Web Scraping
Arslan
ArslanOct 7, 2026

Compare the best web scraping MCP servers in 2026 for scraping, crawling, browser automation, structured data, scale, context cost, and self-hosting.

Best Web Scraping MCP Servers in 2026: Top 10 Compared

Built-in web fetch in Claude, Cursor and similar clients often fails on JavaScript-heavy pages or floods the context window with raw HTML. A web scraping MCP server fixes this with dedicated tools, using the Model Context Protocol (MCP), the open standard AI apps use to call tools. According to MCP's 2026-07-28 release notes, "Across our Tier 1 SDKs, we're seeing close to half-a-billion downloads a month, with both TypeScript and Python SDKs crossing the 1 billion total downloads threshold."

There is no single best web scraping MCP server. Pick the type of server your job needs first, then compare options on context cost, block-page handling, scale and setup.

TL;DR

  • Reading and extracting pages at scale: Olostep or Firecrawl, two managed scraping-API servers with batch and crawl tools.
  • Interactive browser tasks: Playwright MCP, which clicks, types and returns accessibility snapshots.
  • Structured data from big sites: Bright Data or Apify, which ship prebuilt extractors and Actors.
  • Free and self-hosted: Scrapling or Crawl4AI, which run on your own hardware.
  • Quick and simple: Fetch, a one-tool reference server for static pages.

What Is a Web Scraping MCP Server?

A web scraping MCP server is a program that gives an AI client web tools, such as scrape, crawl, search or screenshot, through the Model Context Protocol. The model calls those tools and gets back Markdown, JSON or page snapshots.

The host is the AI app you use, such as Claude Code, and the client is the connector it runs for each server. Our explainer on what an MCP server is covers these roles in more depth.

A tool call follows the same steps on every server:

  • Listing: When the client connects, the server sends its tool definitions. A tool definition is the tool's name, a plain-language description and a JSON schema for its inputs.
  • Loading: The host places those definitions in the model's context window. That is how the model knows which tools it can call.
  • Calling: The model picks a tool and fills in inputs, such as a URL and an output format. The server runs the request and returns the result as text.
  • Cost: Every definition and every result uses tokens. A server with dozens of tools uses more context before the first page is read than a server with a handful.

Local (stdio) vs. Hosted (Streamable HTTP) Servers

MCP servers connect over a transport, the channel that carries messages between client and server. The current MCP transport specification defines the two standard options as "stdio: newline-delimited messages over the standard streams of a client-launched subprocess" and "Streamable HTTP: each message is an HTTP POST to a single MCP endpoint."

The choice comes down to four trade-offs:

  • Local stdio: The client starts the server as a subprocess on your machine. Your API key sits in an environment variable, and you need the runtime, such as Node 18+ for npx packages.
  • Hosted Streamable HTTP: The server runs at a remote URL, so there is nothing to install. The key travels in a request header or, on some servers, in the URL.
  • Corporate networks: A proxy or firewall can block a remote endpoint. Olostep's docs recommend switching to stdio when that happens.
  • Maintenance: The vendor updates a hosted server for you. A local server changes only when you update the package.

Some older servers still use HTTP with Server-Sent Events (SSE), an earlier remote transport. The MCP maintainers' release post confirms that HTTP+SSE is officially deprecated, stating, "The legacy HTTP+SSE transport is also considered to be officially deprecated, with a year-long offramp." An SSE-only server means a migration you will have to make later.

The Three Types of Web Scraping MCP Servers

Web scraping MCP servers fall into three types, and the type decides what your agent can do. Pick the type before comparing servers.

TypeWhat it returnsStrengthsWeak spotsExamples
Browser controlAccessibility snapshots, screenshots, extracted textClicks, types, fills forms, keeps sessionsCosts more tokens for plain reading; someone has to run the browserPlaywright MCP, Browserbase
Managed scraping APIMarkdown, HTML, JSON, parser outputVendor handles rendering and infrastructure; batch and crawl jobsPaid after the free tier; little or no interactionOlostep, Firecrawl, Bright Data, Apify, Scrapfly, Oxylabs
Self-hosted libraryMarkdown, HTML, screenshotsFree and open source; runs on your hardwareYou manage browsers, IP addresses and updatesCrawl4AI, Scrapling, Fetch

An accessibility snapshot is a text outline of a page's elements, such as buttons, links and headings, with their labels. Some servers cross types: Firecrawl has an interact tool, and Scrapfly and Bright Data include browser tools. If you only need search results (links and snippets) and not page content, our guide to the best web search MCP servers covers that job.

How We Evaluated These Servers

We judged each server on five criteria, using only official docs, repos and pricing pages checked on October 7, 2026. No vendor paid for placement.

Olostep publishes this article, and its own server is held to the same criteria. We ran no speed or accuracy benchmarks, so none are claimed.

Output Format and Context Cost

Output format decides how many tokens each page costs. Raw HTML carries scripts and navigation the model must read past, while Markdown or schema JSON keeps only text and fields. Olostep's glossary compares the formats AI models read best for different jobs.

Tool definitions add a second cost, because they load into context before any page is read. According to Anthropic's engineering team, "In cases where agents are connected to thousands of tools, they'll need to process hundreds of thousands of tokens before reading a request."

Tool counts vary widely across scraping servers, as stated in each vendor's docs and checked October 2026. Bright Data exposes 69 tools in total, HasData exposes 57 by default, and Olostep exposes 10.

Large results need the same care. They should come back in pages or as async jobs, so one call does not fill the context window.

Rendering and Block-Page Handling

JavaScript-heavy pages need a rendered DOM, the page structure after scripts run, before their content exists. A plain HTTP fetch of such a page can return an empty shell.

Many sites also push back on automated traffic. According to Thales' 2026 bad bot research, "Bots now account for 53% of all internet traffic. Bad bots alone account for 40%, rising 3% from last year."

The most damaging failure is a response that looks successful but contains a challenge page, such as "Just a moment" or a CAPTCHA prompt. The agent then summarizes that page as if it were real content.

We scored servers on rendering publicly accessible pages reliably and helping you spot blocks. Olostep's terms prohibit bypassing access controls or CAPTCHAs without authorization, so evasion is not a criterion here.

Scale: Batch and Crawl Jobs

Single-URL tools are fine for a few pages. At hundreds or thousands of URLs, you need async batch or crawl tools.

An async tool starts the job and returns a job ID right away. The agent then polls, meaning it calls a results tool again until the job finishes, and it receives results in pages.

This keeps the full result set out of any single tool response. Among the servers reviewed, Olostep, Firecrawl, Crawl4AI and Oxylabs list crawl tools, while Fetch handles one URL per call.

Setup, Authentication, and Maintenance

Send your API key in a request header instead of the URL query string. Query strings such as ?key= or ?token= get written to logs and copied into shared config files.

Bright Data's hosted URL carries a token parameter, and Scrapfly's carries a key parameter by default. Scrapfly's docs allow an Authorization header instead, and Olostep's hosted endpoint uses an Authorization: Bearer header.

Check the repo's status before you install. The Puppeteer reference server now sits in the archived servers repository, and Browserbase marks its self-host repo as archived and recommends its hosted server.

Security and Cost

Treat every scraped page as untrusted input. A page can carry prompt injection, meaning hidden text that tries to give the model new instructions.

Local fetch servers add a second risk. The Fetch reference server can reach internal IP addresses, which creates SSRF (server-side request forgery) risk. In SSRF, a request gets steered at private systems.

In one academic scan of MCP servers, the researchers report, "Additionally, 7.2% of servers contain general vulnerabilities and 5.5% exhibit MCP-specific tool poisoning." Tool poisoning means harmful instructions hidden inside a tool's own definition.

For cost, compare the free tier and the price after it. The Olostep pricing plans page lists 500 free requests with no card, then Starter at $9 per month for 5,000 requests (checked October 2026). Olostep's docs say a standard scrape costs 1 credit and failed requests are not charged.

Best Web Scraping MCP Servers Compared

The table applies the same columns to every server on the shortlist. All counts, prices and free tiers are as stated on each vendor's pricing page or docs, checked October 2026.

ServerTypeHostingDefault toolsOutputFree tier (as stated)Best for
OlostepManaged scraping APIHosted Streamable HTTP (Bearer header), local stdio, Docker10Markdown, HTML, JSON, text; parser JSON; cited answers500 requests, no cardScrape, batch, crawl and search in one server
FirecrawlManaged scraping APIHosted (keyless tier limited), local npxNot counted; includes scrape, map, crawl, search, parse, interact, agent, monitorMarkdown, schema JSON, screenshots1,000 credits/month, no cardBroad toolset with light interaction
Bright DataManaged scraping APIHosted (token in URL), local npx69 (filter with GROUPS or TOOLS)Markdown, HTML, structured JSON5,000 requests/month, no cardStructured data from big sites
ApifyManaged scraping API (Actors)Hosted Streamable HTTP, local stdioVaries; Actor search and call, docs, RAG web browser, web fetchActor dataset JSON, Markdown$5 of usage/monthPrebuilt scrapers for specific sites
Playwright MCPBrowser controlLocal stdio (optional HTTP port, Docker)Not counted; includes navigate, snapshot, click, typeAccessibility snapshots, screenshots, PDFsFree, open source (Apache-2.0)Interactive browser tasks
BrowserbaseBrowser controlHosted Streamable HTTP (recommended); self-host repo archived6Extracted data and textNot stated on pages checkedInteraction without a local browser
ScrapflyManaged scraping APIHosted (key in URL by default, header supported), local CLI10 (per docs)Markdown, text, JSON, clean HTML, raw; screenshots1,000 credits, no cardScraping with block detection
OxylabsManaged scraping APIHosted (Basic auth), local uvx10Markdown, HTML, links; JSON and CSV via AI Studio1-week Web Scraper API trial; 1,000 AI Studio creditsGoogle and Amazon data plus AI crawling
Crawl4AISelf-hosted librarySelf-hosted Docker; MCP over SSE or WebSocket7Markdown, HTML, screenshots, PDFsFree, open sourceFree crawling on your hardware
ScraplingSelf-hosted libraryLocal stdio, Streamable HTTP (auth token), Docker13Markdown, text, HTML; prompt-injection content strippedFree, open source (BSD-3)Free scraping with sessions
FetchSelf-hosted libraryLocal (uvx, pip, Docker)1Markdown from HTML, or rawFree, open source (MIT)Quick reads of static pages

How the servers compare on the criteria above:

  • Tool footprint: Fetch has 1 tool, Browserbase has 6, and Olostep, Scrapfly and Oxylabs have 10 each. Bright Data's 69 tools need filtering to stay lean.
  • Interaction: Playwright MCP and Browserbase are built for clicking and typing. Firecrawl, Scrapfly and Bright Data add browser tools to scraping, and Olostep has none.
  • Scale: Olostep's batch tool takes up to 10,000 URLs per job, though new accounts start at 100 per batch. Olostep, Firecrawl, Crawl4AI and Oxylabs list crawl tools.
  • Transport risk: Crawl4AI's MCP endpoint uses deprecated SSE, and Browserbase's self-host repo is archived.

The Shortlist: 10 Web Scraping MCP Servers Reviewed

Each review covers what the server is, its main tools, hosting and install, limits, and the job it fits. Vendor numbers are as stated in their docs or pricing pages, checked October 2026.

1. Olostep MCP Server: Scrape, Crawl, Batch, and Search in One Hosted Server

Olostep's MCP server connects MCP clients to Olostep's web data API. Its docs list 10 tools: scrape_website, get_webpage_content, search_web, answers, batch_scrape_urls, get_batch_results, create_crawl, get_crawl_results, create_map and get_website_urls.

scrape_website returns Markdown, HTML, JSON or text, and its parser option returns structured JSON, for example with @olostep/amazon-product. Olostep's docs say it "returns LLM-ready markdown by default with automatic boilerplate removal." The answers tool returns cited answers, and search_web returns search results as JSON.

batch_scrape_urls handles batch scraping thousands of URLs: it accepts 2 to 10,000 URLs per job (Olostep's batch docs note new accounts start at 100 items per batch until the limit is raised) and returns a batch_id. The agent polls get_batch_results until the job reads completed.

create_crawl returns a crawl_id for crawling an entire website, with URL patterns such as /blog/**. get_crawl_results pages through results, up to 100 items per call.

The hosted endpoint is https://mcp.olostep.com/mcp with an Authorization: Bearer header, so the key stays out of the URL. Locally, npx -y olostep-mcp runs over stdio with OLOSTEP_API_KEY and Node 18+, and a Docker image (olostep/mcp-server) exists. Install steps for each client are on the Olostep MCP server setup page.

Olostep has no click or type browser control, so it cannot fill forms or step through a UI. It also has no dedicated block-check tool. The trial includes 500 free requests, then Starter costs $9 per month, as stated on its pricing page, checked October 2026.

Best for: agents that need scraping, search, batch jobs and crawls from one hosted server.

2. Firecrawl MCP Server: Broad Toolset With a Keyless Hosted Tier

Firecrawl's MCP server connects agents to Firecrawl's scraping API. Tools include scrape, map, crawl, search, parse, interact, agent and monitor, with Markdown, schema JSON and screenshot output.

The hosted endpoint is https://mcp.firecrawl.dev/v2/mcp, with a rate-limited keyless tier that covers only scrape, search and parse. The local option is npx -y firecrawl-mcp with FIRECRAWL_API_KEY.

The free plan is 1,000 credits per month with no card, as stated on its pricing page, checked October 2026. The longer tool list costs more context than a lean server, and the free plan allows 10 scrape requests per minute.

Best for: teams that want a broad toolset, including light interaction, from one vendor.

3. Bright Data MCP: Largest Catalog of Structured Extractors

Bright Data's MCP server connects agents to Bright Data's scraping infrastructure. It has 69 tools in total, including search_engine, scrape_as_markdown, browser tools and 45 web_data_* structured extractors.

Use the hosted URL, which carries a token query parameter, or run npx @brightdata/mcp locally with API_TOKEN. Set the GROUPS or TOOLS variables to load only the tools you need.

The free tier is 5,000 requests per month with no card, then $1.50 per 1,000 results, as stated on its pricing page, checked October 2026. With all 69 tools loaded, it has the largest tool list in this review.

Best for: structured data from big sites through prebuilt extractors.

4. Apify MCP Server: Thousands of Prebuilt Actors as Tools

Apify's MCP server turns Apify Actors, its prebuilt cloud scrapers, into tools the agent can find and run. Default tools include search-actors, fetch-actor-details, call-actor, docs tools and the RAG web browser, and you can add any Actor as its own tool.

The hosted server at https://mcp.apify.com uses Streamable HTTP, and Apify has removed SSE support. The local option is npx @apify/actors-mcp-server with APIFY_TOKEN.

The Free plan includes $5 of usage per month, as stated on its pricing page, checked October 2026. Output quality varies from Actor to Actor, so test the one you pick on your target pages.

Best for: site-specific scraping where a prebuilt Actor already exists.

5. Playwright MCP: Best for Interactive Browser Tasks

Playwright MCP is Microsoft's open-source browser-control server, licensed under Apache-2.0. Tools such as browser_navigate, browser_click and browser_type drive a real browser, and browser_snapshot returns an accessibility snapshot instead of HTML.

Install it with npx @playwright/mcp@latest. It runs locally over stdio, with an optional HTTP port and a Docker image limited to headless Chromium.

The browser runs on your machine, with no managed handling for sites that push back on automation. For plain reading, it costs more tokens than a Markdown scrape.

Best for: interactive browser tasks such as forms and multi-step UI flows.

6. Browserbase MCP: Hosted Browsers With Stagehand Actions

Browserbase's MCP server drives hosted cloud browsers with Stagehand actions. Its six tools cover starting and ending a session plus navigate, act, observe and extract.

Browserbase recommends the hosted endpoint at https://mcp.browserbase.com/mcp. The self-host repo, which ran npx @browserbasehq/mcp over stdio, is marked archived.

The pages we checked do not state a free tier. Like Playwright MCP, it is a heavier choice for plain reading.

Best for: interactive tasks when you do not want to run a local browser.

7. Scrapfly MCP Cloud: Scraping Plus Block Detection

Scrapfly MCP Cloud is a hosted server on top of Scrapfly's scraping API. Its docs list 10 tools, including web_get_page, web_scrape, screenshot, check_if_blocked and cloud browser session tools, though the product page still says 5.

The endpoint is https://mcp.scrapfly.io/mcp with the key passed as ?key= by default. Scrapfly's docs also allow an Authorization header or OAuth2, which keeps the key out of the URL.

The free tier is 1,000 credits with no card, as stated on its pricing page, checked October 2026. check_if_blocked targets the challenge-page problem, which most servers here leave to you.

Best for: scraping where you want a built-in block check.

8. Oxylabs MCP: Scraper APIs and AI Studio Tools

Oxylabs' MCP server exposes Oxylabs' Scraper APIs and AI Studio tools. Its 10 tools include universal_scraper, google_search_scraper, amazon_product_scraper, ai_crawler and ai_browser_agent.

The hosted endpoint https://mcp.oxylabs.io/mcp uses Basic auth plus an AI Studio header, and uvx oxylabs-mcp runs it locally. Output is Markdown, HTML or links, plus JSON, CSV and other formats through AI Studio.

Oxylabs lists a 1-week Web Scraper API trial and 1,000 free AI Studio credits, as stated on its site, checked October 2026. Setup needs credentials for both products.

Best for: Google and Amazon data plus AI-driven crawling.

9. Crawl4AI and Scrapling: Free, Self-Hosted Options

Crawl4AI and Scrapling are open-source scrapers you run on your own hardware. You manage the browsers, IP addresses and updates yourself.

Crawl4AI runs as a Docker server (unclecode/crawl4ai) with tools such as md, html, screenshot, pdf, execute_js and crawl. Its MCP endpoint uses SSE at /mcp/sse or WebSocket, so plan for the SSE deprecation covered earlier.

Scrapling installs with pip install "scrapling[ai]" and scrapling install, then runs as scrapling-mcp over stdio. Its 13 tools include stealthy_fetch, bulk fetch tools and session tools, and it strips prompt-injection content from output. HTTP mode requires an auth token, and the project is licensed BSD-3.

Best for: $0 scraping and crawling when you can run the infrastructure.

10. Fetch Reference Server: Simplest Baseline

Fetch is described as a reference implementation by the MCP project. It has one tool, fetch, which converts HTML to Markdown and has a default max_length of 5,000.

Install it with uvx mcp-server-fetch, pip or the mcp/fetch Docker image. It does not render JavaScript, and it can reach internal IP addresses, so keep it away from private services.

The Puppeteer reference server, still recommended in some guides, has moved to the archived servers repository. Use Playwright MCP for new browser setups.

Best for: quick reads of static pages during development.

Which Web Scraping MCP Server Should You Use?

Match the server to the job, then confirm it fits your budget and network. The Why column gives the deciding factor for each pick.

If your job is…Use…Why
Read docs or articles inside an IDEOlostep (get_webpage_content) or Fetch for static pagesOne Markdown result per page keeps context small
Scrape 1,000+ URLs for a datasetOlostep (batch_scrape_urls)Async job ID and polling for up to 10,000 URLs per job (new accounts start at 100)
Crawl a whole docs site into RAGOlostep (create_crawl) or Firecrawl; Crawl4AI to self-hostCrawl tools follow links and return results in pages
Fill forms and click through a UIPlaywright MCP (local) or Browserbase (hosted)Browser control with click, type and snapshots
Get structured product dataBright Data, Apify or Olostep parsersPrebuilt extractors return JSON fields
Spend $0 and self-hostScrapling or Crawl4AIOpen source, runs on your hardware
Get sourced answersOlostep (answers)Returns an answer with source citations

How to Install a Web Scraping MCP Server in Claude Code and Cursor

Olostep's hosted server is the worked example. Other servers follow the same steps with their own endpoint, package and key.

  1. Get an API key. Create one on the API keys page of the Olostep dashboard. The trial needs no card.
  2. Restart the client. Fully quit and reopen it so it reloads the tool list.
  3. Verify the tools. In Claude Code, run claude mcp list; in Cursor, open the MCP settings. You should see the 10 Olostep tools.
  4. Run a test scrape. Ask: Use scrape_website to get https://example.com as Markdown and quote the title. A working setup returns the title "Example Domain."

Add the server. In Claude Code, run this command:

bash
claude mcp add --transport http olostep https://mcp.olostep.com/mcp --header "Authorization: Bearer YOUR_API_KEY"

In Cursor, add this to .cursor/mcp.json in your project:

json
{
  "mcpServers": {
    "olostep": {
      "url": "https://mcp.olostep.com/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_API_KEY"
      }
    }
  }
}

If a proxy blocks the hosted endpoint, run the server locally over stdio. In Claude Code, use this command:

bash
claude mcp add --transport stdio --env OLOSTEP_API_KEY=YOUR_API_KEY olostep -- npx -y olostep-mcp

In clients that use a JSON config, use this block:

json
{
  "mcpServers": {
    "olostep": {
      "command": "npx",
      "args": ["-y", "olostep-mcp"],
      "env": {
        "OLOSTEP_API_KEY": "YOUR_API_KEY"
      }
    }
  }
}

If the setup fails, check these causes first:

  • Connected, 0 tools: This usually means a bad key or the wrong auth mode. Olostep's docs say "wrong auth mode is the #1 onboarding error": hosted needs the Bearer header, stdio needs the variable.
  • 401 Missing Authorization: The header is missing or misspelled. Check the Bearer prefix and the key.
  • Endpoint blocked: A corporate proxy or DNS filter blocks mcp.olostep.com. Switch to the stdio setup above.
  • npx not found: Install Node 18 or later. Then restart the client.
  • Windows: Run the command as cmd /c npx. In JSON, set "command" to "cmd" and put "/c", "npx" first in args.
  • Old tool list: The client cached an earlier list, so fully relaunch it.

How to Keep Scraped Content Reliable Inside an Agent

A scrape can succeed and still return the wrong page. Add these checks to your agent's instructions or code:

  • Check for challenge pages: Look for "Just a moment", CAPTCHA text or very short content before trusting output. Retry or flag the URL instead of summarizing it.
  • Ask for the title and URL: Have the agent quote the page title and the URL it read. A mismatch points to a redirect or block page.
  • Use async jobs for large sets: Send hundreds of URLs to a batch or crawl tool and poll for results. Each response stays small.
  • Treat page text as untrusted: Tell the agent that instructions inside scraped pages are content to report, not commands to follow. Review any action it takes after reading a page.
  • Respect site rules: Follow each site's terms, robots rules and rate limits. Scrape only pages you are allowed to access.

Frequently Asked Questions

Each answer below draws on the vendor docs and repos reviewed in October 2026.

What Is the Best MCP Server for Web Scraping?

It depends on the job: for reading and extracting public pages at scale, a managed scraping-API server such as Olostep or Firecrawl fits best. For clicking through pages and filling forms, use Playwright MCP.

Is There a Free Web Scraping MCP Server?

Fetch, Playwright MCP, Crawl4AI and Scrapling are free and open source, though you run them on your own hardware. Most hosted servers also have free tiers, such as Olostep's 500 free requests and Firecrawl's 1,000 monthly credits.

Can I Use a Web Scraping MCP Server With Claude Code, Cursor, or VS Code?

Yes. Claude Code, Cursor and VS Code all support MCP servers, through local stdio servers or remote Streamable HTTP servers.

What Is the Difference Between Playwright MCP and a Scraping-API MCP Server?

Playwright MCP drives a browser on your machine so the agent can click, type and read accessibility snapshots. A scraping-API server returns cleaned page content, such as Markdown or JSON, from the vendor's managed infrastructure.

Why Does My MCP Server Fail on Pages That Load in My Browser?

The page probably needs JavaScript rendering, which a plain fetch skips, or the site served a bot challenge to automated traffic. Use a server that renders pages, and check the output for challenge-page text.

Can I Run More Than One MCP Server at Once?

Yes, clients can connect to several servers at once. Every server's tool definitions add to the context, so enable only the servers and tools the task needs.

Can an MCP Server Scrape Pages Behind a Login?

Servers with browser sessions, such as Browserbase or Scrapfly's cloud browser tools, can reach login pages when you are authorized to use the account. Follow the site's terms, and keep credentials out of prompts and shared configs.

Is Puppeteer MCP Still Maintained?

No. The Puppeteer reference server was moved to the archived servers repository, so use Playwright MCP for new browser setups.
e Playwright MCP for new browser setups.

About the Author

Arslan Ali

Co-Founder, Olostep · San Francisco, CA

Arslan is the co-founder of Olostep, a web data infrastructure platform that helps developers and teams access, extract, and structure web data at scale. He works closely on the product and technology behind Olostep, with a focus on building reliable infrastructure for web scraping, search APIs, and structured web data.

On this page

Read more