What is web crawling?
Web crawling is the automated process of discovering and retrieving pages from the web.
A crawler starts with one or more URLs, downloads a page, identifies links on that page, decides which links it is allowed or configured to visit, and repeats the process with newly discovered URLs.
This is how search engines discover much of the web. Google describes Search as using automated programs called web crawlers to find pages and download content before indexing it.
Crawling is also used outside search engines. A developer might crawl a documentation site to build a RAG knowledge base, an SEO platform might crawl a domain to inspect its internal structure, or a migration tool might collect the pages that need to be moved into a new CMS.
The defining feature is discovery: the crawler does not need the complete list of URLs before it starts.
How web crawling works
A basic crawler repeats the same process until it reaches its configured limits.
- Start with a seed URL. The crawler receives a starting address such as https://example.com/docs.
- Fetch the page. It sends a request for the page. If the site depends on client-side JavaScript, the crawler may also need to render the page in a browser environment.
- Extract links. It finds links to other pages in the retrieved document.
- Filter and queue URLs. It checks each discovered URL against rules such as domain restrictions, allowed paths, robots.txt, crawl depth, and previously visited URLs.
- Repeat. Eligible URLs are fetched, their links are discovered, and the process continues until the queue is empty or a page, depth, time, or other crawl limit is reached.
The set of URLs waiting to be crawled is commonly called the crawl frontier.
A crawler also needs to remember what it has already processed. Without deduplication, it could request the same page repeatedly through slightly different URLs.
For example:
https://example.com/product?utm_source=email
Those URLs may point to the same underlying content. A crawler needs normalization and filtering rules to decide whether they should be treated as separate resources.
A simple web crawling example
Suppose you want every page in a documentation section, but you only know its starting URL:
The first page contains links to:
/docs/getting-started
/docs/api
/pricing
/blog
If the crawl is restricted to /docs/, the crawler adds the first two URLs to its queue and ignores the others.
It visits /docs/api and discovers:
/docs/api/authentication
/docs/api/search
Those pages are added to the queue as well.
The crawl continues until there are no more eligible documentation pages or until a configured limit is reached.
You supplied one URL. The crawler discovered the rest.
Web crawling vs. web scraping
Crawling and scraping are related, but they describe different parts of a web-data workflow.
| Web crawling | Web scraping | |
|---|---|---|
| Main purpose | Discover and traverse pages | Extract information from pages |
| Typical input | One or more starting URLs | Known target URL or URLs |
| Follows links | Usually | Not required |
| Typical output | Discovered pages and their content | HTML, text, Markdown, or structured fields |
| Example | Find every product page on a store | Extract the price from a product page |
If you already have 5,000 URLs and need data from each of them, discovery is not your main problem. That is primarily a scraping or batch-processing task.
If you only know the site's homepage and need to find the relevant pages first, crawling is the appropriate operation.
A system can do both. A crawler can discover product pages, then an extraction step can pull the product name, SKU, price, description, and availability from each page.
Web crawling vs. indexing
Crawling retrieves pages. Indexing processes and stores information from those pages so it can be retrieved later.
Google separates these stages in its description of how Search works. During crawling, Google discovers and downloads pages. During indexing, it analyzes the retrieved content, including text, titles, images, videos, and canonical information.
A simplified search workflow is:
Discover URL → crawl page → process content → add eligible information to an index → retrieve it for searches
Being crawled does not mean a page will necessarily be indexed.
Crawling vs. website mapping
Sometimes you need the URLs on a site without needing the full content of every page.
That is closer to website mapping.
For example:
Mapping request: Find the URLs under example.com/docs.
Crawling request: Find those URLs, visit them, and retrieve their content.
Separating these operations can save requests when your application only needs URL discovery.
Olostep exposes these as separate operations: Maps for URL discovery and Crawls for recursively processing pages from a starting URL.
What is crawl depth?
Crawl depth describes how many link steps a page is from the starting page.
If the crawler starts at:
that page is depth 0.
A page linked directly from it, such as:
is depth 1.
A page reached from the documentation page:
is depth 2.
Crawl depth is based on the path through links, not the number of folders visible in the URL.
A deeply nested URL can still be depth 1 if the starting page links directly to it.
Setting a maximum depth is one way to stop a crawler from moving too far away from the part of a site that matters to the job.
How robots.txt affects web crawling
robots.txt is a standardized mechanism that site owners use to provide instructions to crawlers about which URL paths they may access.
The file normally lives at:
https://example.com/robots.txt
RFC 9309 defines crawlers as automated clients and specifies the Robots Exclusion Protocol used to communicate crawling rules. It also states that these rules are not a form of access authorization.
That distinction matters.
A robots.txt rule can tell compliant crawlers not to request a path. It does not password-protect the resource.
For search engines, crawling and indexing controls are also different. Google says a URL blocked through robots.txt can still appear in search results if Google discovers the URL elsewhere, even though Google cannot crawl the page to read its content.
If the goal is to prevent a page from appearing in Google Search, Google recommends controls such as noindex or authentication rather than using robots.txt for that purpose.
Why web crawling becomes difficult at scale
Writing a crawler that follows a few HTML links is straightforward. Crawling real websites reliably introduces several engineering problems.
JavaScript-rendered pages
Some sites return little useful content in the initial HTML response. The browser loads the rest after JavaScript executes.
Google itself renders pages and runs JavaScript during its crawling and indexing process because some website content would otherwise be unavailable to the crawler.
A crawler that needs the rendered version of a page must operate browser infrastructure in addition to making HTTP requests.
That increases processing time and resource usage.
Duplicate URL spaces
Tracking parameters, faceted navigation, search pages, filters, session parameters, and pagination can produce large numbers of URLs.
For example:
/shoes?color=black
/shoes?color=black&sort=price
/shoes?sort=price&color=black
A crawler needs rules for normalization and deduplication or it can repeatedly retrieve equivalent pages.
Google's crawler documentation specifically discusses managing duplicate and faceted-navigation URLs because they can consume unnecessary crawling resources.
Crawl traps
Some URL structures have no natural end.
A calendar could generate:
/calendar/2026
/calendar/2027
/calendar/2028
and continue indefinitely.
Search-result pages, combinations of filters, dynamically generated pagination, and session URLs can create similar spaces.
Page limits, depth limits, URL-pattern restrictions, and deduplication protect a crawler from spending its entire job inside one such pattern.
Request rates and retries
A crawler can send many requests to one server.
It therefore needs controls for concurrency, rate limits, failed requests, redirects, server errors, timeouts, and retries.
At larger scale, multiple workers may process URLs simultaneously. They need shared state so two workers do not continually pick up the same page.
This is where a crawler becomes infrastructure rather than a simple loop over links.
What web crawling is used for
Search engines are the most familiar example. Googlebot discovers URLs primarily from links on pages it has already crawled and then retrieves those pages for Google's search-processing systems.
AI applications use crawling for a different reason: they need usable website content. A RAG application might crawl a company's documentation, retrieve each page, convert the pages into clean text or Markdown, create embeddings, and store them in a retrieval system.
SEO crawlers collect page-level and site-level information such as status codes, links, redirects, titles, canonical elements, crawlability signals, and website structure.
Website migrations use crawling to create an inventory of existing pages before URLs or content are moved. The same site can be crawled again after deployment to identify missing pages or unexpected redirects.
Web-data products can also use crawling to collect public product pages, documentation, company information, or other datasets where the complete set of URLs is not known in advance.
What is a web crawling API?
A web crawling API lets an application start and control a crawl without operating the crawling infrastructure itself.
Instead of building the URL queue, browser workers, retry system, deduplication logic, job state, and concurrency management internally, the application sends a crawl configuration to an API.
The result might contain discovered URLs, page metadata, extracted content, or identifiers that can be used to retrieve the content later.
The exact behavior depends on the provider, so a crawling API should be evaluated on what it actually returns, how it handles JavaScript, how crawl boundaries are configured, and how usage is billed.
How to crawl a website with Olostep
Olostep's Crawls API starts from one URL and recursively follows eligible links across the site.
A crawl is created through:
POST https://api.olostep.com/v1/crawls
A minimal request can define the starting URL, the number of pages to process, and which URL patterns are in scope:
{
"start_url": "https://example.com/docs",
"max_pages": 100,
"include_urls": ["/docs/**"]
}
The crawl runs asynchronously. The application receives a crawl ID, checks the crawl's status, and then lists the pages discovered by the job.
Olostep also exposes controls for max_depth, include_urls, and exclude_urls. Crawls respect robots.txt by default, and the crawler can render JavaScript-dependent pages.
Each successfully crawled page can return a retrieve_id. That identifier can be used to retrieve the page's full HTML or Markdown content.
This creates a workflow such as:
Start URL → discover pages → crawl pages → retrieve page content → send the content to your application
For a documentation ingestion pipeline, for example, the starting URL might be /docs, the crawl could be restricted to /docs/**, and the resulting pages could be retrieved as Markdown for downstream indexing.
Olostep currently gives new accounts 500 free requests. Crawl usage is based on successfully processed pages: a crawl that successfully processes 50 pages consumes 50 requests. Failed pages are not counted as successful requests.
When should you use crawling?
Use crawling when you know where the relevant website or section starts but do not already have every URL you need.
If you know the exact URL and only need its content, use a scrape.
If you only need to discover the URLs on a domain, use a mapping operation.
If you already have a large URL list, process that list directly rather than rediscovering the same URLs through a crawl.
The difference is the input you have before the job starts:
Known page → scrape it
Known URL list → process the list
Known domain or section, unknown pages → crawl it
That distinction keeps the workflow smaller and avoids spending crawl requests on discovery that your application has already completed.
Frequently asked questions
What is web crawling in simple terms?
Web crawling means automatically visiting web pages and following their links to discover additional pages.
A crawler can start with one URL and progressively find hundreds or thousands of related URLs without receiving the complete URL list in advance.
What is a web crawler?
A web crawler is software that automatically retrieves web resources and follows links according to a set of rules. Search-engine crawlers such as Googlebot are one example, but crawlers are also used for website analysis, data collection, RAG ingestion, and migrations.
Is web crawling the same as web scraping?
No. Crawling is primarily concerned with finding and traversing pages. Scraping is concerned with extracting information from a page.
They are frequently combined in the same pipeline.
Can web crawlers process JavaScript?
Some can.
A simple crawler may only fetch the server's HTML response. A browser-based crawler can execute JavaScript and inspect the rendered page. JavaScript rendering is necessary when links or content are created only after scripts run.
Do web crawlers follow robots.txt?
Compliant crawlers can use robots.txt to determine which URL paths the site owner has allowed or disallowed for that crawler.
Olostep's crawler respects robots.txt by default.
What is the difference between crawling and indexing?
Crawling retrieves pages. Indexing analyzes and stores information from those pages so that it can later be searched or retrieved.
Search engines normally crawl pages before deciding whether and how to index them.
Can I crawl an entire website?
A crawler can recursively traverse eligible pages across a site, but the actual coverage depends on accessible links, crawl restrictions, robots.txt, authentication, JavaScript behavior, duplicate URLs, page limits, and crawl depth.
For large sites, setting explicit boundaries is usually preferable to asking a crawler to follow every URL it can find.
Ready to get started?
Start using the Olostep API to implement what is web crawling? in your application.