Is there a scraper that can navigate subpages and find all the links for me?

Yes. What you need is usually called a web crawler or recursive scraper.

Instead of scraping only the URL you provide, a crawler starts from that page, finds internal links, visits those subpages, discovers more links, and continues until it reaches the limits you set.

With Olostep, there are two ways to do this depending on what you actually need:

  • Use Maps when you want a list of discoverable URLs from a website.
  • Use Crawls when you want Olostep to navigate those subpages and collect their content as well.

Olostep also returns links_on_page when scraping an individual URL, so a normal Scrape works when you only need the links present on one known page.

A scraper and a crawler are not quite the same thing

The distinction matters when your starting point is a single URL.

A basic scraper receives something like:

https://example.com

It loads that page and extracts what is present there: HTML, text, structured fields, or links.

Suppose the homepage contains these links:

https://example.com/pricing
https://example.com/docs
https://example.com/blog

A one-page scraper can return those URLs, but its job may end there.

A crawler adds another step. It puts the discovered internal URLs into a queue, visits them, extracts the links found on each page, adds previously unseen URLs to the queue, and repeats.

The process looks roughly like this:

Homepage
├── /pricing
├── /docs
│   ├── /docs/api
│   ├── /docs/authentication
│   └── /docs/webhooks
└── /blog
    ├── /blog/web-scraping
    └── /blog/web-crawling

That recursive navigation is what you want when the request is: start with this website and find the pages underneath it for me.

This is the standard crawler model: begin with a seed page, extract links, visit eligible links, and repeat until no new URLs remain or a configured boundary is reached.

Which Olostep endpoint should you use?

The answer depends on whether you need URLs, page content, or links from only one page.

What you needOlostep endpointWhat it returns
Links found on one known pageScrapesPage content plus links_on_page
Discoverable URLs across a websiteMapsWebsite URL inventory
Navigate subpages and collect their contentCrawlsCrawled pages plus IDs for retrieving their content

Olostep separates these operations so you do not have to crawl and process hundreds of pages when all you wanted was the URL list.

If you only want all the URLs, use Maps

Imagine you are building an SEO auditing tool. You receive:

https://example.com

Your first task is not necessarily to download the content of every page. You may simply need to determine what pages exist.

Olostep's Maps endpoint is built for that job.

It discovers URLs using website signals including sitemaps and links, and lets you control the scope with include and exclude patterns. Large URL inventories can be paginated rather than forcing the entire result into one response.

A simple Python request looks like this:

from olostep import Olostep

client = Olostep(api_key="YOUR_API_KEY")

site_map = client.maps.create(
    url="https://example.com",
    include_urls=["/**"],
    top_n=1000,
)

for url in site_map.urls():
    print(url)

You could narrow discovery to a section of the website instead:

site_map = client.maps.create(
    url="https://example.com",
    include_urls=["/blog/**"],
)

Now URLs outside /blog/ do not need to become part of the job.

This is useful when your downstream system needs a clean URL inventory before deciding which pages are worth processing.

If you need the subpage content too, use Crawls

Sometimes finding the URL is only half the job.

You may want to:

  • build a RAG dataset from an entire documentation site;
  • collect product pages for analysis;
  • ingest a help center into an AI agent;
  • audit pages across a domain;
  • monitor content across a specific website section.

In that case, use a crawl.

Olostep's Crawl endpoint starts from a URL and recursively follows links across the website. You can restrict the number of pages, maximum crawl depth, URL patterns, subdomains, and whether external links are included. Each discovered page can then be retrieved as content.

For example:

from olostep import Olostep

client = Olostep(api_key="YOUR_API_KEY")

crawl = client.crawls.create(
    start_url="https://example.com",
    max_pages=500,
    max_depth=4,
    include_urls=["/**"],
    exclude_urls=[
        "/account/**",
        "/cart/**"
    ],
    include_external=False,
)

for page in crawl.pages():
    print(page.url)

    content = page.retrieve(["markdown"])
    print(content.markdown_content[:500])

The important difference is that the crawler does not require you to know those 500 URLs beforehand.

You provide the starting point and boundaries. The crawler handles discovery.

How does a crawler navigate from one page to another?

A typical recursive crawl has a few operations happening repeatedly.

The crawler first fetches the starting URL. It parses the rendered page and extracts URLs from links it can discover.

Those URLs are normalized so obvious duplicates are not repeatedly requested. URLs outside the allowed domain or path can be discarded.

New internal URLs are placed into the crawl queue.

The crawler then visits the next URL and runs the same process again.

For example:

1. Visit /
2. Find /products and /blog
3. Visit /products
4. Find /products/a and /products/b
5. Visit /blog
6. Find /blog/article-1
7. Continue until the crawl limit is reached

You do not need to manually write a separate scraping request for every URL encountered.

Yes. That is what crawl depth controls.

Consider this structure:

example.com/
    ↓
example.com/docs/
    ↓
example.com/docs/api/
    ↓
example.com/docs/api/search/

The last URL is several link hops away from the homepage.

A homepage-only link extractor will not see it unless the homepage links directly to that URL.

A recursive crawler can reach it by navigating through the intermediate pages.

Olostep exposes max_depth specifically for bounding how far the crawler should follow the link graph. It also exposes max_pages, so a website with thousands of pages cannot accidentally turn into an unbounded crawl.

Can it handle websites that use JavaScript?

This is where a simple HTML link extractor and a browser-based crawler can produce different results.

Some websites return most of their useful HTML directly from the server. Fetching the document and parsing <a href> elements may be enough.

Other sites generate navigation or page content after JavaScript runs.

Olostep's crawling infrastructure renders pages using a browser before processing them, which allows it to work with JavaScript-rendered websites rather than relying only on the initial HTML response.

There is still a distinction between rendering a page and performing arbitrary user interactions.

If a URL appears only after selecting filters, submitting a form, opening an application state, or clicking a control that does not initially expose a navigable URL, that may require browser actions rather than ordinary recursive link following.

Olostep's Scrape endpoint supports actions including waiting, clicking, filling inputs, and scrolling for those interaction-dependent cases.

No crawler can discover a URL for which there is no discoverable path or signal.

A crawler can find discoverable URLs within its configured scope.

That may include URLs exposed through:

  • internal links;
  • XML sitemaps;
  • rendered page content;
  • subpages found during recursive traversal.

But consider a page such as:

https://example.com/private-old-page

If no page links to it, it is absent from the sitemap, and there is no other discovery signal, a crawler starting from the homepage has no URL to follow.

The same issue can apply to authenticated pages or URLs that only exist after particular application states.

This is one reason Olostep separates mapping from crawling. Maps can use both sitemap and link signals to build the URL inventory instead of relying on the homepage alone.

You should usually limit what the crawler follows

Following every URL without rules can produce a lot of unnecessary requests.

Consider an ecommerce site that generates URLs such as:

/products?sort=price
/products?sort=newest
/products?color=red
/products?color=blue
/products?page=2
/products?page=3

Or a documentation site containing account, authentication, language, search, and version-specific paths.

What looks like one site can quickly become thousands of crawlable URL variants.

A crawler intended for production should therefore let you define boundaries.

Olostep Crawls supports controls including:

max_pages
max_depth
include_urls
exclude_urls
include_external
include_subdomain

It also follows robots.txt by default according to the current Crawl API documentation.

If you only care about documentation, for example, targeting:

/docs/**

is better than crawling an entire company website and later throwing most of the results away.

Maps first, Crawl second can be more efficient

You do not always have to start with a full crawl.

For many web-data pipelines, a useful sequence is:

Website
   ↓
Maps
   ↓
URL inventory
   ↓
Filter relevant URLs
   ↓
Crawl / Scrape / Batch
   ↓
Page content

Suppose a site has 20,000 discoverable URLs but your application only needs /docs/, which contains 1,400 pages.

Mapping first lets you understand the URL space and restrict subsequent processing instead of extracting content from pages your application will discard.

Olostep explicitly supports using mapped URLs as inputs for later scrapes, batches, or crawling workflows.

This separation is particularly useful for AI agents. The agent can first obtain a bounded set of URLs, select the pages relevant to its task, and retrieve content only where necessary.

Then you do not need a site-wide crawl.

Olostep's Scrape endpoint returns links_on_page alongside the requested page output.

For example, you could scrape:

https://example.com/resources

and receive the links exposed on that particular page without recursively following them across the domain.

The decision is therefore straightforward:

One page → Scrape
Whole-site URL discovery → Map
Navigate subpages + retrieve content → Crawl

Where subpage crawling is useful

Building RAG and AI knowledge bases

An AI application trained or grounded on a documentation website cannot rely only on the homepage.

The useful information is normally distributed across documentation routes, guides, reference pages, changelogs, and support content.

A crawler can start from the documentation root and collect those subpages for indexing. Olostep specifically lists RAG knowledge bases and site ingestion among the uses for its Crawl API.

SEO crawls and URL inventories

A crawler can expose which pages are reachable through the site's internal link structure.

A map can provide the broader URL inventory.

That data can then be compared with sitemaps, indexing data, analytics, or your own page database instead of manually collecting URLs.

Competitor and market research

If you are studying a competitor's product documentation, integrations, categories, or resource library, recursive discovery prevents the research process from depending on a manually assembled URL list.

Path filters also let you stay inside the section relevant to the research.

Website migrations

A URL inventory is useful before a migration because existing paths need to be identified before redirect rules and post-launch checks can be prepared.

Olostep lists URL inventory and site migration workflows among the use cases for Maps.

Do you need to build your own recursive scraper?

You can.

The basic implementation is conceptually simple:

queue = [starting_url]
visited = set()

while queue:
    url = queue.pop()

    if url in visited:
        continue

    page = fetch(url)
    visited.add(url)

    links = extract_links(page)

    for link in links:
        if link belongs to target site:
            queue.append(link)

The implementation gets more involved once the target includes JavaScript rendering, redirects, duplicate URL variants, retry logic, concurrency, crawl limits, timeouts, robots rules, and large queues.

If crawling is part of a production application, using an API means your application can work at the job level instead:

start URL → crawl job → discovered pages → retrieve content

Olostep's Crawl API uses that asynchronous job model. A request creates the crawl, the crawl ID tracks its status, and the pages can be listed once they are processed.

FAQ

Yes, but once it recursively follows discovered URLs it is generally functioning as a web crawler. A basic scraper normally processes a URL it has already been given.

Use a website mapping or crawling tool rather than a single-page link extractor. With Olostep, use Maps when you need the URL inventory and Crawls when you also need the content from the discovered pages.

Can I crawl only one section of a website?

Yes. Olostep supports URL include and exclude patterns. For example, /blog/** or /docs/** can be used to limit the crawl to a particular part of the site.

Can a crawler find pages that are not linked from the homepage?

Yes, if it can discover a path to them through deeper internal links or another URL-discovery signal such as a sitemap. A page that has no discoverable link or sitemap entry cannot be guaranteed to appear.

Does Olostep crawl JavaScript websites?

Olostep's Web Crawling API renders pages with a browser before processing them, so it can capture links and content generated by JavaScript.

Yes. That is the purpose of the Olostep Maps endpoint. It returns discovered URLs without requiring you to treat every page as a full content-extraction job.

What is the difference between Maps and Crawls in Olostep?

Maps is for URL discovery. Crawls recursively walks through pages and creates page records whose content can be retrieved. If your question is simply "what URLs exist?", use Maps. If the next question is "what is on those pages?", use Crawls.

If you start with one website and need to discover its subpages automatically, you do not need to maintain a manual URL list.

Use Olostep Maps when the desired output is a list of discoverable website URLs.

Use Olostep Crawls when the application needs to follow those URLs and retrieve page content.

Use Olostep Scrapes when you already know the page and only need its content or the links present on that page.

Ready to get started?

Start using the Olostep API to implement is there a scraper that can navigate subpages and find all the links for me? in your application.