What is web extraction API?
A web extraction API is an API that takes a webpage or URL and returns useful data extracted from it. Instead of downloading a page and writing your own code to find the information inside its HTML, you send the URL to the API and request the output you need, such as clean text, Markdown, links, metadata, or structured JSON.
For example, suppose a product page contains a product name, price, description, availability status, images, navigation, scripts, tracking code, and hundreds of HTML elements.
Your application may only need:
{
"product_name": "Wireless Headphones",
"price": "$129",
"availability": "In stock"
}
A web extraction API handles the work between the URL and that usable output.
Modern extraction APIs may also render JavaScript, interact with page elements, remove unwanted content, interpret the page semantically, and map extracted information to a schema defined by the developer.
How does a web extraction API work?
The exact implementation varies by provider, but the basic process is straightforward.
Your application sends an API request containing a URL and instructions about what data or format it wants.
The extraction service then retrieves the page. For a simple server-rendered website, an HTTP request may be enough. For pages where important content appears only after JavaScript executes, the service may need to load the page in a browser environment first.
Once the page is available, the extraction layer identifies the useful content. That can happen through deterministic parsing rules, CSS selectors, XPath expressions, prebuilt parsers, machine-learning models, or an LLM.
The result is converted into the requested format and returned to the application.
A typical flow looks like this:
URL
↓
Fetch or render webpage
↓
Read page content
↓
Extract requested information
↓
Structure or transform the result
↓
Return Markdown, text, HTML, JSON, links, or another format
That removes much of the retrieval and parsing infrastructure developers would otherwise need to build themselves.
Web extraction API vs web scraping API
The terms overlap, and many APIs provide both functions.
Web scraping generally refers to retrieving information from websites programmatically. Extraction describes the part of that process where useful information is identified and separated from the rest of the webpage.
If you request the complete HTML of a page, you are primarily retrieving or scraping the page.
If you request:
{
"company_name": "...",
"industry": "...",
"headquarters": "..."
}
you are asking for structured extraction.
A web scraping API can therefore also be a web extraction API when it includes parsing or structured extraction capabilities.
The distinction becomes more useful when comparing the different jobs involved in working with web data:
| API type | Main job |
|---|---|
| Web Search API | Find relevant webpages from a query |
| Web Scraping API | Retrieve content from a known webpage |
| Web Extraction API | Identify and return useful data from webpage content |
| Web Crawling API | Follow links and process multiple pages across a site |
| Website Mapping API | Discover URLs available on a website |
These functions are often combined. You might search for relevant pages, scrape them, extract specific fields, and send the resulting JSON to another application.
What can a web extraction API return?
Extraction does not always mean structured JSON.
Sometimes the useful output is simply a cleaned version of the page.
Clean text
Text extraction removes HTML markup and returns readable text. This can work well for indexing, text analysis, summarization, and other downstream processing.
Markdown
Markdown keeps useful document structure such as headings, paragraphs, lists, and links without carrying most of the HTML surrounding the content.
This makes it useful when webpage content will be passed to an LLM, RAG pipeline, search index, or knowledge base.
Structured JSON
JSON extraction maps information from a page into specific fields.
For an event page, for example:
{
"event": {
"title": "Developer Conference",
"date": "2026-10-15",
"venue": "Convention Center",
"start_time": "09:00"
}
}
Structured extraction is useful when another application needs predictable fields rather than an entire page.
HTML
Raw or processed HTML can be useful when the application needs access to the page structure or intends to run its own parsing logic.
Links and metadata
Extraction APIs can also return page titles, descriptions, canonical URLs, links found on the page, status information, or similar metadata.
Some APIs additionally support screenshots, PDFs, or raw page assets.
Structured vs unstructured web extraction
One useful distinction is whether you need page content or predefined fields.
Unstructured extraction returns content such as text or Markdown:
Product name
The Wireless Headphones provide...
Price: $129
Available now
Structured extraction converts that information into a predictable data model:
{
"name": "Wireless Headphones",
"price": 129,
"currency": "USD",
"available": true
}
The second format is usually easier to insert directly into databases, spreadsheets, CRMs, monitoring systems, applications, or data pipelines.
The first can be better when the next step involves semantic search, summarization, question answering, embeddings, or another model that needs the surrounding context.
Neither is universally better. The right output depends on what will consume the data next.
How structured web extraction works
There are several ways to turn a webpage into structured data.
CSS selectors and XPath
Traditional extraction logic identifies elements based on the HTML structure.
For example:
price = page.select_one(".product-price").text
This works well when page structure is predictable.
The downside is maintenance. If the website changes .product-price to another element or reorganizes the page, the extraction rule may need to be updated.
Site-specific parsers
A parser contains extraction logic designed for a particular website or page type.
For recurring jobs against predictable websites, parsers can produce consistent JSON without asking a language model to interpret every page individually.
This is useful for workloads such as search result extraction, product catalogs, directory pages, or other high-volume sources with repeatable structures.
Schema-based AI extraction
Instead of describing where the information exists in the HTML, the developer describes what the output should contain.
For example:
{
"type": "object",
"properties": {
"company_name": {
"type": "string"
},
"pricing": {
"type": "string"
}
}
}
The extraction system interprets the page and attempts to populate those fields.
This is useful when pages have different layouts but contain semantically similar information.
Prompt-based extraction
Some extraction APIs also accept natural-language instructions.
For example:
Extract the company name, pricing information, and supported integrations.
This is useful for exploratory or one-off extraction where defining a full schema is unnecessary.
For production systems that depend on a fixed response structure, a schema is usually easier to validate downstream.
Why JavaScript rendering matters for extraction
The HTML returned by an initial HTTP request does not always contain everything visible to a visitor.
Modern websites may load product information, search results, reviews, pricing, tables, or application content through JavaScript after the initial document loads.
An extraction API that supports browser rendering can wait for that content before reading the page.
For some workflows, the extractor may also need to interact with the page by:
- waiting for content to appear;
- scrolling;
- clicking an element;
- filling an input;
- opening content hidden behind a public interface.
Without that step, an extraction process can return an apparently successful page while still missing the information the application needs.
What are web extraction APIs used for?
The same extraction mechanism can support many different products because the output is determined by the source page and requested fields.
Product and pricing data
A product page can be converted into fields such as:
{
"name": "...",
"price": "...",
"currency": "...",
"availability": "...",
"rating": "..."
}
That data can then feed product intelligence, catalog analysis, or permitted monitoring workflows.
Company and lead enrichment
An extraction API can process company websites and collect publicly available fields needed by an enrichment pipeline, such as company descriptions, locations, contact information, or product information.
Articles and research
News articles, documentation, blogs, and research pages can be converted into clean text or Markdown without sending navigation menus, scripts, and unrelated page elements further through the pipeline.
RAG and knowledge bases
Retrieval-augmented generation systems need usable documents before they can chunk, embed, index, and retrieve them.
A web extraction API can convert webpages into Markdown, text, or JSON before those documents enter the knowledge system.
AI agents
An agent may know which URL it needs to read but still require a way to access the current page.
Instead of giving the agent raw HTML, an extraction API can return cleaner content or a schema containing only the information required for the current task.
Monitoring
Repeated extraction makes it possible to compare a selected field over time.
For example, instead of comparing two complete HTML documents, a monitoring system could compare:
{
"plan": "Pro",
"monthly_price": "$49"
}
against the result from its previous run.
Using Olostep as a web extraction API
Olostep's Scrapes endpoint handles extraction from a known URL.
A request is sent to:
POST https://api.olostep.com/v1/scrapes
The request can specify formats such as Markdown, HTML, text, JSON, or screenshots.
For example, a basic request for page content looks like:
curl --request POST \
--url https://api.olostep.com/v1/scrapes \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"url_to_scrape": "https://example.com",
"formats": ["markdown", "text"]
}'
If your application needs fields instead of the complete page, Olostep can perform structured extraction.
For example:
curl --request POST \
--url https://api.olostep.com/v1/scrapes \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"url_to_scrape": "https://example.com/product",
"formats": ["json"],
"llm_extract": {
"schema": {
"type": "object",
"properties": {
"product_name": {
"type": "string"
},
"price": {
"type": "string"
},
"availability": {
"type": "string"
}
}
}
}
}'
Instead of writing a selector for every field, the requested structure defines what the extraction should return.
Olostep also supports prompt-based LLM extraction when a schema is unnecessary:
{
"llm_extract": {
"prompt": "Extract the product name, price, and availability."
}
}
The output is returned through the scrape result as structured JSON.
Parser or LLM extraction: which should you use?
Olostep supports both approaches because they solve different extraction problems.
Use LLM extraction when the page structure varies, the task is exploratory, or you need to extract fields from websites that do not share the same HTML structure.
Use a parser when you repeatedly extract predictable data from the same source and want deterministic fields without interpreting the page with an LLM on every request.
For example:
| Requirement | Better starting point |
|---|---|
| Extract data once from an unfamiliar page | LLM extraction |
| Extract the same fields from many differently designed sites | Schema-based LLM extraction |
| Repeatedly extract the same data from one predictable source | Parser |
| Send readable page content to an LLM | Markdown extraction |
| Keep the original document structure for custom processing | HTML |
The extraction method should match the stability of the source and the structure expected by the downstream system.
Web extraction API vs web crawling API
A web extraction API does not necessarily explore a website.
If you give an extraction API:
https://example.com/product/123
it processes that page.
A crawler starts from a URL and follows links to discover additional pages.
For example:
https://example.com
├── /products
│ ├── /product/1
│ ├── /product/2
│ └── /product/3
└── /blog
├── /post/1
└── /post/2
If you already have the URLs, extraction is enough.
If you need to discover and process a site's subpages, crawling and extraction are usually combined.
Olostep separates these jobs. Scrapes processes known URLs, Maps discovers URLs, and Crawls follows pages across a website. That lets an application choose the amount of navigation it actually needs instead of running a full crawl for every extraction task.
Web extraction API vs web search API
Search answers a different question.
A Web Search API starts with something such as:
best database for vector search
and returns relevant webpages.
A web extraction API normally starts with:
https://example.com/database-comparison
and returns information from that page.
A research agent may use both:
Search query
↓
Find relevant URLs
↓
Extract content from those URLs
↓
Analyze or structure the information
When the source URL is already known, the search stage can be skipped.
What should you look for in a web extraction API?
The right criteria depend on the pages you need to process.
If your sources rely heavily on client-side rendering, JavaScript support matters. If the data feeds a database, schema-based JSON extraction may matter more than Markdown quality.
For AI applications, useful capabilities often include clean Markdown, structured JSON, link extraction, browser rendering, and control over which page content is retained.
For recurring data pipelines, the questions change. You may need parsers, batch processing, predictable schemas, retry handling, and the ability to process large URL lists.
There is no single extraction mode that works best for every workflow. Start from the output your application needs, then work backward to the retrieval and parsing method required to produce it.
Frequently asked questions about web extraction APIs
Is a web extraction API the same as a web scraping API?
Not exactly, although the terms are often used interchangeably. Scraping covers retrieving web data more broadly. Extraction focuses on identifying and transforming useful information from the retrieved page. Many web scraping APIs include extraction features, so one API can perform both jobs.
Can a web extraction API return JSON?
Yes, if the API supports structured extraction. The fields may be produced through predefined parsers, custom parsing logic, a JSON schema, or an AI extraction system.
Can a web extraction API process JavaScript websites?
Only if it supports browser or JavaScript rendering. Pages that load important information client-side may return incomplete data when processed through a basic HTTP request alone.
Does a web extraction API crawl an entire website?
Not necessarily. Extraction normally operates on supplied URLs. A web crawling API discovers and follows additional URLs. Platforms that provide both capabilities can crawl pages first and then run extraction on the discovered content.
Can web extraction APIs be used with LLMs?
Yes. Clean text and Markdown can be used as context for LLMs, while JSON extraction can give agents predictable fields that are easier to pass between tools or store in databases.
What is the difference between parsing and extraction?
Parsing interprets a document's structure so software can identify elements and values inside it. Extraction is the broader task of selecting the information you actually want from that document. Parsing is one method used to perform extraction.
When should I use structured extraction instead of Markdown?
Use structured extraction when your application expects known fields such as price, company_name, or publication_date. Use Markdown when retaining the surrounding page content and document structure is useful for search, RAG, summarization, or other language-model workflows.
From a webpage to usable data
A web extraction API turns webpage content into something another system can use without requiring the application to manage every retrieval, rendering, and parsing step itself.
For a known URL, that may be as simple as converting the page into clean Markdown. For a data pipeline, it may mean extracting a fixed JSON schema. For changing websites, it may involve semantic or LLM-based extraction. For recurring predictable sources, a parser may be the better fit.
Olostep exposes these options through the Scrapes API, so the same endpoint can be used to retrieve page content or turn a webpage into structured data depending on what the downstream application needs.
Ready to get started?
Start using the Olostep API to implement what is web extraction api? in your application.