How Web Extraction APIs Handle Structured Output Formats (JSON, CSV, XML)

Web extraction APIs take the messy HTML behind a web page and return it in a structured, machine-readable format your code can use right away. The most common format is JSON, followed by CSV, XML, and Markdown. Instead of writing custom parsing logic for every site, you tell the API what data you want and which shape you want it in, and it hands back clean, structured output.

This matters more than it used to. The web is now a primary data source for software and AI, and the tooling around it is growing fast. Mordor Intelligence values the global web scraping market at USD 1.56 billion in 2026, growing to USD 3.49 billion by 2031 at a 17.39% CAGR (commercial software and managed services).

A big reason for that growth is AI. According to the 2024 Web Almanac, structured data provides a parsable source for AI systems, enabling them to efficiently extract, interpret, and validate information. In short, models work better when the data arrives in a predictable shape.

From Unstructured HTML to Structured Data

Raw HTML is unstructured data. The information you want is buried inside tags, attributes, and nested elements, mixed with navigation, ads, and scripts.

Structured data is the opposite. It follows a fixed shape with named fields, so a program can query it, store it, or feed it to a model without guessing.

A web scraping API does the conversion for you. You send a URL, and it returns the page as fields you can use, instead of a wall of markup you have to untangle yourself.

The Four Formats You'll See Most: JSON, CSV, XML, and Markdown

Extraction APIs usually offer the same short list of output formats. The right choice depends on where the data is going, not on habit.

FormatBest forStructure preservedTypical destination
JSONTyped, nested recordsObjects and arraysAPIs, databases, AI schemas
CSVFlat, tabular dataRows and columnsSpreadsheets, BI tools
XMLHierarchical markupNested elementsLegacy and enterprise systems
MarkdownContent-heavy textHeadings, lists, tablesLLMs, RAG pipelines

JSON: The Default for APIs, Databases, and Typed Records

JSON is the default output for most extraction APIs. As defined in RFC 8259, JavaScript Object Notation (JSON) is a lightweight, text-based, language-independent data interchange format that defines a small set of formatting rules for the portable representation of structured data.

JSON can hold both simple values and nested structures. That is why it fits typed records, APIs, and databases so well: a price stays a number, a flag stays a boolean, and related items nest inside arrays.

The format is also governed by a second standard. ECMA-404 (2nd edition, December 2017) is co-normative with RFC 8259, so the two specifications describe the same syntax.

CSV: Flat, Tabular Data for Spreadsheets and Analytics

CSV is the right pick when your data is flat, meaning simple rows and columns with no nesting. Think product catalogs, price lists, or anything headed straight for a spreadsheet or a BI tool.

RFC 4180 (October 2005) documents the format of comma separated values (CSV) files and formally registers the "text/csv" MIME type. It is an Informational document rather than a formal Internet standard, and it notes there is no single "master" specification for CSV.

That lack of one strict standard is worth remembering. Different tools quote and escape commas, quotes, and line breaks in slightly different ways, so watch those edge cases when your fields contain punctuation.

XML: Hierarchical Markup for Legacy and Enterprise Systems

XML is a good fit when a system on the other end requires it. It shows up in older enterprise stacks, SOAP services, RSS feeds, and many document and configuration standards.

XML is defined by the W3C XML 1.0 specification, a W3C Recommendation from 26 November 2008 that describes XML as a subset of SGML. It represents data as a tree of nested elements, which makes it expressive but more verbose than JSON for the same content.

Markdown: Token-Efficient Text for LLMs and RAG

Markdown is the format most guides skip, yet it is often the best choice for AI. It keeps headings, lists, and tables while stripping the boilerplate that a model does not need, which makes it ideal for RAG context and fine-tuning corpora.

It is also cheaper to process. Olostep's own figures show that a webpage's HTML can consume about 50,000 tokens, while the same content as Markdown uses only about 5,000 tokens, a roughly 10x reduction, which is one reason LLM-ready data is usually delivered as Markdown or JSON.

How Web Extraction APIs Actually Produce Structured Output

Under the hood, the API renders the page, reads its content, and maps that content into the shape you asked for. There are three common ways it does the mapping.

Schema-Based Extraction (Strict, Typed JSON)

With schema-based extraction, you declare the fields and types you want before extraction runs, and the API returns data that matches. You get type safety, a consistent shape across pages, and built-in validation.

This is the most reliable method for production. Because the model maps by meaning rather than by CSS class, one schema works across different layouts, and it keeps working after a site redesign.

Olostep's JSON mode accepts schemas in OpenAI's JSON Schema format. A request looks like this:

curl --request POST \
  --url https://api.olostep.com/v1/scrapes \
  --header 'Authorization: Bearer YOUR_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
    "url_to_scrape": "https://example.com/product",
    "formats": ["json"],
    "llm_extract": {
      "schema": {
        "type": "object",
        "properties": {
          "name": { "type": "string" },
          "price": { "type": "number" },
          "in_stock": { "type": "boolean" }
        },
        "required": ["name", "price"]
      }
    }
  }'

The response comes back shaped to that schema, with the price as a number and the stock flag as a boolean.

Prompt-Based Extraction (Flexible, No Schema)

Sometimes you do not know the exact shape in advance. With prompt-based extraction, you describe the data in plain language, such as "extract the company name, revenue, and employee count," and the AI decides how to structure the JSON.

This method is faster to start and handles pages whose structure varies. The trade-off is that you give up the strict guarantees a schema provides, so it suits exploratory work more than locked-down production fields.

Parsers vs LLM Extraction (Stable Contracts vs Fuzzy Fields)

There are two ways to turn a page into typed JSON, and they fit different jobs. The comparison of parsers and LLM extraction comes down to how stable your fields need to be.

ApproachBest forWhat to watch
ParsersRecurring runs and stable fieldsLimited to supported extractors unless you build a custom one
LLM extractionOne-off jobs and changing schemasDrift risk; cost and latency can vary
HybridStable fields plus fuzzy enrichmentYou manage two paths in one pipeline

Parsers return a stable JSON contract with provenance, so downstream systems can depend on the same keys every run, and self-healing parsers absorb small layout changes. Olostep offers pre-built parsers for common sources, plus the option to build your own.

A useful way to frame reliability is this: pages are not the product, fields are. A job can look healthy while a field quietly goes null, so monitor field drift rather than just checking that the request succeeded.

Handling Nested Data, Tables, and Lists

Real pages rarely hold one flat record. A product page might list variants, sizes, and a table of specifications all at once.

You describe that shape with nested objects and arrays in your schema, and the API preserves the hierarchy. For example, one schema can capture a product and its variants:

{
  "name": "string",
  "features": ["string"],
  "variants": [
    { "color": "string", "sizes": ["string"], "price": "number" }
  ]
}

Tables become arrays of typed row objects, and lists become arrays of items. You define the structure once, and the extraction fills it in regardless of the underlying markup.

Getting Structured Output at Scale

One URL is easy. The real work starts when you need thousands of pages in the same shape.

A web crawling API turns a single starting URL into a dataset of clean, structured pages by following links across a site. For known URL lists, batch workflows process large volumes at once: Olostep's Batches handle up to 10,000 URLs per batch, commonly in about five to eight minutes, and hosted result URLs expire after seven days, so persist the JSON you need in your own storage.

Cost is part of the format decision too. With Olostep, a standard scrape costs 1 credit while an LLM extraction costs 20 credits, so at high volume it is worth using deterministic parsers for stable fields and reserving LLM extraction for the fuzzy ones.

Feeding Structured Output to AI Agents

Typed JSON is exactly what AI agents need to act reliably. A predictable schema lets an agent read a value, store it, or pass it to the next step without guessing.

You can also skip the URL entirely. Olostep's real-time web search API lets an agent describe the data it needs, then searches the live web, reads pages, and returns a validated answer with an optional JSON schema and cited sources.

Choosing the Right Format for Your Pipeline

The format follows the job. Here is the short version:

  • Key point: JSON for typed, nested records headed to APIs, databases, or AI schemas.
  • Key point: CSV for flat, tabular data going into spreadsheets or BI tools.
  • Key point: XML when a legacy or enterprise system requires it.
  • Key point: Markdown for content-heavy text feeding LLMs and RAG.
  • Key point: JSONL for streaming records or training data, where every line shares the same fields.

Whichever format you pick, the reliability comes from the same place: a defined schema, validation, and stored provenance so you can trust and audit the output later.

Frequently Asked Questions

When Should I Use JSON vs CSV vs XML?

Use JSON for nested, typed records and APIs, CSV for flat tabular data bound for spreadsheets or BI tools, and XML when a legacy or enterprise system requires it. Match the format to the destination rather than defaulting to one for every job.

How Do Extraction APIs Convert HTML Into JSON?

The API renders the page, reads its content, and maps that content to your fields using schema-based or prompt-based AI extraction, or a parser. You get structured JSON back without writing or maintaining CSS selectors.

What's the Difference Between Schema-Based and Prompt-Based Extraction?

Schema-based extraction declares exact fields and types up front for strict, validated JSON, while prompt-based extraction describes the data in plain language and lets the AI infer the structure. Schema-based is best for production stability, and prompt-based is best for exploratory or variable pages.

What Happens When a Website Changes Its HTML?

Because semantic AI extraction and self-healing parsers map data by meaning rather than by CSS class, extraction can keep working through many redesigns. It is still wise to monitor field drift so you catch quiet failures where a field goes empty.

Can I Get CSV or XML Instead of JSON?

Yes. Extraction usually produces JSON first, which converts directly into CSV for spreadsheets or XML for legacy systems, so you can deliver whichever format your application needs.

Can I Extract Structured Data From Many URLs at Once?

Yes. Crawl and batch workflows return consistent structured output across thousands of pages in a single job, which suits RAG knowledge bases, dataset building, and large content migrations.

Ready to get started?

Start using the Olostep API to implement how web extraction apis handle structured output formats (json, csv, xml) in your application.