Semantic chunking is a text-splitting method that groups sentences by meaning so each chunk is a topically coherent unit, instead of cutting text at a fixed length. It uses embedding similarity to decide where one idea ends and the next begins.
A few terms make this easier to follow. A chunk is a small piece of a document that you store and later retrieve. An embedding is a list of numbers that represents the meaning of a piece of text, so that texts with similar meaning have similar numbers. RAG, or retrieval-augmented generation, is a pattern where a system retrieves relevant chunks and passes them to a language model as context before it answers.
The contrast with fixed-size chunking is the whole point. Fixed-size chunking cuts every N characters or tokens regardless of meaning, so it can split a single idea across two chunks. Semantic chunking tries to keep related sentences together instead.
Why Chunking Matters for RAG Retrieval Quality
A RAG system can only answer from the chunks it retrieves, so chunk boundaries directly shape answer quality. If a chunk mixes two topics or splits one idea in half, retrieval returns muddled context and the model has less to work with.
This makes chunking a first-order design decision, not an implementation detail you settle later. The way you draw boundaries affects what gets embedded, what matches a query, and what the model finally sees.
The business stakes follow from data readiness. According to Gartner's AI-ready data research, the firm predicts that through 2026, organizations will abandon 60% of AI projects that are not supported by AI-ready data. Chunking is one step in getting data into that ready state, which is a large reason web scraping for RAG emphasizes clean text, logical boundaries, and preserved sources.
How Semantic Chunking Works
Semantic chunking decides boundaries by measuring how similar adjacent sentences are in meaning, then cutting where that similarity drops. It needs an embedding model and adds compute at ingestion time.
The mechanism runs as an ordered sequence:
- Split the document into individual sentences.
- Embed each sentence into a vector using an embedding model.
- Measure cosine similarity between neighboring sentences.
- Insert a boundary where similarity falls below a chosen threshold.
- Group the sentences between boundaries into one chunk.
Cosine similarity is a score, usually from -1 to 1, that measures how closely two vectors point in the same direction. Sentences with similar meaning produce a high score, while a sharp drop suggests the topic has shifted. That drop is what triggers a new chunk.
This extra work has a cost. Every sentence must be embedded before you can even decide where chunks go, which adds latency and compute to ingestion compared with splitting on character counts.
Choosing a Similarity Threshold
The threshold controls how sensitive the boundaries are. A lower threshold triggers more cuts and produces more, smaller chunks, while a higher threshold produces fewer, larger ones.
Test the threshold on your own data rather than trusting a default, because the right value depends on your content and embedding model. Technical documentation, support tickets, and long articles can all behave differently.
LangChain and LlamaIndex both expose this as a breakpoint setting, often based on a percentile of the similarity distribution rather than a raw number. A percentile approach adapts to each document, which can be more stable than a single fixed cutoff across a varied corpus.
Semantic Chunking vs Other Chunking Strategies
The main chunking strategies differ mostly in how they set boundaries and how much control you get over chunk size. The table below compares five common approaches across the same criteria.
| Strategy | How boundaries are set | Chunk size control | Pros | Cons | Best fit |
|---|---|---|---|---|---|
| Fixed-size | Cut every N characters or tokens | Precise and uniform | Simple, fast, cheap, predictable | Splits ideas mid-thought; ignores meaning | High-volume pipelines where speed and cost dominate |
| Recursive | Split on a priority list of separators (paragraphs, then sentences, then words) until chunks fit a size limit | Good, within a target range | Respects some structure; widely available defaults | Still size-driven, so it can break topics | General-purpose starting point for mixed text |
| Semantic | Cut where embedding similarity between adjacent sentences drops | Indirect; size varies by content | Keeps related sentences together | Adds embedding compute; sizes are uneven | Static, high-value corpora where retrieval quality dominates |
| Structure-aware / page-level | Cut on document structure such as headings, sections, or pages | Depends on the source layout | Uses real author-defined boundaries; cheap once structure exists | Requires clean, preserved structure to work | Well-structured docs, Markdown, and paginated sources |
| Hybrid | Combine structure-aware splitting first, then refine with semantic or size limits | Tunable at each stage | Balances structure and meaning | More moving parts to configure and maintain | Large, mixed corpora that need both order and coherence |
Two findings help set expectations. A 2026 cross-domain study found that content-aware chunking significantly improves retrieval effectiveness over naive fixed-length splitting, with the top method reaching about 24% Precision@1 versus about 2 to 3% for fixed-character chunking, as reported in a 2026 chunking benchmark. Structure also competes well with pure semantic methods: in NVIDIA's 2025 tests, page-level chunking reached the highest average end-to-end RAG accuracy of 0.648, while 128-token chunks scored worst on one dataset at 0.421, according to NVIDIA's chunking benchmark.
The practical takeaway is that content-aware strategies tend to beat naive fixed-size splitting, but semantic chunking is not automatically the winner. Structure-aware and page-level chunking often rival or beat it when the source has usable structure.
Is Semantic Chunking Worth the Cost?
Sometimes, but not always. Semantic chunking adds embedding compute at ingestion and can produce uneven chunk sizes that complicate index tuning, and the retrieval gains are inconsistent against simpler methods.
The evidence supports caution. A peer-reviewed NAACL 2025 study found that the computational costs associated with semantic chunking are not justified by consistent performance gains, as documented in a peer-reviewed study.
A workable decision rule follows from that trade-off. Semantic chunking can be worth it for static, high-value corpora where retrieval quality matters more than ingestion cost, and it is harder to justify for frequently updated or cost-sensitive pipelines where you re-embed often.
Context enrichment is a separate lever worth considering. Anthropic reported that adding context to chunks before embedding, an approach it calls contextual embeddings, reduced the top-20-chunk retrieval failure rate by 35%, from 5.7% to 3.7%, in Anthropic's contextual retrieval research. That gain comes from enriching each chunk with surrounding context, not from the semantic splitting method itself, so the two techniques address different problems.
Garbage In, Garbage Out: Chunk Quality Starts With Clean Input
Every strategy above assumes clean text, and most real source content does not arrive clean. Web pages, PDFs, and raw HTML carry markup, navigation, and layout noise that no boundary algorithm can fix after the fact.
Clean the Boilerplate Before You Chunk
Strip navigation, ads, footers, and scripts before you chunk, because they get embedded as noise and retrieved alongside real content. When boilerplate lands in your index, queries can match a cookie banner instead of the answer.
Cleaning also cuts token waste. By Olostep's own self-reported internal testing, a web page's raw HTML might consume around 50,000 tokens while the same content in Markdown uses only about 5,000 tokens, a figure Olostep reports rather than an independent benchmark. Fewer tokens means less wasted embedding cost and cleaner text to split.
This is why converting HTML to clean Markdown helps before chunking: it removes markup and keeps the readable content. Olostep's scrape endpoint takes a URL or PDF and returns clean, LLM-ready extraction as Markdown or JSON, which gives chunking a better starting point.
Preserve Document Structure for Better Boundaries
Headings, lists, and paragraph breaks are the strongest natural chunk boundaries, but only when extraction preserves them. An author already grouped related ideas under each heading, so that structure is a free signal for where chunks should start and end.
Structure-aware and page-level chunking depend directly on this. If parsing flattens a document into one undifferentiated block of text, those methods lose the boundaries they rely on, and even a well-tuned splitter degrades.
That dependency is why preserving structure for chunking matters as much as removing noise. Clean Markdown that keeps its heading hierarchy lets the model and the splitter both read the document the way its author organized it.
Chunking Web-Scraped and Crawled Content
Web corpora add problems that plain-text tutorials skip: JavaScript-rendered pages, duplicate navigation repeated across every page, and the need to keep each chunk traceable to its source. A page whose content loads through JavaScript can return nearly empty HTML unless it is rendered first.
Duplicate boilerplate is its own retrieval hazard. When the same menu and footer appear on hundreds of crawled pages, those near-identical chunks crowd your index and can outrank the unique content you actually want.
A reliable order of operations handles this. Fetch clean Markdown per page, then chunk, then store metadata such as the source URL, page title, and crawl date alongside each chunk. To gather the pages, you can crawl entire sites for RAG and receive per-page Markdown ready to split. Keeping that per-page provenance is what makes answers auditable, which is the point of preserving source citations so each retrieved chunk can point back to where it came from.
Putting It Together: A Practical RAG Chunking Pipeline
A production RAG chunking pipeline runs as a repeatable sequence from acquisition to storage:
- Crawl or scrape your sources.
- Convert each page to clean, structure-preserving Markdown.
- Chunk with a content-aware strategy suited to the content.
- Embed each chunk with your chosen model.
- Store the vectors in a vector database with source metadata attached.
A vector database is a store built to hold embeddings and find the nearest matches to a query vector quickly. It is where your chunks live once they are embedded, and it is what retrieval searches at query time.
Scale changes the work at the top of this pipeline. Real corpora mean thousands of pages, failed requests that need retries, and periodic re-crawls to keep the index fresh. Handling that volume is where you process URLs at scale through asynchronous batches rather than one request at a time. The acquisition and cleaning steps determine everything downstream, because clean, well-structured input is what makes any chunking strategy perform.
Frequently Asked Questions
What is semantic chunking in RAG?
Semantic chunking is a method that groups sentences by meaning using embedding similarity, so each chunk stays topically coherent instead of being cut at a fixed length.
Is semantic chunking better than fixed-size chunking?
Content-aware methods, including semantic chunking, generally retrieve better than naive fixed-size splitting, but structure-aware and page-level chunking often match or beat pure semantic chunking on well-structured content.
What chunk size or threshold should I use?
There is no universal value, so test the similarity threshold and target chunk size on your own data and embedding model rather than trusting a default.
Is semantic chunking worth the extra cost?
It can be worth the added embedding compute for static, high-value corpora where retrieval quality dominates, but it is harder to justify for frequently updated or cost-sensitive pipelines.
How do I chunk PDFs or web pages for RAG?
Convert the source to clean, structure-preserving Markdown first, then chunk on that clean text and store the source URL and title with each chunk so it stays citable.
Which tools do semantic chunking?
LangChain and LlamaIndex both provide semantic chunking with configurable breakpoint thresholds, usually based on a percentile of the similarity distribution between adjacent sentences.



