RAG & Knowledge Systems · 6 min
Ingestion and Chunking: Turning Documents into Retrievable Units
How raw PDFs, HTML, code, and tickets become indexable chunks, and why the boundaries you draw at ingest quietly cap everything retrieval can ever do.
Retrieval can only return units you created at index time. That single fact is the reason this lesson exists. A vector search over your corpus never reads documents. It reads chunks: the pre-cut segments you decided to store. If a fact is split awkwardly across chunks, you've made the retriever's job substantially harder. Unless both pieces are retrieved — or your system deliberately expands neighboring or parent context — the model may never see the complete fact. The boundary you drew during ingestion becomes a hard ceiling on what the system can recall. Chunking is not a preprocessing detail you leave on defaults. It is the schema of your knowledge base.
The pipeline: parse, clean, split, enrich
Ingestion runs in four stages, and most failures happen in the first one, long before anybody has thought about the third.
raw source parse clean split enrich
┌────────┐ ┌────────────┐ ┌──────────┐ ┌───────────┐ ┌──────────────┐
│ PDF │──▶│ extract │──▶│ strip │──▶│ chunk into│──▶│ attach │
│ HTML │ │ text + │ │ boiler- │ │retrievable│ │ metadata: │
│ code │ │ structure │ │ plate, │ │ units + │ │ source, │
│ tickets│ │ (headings, │ │ fix │ │ overlap │ │ section, │
└────────┘ │ tables) │ │ encoding │ └───────────┘ │ timestamp... │
└────────────┘ └──────────┘ └──────────────┘Parsing is where PDFs betray you. Extract a two-column research paper naively and the columns interleave line by line, producing text that is grammatically shredded. Tables collapse into digit soup. Scanned pages need OCR, which adds its own errors. HTML arrives wrapped in nav bars, cookie banners, and footers that, left in place, become high-frequency noise: every chunk starts to look a little like every other one. Code and tickets carry real structure (function boundaries, comment threads, ticket fields) that a text-blind splitter discards.
Cleaning removes that boilerplate, normalizes whitespace and encoding, and preserves the structural signals parsing recovered: heading levels, list boundaries, table cells, code blocks. Those signals are what the next stage splits on. Throw them away here and you have forced yourself into blind fixed-size slicing downstream.
The granularity tradeoff
Now the core decision. Chunks must be small enough to be precise. A tight chunk embeds to a focused vector, so a query about one narrow fact matches it cleanly instead of being diluted by four unrelated paragraphs. But chunks must also be large enough to be self-contained. A fragment like "It raises the limit to 90 days" is useless when the "it" (which policy, which limit?) lives in the chunk before it.
Pinecone's guide offers a heuristic worth keeping: if a chunk makes sense to a human reading it with no surrounding context, it will make sense to the model too. So read a few of your chunks cold. If you cannot tell what they are about, neither can the retriever.
Some numbers to anchor on rather than copy blindly. Recursive splitting at 256 to 512 tokens with 10 to 20% overlap is the common starting point for prose. Smaller chunks around 128 to 256 tokens buy granularity for fact-lookup workloads; larger ones in the 512-plus range keep context for narrative or reasoning-heavy sources. Chunk size also trades against your generation budget. When the model has a small context window, a handful of long chunks fill the prompt and you retrieve fewer distinct sources, which costs you coverage. As the window grows, that pressure eases and larger chunks start to pay off. The 2024 "Searching for Best Practices in RAG" survey lands on a few hundred tokens as a sensible default, which is roughly where most production systems sit.
Three splitting strategies
Fixed-token windows with overlap. Cut every N tokens and slide the window back by an overlap fraction, so each boundary fact lands in two adjacent chunks. Dead simple, cheap, and it will cheerfully guillotine a sentence, or a code function, straight down the middle. Overlap is the damage control. A 10 to 20% window catches most boundary-straddling facts for little storage cost. Push past 25% and you mostly buy duplicated text and index bloat with nothing to show for it.
Fixed windows, 20% overlap:
[--------- chunk 1 ---------]
[--------- chunk 2 ---------]
[------- chunk 3 -----]
└─ overlap ─┘ └─ overlap ─┘
A fact landing on a boundary survives because it appears in both neighbors.Recursive and structural. Split on a priority list of separators: paragraph breaks first (\n\n), then lines, then sentences, then whitespace, then raw characters, dropping to the next level only when a piece still exceeds the size limit. This respects natural boundaries, so chunks tend to be whole thoughts. It is the sensible default for prose. For structured sources you go further and split code on function or class boundaries, Markdown on heading hierarchy, tickets on field structure. Same idea: cut where the document already has seams.
Semantic. Embed the sentences, walk through them in order, and open a new chunk when topic similarity between neighbors drops past a threshold, so boundaries fall on genuine topic shifts. This helps on heterogeneous documents that wander across subjects, and it costs you an embedding pass over every sentence at ingest time. For uniform corpora it is overkill.
Builder. Start with recursive at about 500 tokens and 15% overlap, then measure. Build a small eval set of real queries with known answer chunks and track retrieval recall@k. Change the splitter only when the numbers move. Reaching for "semantic chunking" without an eval harness is cargo-culting.
Metadata is half the point
Every chunk should carry structured fields next to its text: source URI, document title, section path, page or line range, author, timestamp, version, access level. This buys two things vectors alone cannot give you. Filtering: "only chunks from docs updated since the last release," or "only tickets tagged billing." Those pre-filters stop the vector search from spending its budget on irrelevant regions. Citation: when the model answers, you can point the user at the exact page and section. That is the line between a trustworthy system and a fluent liar.
Defender. The access-level field is a security boundary, not a nicety. A RAG index is a fast path to aggregate exfiltration: one crafted query can surface fragments a user was never cleared to read, because the retriever knows nothing about your app's authz. Enforce document-level permissions as a hard metadata pre-filter at query time, and re-check at generation. Treat ingested content as untrusted input, too. A PDF or a ticket can smuggle an indirect prompt-injection payload that only fires once it is retrieved into a prompt.
Where the fact-split problem bites
Say a policy doc reads: "Enterprise accounts qualify for extended retention. This raises the log retention limit to 90 days." Split between those two sentences and you get one chunk that says who qualifies and another that says what they get, and neither one embeds anywhere near the query "how long do enterprise accounts keep logs?" Overlap might rescue you. Structural splitting that keeps the paragraph whole definitely would.
When better boundaries are not enough, the strongest known fix is contextual retrieval: have an LLM write a two- or three-sentence blurb situating each chunk inside its parent document, and prepend that before embedding. Anthropic reported this cut top-20 retrieval failures sharply. Combined with contextual BM25 and reranking, the failure rate fell from 5.7% to 1.9%, a 67% reduction. It works because it re-injects the context the boundary stripped out.
So chunking is an engineering decision with a measurable ceiling, not a default to inherit. In the next lesson, Embeddings and Vector Indexes, we turn these units into searchable vectors and hit another quiet cap on chunk size: the embedding model's own context limit.
Sources
- Chunking Strategies for LLM Applications — Pinecone
- Chunking Strategies to Improve LLM RAG Pipeline Performance — Weaviate
- Introducing Contextual Retrieval — Anthropic Engineering
- Searching for Best Practices in Retrieval-Augmented Generation (arXiv 2407.01219)
- Beyond Chunk-Then-Embed: A Comprehensive Taxonomy and Evaluation of Document Chunking Strategies for Information Retrieval (arXiv 2602.16974)
- Best Chunking Strategies for RAG — Firecrawl
- The Ultimate Guide to Chunking Strategies for RAG — Databricks Community
- LangChain RecursiveCharacterTextSplitter documentation