
Learned Contextual Chunking: From LLM Supervision to Fast CPU Inference
TL;DR
Learned Contextual Chunking uses an LLM to label document boundaries, then trains a small model to make those decisions on CPU. We evaluated Simple XGBoost and MiniLM + XGBoost, selecting the approach for each dataset through retrieval validation. Simple XGBoost chunked documents 15–29× faster than hosted Gemma 12B in matched runtime measurements. MiniLM + XGBoost achieved the strongest evidence-retrieval results across the five methods tested in both reported evaluations, with additional encoding cost. The broader lesson: LLMs can provide useful supervision while traditional ML handles narrow, repeated tasks at runtime.
At Valyu, we manage a petabyte-scale corpus spanning varied unstructured and semi-structured documents. That scale makes every optimisation to our indexing pipeline an improvement that compounds as we scale. A model call that is manageable in a small experiment can become a significant operational resource and throughput bottleneck when applied across a large, changing corpus.
Chunking is one of those operations. Before a document can be retrieved, we need to decide which parts belong together. A chunk cannot be so large that it consumes most of the retrieval budget or makes its embedding too noisy, nor so small that it lacks enough context. Poorly placed chunk boundaries can separate relevant information from its surrounding context.
Language models offer an appealing way to make these decisions. They can read the surrounding text and identify meaningful boundaries across different document formats. But that leads to another engineering question: if the decision is narrow and repeated within a specific dataset, can we learn it once and execute it with a much smaller model?
We call this approach Learned Contextual Chunking: use an LLM to generate chunking supervision, then train a small model to make boundary decisions on CPU. We tested two approaches: a simple XGBoost model and another augmented with a small embedding model (MiniLM + XGBoost), alongside experiments with richer contextual features.
LLM-Based Chunking
The starting point was ZeroEntropy's LlamaChunk work, published alongside the zChunk implementation.
Their approach asks Llama 3.1 70B to identify semantic boundaries using a special marker. A straightforward implementation would have the model reproduce the document with that marker inserted at appropriate positions.
Their optimization is more interesting: inspect the probability of the boundary marker at positions in the source, avoiding the need to generate another copy of the document. As their article puts it, they can “check the logprobs to see the probability” of the marker. They also normalize a position-dependent drift in the boundary scores.
That provides a useful way to treat chunking as a learned decision. However, if we were to put this into production, the cost of running an LLM to chunk billions of documents across hundreds of different datasets would be too expensive. This led us to test whether we could distil this behaviour into a compact student that approximates useful boundary choices using inexpensive signals for each dataset.
Learning Chunk Boundaries
A chunker ultimately needs to decide where to stop the current chunk. We make that decision over candidate gaps produced by a document parser.
First, we parse training documents into source units and preserve their offsets. The units depend on the document profile: for example, paragraphs in a scientific paper. Candidate boundaries are gaps between these units.
We then ask the teacher to choose among numbered endpoints. It receives accumulated source text, nearby following text, and a soft token target. We validate its output and connect the selected endpoints to feature rows for student training.
The Simple XGBoost variant uses 13 numeric features. These describe things such as the lengths of neighboring units, position within the document, section transitions, heading flags, punctuation, nearby word overlap, and adjacent TF-IDF similarity.
These features express a limited view of the text. A heading transition can indicate a change of subject. Strong lexical overlap can suggest that neighboring passages belong together. Punctuation and list signals help describe how the source is organized. XGBoost learns how to combine these signals from the teacher's choices.
The student variants use 120 trees of depth three. The lexical student does not need a language model or semantic encoder to score new boundaries.
The MiniLM + XGBoost variant adds semantic similarity features from a small text encoder to the same lexical and structural signals.
A dataset-specific parser can still require engineering work to define useful source units. Once that parser is in place, teacher labeling and student training can run automatically, without manually annotating chunk boundaries for each dataset. This shifts the repeated work from handcrafting chunking rules to running a training and validation pipeline.
Figure 1. Boundary scoring. Candidate gaps are described by lexical and structural signals, with optional MiniLM semantic similarity. The CPU student scores those gaps, and decoding applies the selected settings while preserving the source text.
Keep Teaching Separate from Execution
The teacher runs during label generation. The student runs on new documents.
Figure 2. Training documents supply endpoint labels and numeric features for the student. The frozen CPU chunker scores boundaries and decodes source-preserving chunks; retrieval embeddings and indexing follow.
The workflow is:
- Generate endpoint labels on training documents.
- Train a student on features associated with those endpoints.
- Select thresholds or decoding settings using validation retrieval.
- Freeze the model and settings before scoring test documents.
- Use the CPU student to score candidate boundaries and decode source-preserving chunks.
See the Boundaries in the Source Text
The columns below align excerpts from the same contract so we can inspect where each chunk starts and ends. Fixed length and Gemma begin in an earlier payment clause; the illustrated student begins later, at duty 5.1.2, leaving room to retain the whole of section 5.2. The starting point matters as much as the endpoint.
Figure 3. Shared passages align across three columns. Dark text is inside each selected chunk; pale text shows adjacent source context. START and END mark saved boundaries; ellipses omit text. Gemma is shown at its native teacher endpoint before cap enforcement. The learned column uses an earlier contextual-feature experiment, distinct from the Simple XGBoost and MiniLM + XGBoost variants in the result tables.
Measure Retrieval Rather than Imitation
Matching the teacher is a useful training diagnostic. However, it doesn’t guarantee that the chunks improve retrieval. A teacher model can still make poor decisions that the student then copies.
We therefore evaluated recovery of independently annotated evidence. The main measures were piece recall, which measures how much of the annotated evidence is recovered, and complete evidence, which asks whether all pieces in a valid annotation are recovered.
We held the retriever and budget policy constant within each comparison. Retrieval selects a ranked prefix of whole chunks and stops when the next chunk would exceed the token budget. Under that policy, chunk length and ranking jointly determine what evidence fits.
This can be seen in a retrieval example. For the question “How is intellectual property ownership assigned in this contract?”, all three methods first return the same contract opening. However, their second chunks differ. Fixed-length chunking and Gemma return incomplete sections of ownership language while the student returns the complete fallback assignment clause, which assigns rights to On2 if ownership does not vest automatically.
Figure 4. Excerpts from actual returned chunks for a selected On2 / Wildform question. Shading marks the complete annotated ownership and fallback assignment span. Each method returns two chunks totaling 1,022 tokens under the same 1,024-token retrieval budget, embedding model, cosine ranking, and 512-token chunk cap. The learned column uses the same earlier contextual-feature experiment as Figure 3.
Since Learned Contextual Chunking uses both Simple XGBoost and MiniLM + XGBoost approaches, the choice of chunker for each dataset comes from validation retrieval results during training. We benchmark both variants below so their quality and runtime tradeoffs remain visible.
We added a structured-rules baseline that prefers a section transition once half the target length has accumulated; otherwise, it cuts at the first source-unit boundary reaching the target. It compares an engineered boundary policy, similar to those we used to build, with the learned boundary model.
The arXiv/QASPER evaluation uses a 4,096-token retrieval budget. It covers 70 questions with conservatively mapped evidence, from 111 test questions across 29 papers. Retrieval is within the known paper, and the cohort was reused during research.
| Chunker | Evidence piece recall | Complete evidence |
|---|---|---|
| Fixed-length chunking | 74.3% | 72.9% |
| Structured rules | 84.3% | 82.9% |
| Gemma 12B endpoints | 77.9% | 77.1% |
| Simple XGBoost | 84.3% | 82.9% |
| MiniLM + XGBoost | 85.7% | 84.3% |
Simple XGBoost matches structured rules on both metrics while exceeding the teacher. MiniLM + XGBoost improves on the rules by 1.4 percentage points on each metric. This is the useful automation result: teacher-generated supervision can produce a competitive boundary policy without manually annotating cuts or handcrafting the boundary decision rule.
The fresh CUAD/MAUD transfer evaluation covers 163 questions across 18 documents, at a 1,024-token retrieval budget with a shared 512-token chunk cap.
| Chunker | Evidence piece recall | Complete evidence |
|---|---|---|
| Fixed-length chunking | 12.3% | 11.0% |
| Structured rules | 17.4% | 16.6% |
| Gemma 12B endpoints | 17.4% | 14.1% |
| Simple XGBoost | 11.8% | 9.2% |
| MiniLM + XGBoost | 17.9% | 17.2% |
MiniLM + XGBoost records the strongest results among these five methods: 0.5 percentage points more piece recall and 0.6 percentage points more complete evidence than structured rules.
Latency Is Key
Another benefit of a trained tree model is that it handles boundary decisions through local feature computation and CPU prediction. In our comparison, the LLM chunker ran on a GPU and was slower.
We created a latency benchmark measuring mean chunking latency per document across a selection of 125 test documents from the benchmark datasets.
| Method | Execution | arXiv papers | Consumer contract passages | Legal documents |
|---|---|---|---|---|
| Fixed chunks | CPU | 1,108.5 ms | 8.2 ms | 136.7 ms |
| Structured rules | CPU | 201.7 ms | 15.6 ms | 90.4 ms |
| Simple XGBoost | CPU | 196.6 ms | 23.4 ms | 93.8 ms |
| MiniLM + XGBoost | CPU | 2,782.3 ms | 289.9 ms | 1,198.6 ms |
| Gemma 12B | GPU | 5,750.1 ms | 350.2 ms | 2,100.2 ms |
Simple XGBoost on CPU completed chunking 29.2× faster on arXiv, 15.0× faster on consumer passages, and 22.4× faster on legal documents than hosted Gemma. Richer features do come with a price. MiniLM introduces an encoder into chunk construction, and the runtime table shows the additional work. That increase can be worthwhile, as the better retrieval performance on certain datasets shows.
Why Traditional ML Still Belongs in LLM Pipelines
I think we have become too quick to reach for a general LLM whenever a problem involves language, unstructured or semi-structured data. The benefits of generalisable models can distract from alternative approaches. The old-school route still has its merits: define a narrow task, collect useful supervision, and train a model with the capacity that task requires.
LLMs can make that option more accessible. They help generate supervision for decisions that would otherwise require substantial manual annotation. We can then test how much of that behavior a smaller model can learn, and whether its output meets the downstream requirements.
As shown above, chunking is just one example: a bounded decision about a source boundary. A tree model can learn useful preferences from lexical and structural signals, execute them on cheap compute, and still leave the source text intact under explicit constraints.
The question I want to keep asking is whether every repeated language operation needs a general model at runtime. Recent decision models such as Jev point to the same question: how much of a narrow task can a smaller model handle? With the larger, more capable teacher models now available and a carefully chosen feature representation, even a small student can be useful. Traditional ML still has a useful role in building modern language systems, and we should make room for it in the experiments we run and the products we build.
FAQ
What is Learned Contextual Chunking?
Learned Contextual Chunking uses an LLM to label useful document endpoints, then trains a small model to make boundary decisions. A parser divides the source into units and preserves their offsets. The student scores the gaps between those units, and a decoder turns the selected boundaries into chunks.
Does the chunker need an LLM or GPU at runtime?
The two student variants tested here run on CPU. The teacher is used to generate training labels, rather than being called for every new document. Simple XGBoost uses lexical and structural features without a language model or semantic encoder. MiniLM + XGBoost also runs a small text encoder during chunk construction.
What is the difference between Simple XGBoost and MiniLM + XGBoost?
Simple XGBoost learns from 13 numeric features describing document structure and neighboring text. MiniLM + XGBoost adds semantic similarity features from a small encoder. Both students use 120 trees of depth three. Simple XGBoost was faster in the runtime comparison; MiniLM + XGBoost had the strongest evidence-retrieval results among the five methods in both reported evaluations.
Does Learned Contextual Chunking rewrite or summarise the document?
No. The student selects boundaries in the original source; it does not generate replacement text. The parser preserves source offsets, and decoding produces source-preserving chunks. This keeps the retrieved passages tied to the document they came from.
How do you choose chunk sizes and boundary settings?
The teacher receives a soft token target when labeling endpoints. Student thresholds or decoding settings are selected using validation retrieval, then frozen before test documents are scored. The CUAD/MAUD evaluation used a shared 512-token chunk cap and a 1,024-token retrieval budget. Those experimental settings are not a universal recommendation for every dataset.
Why measure evidence retrieval instead of agreement with the teacher?
A student can reproduce the teacher’s mistakes. We therefore measure recovery of independently annotated evidence: piece recall measures how much evidence is recovered, while complete evidence checks whether all pieces in a valid annotation are recovered. Holding the retriever and budget policy constant lets us compare how boundary choices affect the evidence that fits.
How much faster was the CPU student than the LLM chunker?
In matched mean per-document chunking measurements, Simple XGBoost was 29.2× faster on arXiv papers, 15.0× faster on consumer contract passages, and 22.4× faster on legal documents than hosted Gemma 12B. These figures describe chunking latency, not end-to-end indexing speed. MiniLM + XGBoost had additional encoding cost and should not be assigned the same speedup.
Will the same student work well on every dataset?
That is not established by these results. Simple XGBoost matched structured rules on arXiv/QASPER but underperformed them in the fresh CUAD/MAUD transfer evaluation. MiniLM + XGBoost performed best among the tested methods there, with additional runtime cost. Dataset-specific parsing and retrieval validation still matter. The arXiv/QASPER cohort was also reused during research, so these results should not be read as a universal generalisation claim.
More from the blog





Valyu Add
Join 12,000+ professionals and knowledge workers.
Valyu Add is a free weekly research briefing for builders, investors and operators. Every issue is sourced, cited and verified with Valyu DeepResearch.
