Chunk¶
A Chunk is a contiguous excerpt of a source's text. Chunks are the units that
get indexed and retrieved, so each one carries enough information to find it again in its source:
the source's origin and the character offset start.
Example¶
from triplum.datatype import Chunk, Source
source = Source(
origin="notes/returns.md",
text="Returns are accepted within 30 days. Refunds take 5 days.",
)
chunk = Chunk(source_id=source.id, origin=source.origin, start=37, text="Refunds take 5 days.")
assert source.text[chunk.start : chunk.start + len(chunk.text)] == chunk.text
- Build a
Sourceas before. - Create a
Chunkpointing at its source:source_idis the source'sid, andoriginis copied from it. startis the offset of the excerpt insource.text, andtextis the excerpt, verbatim.- The assertion is the relationship every chunk has to its source: slicing the source text at
startgives back the chunk text.
What it guarantees¶
startis a Python character offset: it counts code points, not bytes, so it matches Python string slicing directly. It must be non-negative; a negative value fails validation.textis copied verbatim from the source text, never rewritten. Text prepared for embedding (descriptions, questions, context) is a separate concern; see chunk preprocessing.- Like
Source,Chunkis a plain Pydantic model. It cannot check the slicing relationship itself, because it does not hold the source; the code that creates chunks is responsible for it.
Reference¶
Chunk in the API reference.
Next: Datasets and the DataLoader, where sources come from.