Source¶
A Source is identified input text: the full text of one document, webpage or article,
together with an origin that says where it came from. It is what chunking consumes and what
every chunk points back to.
Upstream steps (reading files, extracting HTML, OCR) exist to produce sources. Everything after them works on plain text and never needs to know the original file format.
Example¶
from triplum.datatype import Source
source = Source(
origin="notes/returns.md",
text="Returns are accepted within 30 days. Refunds take 5 days.",
)
print(source.origin, len(source.text)) # notes/returns.md 57
- Import
Sourcefromtriplum.datatype. - Set
originto something that identifies the text. Here it is a file path; a URL or another identifier works too. It is always a string. - Set
textto the full text, already decoded. The source holds it in memory.
What it guarantees¶
Sourceis a plain Pydantic model, so construction validates the field types. It has no storage behavior of its own.textmay be empty; chunking then produces no chunks.- Every source gets an
id(a time-ordered UUIDv7) when it is created. Chunks refer back to their source throughsource_id, which is thatid. Creating a secondSourcewith the same text gives a differentid. fingerprintis a hash oforiginandtext: two sources with the same content have the same fingerprint, whatever theirid. See Fingerprints.originis for people and for matching back to the original document; it is not checked for uniqueness.
Reference¶
Source in the API reference.
Next: Chunk, an excerpt of a source.