Fingerprints¶
A fingerprint identifies the content you chose to treat as equivalent. For caching, two questions matter: is this the same input data? and is this the same computation? The cache matches both. A record's storage ID answers a third question: which persisted record does this refer to?
from triplum.datatype import Source
from triplum.steps.chunking import FixedSize
first = Source(origin="notes/returns.md", text="Returns within 30 days.")
second = Source(origin="notes/returns.md", text="Returns within 30 days.")
print(first.id == second.id, first.fingerprint == second.fingerprint)
print(FixedSize(20).fingerprint() == FixedSize(20).fingerprint())
print(FixedSize(20).fingerprint() == FixedSize(30).fingerprint())
Output:
False True
True
False
The sources contain equivalent data but have distinct UUIDv7 storage IDs. The chunkers identify configured computations: changing the chunk size changes their identity.
Choose the identity question¶
For Pydantic data, inherit FingerprintedDataModel to get a
fingerprint() method based on actual field values and qualified type name. Use @cached on the
computation. Start with Value fingerprints for the smallest example,
bookkeeping exclusions and custom projections. External value types can implement the method
themselves. Computation fingerprints explains configured
steps and automatic function inference. Those pages include complete extension examples.
This page describes existing record, object and dataset identities. They use related digest helpers but have different equality promises; the common name does not make them interchangeable.
Existing record properties¶
Source and Chunk expose
fingerprint as a property, recomputed on access:
| Record | Included fields | Excluded identifiers |
|---|---|---|
| Source | origin, text |
id |
| Chunk | origin, start, text |
id, source_id |
Equal record fingerprints therefore do not mean equal storage identity or provenance. Two chunks
can have equal fingerprints while referring to different source records. The computed property
is also included in model_dump(); SQL storage retains it in an indexed column.
These properties are not compatible with the new cache's fingerprint() method requirement.
Passing a Source or Chunk directly to a cached function does not make it a valid cache value.
Their identity migration remains under review. Define an explicit value model for the data a
computation needs; do not assume an entire record is interchangeable merely because
its content property matches.
Configured-object identity¶
FingerprintedComputationMixin in triplum.utils.fingerprint supplies
a fingerprint() implementation used by steps such as
FixedSize. It combines the qualified class name, loaded application method definitions,
immutable class constants, selected configuration and statically resolved application helpers. It does not scan instance attributes.
The initial FixedSize example uses this implementation: equal sizes identify equivalent
chunkers, while different sizes identify different computations. Settings are selected explicitly;
runtime counters and clients stay out. For cached functions with settings, start with a
function factory. The same guide shows CachedStep for existing classes.
The computation guide explains
configuration selection and supported class constants. Decorators and computation classes share
loaded-definition hashing and bounded helper discovery. See the
automatic boundary for external resources,
dynamic dispatch and complete process_id overrides.
Dataset identity¶
Dataset and
IterableDataset require fingerprint() methods.
Their contract describes logical output: equal identities promise the same ordered logical data,
including source revision and transformations. Computing identity must not consume iteration
state. Physical loader batch size is not content, and identifying a one-shot stream does not
make it replayable.
Concrete implementations choose their own projection:
RecordDatasethashes the ordered, completemodel_dump(mode="json")of every record, recomputing on each call. It does not delegate to each record's fingerprint property or method. Freshly created Source records with matching content but different IDs therefore produce different RecordDataset identities.MarkdownFolderhashes its fixed, sorted file list by file name and bytes. Reading an unchanged folder again gives the same dataset identity, although accessed Source records receive fresh IDs. Its fingerprint omits the absolute folder path, while emittedSource.originincludes that path. Equal fingerprints across copied folders therefore do not promise equal origins; account for location separately if a computation uses it.
For example, record content equivalence and full-record dataset equivalence differ:
from triplum.datatype import Source
from triplum.utils.data import RecordDataset
first = Source(origin="notes.md", text="Hello")
second = Source(origin="notes.md", text="Hello")
print(first.fingerprint == second.fingerprint)
print(RecordDataset([first]).fingerprint() == RecordDataset([second]).fingerprint())
Output:
True
False
Choose the dataset's identity according to its concrete contract; a method named fingerprint
alone does not establish that it excludes storage IDs or bookkeeping.
Next: Cache.