Skip to content

Fingerprints

A fingerprint identifies the content you chose to treat as equivalent. For caching, two questions matter: is this the same input data? and is this the same computation? The cache matches both. A record's storage ID answers a third question: which persisted record does this refer to?

from triplum.datatype import Source
from triplum.steps.chunking import FixedSize

first = Source(origin="notes/returns.md", text="Returns within 30 days.")
second = Source(origin="notes/returns.md", text="Returns within 30 days.")

print(first.id == second.id, first.fingerprint == second.fingerprint)
print(FixedSize(20).fingerprint() == FixedSize(20).fingerprint())
print(FixedSize(20).fingerprint() == FixedSize(30).fingerprint())

Output:

False True
True
False

The sources contain equivalent data but have distinct UUIDv7 storage IDs. The chunkers identify configured computations: changing the chunk size changes their identity.

Choose the identity question

For Pydantic data, inherit FingerprintedDataModel to get a fingerprint() method based on actual field values and qualified type name. Use @cached on the computation. Start with Value fingerprints for the smallest example, bookkeeping exclusions and custom projections. External value types can implement the method themselves. Computation fingerprints explains configured steps and automatic function inference. Those pages include complete extension examples.

This page describes existing record, object and dataset identities. They use related digest helpers but have different equality promises; the common name does not make them interchangeable.

Existing record properties

Source and Chunk expose fingerprint as a property, recomputed on access:

Record Included fields Excluded identifiers
Source origin, text id
Chunk origin, start, text id, source_id

Equal record fingerprints therefore do not mean equal storage identity or provenance. Two chunks can have equal fingerprints while referring to different source records. The computed property is also included in model_dump(); SQL storage retains it in an indexed column.

These properties are not compatible with the new cache's fingerprint() method requirement. Passing a Source or Chunk directly to a cached function does not make it a valid cache value. Their identity migration remains under review. Define an explicit value model for the data a computation needs; do not assume an entire record is interchangeable merely because its content property matches.

Configured-object identity

FingerprintedComputationMixin in triplum.utils.fingerprint supplies a fingerprint() implementation used by steps such as FixedSize. It combines the qualified class name, loaded application method definitions, immutable class constants, selected configuration and statically resolved application helpers. It does not scan instance attributes.

The initial FixedSize example uses this implementation: equal sizes identify equivalent chunkers, while different sizes identify different computations. Settings are selected explicitly; runtime counters and clients stay out. For cached functions with settings, start with a function factory. The same guide shows CachedStep for existing classes.

The computation guide explains configuration selection and supported class constants. Decorators and computation classes share loaded-definition hashing and bounded helper discovery. See the automatic boundary for external resources, dynamic dispatch and complete process_id overrides.

Dataset identity

Dataset and IterableDataset require fingerprint() methods. Their contract describes logical output: equal identities promise the same ordered logical data, including source revision and transformations. Computing identity must not consume iteration state. Physical loader batch size is not content, and identifying a one-shot stream does not make it replayable.

Concrete implementations choose their own projection:

  • RecordDataset hashes the ordered, complete model_dump(mode="json") of every record, recomputing on each call. It does not delegate to each record's fingerprint property or method. Freshly created Source records with matching content but different IDs therefore produce different RecordDataset identities.
  • MarkdownFolder hashes its fixed, sorted file list by file name and bytes. Reading an unchanged folder again gives the same dataset identity, although accessed Source records receive fresh IDs. Its fingerprint omits the absolute folder path, while emitted Source.origin includes that path. Equal fingerprints across copied folders therefore do not promise equal origins; account for location separately if a computation uses it.

For example, record content equivalence and full-record dataset equivalence differ:

from triplum.datatype import Source
from triplum.utils.data import RecordDataset

first = Source(origin="notes.md", text="Hello")
second = Source(origin="notes.md", text="Hello")
print(first.fingerprint == second.fingerprint)
print(RecordDataset([first]).fingerprint() == RecordDataset([second]).fingerprint())

Output:

True
False

Choose the dataset's identity according to its concrete contract; a method named fingerprint alone does not establish that it excludes storage IDs or bookkeeping.

Next: Cache.