Value fingerprints¶
A value fingerprint answers “would this be the same input to the computation?” Equal fingerprints let a cached computation reuse a previous result. The value's producer, cache location and observation history do not have to match.
Start with a data model¶
Inherit FingerprintedDataModel to get fingerprint() and
Pydantic serialization. By default, the fingerprint includes the model's qualified type name and
actual field values. You do not write a hash function or maintain a version string.
from triplum.datatype import FingerprintedDataModel
class Text(FingerprintedDataModel):
text: str
value = Text(text="A")
original = value.fingerprint()
value.text = "B"
print(original == value.fingerprint())
value.text = "A"
print(original == value.fingerprint())
False
True
The digest is recomputed on each call. Returning to A restores its fingerprint, without including producer identity or processing history. Model code, validators and the full schema do not enter data identity. Changing a selected value changes its fingerprint; changing only unselected structure does not. Different qualified model names identify different kinds of values.
Add identity to an existing model base¶
Use FingerprintedDataModelMixin when your
application already has a Pydantic base. Put the mixin first so its fingerprint hooks participate:
from pydantic import BaseModel
from triplum.datatype import FingerprintedDataModelMixin
class ExistingModel(BaseModel):
pass
class Text(FingerprintedDataModelMixin, ExistingModel):
text: str
print(Text(text="Hello").fingerprint() == Text(text="Hello").fingerprint())
This prints True. The mixin supplies data identity; ExistingModel supplies validation and
serialization. FingerprintedDataModel combines this same mixin with BaseModel for the usual
case. Both support the exclusions and custom projections below. Other bases that define
__init_subclass__ or __pydantic_init_subclass__ must call super() so these hooks cooperate.
Exclude bookkeeping¶
Declare fields that consumers do not use in fingerprint_exclude. Exclusions affect identity,
not serialization, so the field remains available in stored results:
from triplum.datatype import FingerprintedDataModel
class Text(FingerprintedDataModel):
fingerprint_exclude = frozenset({"observed_at"})
text: str
observed_at: int
a = Text(text="A", observed_at=1)
b = Text(text="A", observed_at=2)
print(a.fingerprint() == b.fingerprint())
print(b.model_dump()["observed_at"])
True
2
FingerprintedDataModel already declares this setting as a class variable; assigning it needs no
new annotation. When first composing FingerprintedDataModelMixin with an existing model base,
an override needs an explicit annotation: import ClassVar from typing and write
fingerprint_exclude: ClassVar[frozenset[str]] = frozenset({"observed_at"}).
Keep that ClassVar annotation whenever you annotate this setting, rather than declaring a data field.
Subclasses inherit exclusions; assigning a new set replaces the inherited set. Exclusions must
name declared model fields; unknown names fail when the class is defined. Allowed extra fields
always contribute unless a custom projection removes them. There is no automatic timestamp filter:
a temporal query constraint can affect the answer, while an observation time may be bookkeeping.
An operation that uses an excluded field needs an input contract that includes it.
A cache hit can return the earlier result's bookkeeping. Attach current-event metadata after reusing the result when it must describe the current call.
Select a custom projection¶
Override fingerprint_data() when field values need a different semantic representation. This
example projects a date explicitly into its ISO spelling:
from datetime import date
from triplum.datatype import FingerprintedDataModel
class DatedText(FingerprintedDataModel):
text: str
day: date
def fingerprint_data(self) -> dict[str, object]:
return {"text": self.text, "day": self.day.isoformat()}
first = DatedText(text="Hello", day=date(2026, 1, 1))
second = DatedText(text="Hello", day=date(2026, 1, 1))
print(first.fingerprint() == second.fingerprint())
True
The model's qualified name still distinguishes the projected value kind. An override defines the whole projection: include every semantic field its consumers need and apply any exclusions there. Subclasses inherit that projection; new subclass fields contribute only when the projection includes them. Equal projections intentionally share identity, regardless of how they were produced.
The default projection separates declared fields under "fields" from allowed extras under
"extras", so an extra with the same name cannot overwrite a declared field. It omits private
attributes and computed fields. fingerprint, fingerprint_data and fingerprint_exclude are
reserved for the identity interface, so they cannot be data-field names. Serialization aliases and Field(exclude=True) do not hide actual
semantic fields from it. That separation lets the default codec detect when serialization loses
semantic data; see serialization.
Supported values are None, booleans, integers, finite floats, strings, bytes, UUIDs, paths,
lists, tuples and string-keyed mappings. Lists and tuples remain distinct; mapping insertion order
does not matter. A path identifies its spelling, not file contents. Nested values must implement fingerprint() themselves, normally by inheriting from
FingerprintedDataModel, or be projected explicitly. A plain nested Pydantic model is not supported
automatically. Neither are arbitrary Pydantic field types: project dates, decimals and other
unsupported values explicitly. Non-finite floats are rejected. Fingerprinting support does not
guarantee a default JSON round trip: arbitrary binary bytes may need Pydantic serialization
configuration or a custom codec.
Do not silently normalize content merely to increase reuse. Perform normalization as an explicit step when intended, then fingerprint its output. Use ordered pairs when mapping order affects the computation. Keep recursive object graphs outside these value models.
Use an external value type¶
Fingerprintable is a structural Protocol. External types can
implement fingerprint() -> str without inheriting from our model. Return a full 64-character
SHA-256 hexadecimal digest covering every semantic field:
from dataclasses import dataclass
from triplum.utils.cache import content_key
@dataclass(frozen=True)
class Text:
text: str
def fingerprint(self) -> str:
return content_key("example.Text", {"text": self.text})
print(Text("Hello").fingerprint() == Text("Hello").fingerprint())
content_key hashes a kind identifier and JSON data. It is not the model's typed encoder:
JSON can collapse distinctions such as tuples versus lists, so represent them explicitly when
needed. The kind distinguishes meaning, not a manually maintained version. Both cached inputs
and outputs need a fingerprint; custom non-Pydantic outputs also need a codec.
The Protocol cannot check whether a projection includes every field its consumers rely on.
Share content without sharing provenance¶
Two sources can contain the same text while referring to different documents. A text-only
embedding request can reuse its vector across both. A result carrying a parent source_id
cannot reuse an arbitrary earlier parent reference merely because the text matches.
Project a record into the data the computation actually consumes. Include parent/context fields when the result depends on them, or attach record references after reusing an identity-free result. A fingerprint does not merge storage identity, provenance or access permissions.
Current Source.fingerprint and Chunk.fingerprint are properties, not the required method.
They also omit storage IDs, and Chunk's property omits source_id. They are not directly usable
as cache values. See record identity for the
current behavior; defining the explicit value above does not require changing record IDs.
Compose cached steps¶
A cache lookup matches the input fingerprint and computation fingerprint together. Its output is another value with its own fingerprint. The next step uses that output's semantic fingerprint, not the upstream computation's fingerprint.
Consequently, different normalizers that produce the same Text can share the next computation's
entry. A plain uncached function can also produce that value. There is no required pipeline
executor or requirement to cache every stage. The repository example demonstrates both reuse
across upstream computations and reuse after reopening:
uv run python examples/cached_pipeline.py