Reuse computed results¶
Cache a step when reusing its result is cheaper than computing it again. A cache hit requires both the same computation and the same immediate input content. Different pipelines can therefore reuse an intermediate result without sharing their whole processing history.
Add the decorator¶
Use @cached on a function with fingerprintable inputs and outputs. The decorator identifies the
computation, chooses serialization from the return annotation and opens a shared SQLite cache
when first called. No cache configuration or manual computation ID is needed.
Save this complete example as cache_example.py and run uv run python cache_example.py.
FingerprintedDataModel supplies data fingerprinting and Pydantic serialization; the decorator
supplies computation identity and caching.
from triplum.cache import cached
from triplum.datatype import FingerprintedDataModel
class Text(FingerprintedDataModel):
text: str
@cached
def lowercase(value: Text) -> Text:
print("Computing")
return Text(text=value.text.lower())
for text in ("Hello", "Hello", "WORLD", "Hello"):
print(lowercase(Text(text=text)).text)
With an empty cache:
Computing
hello
hello
Computing
world
hello
The second call reuses the result. Changing the input to WORLD computes another result;
returning to Hello reuses its earlier result. Results persist between runs, so subsequent runs
may print no Computing lines. The -> Text annotation selects Pydantic serialization, while
the inherited Text.fingerprint() identifies its field values.
Value fingerprints explains that separate data contract.
The cache opens on the first call and is shared by decorated functions. For its location, explicit cleanup, or write policy, see ownership and write policy.
Function identity covers loaded code, defaults, captures and statically resolved application helpers. External resources have their own identity boundaries; see computation fingerprints.
Specialize only the part you need¶
Ordinary functions remain uncached, and several cached functions share the same default cache.
Each computation has its own namespace. Use @cache.cached with an explicitly owned Cache
when choosing storage or lifetime. A function factory can capture settings;
CachedStep is available for operations already implemented as classes.
Choose the next concept¶
Each page explains one boundary, with a small runnable example and the rules for supplying your own implementation. You can use each extension without adopting the others.
| What you want to do | Read | Interface |
|---|---|---|
| Define which input/output fields mean the same thing | Value fingerprints | FingerprintedDataModel, Fingerprintable |
| Identify an operation or write a configured cached step | Computation fingerprints | cached, CachedStep |
| Follow application methods used at runtime | Runtime dependency tracing | dependency_mode="traced" |
| Use Pydantic output or supply another encoding | Serialization | Codec, PydanticCodec |
| Supply storage or use keys/bytes directly | Storage backends | CacheBackend, SQLiteBackend, CacheKey |
| Choose skipping/blocking and manage lifetime | Ownership and write policy | Cache, CachePolicy, shared-default helpers |
| Inspect or clear persisted results | Administration | CLI, cache_stats, clear_cache |
For a two-step composition demonstrating reuse across upstream computations and after reopening,
run uv run python examples/cached_pipeline.py. Ordinary functions can sit between cached steps;
there is no requirement to cache every operation or use a pipeline executor.
Cache versus record storage¶
The cache holds recomputable results. A record store holds selected records and their references. Equal text may reuse a vector computation without merging source identity or copying another source's provenance. Value fingerprints explain the boundary; current Source/Chunk fingerprint properties cannot be used directly as cache values.