Skip to content

Reuse computed results

Cache a step when reusing its result is cheaper than computing it again. A cache hit requires both the same computation and the same immediate input content. Different pipelines can therefore reuse an intermediate result without sharing their whole processing history.

Add the decorator

Use @cached on a function with fingerprintable inputs and outputs. The decorator identifies the computation, chooses serialization from the return annotation and opens a shared SQLite cache when first called. No cache configuration or manual computation ID is needed.

Save this complete example as cache_example.py and run uv run python cache_example.py. FingerprintedDataModel supplies data fingerprinting and Pydantic serialization; the decorator supplies computation identity and caching.

from triplum.cache import cached
from triplum.datatype import FingerprintedDataModel


class Text(FingerprintedDataModel):
    text: str


@cached
def lowercase(value: Text) -> Text:
    print("Computing")
    return Text(text=value.text.lower())


for text in ("Hello", "Hello", "WORLD", "Hello"):
    print(lowercase(Text(text=text)).text)

With an empty cache:

Computing
hello
hello
Computing
world
hello

The second call reuses the result. Changing the input to WORLD computes another result; returning to Hello reuses its earlier result. Results persist between runs, so subsequent runs may print no Computing lines. The -> Text annotation selects Pydantic serialization, while the inherited Text.fingerprint() identifies its field values. Value fingerprints explains that separate data contract.

The cache opens on the first call and is shared by decorated functions. For its location, explicit cleanup, or write policy, see ownership and write policy.

Function identity covers loaded code, defaults, captures and statically resolved application helpers. External resources have their own identity boundaries; see computation fingerprints.

Specialize only the part you need

Ordinary functions remain uncached, and several cached functions share the same default cache. Each computation has its own namespace. Use @cache.cached with an explicitly owned Cache when choosing storage or lifetime. A function factory can capture settings; CachedStep is available for operations already implemented as classes.

Choose the next concept

Each page explains one boundary, with a small runnable example and the rules for supplying your own implementation. You can use each extension without adopting the others.

What you want to do Read Interface
Define which input/output fields mean the same thing Value fingerprints FingerprintedDataModel, Fingerprintable
Identify an operation or write a configured cached step Computation fingerprints cached, CachedStep
Follow application methods used at runtime Runtime dependency tracing dependency_mode="traced"
Use Pydantic output or supply another encoding Serialization Codec, PydanticCodec
Supply storage or use keys/bytes directly Storage backends CacheBackend, SQLiteBackend, CacheKey
Choose skipping/blocking and manage lifetime Ownership and write policy Cache, CachePolicy, shared-default helpers
Inspect or clear persisted results Administration CLI, cache_stats, clear_cache

For a two-step composition demonstrating reuse across upstream computations and after reopening, run uv run python examples/cached_pipeline.py. Ordinary functions can sit between cached steps; there is no requirement to cache every operation or use a pipeline executor.

Cache versus record storage

The cache holds recomputable results. A record store holds selected records and their references. Equal text may reuse a vector computation without merging source identity or copying another source's provenance. Value fingerprints explain the boundary; current Source/Chunk fingerprint properties cannot be used directly as cache values.