Store¶
A store keeps the records a pipeline produces, so you can read them back later or in another process. Today the store holds sources and their chunks. Vectors, files and graph data are not stored yet.
There is one contract, RecordStore, and two
implementations:
MemoryStorekeeps records in Python dictionaries.SQLAlchemyStorekeeps them in a relational database through SQLModel (which is built on SQLAlchemy). By default it uses SQLite, which ships with Python and needs no server.
All three are importable from triplum.store. SQLAlchemyStore uses SQLAlchemy for database
access; only SQLite is covered by triplum's tests. Backend-specific stores (for
SQLite or PostgreSQL) are planned where database-specific SQL is faster; they will share its
tables, SourceRow and
ChunkRow in triplum.store.sql.tables.
A CollectionRow mapping is also available, with the
id and name of a collection. SQLAlchemyStore does not yet create
its table or manage collection records; RecordStore has no collection methods. Source and chunk
rows do not yet carry collection membership.
The current read methods take record IDs, with no Viewer or temporal-view argument. They do not filter access by principal or reconstruct historical state. Those are requirements of the intended retrieval flow, not guarantees of these stores. Adding a record with an existing ID replaces it rather than preserving history.
Example¶
This chunks one source, writes both to a SQLite file, and reads them back through a second store opened on the same file, as a later run of your program would.
from triplum.datatype import Source
from triplum.steps.chunking import FixedSize
from triplum.store import SQLAlchemyStore
source = Source(origin="notes/returns.md", text="Returns are accepted within 30 days.")
chunks = FixedSize(20)(source)
store = SQLAlchemyStore("sqlite:///records.db")
store.add_sources([source])
store.add_chunks(chunks)
reopened = SQLAlchemyStore("sqlite:///records.db")
print(reopened.source(source.id) == source)
for chunk in reopened.chunks(source.id):
print(chunk.start, repr(chunk.text))
Output:
True
0 'Returns are accepted'
20 ' within 30 days.'
FixedSize(20)(source)splits the source into chunks. Each chunk'ssource_idis the source'sid, which is how the store connects them.SQLAlchemyStore("sqlite:///records.db")opens (or creates) the SQLite filerecords.dbin the current working directory and creates itssourceandchunktables if they are missing.add_sourcesmust come beforeadd_chunks: a chunk whose source is not stored is rejected.SQLAlchemyStore(...)on the same URL sees everything the first store committed.source(id)returns the source with thatid;chunks(source_id)returns its chunks ordered bystart.
Choosing where records go¶
| You write | What you get |
|---|---|
MemoryStore() |
Python dictionaries. Gone when the process ends. |
SQLAlchemyStore() |
An in-memory SQLite database. Gone when the process ends, and private to this store instance. |
SQLAlchemyStore("sqlite:///records.db") |
The file records.db, relative to the current working directory (three slashes). |
SQLAlchemyStore("sqlite:////data/triplum/records.db") |
The absolute path /data/triplum/records.db (four slashes: three for the URL, one for the path). |
SQLAlchemyStore("postgresql+psycopg://user:password@host/dbname") |
A PostgreSQL database; see below. |
Practical details:
- The file is created for you, the folder is not. SQLite creates
records.dbif it does not exist, but the directory it lives in must already exist. - Tables are created when the store is constructed.
SQLAlchemyStore(url)runsCREATE TABLEfor its ownsourceandchunktables if they do not exist yet, and for no other table. Existing tables and rows are left alone; there are no migrations, so a table created by an older layout is not updated. - The in-memory default is shared across threads.
SQLAlchemyStore()keeps its in-memory SQLite database on a single connection, so every thread using that store instance sees the same records. Another store instance, even anotherSQLAlchemyStore(), gets its own empty database. - To start over, delete the SQLite file (
rm records.db). There is no method that clears a store. - Inspecting the data: the SQLite file is an ordinary database. Any SQLite client can open it,
for example
sqlite3 records.db "select origin, length(text) from source"if thesqlite3command-line tool is installed.
PostgreSQL and other databases¶
SQLAlchemyStore hands the URL to SQLAlchemy's create_engine unchanged, so any database SQLAlchemy
supports can be named. SQLAlchemy does not include database drivers other than SQLite's, and
triplum does not depend on one. For PostgreSQL, install a driver into the same environment and
name it in the URL:
uv add psycopg(psycopg 3), thenpostgresql+psycopg://user:password@host/dbname.uv add psycopg2, thenpostgresql+psycopg2://user:password@host/dbname. A barepostgresql://URL also selects psycopg2, SQLAlchemy's default PostgreSQL driver.
The database itself must already exist; SQLAlchemyStore creates tables, not databases. Only SQLite is
covered by triplum's tests. PostgreSQL is expected to work because the tables use only portable
column types, but this has not been verified.
What it guarantees¶
- Records are keyed by
id. Every source and chunk carries anid, a UUIDv7 generated when the record is created. It is not derived from the content: two sources with identical text get different ids and are stored as two records. Their fingerprints are equal, which is how you detect such duplicates. - Adding an
idthat is already stored replaces the record. Chunks stored for a replaced source are not removed. - A chunk needs its source. If any chunk in
add_chunksrefers to a source that is not stored,add_chunksraisesKeyErrorand stores none of the batch. - Reads:
source(id)raisesKeyErrorfor an unknown id.sources()yields every source in no particular order.chunks(source_id)returns a list ordered bystart, empty if the source has no chunks. - Round trip: a record read back equals the one you stored, including whether a chunk's
originwas aPathor astr.
Capabilities, not one big store¶
Storage is split by capability: one Protocol for each kind of thing kept. RecordStore covers
records; vectors, blobs (files) and graph data will each get their own Protocol when they are
built. A single backend may implement several of them, and a pipeline asks only for the
capabilities it uses. Code that only reads chunks therefore accepts a RecordStore and works with
either implementation above, or with your own class that has the same methods.
The store is also separate from the cache. The cache remembers results of computations so they need not be repeated; the store holds the records you decided to keep.
Reference¶
Next: Fingerprints.