Embedding¶
Embedding turns the text prepared for each chunk into a vector, so that chunks can be found by similarity to a query's vector. An embedder works on a batch of texts at once, because real models are much faster that way.
The library provides only a placeholder, ZeroEmbedder, which returns zero vectors. It fixes
the shape of the step; it does not embed anything meaningful.
Example¶
from triplum.steps.embedding import ZeroEmbedder
embed = ZeroEmbedder(dimensions=4)
vectors = embed(["Returns are accepted within 30 days.", "Refunds take 5 days."])
print(vectors.shape, vectors.dtype) # (2, 4) float32
print(vectors[0]) # [0. 0. 0. 0.]
- The vector size is configured in
__init__and exposed asembed.dimensions(default 1536). A real embedder would load its model or create its client here, once. - Calling it with a list of texts returns one matrix: row
iis the vector of texti. - The result is a numpy
float32array of shape(len(texts), dimensions). The type aliasVectorsnames that shape.
The contract¶
Embedder takes list[str] and returns
Vectors, a NumPy float32 matrix. Its dimensions
attribute gives the vector size without running the model.
- Input: one batch of texts, usually produced by chunk preprocessing.
Forming batches (for example with a
DataLoader) is the pipeline's job. - Output: a
float32matrix with one row per text, in input order, each rowdimensionslong. dimensionsis an attribute, so the vector size is known without embedding anything.
Benchmarks hold the embedding configuration fixed across the pipelines they compare, so that differences in results come from the pipelines, not the embedder. Encoding a user's question at retrieval time is a separate use of the same model and is not part of indexing.
Writing your own¶
Subclass the protocol, set dimensions and implement __call__. This toy embedder counts
letters:
import numpy as np
from triplum.steps.embedding import Embedder, Vectors
class LetterCounts(Embedder):
"""Count the letters a-z in each text; a toy embedder with 26 dimensions."""
dimensions = 26
def __call__(self, texts: list[str], /) -> Vectors:
vectors = np.zeros((len(texts), self.dimensions), dtype=np.float32)
for row, text in enumerate(texts):
for char in text.lower():
if "a" <= char <= "z":
vectors[row, ord(char) - ord("a")] += 1
return vectors
vectors = LetterCounts()(["abba", "Cab"])
print(vectors.shape, vectors.dtype) # (2, 26) float32
print(vectors[:, :3])
# [[2. 2. 0.]
# [1. 1. 1.]]
Unlike the other steps, a plain function does not fit Embedder: the protocol also requires the
dimensions attribute, so the type checker rejects a bare function.
Open questions¶
How an embedder identifies itself (model, revision, settings) so that its vectors can be cached and recorded in a run's identity is not decided yet.
Next: Decision points, every place you choose an implementation.