Skip to content

Core Concepts

Follow these pages for small examples in pipeline order. Decision points compares the available implementations and explains how to plug in your own.

Indexing

The examples work directly with Sources. A Collection identifies a corpus scope, but is not an executable prerequisite: its baseline record exists while source membership and scope enforcement remain unimplemented.

Indexing turns input material into searchable chunks. The concepts in pipeline order:

  1. Source: identified input text.
  2. Chunk: an excerpt of a source, located by origin and offset.
  3. Datasets and the DataLoader: where records come from, and how they are consumed one at a time or in batches.
  4. Conversion: dataset records to sources.
  5. Chunking: sources to chunks.
  6. Chunk preprocessing: chunks to the text that gets embedded.
  7. Embedding: that text to one vector per chunk.

Indexing and retrieval shows the larger flow; Decision points explains the shared Protocol and configuration conventions.

Retrieval

Retrieval concepts are not written yet; see Indexing and retrieval for the intended flow.