Conversion¶
Conversion turns one dataset item into sources. A dataset may yield file paths,
raw bytes, HTML pages or scanned images; chunking only understands Source. A converter is the
step in between: reading and decoding a file, extracting text from HTML, or running OCR.
The library provides one naive converter, Utf8File, which
reads a file as text.
Example¶
from pathlib import Path
from triplum.steps.conversion import Utf8File
path = Path("returns.md")
path.write_text("# Returns\nReturns are accepted within 30 days.\n", encoding="utf-8")
convert = Utf8File()
sources = convert(path)
print(len(sources)) # 1
print(sources[0].origin, repr(sources[0].text))
# returns.md '# Returns\nReturns are accepted within 30 days.\n'
- Create the converter once. Converters that need configuration or resources (an OCR model, an
HTTP client) take them in
__init__;Utf8Fileneeds none. - Call it with one item, here a
Path. It returns a list of sources, because one item may hold several documents (an archive, a JSON Lines file).Utf8Filealways returns one. - The source's
originis the path as given, as a string, andtextis the file decoded as UTF-8 with its newlines unchanged. Invalid UTF-8 raises an error.
The contract¶
Converter[A] takes one item of type A and
returns list[Source]. Utf8File is a Converter[Path].
- Input: one item, of whatever type the dataset yields.
Utf8Fileis aConverter[Path]. - Output: zero or more
Sources. Eachoriginshould identify where its text came from, and be unique within what you index together. - A converter handles one item. Mapping it over a dataset or
DataLoaderis the pipeline's job, and the loader never converts.
Datasets whose items are already sources skip this step. The MultiHop-RAG corpus is one.
Writing your own¶
Subclass the protocol and implement __call__. This converter turns one JSON Lines file into
several sources:
import json
from pathlib import Path
from triplum.datatype import Source
from triplum.steps.conversion import Converter
class JsonLines(Converter[Path]):
"""One source per line of a JSON Lines file with `id` and `text` fields."""
def __call__(self, item: Path, /) -> list[Source]:
records = [json.loads(line) for line in item.read_text("utf-8").splitlines() if line]
return [Source(origin=f"{item}#{r['id']}", text=r["text"]) for r in records]
path = Path("faq.jsonl")
path.write_text('{"id": "returns", "text": "30 days."}\n{"id": "refunds", "text": "5 days."}\n')
for source in JsonLines()(path):
print(source.origin, source.text)
# faq.jsonl#returns 30 days.
# faq.jsonl#refunds 5 days.
A matching plain function also fits. See Decision points for the shared Protocol and configuration conventions.
Next: Chunking.