Datasets and the DataLoader¶
A dataset gives access to records of one type, such as Source. A DataLoader consumes
a dataset one record or one batch at a time. The split is deliberate: the dataset decides what
the records are, the loader only decides how many at once. A loader never converts records.
Batch records without a download¶
Start with an in-memory dataset to see the boundary between records and batching:
from triplum.datatype import Source
from triplum.utils.data import DataLoader, RecordDataset
records = RecordDataset(
[
Source(origin="first", text="Hello"),
Source(origin="second", text="World"),
Source(origin="third", text="Again"),
]
)
for batch in DataLoader(records, batch_size=2):
print([source.text for source in batch])
['Hello', 'World']
['Again']
The loader keeps record order and includes the final, smaller batch. With batch_size=None, it
yields individual records instead. Neither form converts or changes the records.
Load a benchmark corpus¶
This loads the MultiHop-RAG corpus of 609 news articles (licence: ODC-BY 1.0) and iterates over it as sources. The first run downloads the corpus file.
from pathlib import Path
from triplum.datasets.multihoprag import MultiHopRAGCorpus, download
from triplum.utils.data import DataLoader
corpus = MultiHopRAGCorpus(download(Path("data")))
print(len(corpus)) # 609
for source in DataLoader(corpus, batch_size=None):
print(source.origin) # https://mashable.com/article/cyber-monday-deals-amazon-2023
break
for batch in DataLoader(corpus, batch_size=256):
print(len(batch)) # 256, 256, 97
download(Path("data"))fetches the pinnedcorpus.jsontodata/multihoprag/corpus.json, checks its SHA-256 hash and returns the path. If a verified copy already exists, it skips the download.MultiHopRAGCorpus(path)parses the file and is aDataset[Source]: oneSourceper article.originis the article URL;textis the title, a blank line, then the body.DataLoader(corpus, batch_size=None)yields the records one by one, unchanged.DataLoader(corpus, batch_size=256)yields lists of up to 256 sources. The last batch keeps the remainder.
A local example: Markdown files¶
MarkdownFolder reads your own notes: one Source
per .md file in a folder. Given a folder notes/ with returns.md and shipping.md:
from pathlib import Path
from triplum.datasets.markdownfolder import MarkdownFolder
from triplum.utils.data import DataLoader
notes = MarkdownFolder(Path("notes"))
print(len(notes))
for source in DataLoader(notes, batch_size=None):
print(Path(source.origin).name, repr(source.text[:20]))
Output:
2
returns.md '# Returns\n\nReturns a'
shipping.md '# Shipping\n\nOrders s'
The example prints only each file's name; source.origin retains its absolute path.
- Only
.mdfiles directly in the folder count; subfolders are not searched. Files come in sorted name order. - The file list is fixed when you create the dataset; a file added later needs a new
MarkdownFolder. Contents are read, as UTF-8, each time a record is accessed. originis the absolute path of the file, as a string.- Its
fingerprint()hashes every file's name and bytes, so it reads the whole folder.
Two kinds of dataset¶
Dataset: indexed. Implement__len__,__getitem__andfingerprint(). Use it when records can be counted and fetched by position, as with the corpus above.RecordDatasetwraps an in-memory list of Pydantic records.IterableDataset: streaming. Implement__iter__andfingerprint(). Use it when there is no length or random access, for example a stream that can only be read once.
Neither class imposes a schema; each dataset chooses its record type.
Fingerprints¶
Every dataset must implement fingerprint(): a stable string identifying its logical content.
Two datasets with equal fingerprints promise the same records in the same order. A name or a URL
that can change is not enough; the MultiHop-RAG corpus uses the pinned file hash. Computing the
fingerprint must not consume the dataset, and batch size is not part of it. Fingerprints are meant
to let expensive later steps be cached by their inputs.
What the loader does¶
batch_size=Nonepasses each record through unchanged.- An integer
batch_sizeyields lists of that size (the default is 1), or whatevercollate_fnbuilds from each list. - It is lazy: it does not prefetch, does not replay a one-shot stream, and lets errors from the dataset propagate.
Turning a dataset's records into sources, when they are not sources already, is conversion, not the loader's job.
Reference¶
Dataset,IterableDatasetandRecordDatasetDataLoadertriplum.datasets.multihopragandMarkdownFolder
Next: Conversion.