Designs by DuhartAll work
Rank
26 of 28
Tier
Tier C 46 of 120

20 — Solomon: Corpus, Tokenizer & Data Governance

Solomon 2 of 6 · Solomon/solomon/data, Solomon/solomon/tokenizer, Solomon/artifacts/corpus Stack: SentencePiece (BPE), custom extractors, no HuggingFace datasets

Where this sits. The unglamorous half of ML that separates people who have shipped a model from people who have run a notebook. Licensing, cleaning, deduplication, contamination checking, and a written data card.


The corpus

WikiText-103 raw, CC BY-SA — license recorded, not assumed. Cleaned, exact-line deduplicated, and split:

SplitTokens
train130M
validation275K
test314K

A documented fallback existed if the mirror was unreachable: Simple English Wikipedia (CC BY-SA, own extractor) plus Project Gutenberg (public domain). Planning the fallback before needing it is the operational habit worth noting.

The data card — artifacts/corpus/DATA_CARD.md

Provenance, license, cleaning steps, split sizes, and contamination notes.

Why contamination checking is the headline

The corpus was checked against the evaluation prompt suites and against LAMBADA and HellaSwag. Without that check, the 9-gate promotion pipeline is measuring memorization and cheerfully promoting a bad model.

Most candidates can describe perplexity. Far fewer will volunteer, unprompted, that their eval number is worthless if the eval set leaked into training — and fewer still have the check written down as an artifact.

The tokenizer

SentencePiece BPE, vocab 8192, byte fallback. Trained here, not downloaded.

MetricValue
Compression4.02 chars/token
UNK tokens0

Recorded in artifacts/tokenizer/tokenizer_stats.json.

Byte fallback is the reason UNK is zero: any codepoint the BPE merges do not cover decomposes into bytes rather than becoming an unknown. That is a design choice with a measurable consequence, and the measurement is in the repo.

Vocabulary sizing

8192 was chosen against a 6.84M-parameter model with tied embedding and output weights. The embedding matrix is vocab × d_model; at vocab 32k it would dominate the parameter budget and starve the layers. This is the vocab/depth trade-off, made explicitly rather than by default.

Interview surface this opens