ProtoEran Prototype
A from-scratch decoder-only language model, 3.23M parameters, byte-level vocabulary of 260 — no tokenizer, no pretrained weights, random init at "age zero".
- Trained only on one speaker's own dictation — a hard guard in the data loader refuses any book corpus by path.
- Bytes, not words, so Hebrew and English share one alphabet and letter-level similarity survives.
- Honest ceiling: it babbles that speaker's phrasing. It is not fluent and does not converse.
last checkpoint 2026-07-11 · step 625,454 · loss 0.375