The result is a 1.38M parameter, 4 layer decoder only Transformer trained on 8M tokens. It has a custom 4K token BPE tokenizer and a 256-token context window, and it’s small enough to train and run l
Why CEREBRO kept it
Deep technical breakdown of trained model architecture and tokenizer.
The text below is an automated extraction of the article at https://x.com/skirano/status/2075283309610078428, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
It has a custom 4K token BPE tokenizer and a 256-token context window, and it’s small enough to train and run l
> Deep technical breakdown of trained model architecture and tokenizer.
The result is a 1.38M parameter, 4 layer decoder only Transformer trained on 8M tokens.
It has a custom 4K token BPE tokenizer and a 256-token context window, and it’s small enough to train and run locally on my Mac.
Backlinks
Appeared in 1 briefing