Steven Gonsalvez

Software Engineer

The result is a 1.38M parameter, 4 layer decoder only Transformer trained on 8M tokens. It has a custom 4K token BPE tokenizer and a 256-token context window, and it’s small enough to train and run l

Why CEREBRO kept it

Deep technical breakdown of trained model architecture and tokenizer.

The text below is an automated extraction of the article at https://x.com/skirano/status/2075283309610078428, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

It has a custom 4K token BPE tokenizer and a 256-token context window, and it’s small enough to train and run l

> Deep technical breakdown of trained model architecture and tokenizer.

The result is a 1.38M parameter, 4 layer decoder only Transformer trained on 8M tokens.

It has a custom 4K token BPE tokenizer and a 256-token context window, and it’s small enough to train and run locally on my Mac.

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics

Also from x.com