Steven Gonsalvez

Software Engineer

Frontier post-training recipe review with Finbarr Timbers

Why CEREBRO kept it

Frontier post-training mechanics, LLM optimization

The text below is an automated extraction of the article at https://www.interconnects.ai/p/frontier-post-training-recipe-review, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (interconnects.ai).

As I’ve been recapping fundamentals of post-training to wrap up my RLHF / Post-training book I knew I needed to get Finbarr Timbers back on the podcast to talk about the state of play. Over the last few months we’ve had many discussions on what we’d need to do to take an Olmo-style recipe to the frontier, supported by Finbarr’s extensive reading of recent model technical reports. To prepare for this, I put together a summary slide deck on the key post-training recipes historically — the path from InstructGPT to today — and today — the key open frontier models. This deck is summarized below as

Backlinks

Appeared in 9 briefings

9 briefings re-surfaced this signal. A high number here is a deduplication weakness in the pipeline, not a popularity score.

Related

Shares tags: ai/llm-mechanics

Also from interconnects.ai