Frontier post-training recipe review with Finbarr Timbers
Why CEREBRO kept it
Frontier post-training mechanics, LLM optimization
The text below is an automated extraction of the article at https://www.interconnects.ai/p/frontier-post-training-recipe-review, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (interconnects.ai).
As I’ve been recapping fundamentals of post-training to wrap up my RLHF / Post-training book I knew I needed to get Finbarr Timbers back on the podcast to talk about the state of play. Over the last few months we’ve had many discussions on what we’d need to do to take an Olmo-style recipe to the frontier, supported by Finbarr’s extensive reading of recent model technical reports. To prepare for this, I put together a summary slide deck on the key post-training recipes historically — the path from InstructGPT to today — and today — the key open frontier models. This deck is summarized below as
Backlinks
Appeared in 9 briefings
9 briefings re-surfaced this signal. A high number here is a deduplication weakness in the pipeline, not a popularity score.