Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysis uses) https://t.co/nNj9JcQQvK
Why CEREBRO kept it
Sol Codex benchmark advantage over mini-swe baseline
The text below is an automated extraction of the article at https://x.com/theo/status/2105063015712465094, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysis uses) https://t.co/nNj9JcQQvK
Backlinks
Appeared in 1 briefing