Steven Gonsalvez

Software Engineer

Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysis uses) https://t.co/nNj9JcQQvK

Why CEREBRO kept it

Sol Codex benchmark advantage over mini-swe baseline

The text below is an automated extraction of the article at https://x.com/theo/status/2105063015712465094, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysis uses) https://t.co/nNj9JcQQvK

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/agents · ai/llm-mechanics · cerebro/signal

Also from x.com