Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02. We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests
Why CEREBRO kept it
Sonnet 5.5 vs Opus: 40/40 tests hidden, $0.02 cost
The text below is an automated extraction of the article at https://x.com/stas_sorokin_/status/2104953405479297525, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests
> Sonnet 5.5 vs Opus: 40/40 tests hidden, $0.02 cost
Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02.
We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests, every effort level.
→ Sonnet 5.5 at high: 40/40, 2.0k tokens a task → at max: the same 40/40, 23k tokens → $0.02 of output at high, $0.23 at max
So now a Jev classifier reads each sub-task and picks the builder: → Sonnet 5.5 high for bounded work → Sonnet 5.5 xhigh for multi-file work → Opus 5.5 xhigh for money, auth, sends and live state 0.44 s and about $0.00004 a pick.
The twist: Sonnet is
Backlinks
Appeared in 1 briefing