Steven Gonsalvez

Software Engineer

Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02. We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests

Why CEREBRO kept it

Sonnet 5.5 vs Opus: 40/40 tests hidden, $0.02 cost

The text below is an automated extraction of the article at https://x.com/stas_sorokin_/status/2104953405479297525, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests

> Sonnet 5.5 vs Opus: 40/40 tests hidden, $0.02 cost

Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02.

We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests, every effort level.

→ Sonnet 5.5 at high: 40/40, 2.0k tokens a task → at max: the same 40/40, 23k tokens → $0.02 of output at high, $0.23 at max

So now a Jev classifier reads each sub-task and picks the builder: → Sonnet 5.5 high for bounded work → Sonnet 5.5 xhigh for multi-file work → Opus 5.5 xhigh for money, auth, sends and live state 0.44 s and about $0.00004 a pick.

The twist: Sonnet is

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/agents · ai/llm-mechanics · cerebro/signal

Also from x.com