This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index leaderboard, but each has a different balance of performance and cost The Artificial An
Why CEREBRO kept it
LLM coding leaderboard: Sonnet vs Opus vs Gemini performance
The text below is an automated extraction of the article at https://x.com/ArtificialAnlys/status/2105814318294114720, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
The Artificial An
> LLM coding leaderboard: Sonnet vs Opus vs Gemini performance
This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index leaderboard, but each has a different balance of performance and cost
The Artificial Analysis Coding Agent Index measures agents (a combination of model and harness) across three agentic coding evaluations.
➤ Claude Sonnet 5.5 (max) in Claude Code takes the top spot at 68, but also has the highest measured cost per task: $14.19
➤ Gemini 4 Argon (high) in Antigravity CLI scores 64 at $5.84 per task - less than half of Sonnet 5.5’s cost. Note, this uses Google’s promotional pricing
Backlinks
Appeared in 1 briefing