Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents: Sonnet 5.5 launch

  • Claude Sonnet 5.5: second model in the 5.5 family, 30%+ faster and up to 30% cheaper than Sonnet 5, aimed at well-scoped everyday work like bug fixes and docs/slides/spreadsheets, positioned as the low-cost complement to Opus 5.5; HN commenters note it matches Opus 5.5 on agentic coding benchmarks but Anthropic's cyber safeguards block authorized security work.
  • Claude Sonnet 5.5: Simon Willison notes same price as Sonnet 5 yet it beats it on every benchmark, so effective cost drops.
  • v2.1.284: Claude Code makes claude-sonnet-5-5 the default Sonnet on the API (1M context, $2/$10 per Mtok, $0.20/Mtok cache reads) and adds a "Yes, but ask again next time" option to auto mode's out-of-directory read prompt.
  • Claude Sonnet 5.5 in GitHub Copilot: now GA in Copilot for feature building and bug fixing.
  • Sonnet 5.5 fixing a bug with Claude Code: Boris Cherny demo of the "30% faster, 30% less usage" claim in Claude Code.
  • Sonnet 5.5 vs Sonnet 5 benchmarks: Anthropic says gains are dramatic in some cases, strongest on scoped tasks, bug fixes, and polished documents.
  • Introducing Claude Sonnet 5.5: official announcement, the 30% speed and cost claims in one post.

Model landscape and usage economics

  • A $200 Claude Code sub gets you ~$9,000 of Opus usage per month: Theo drained three accounts since Opus 5.5 dropped, averaging about $2,200/week of API-equivalent usage, so the flat-rate plan is heavily subsidized and limits may tighten.
  • Opus 5.5 is good at explainer videos: AINews recap says Opus 5.5 vibes are overwhelmingly positive and explainer videos took over the timeline.
  • Fable class models: Simon Willison's working definition groups Claude Fable 5, Opus 5.5, GPT-Astra 6 and maybe GPT-5.6 Sol as the current top tier.
  • Claude Code's Next Era: Thariq Shihipar (Anthropic) on Claude Code's direction, plus a rundown of Anthropic's summer shipping cadence (Sonnet 5, Fable 5, Opus 5, /checkup, $65B ARR); AI Engineer New York is in two weeks.

Routing, cost, and agent architecture

Computer use and agent tooling

  • trycua/cua: open-source computer-use stack with desktop drivers, cloud and local macOS VMs, cross-OS fleets, small CUA-S1 decision models, and benchmarks.
  • Cf: The Agentic CLI for the Cloudflare API: Cloudflare says agents now drive 48% of Wrangler use (up from 25% in March), hence a CLI built for agents; commenters argue plain REST already works and the CLI adds a TypeScript dependency.
  • Launch HN: Vespper: a Docx MCP for agents editing Word documents in legal and finance workflows; commenters point to competing open-source Word MCPs and prefer local tools.

Agent security and observability

Low signal

Signals

  1. Claude Code’s Next Era — Thariq Shihipar, Anthropic
  2. Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, an…
  3. trycua/cua: Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
  4. v2.1.284
  5. Sonnet 5.5 fixing a bug with Claude Code. 30% faster and 30% less usage. https://t.co/Vt1K1ubcFd
  6. Sonnet 5.5 improves on Sonnet 5 across benchmarks, in some cases dramatically. It’s a faster, lower-cost complement to Claude Opus 5.5, str…
  7. Sonnet 5.5
  8. Claude Sonnet 5.5
  9. Claude Sonnet 5.5 in GitHub Copilot
  10. VibeDefend by CybeDefend
  11. RT @ClaudeDevs: Claude can now help you build evaluations and hillclimb on them. In this article, we share guidance on eval design & sk…
  12. Cf: The Agentic CLI for the Cloudflare API
  13. Zerg Router
  14. vantage.ai
  15. @WeAreDevs Here's how I define "Fable class models" - first Claude Fable 5, now Claude Opus 5.5 and GPT-Astra 6 and maybe GPT-5.6 Sol as we…
  16. I routed Claude Code tasks with a cheap decision model. A cost saving is not established.
  17. 400 LLM agents living together in an MMO server: what I learned about perception lag, fire-and-forget actions, and shedding load
  18. Coding Agents Build for the Grader They Imagine, Not the User: Speculative Reward Hacking in DeepSWE
  19. I built a security testing platform for AI agents that can move money. Looking for real-world feedback.
  20. Launch HN: Vespper (YC F24) – SOTA Docx MCP
  21. Are you a Codex Original?
  22. Statable Analytics
  23. A $200 Claude Code sub gets you ~$9,000 of Opus usage per month. I have managed to run 3 Claude accounts down to 0% since Opus 5.5 dropped.…
  24. We split one agent between a cloud planner and a local coder. Here's what it cost:
  25. (AINews) Opus 5.5 is good at explainer videos

Sources & citations

rss14010
x2276
reddit505
hackernews302
github341
yc-launch31
show_hn400
github_search200
swipe200
yc-rfs130
gmail90
seed_urls10
crackscan00
ossinsight00
reddit_users00
total58725