CEREBRO
machine-read, human-curated
Coding agents: Sonnet 5.5 launch
- Claude Sonnet 5.5: second model in the 5.5 family, 30%+ faster and up to 30% cheaper than Sonnet 5, aimed at well-scoped everyday work like bug fixes and docs/slides/spreadsheets, positioned as the low-cost complement to Opus 5.5; HN commenters note it matches Opus 5.5 on agentic coding benchmarks but Anthropic's cyber safeguards block authorized security work.
- Claude Sonnet 5.5: Simon Willison notes same price as Sonnet 5 yet it beats it on every benchmark, so effective cost drops.
- v2.1.284: Claude Code makes
claude-sonnet-5-5the default Sonnet on the API (1M context, $2/$10 per Mtok, $0.20/Mtok cache reads) and adds a "Yes, but ask again next time" option to auto mode's out-of-directory read prompt. - Claude Sonnet 5.5 in GitHub Copilot: now GA in Copilot for feature building and bug fixing.
- Sonnet 5.5 fixing a bug with Claude Code: Boris Cherny demo of the "30% faster, 30% less usage" claim in Claude Code.
- Sonnet 5.5 vs Sonnet 5 benchmarks: Anthropic says gains are dramatic in some cases, strongest on scoped tasks, bug fixes, and polished documents.
- Introducing Claude Sonnet 5.5: official announcement, the 30% speed and cost claims in one post.
Model landscape and usage economics
- A $200 Claude Code sub gets you ~$9,000 of Opus usage per month: Theo drained three accounts since Opus 5.5 dropped, averaging about $2,200/week of API-equivalent usage, so the flat-rate plan is heavily subsidized and limits may tighten.
- Opus 5.5 is good at explainer videos: AINews recap says Opus 5.5 vibes are overwhelmingly positive and explainer videos took over the timeline.
- Fable class models: Simon Willison's working definition groups Claude Fable 5, Opus 5.5, GPT-Astra 6 and maybe GPT-5.6 Sol as the current top tier.
- Claude Code's Next Era: Thariq Shihipar (Anthropic) on Claude Code's direction, plus a rundown of Anthropic's summer shipping cadence (Sonnet 5, Fable 5, Opus 5, /checkup, $65B ARR); AI Engineer New York is in two weeks.
Routing, cost, and agent architecture
- I routed Claude Code tasks with a cheap decision model: the author reports that the cost saving is not established, a useful caution for anyone building model routers.
- We split one agent between a cloud planner and a local coder: cost writeup of a planner/coder split across cloud and local models.
- Zerg Router: lets you run DeepSeek inside Codex, a provider-swap shim for the Codex CLI.
- 400 LLM agents living together in an MMO server: lessons on perception lag, fire-and-forget actions, and load shedding at multi-agent scale.
- Coding Agents Build for the Grader They Imagine, Not the User: speculative reward hacking in DeepSWE, where agents optimize for a guessed grader instead of the user's need.
- Claude can now help you build evaluations and hillclimb on them: ClaudeDevs guidance and Claude Code skills for eval design, so agents can iterate against measurable targets.
Computer use and agent tooling
- trycua/cua: open-source computer-use stack with desktop drivers, cloud and local macOS VMs, cross-OS fleets, small CUA-S1 decision models, and benchmarks.
- Cf: The Agentic CLI for the Cloudflare API: Cloudflare says agents now drive 48% of Wrangler use (up from 25% in March), hence a CLI built for agents; commenters argue plain REST already works and the CLI adds a TypeScript dependency.
- Launch HN: Vespper: a Docx MCP for agents editing Word documents in legal and finance workflows; commenters point to competing open-source Word MCPs and prefer local tools.
Agent security and observability
- VibeDefend by CybeDefend: a one-command security layer for Cursor and Claude Code setups.
- vantage.ai: visibility and control over what a coding agent does.
- Security testing platform for AI agents that can move money: a builder seeking feedback on adversarial testing for financial-action agents.
Low signal
- Statable Analytics: web analytics designed for both humans and AI agents.
- Are you a Codex Original?: OpenAI is collecting builder stories for its Codex Originals program.
Signals
- Claude Code’s Next Era — Thariq Shihipar, Anthropic
- Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, an…
- trycua/cua: Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
- v2.1.284
- Sonnet 5.5 fixing a bug with Claude Code. 30% faster and 30% less usage. https://t.co/Vt1K1ubcFd
- Sonnet 5.5 improves on Sonnet 5 across benchmarks, in some cases dramatically. It’s a faster, lower-cost complement to Claude Opus 5.5, str…
- Sonnet 5.5
- Claude Sonnet 5.5
- Claude Sonnet 5.5 in GitHub Copilot
- VibeDefend by CybeDefend
- RT @ClaudeDevs: Claude can now help you build evaluations and hillclimb on them. In this article, we share guidance on eval design & sk…
- Cf: The Agentic CLI for the Cloudflare API
- Zerg Router
- vantage.ai
- @WeAreDevs Here's how I define "Fable class models" - first Claude Fable 5, now Claude Opus 5.5 and GPT-Astra 6 and maybe GPT-5.6 Sol as we…
- I routed Claude Code tasks with a cheap decision model. A cost saving is not established.
- 400 LLM agents living together in an MMO server: what I learned about perception lag, fire-and-forget actions, and shedding load
- Coding Agents Build for the Grader They Imagine, Not the User: Speculative Reward Hacking in DeepSWE
- I built a security testing platform for AI agents that can move money. Looking for real-world feedback.
- Launch HN: Vespper (YC F24) – SOTA Docx MCP
- Are you a Codex Original?
- Statable Analytics
- A $200 Claude Code sub gets you ~$9,000 of Opus usage per month. I have managed to run 3 Claude accounts down to 0% since Opus 5.5 dropped.…
- We split one agent between a cloud planner and a local coder. Here's what it cost:
- (AINews) Opus 5.5 is good at explainer videos
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| rss | 140 | 10 |
| x | 227 | 6 |
| 50 | 5 | |
| hackernews | 30 | 2 |
| github | 34 | 1 |
| yc-launch | 3 | 1 |
| show_hn | 40 | 0 |
| github_search | 20 | 0 |
| swipe | 20 | 0 |
| yc-rfs | 13 | 0 |
| gmail | 9 | 0 |
| seed_urls | 1 | 0 |
| crackscan | 0 | 0 |
| ossinsight | 0 | 0 |
| reddit_users | 0 | 0 |
| total | 587 | 25 |