CEREBRO
machine-read, human-curated
Opus 5.5 and model economics
- Claude Opus 5.5 is out: Anthropic claims about 40% lower cost than Opus 5 at similar capability, with external evals run before release.
- Theo on why Opus 5.5 feels unlimited: Fable was capped at 50% of a Claude Code subscription's usage, and Opus 5.5 High costs over 2x less than Fable 5.1 High, so the headroom comes from both quota and per-task cost.
- Arav Srinivas on Opus 5.5 vs Fable 5.1: after moving 50-100 orchestrator workflows from Fable 5.1 to Opus 5.5 he saw minimal differences, and he is asking which tasks still favour Fable.
- Claude Opus 5.5 official prompting guide: the vendor's own guidance for the new default model, worth diffing against your current prompts and CLAUDE.md.
- How to choose which model for Planning VS Implementing VS Review: a thread on per-phase model routing, which matters more now that the Opus/Fable cost gap is this large.
- I let Opus 5.5 run a business on its own for a week: an unsupervised run that produced €0.00 and 245 dead ideas, with the AI manager ghosting the team on day 1, so it works as a failure-mode catalogue for autonomy.
Multi-agent harnesses and parallel agents
- openrig: defines a Claude Code plus Codex team in YAML and boots it with one command, with a lead agent coordinating specialists across teams.
- Microsoft Research paper on Agensh (via omarsar0): runs 1K+ coding agents with no central orchestrator, coordinating through shared state instead, which is a direct test of flat versus hierarchical agent design.
- I run up to 6 Claude Code sessions on the same codebase: an open-source control room for keeping parallel sessions from colliding.
- Git worktrees solved our parallel-agent file conflicts. The test environment was harder.: worktrees isolate files but not ports, databases or fixtures, so per-agent test environments are the next bottleneck.
- Open-source workspace where coding agents control the app through a CLI: exposes the app's own actions as a CLI so agents drive it directly, a pattern for making SaaS agent-operable.
Agent quality and safety
- Is Vibe Coding Safe? (via thesupermannx): a paper testing 12 coding agents on real tasks finds they write code that works but is often insecure, so functional pass rates overstate readiness.
- Coding agents fix bugs where the error shows up, not where it starts: agents patch the symptom site instead of tracing to the root cause, so review fixes for locality.
- How do you keep large test suites from overwhelming AI coding agents?: a context-budget problem, because full-suite output floods the window and buries the failing signal.
- "You are a senior engineer" changes how sure the answer sounds, not how right it is: persona prompts shift tone and confidence rather than correctness, so don't rely on them as a quality lever.
- Simon Willison on agents making software engineering harder: the gains are real but unlocking them demands extraordinary discipline and knowledge.
- Simon Willison on dropping CLAUDE.md stubs: he can stop creating CLAUDE.md files that only contain
@AGENTS.md, which suggests AGENTS.md is now read natively.
CLI/TUI
- Codex 0.158.0: the fullscreen TUI gains configurable copy-on-select and right-click paste, and copied transcript selections now keep Markdown formatting.
LLM mechanics and training data
- Transformer Inference Arithmetic, part 6 (wafer_ai): kipply's approximate cost model built from per-token work, bytes moved by the GPU and inter-GPU communication, a starting point for reasoning about latency and cost.
- Xiaomi MiMo-V2.6-RL-oss: an open RL dataset of 7,780 samples (12.1 GB) across software engineering, cybersecurity, web dev and knowledge work, with coding tasks verified by executable tests.
- 2026 in LLMs (so far): Simon Willison's annotated WeAreDevelopers keynote, a chronological pass over the year's key trends, useful as a catch-up read.
Agentic SaaS and vibe-coding in the wild
- Harmony: AI agents that resolve IT and HR tickets inside Slack and Teams.
- Claude-powered field operator for a solar company: an operator run through four tests on a real inverter swap, a concrete non-software agent deployment.
- Vibe-coded live AI party installation: a non-coder built it with Claude Code, with a local LLM listening to the room and deciding what a laser wall shows.
Signals
- mvschwarz/openrig: Multi-agent harness that runs Claude Code and Codex together as one system
- Banger paper from Microsoft Research. (bookmark it) They run 1K+ coding agents at once to test a scalable self-organized multi-agent harnes…
- Researchers published the most uncomfortable paper on vibe coding. It’s called "Is Vibe Coding Safe?" They tested 12 of the most advanced A…
- I run up to 6 Claude Code sessions on the same codebase. I built a control room so they stop stepping on each other (open source)
- 0.158.0
- Last few days, I ran 50-100 of my workflows (that earlier used Fable 5.1 as the orchestrator) on Opus 5.5 and found minimal differences. Bu…
- How to choose which model for Planning VS Implementing VS Review
- Why does Opus 5.5 feel practically unlimited when Fable 5.1 was so heavily limited? It's a combination of two things: Opus's efficiency, an…
- Claude Opus 5.5 is out — Anthropic says ~40% cheaper than Opus 5 at similar capability, with external evals before release
- Git worktrees solved our parallel-agent file conflicts. The test environment was harder.
- I run a solar company and built a Claude-powered operator that works in the field with me. Four tests on a real inverter swap (customer dat…
- Claude Opus 5.5 official prompting guide
- 2026 in LLMs (so far)
- Harmony
- Xiaomi just dropped MiMo-V2.6-RL-oss 👀 An open-source RL dataset with 7,780 samples / 12.1 GB covering agentic tasks like software enginee…
- we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread…
- upcoming playlist youtube:- building a production grade generalized ai agent :- omega tentative stack:- golang, postgres, redis, kafka, AWS…
- I can't code, but I vibe-coded a live AI party installation with Claude Code: a local LLM listens to the room and decides what a laser wall…
- Coding agents fix bugs where the error shows up, not where it starts
- How do you keep large test suites from overwhelming AI coding agents?
- "You are a senior engineer" changes how sure the answer sounds, not how right it is
- I built an open-source workspace where coding agents can control the app through a CLI
- Finally! This is great, now I can stop dropping CLAUDEmd files which just contain "@AGENTSmd"
- I let Opus 5.5 run a business on its own for a week: €0.00, 245 dead ideas, and its AI manager ghosted the team on day 1
- The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing…
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| 75 | 12 | |
| x | 231 | 9 |
| rss | 140 | 3 |
| github | 35 | 1 |
| show_hn | 40 | 0 |
| hackernews | 30 | 0 |
| github_search | 20 | 0 |
| swipe | 20 | 0 |
| yc-rfs | 13 | 0 |
| yc-launch | 2 | 0 |
| seed_urls | 1 | 0 |
| crackscan | 0 | 0 |
| gmail | 0 | 0 |
| ossinsight | 0 | 0 |
| reddit_users | 0 | 0 |
| total | 607 | 25 |