CEREBRO
machine-read, human-curated
Coding agents
- Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02.: 72 hidden-test runs show Sonnet 5.5 at high effort hits 40/40 at ~2k tokens/task ($0.02), versus 23k tokens ($0.23) at max for the same score, so route sub-agents by task and cap effort.
- This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index: Sonnet 5.5 (max) in Claude Code leads at 68, but the three trade performance against cost differently, and the index scores model plus harness together.
- Stripe cut payment integration timelines from up to 6 months to 2–6 weeks using coding agents: their LLM orchestrator was slow, costly and did tasks itself, burning its own context, so they moved predictable coordination into code, which matches the sub-agent routing lesson above.
- Muse Spark 1.3: Meta's update targets agentic and coding work, with longer-horizon threads, clarifying questions, and confirmation before consequential actions.
- Asked Fable 5.5 for the three-body problem. One-shot.: about 1000 lines of Bend2 with GPU-computed pixels at 60 FPS and four formally proven properties, a sign that proof-carrying output is reachable one-shot.
- The future is verification-engineering.: Rauch argues proofs, e2e tests, benchmarks and linters (deterministic and agentic) become the main engineering work as agents write the code.
- I gave Claude Code the CEO job for my free Chrome extension: week-1 anecdote of an agent running a product, with the human still approving every public post.
Claude Code mods
- v2.1.287: ships Claude Mods (plugins can modify deeper behavior) plus a built-in "You should know" mod, a side agent that flags what you or Claude missed, enabled via
/plugin enable cc-plugin-you-should-know@builtin. - You can now mod Claude Code: behavior, UI and features are customizable with a few lines of TypeScript or by asking Claude to write the mod, installed as plugins via
/plugin. - Mods are absolutely insane: Boris Cherny pitches prompt-driven customization of Claude's look and workflow, with mods shared as plugins.
Other agent harnesses and tooling
- earendil-works/pi: open toolkit with unified LLM API, agent runtime with tool calling, TUI and a self-extensible coding agent CLI; new-contributor issues and PRs are auto-closed by default.
- DeepSeek Harness: DeepSeek's open-source agent harness; HN commenters say its one-shot benchmark misses long-running use, so published metrics mislead.
- DSH Desktop: official desktop app for the DeepSeek harness.
- OpenCompanion: one desktop app to start, watch and answer multiple AI coding CLIs.
- Crontick: local scheduler for running prompts on a cron-style schedule.
- Effect v4: zero-dependency TypeScript ecosystem pitched as the base for reliable software and AI agents.
Copilot
- GitHub Copilot in VS Code, September 2026 releases: VS Code v1.136–1.140 add automations for repeatable tasks and agent merge, covering the path from implementation to PR merge.
- Dynamic workflows in Copilot CLI and the Copilot app: define orchestration in code across the CLI, app and SDK for more reliable multi-step runs.
- GitHub Copilot can now interact with desktop apps with computer use: public preview of computer use in Copilot CLI and the app on macOS and Windows.
Runtimes, sandboxes and security
- agent-substrate/substrate: secure-by-default agent runtime for millions of sandboxes, with sub-500ms resume, 500+ suspend/resume activations per second, and zero-trust kernel and network isolation.
- Private AI Proxy: checks the AI service with hardware attestation and TLS key pinning, and forwards fail-closed, before your codebase context leaves your machine.
- Polylane: agents that fix production issues unattended.
- Semitexa: PHP framework designed so an AI agent can inspect it.
Models and local inference
- Claude Opus 5.5 on NEAR AI Cloud: now hosted there with 1M context, 128K output and adaptive reasoning for long agentic coding.
- Qwen3.8-Flash-Next on a 64GB Mac: the Slipstream release runs the 95.5 GiB model at 41–52 tok/s, 1.76x faster than llama.cpp, with 130k context scaling and a Swift variant.
Signals
- Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02. We never compromise on quality, so Opus built ever…
- You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScrip…
- agent-substrate/substrate: Agent Substrate: the core system
- v2.1.287
- This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index leaderboard, but each has a…
- earendil-works/pi: AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
- GitHub Copilot in VS Code, September 2026 releases
- Dynamic workflows in Copilot CLI and the Copilot app
- OpenCompanion
- Mods are absolutely insane. You can now customize Claude to work and look the way you want by just prompting it. Each person works differen…
- RT @EffectTS_: Effect v4 is here. One ecosystem. Zero dependencies. The next chapter of Effect and the foundation for building reliable sof…
- DeepSeek Harness
- DSH Desktop
- Stripe cut payment integration timelines from up to 6 months to 2–6 weeks using coding agents. The useful detail: their first orchestrator…
- GitHub Copilot can now interact with desktop apps with computer use
- Your coding agent can read your entire codebase. Private AI Proxy verifies the AI service before that context leaves your machine, using ha…
- We’re excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability. Key cap…
- Polylane
- Asked Fable 5.5 for the three-body problem. One-shot. It gave back ~1000 lines of Bend2: symplectic physics on the CPU, every pixel compute…
- 🤖 $NEAR JUST ADDED CLAUDE OPUS 5.5 TO ITS AI CLOUD Claude Opus 5.5 is now live on $NEAR AI Cloud with a 1M context window, 128K output and…
- Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, +…
- Semitexa
- The future is verification-engineering. Proofs, (e2e) tests, benchmarks, linters… Some tests will be deterministic, some agentic. This look…
- Crontick - Local prompt scheduler
- I gave Claude Code the CEO job for my free Chrome extension. Week 1: it made me rename it, built the next version, and still needs my "go"…
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| x | 227 | 11 |
| rss | 140 | 8 |
| 50 | 3 | |
| github | 42 | 2 |
| hackernews | 30 | 1 |
| show_hn | 40 | 0 |
| github_search | 20 | 0 |
| swipe | 20 | 0 |
| yc-rfs | 13 | 0 |
| gmail | 8 | 0 |
| yc-launch | 3 | 0 |
| seed_urls | 1 | 0 |
| crackscan | 0 | 0 |
| ossinsight | 0 | 0 |
| reddit_users | 0 | 0 |
| total | 594 | 25 |