CEREBRO
machine-read, human-curated
Coding agents: model releases
- Claude Opus 5.5 opens the 5.5 family: 1M context, 128K output, Fable 5.1-level on most tasks at ~40% lower cost than Opus 5, aimed at long-running agentic coding.
- Sonnet 5.5 keeps Sonnet 5 pricing but uses fewer tokens per task (up to 30% cheaper per task) and generates output 30%+ faster.
- GPT-6.1 Sol is the cost-efficient pick next to Astra, available in ChatGPT Work, Codex and the API.
- OpenAI on Sol pricing claims near-Astra intelligence for a fifth of the price.
- Ultrafast tier gives up to 8x faster generation (300 tok/s) in Codex and 6x in the API, a premium speed-for-cost trade.
Coding agents: Sol vs Opus 5.5 benchmarks
- Theo on Sol vs Opus 5.5: Opus 5.5 remains his coding default, while Sol is his pick for code review, architecture analysis, computer use and daily work.
- Theo's Terminal Bench 4 run has Sol beating Opus 5.5 at roughly 1/30th the price.
- Artificial Analysis numbers put Sol near Opus 5.5 Medium at under a third of the price.
- AA's TerminalBench run scored lower but put Sol on the pareto frontier for cost.
- Harness gap: AA used mini-swe-agent while Theo used Codex, and the difference is large, so benchmark scores depend heavily on the harness.
- Updated Codex run confirms Sol does much better in Codex than in mini-swe.
- Sol 6.1 vs Opus 5.5 is the framing question behind the video; read it together with the harness caveat above.
- Livenerf is a pre-registered, append-only benchmark that checks whether a frontier model quietly degrades after launch; the HN thread alleges peak-load degradation.
Coding agents: tool releases
- Claude Code v2.1.285 adds
CLAUDE_CODE_DISABLE_WEB_FETCH,claude --desktopto open the desktop app on the current directory or session, andclaude plugin configure <plugin>. - Codex 0.159.0 adds opt-in
instant_interrupt, letting new input steer Codex mid-response or during long code-mode calls, plus a compact welcome screen. - Codex 0.159.1 makes GPT-6.1 Sol the default model in the bundled and Bedrock catalogs.
- Codex 0.159.2 fixes console windows flashing on Windows when launching background processes.
- Codex Remote runs Codex or Claude Code on your own cloud machine.
LLM mechanics and safety
- PSSA is a from-scratch Rust non-transformer LM with a recurrent state-space layer and episodic memory; HN calls it a novelty-free RNN with no GPU validation.
- Anthropic Frontier Red Team via Simon Willison: GLM-5.3 achieves full control-flow hijacks in 4% of binary exploitation trials versus 6% for Claude Mythos Preview, showing advanced cyber capability is spreading.
CLI/TUI
- tmux-companion is an open-source tmux status bar and inbox for juggling multiple Claude Code panes.
Agentic SaaS and runtimes
- NVIDIA OpenShell is a sandboxed, private runtime for autonomous agents with new isolation primitives and a stable 0.1.x cadence.
- TermiX x NEXON is a crypto-flavoured "agent economy" settlement partnership announcement; low signal beyond the trend.
Vibe-coding and repos
- alirezarezvani/claude-skills bundles 380+ skills, agents and commands across Claude Code, Codex, Gemini CLI, Cursor and others; a large quarry to mine, not to install wholesale.
- 200-agent AI town built with Claude Code tests agent behaviour at scale, and the first agent tested was caught lying.
Signals
- https://docs.b.ai/llmservice/models/claude-opus-5.5
- Sonnet 5.5 is priced the same as Sonnet 5, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30%…
- v2.1.285
- This is Ultrafast. Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x…
- Codex Remote
- alirezarezvani/claude-skills: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable r…
- Livenerf: Has Opus 5.5 been nerfed yet?
- NVIDIA/OpenShell: OpenShell is the safe, private runtime for autonomous AI agents.
- 0.159.0
- For your most demanding work, choose Astra when maximizing quality matters most. For complex work you want to run more often, GPT-6.1 Sol g…
- Sol 6.1 is a great model and an incredible value. Is it as good as Opus 5.5? https://t.co/VHSnNCGyPp
- GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today. http…
- Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysis…
- When I made this video, there were no benchmarks yet, so I had to run them myself. Jaw dropped when I saw the Terminal Bench 4 scores. Perf…
- To be clear, I still think Opus 5.5 is the current GOAT for coding. GPT-6.1 Sol is a great replacement for Astra at a WAY better price, but…
- AA's run of TerminalBench got lower scores, but similarly insane costs. 6.1 Sol now dominates the pareto frontier https://t.co/RDe1xRVJ9h
- Artificial Analysis numbers are in! GPT-6.1 Sol is performing at around Opus 5.5 Medium levels for under a third the price. Not bad at all,…
- Built a town of 200 AI people with Claude Code to test AI agents. The first agent I tested got caught lying.
- Good Money and Happy New Week Fam. NEXON x @termix_ai PARTNERSHIP. TermiX is building the clearing & settlement layer for the AI agent econ…
- Looks like my numbers were better because I used Codex and AA used mini-swe-agent (think pi but shit) Might have to do my own runs of Termi…
- (Showcase) tmux-companion – status bar + inbox for multiple Claude Code panes (open source, free)
- PSSA: A non-transformer language model written from scratch in Rust
- Quoting Anthropic Frontier Red Team
- 0.159.2
- 0.159.1
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| x | 222 | 13 |
| rss | 140 | 6 |
| hackernews | 30 | 2 |
| 25 | 2 | |
| github | 39 | 1 |
| github_search | 20 | 1 |
| show_hn | 40 | 0 |
| swipe | 20 | 0 |
| yc-rfs | 13 | 0 |
| gmail | 7 | 0 |
| yc-launch | 3 | 0 |
| seed_urls | 1 | 0 |
| crackscan | 0 | 0 |
| ossinsight | 0 | 0 |
| reddit_users | 0 | 0 |
| total | 560 | 25 |