Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents: model releases

  • Claude Opus 5.5 opens the 5.5 family: 1M context, 128K output, Fable 5.1-level on most tasks at ~40% lower cost than Opus 5, aimed at long-running agentic coding.
  • Sonnet 5.5 keeps Sonnet 5 pricing but uses fewer tokens per task (up to 30% cheaper per task) and generates output 30%+ faster.
  • GPT-6.1 Sol is the cost-efficient pick next to Astra, available in ChatGPT Work, Codex and the API.
  • OpenAI on Sol pricing claims near-Astra intelligence for a fifth of the price.
  • Ultrafast tier gives up to 8x faster generation (300 tok/s) in Codex and 6x in the API, a premium speed-for-cost trade.

Coding agents: Sol vs Opus 5.5 benchmarks

  • Theo on Sol vs Opus 5.5: Opus 5.5 remains his coding default, while Sol is his pick for code review, architecture analysis, computer use and daily work.
  • Theo's Terminal Bench 4 run has Sol beating Opus 5.5 at roughly 1/30th the price.
  • Artificial Analysis numbers put Sol near Opus 5.5 Medium at under a third of the price.
  • AA's TerminalBench run scored lower but put Sol on the pareto frontier for cost.
  • Harness gap: AA used mini-swe-agent while Theo used Codex, and the difference is large, so benchmark scores depend heavily on the harness.
  • Updated Codex run confirms Sol does much better in Codex than in mini-swe.
  • Sol 6.1 vs Opus 5.5 is the framing question behind the video; read it together with the harness caveat above.
  • Livenerf is a pre-registered, append-only benchmark that checks whether a frontier model quietly degrades after launch; the HN thread alleges peak-load degradation.

Coding agents: tool releases

  • Claude Code v2.1.285 adds CLAUDE_CODE_DISABLE_WEB_FETCH, claude --desktop to open the desktop app on the current directory or session, and claude plugin configure <plugin>.
  • Codex 0.159.0 adds opt-in instant_interrupt, letting new input steer Codex mid-response or during long code-mode calls, plus a compact welcome screen.
  • Codex 0.159.1 makes GPT-6.1 Sol the default model in the bundled and Bedrock catalogs.
  • Codex 0.159.2 fixes console windows flashing on Windows when launching background processes.
  • Codex Remote runs Codex or Claude Code on your own cloud machine.

LLM mechanics and safety

  • PSSA is a from-scratch Rust non-transformer LM with a recurrent state-space layer and episodic memory; HN calls it a novelty-free RNN with no GPU validation.
  • Anthropic Frontier Red Team via Simon Willison: GLM-5.3 achieves full control-flow hijacks in 4% of binary exploitation trials versus 6% for Claude Mythos Preview, showing advanced cyber capability is spreading.

CLI/TUI

  • tmux-companion is an open-source tmux status bar and inbox for juggling multiple Claude Code panes.

Agentic SaaS and runtimes

  • NVIDIA OpenShell is a sandboxed, private runtime for autonomous agents with new isolation primitives and a stable 0.1.x cadence.
  • TermiX x NEXON is a crypto-flavoured "agent economy" settlement partnership announcement; low signal beyond the trend.

Vibe-coding and repos

Signals

  1. https://docs.b.ai/llmservice/models/claude-opus-5.5
  2. Sonnet 5.5 is priced the same as Sonnet 5, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30%…
  3. v2.1.285
  4. This is Ultrafast. Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x…
  5. Codex Remote
  6. alirezarezvani/claude-skills: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable r…
  7. Livenerf: Has Opus 5.5 been nerfed yet?
  8. NVIDIA/OpenShell: OpenShell is the safe, private runtime for autonomous AI agents.
  9. 0.159.0
  10. For your most demanding work, choose Astra when maximizing quality matters most. For complex work you want to run more often, GPT-6.1 Sol g…
  11. Sol 6.1 is a great model and an incredible value. Is it as good as Opus 5.5? https://t.co/VHSnNCGyPp
  12. GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today. http…
  13. Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysis…
  14. When I made this video, there were no benchmarks yet, so I had to run them myself. Jaw dropped when I saw the Terminal Bench 4 scores. Perf…
  15. To be clear, I still think Opus 5.5 is the current GOAT for coding. GPT-6.1 Sol is a great replacement for Astra at a WAY better price, but…
  16. AA's run of TerminalBench got lower scores, but similarly insane costs. 6.1 Sol now dominates the pareto frontier https://t.co/RDe1xRVJ9h
  17. Artificial Analysis numbers are in! GPT-6.1 Sol is performing at around Opus 5.5 Medium levels for under a third the price. Not bad at all,…
  18. Built a town of 200 AI people with Claude Code to test AI agents. The first agent I tested got caught lying.
  19. Good Money and Happy New Week Fam. NEXON x @termix_ai PARTNERSHIP. TermiX is building the clearing & settlement layer for the AI agent econ…
  20. Looks like my numbers were better because I used Codex and AA used mini-swe-agent (think pi but shit) Might have to do my own runs of Termi…
  21. (Showcase) tmux-companion – status bar + inbox for multiple Claude Code panes (open source, free)
  22. PSSA: A non-transformer language model written from scratch in Rust
  23. Quoting Anthropic Frontier Red Team
  24. 0.159.2
  25. 0.159.1

Sources & citations

x22213
rss1406
hackernews302
reddit252
github391
github_search201
show_hn400
swipe200
yc-rfs130
gmail70
yc-launch30
seed_urls10
crackscan00
ossinsight00
reddit_users00
total56025