Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Opus 5.5 and model economics

Multi-agent harnesses and parallel agents

Agent quality and safety

CLI/TUI

  • Codex 0.158.0: the fullscreen TUI gains configurable copy-on-select and right-click paste, and copied transcript selections now keep Markdown formatting.

LLM mechanics and training data

  • Transformer Inference Arithmetic, part 6 (wafer_ai): kipply's approximate cost model built from per-token work, bytes moved by the GPU and inter-GPU communication, a starting point for reasoning about latency and cost.
  • Xiaomi MiMo-V2.6-RL-oss: an open RL dataset of 7,780 samples (12.1 GB) across software engineering, cybersecurity, web dev and knowledge work, with coding tasks verified by executable tests.
  • 2026 in LLMs (so far): Simon Willison's annotated WeAreDevelopers keynote, a chronological pass over the year's key trends, useful as a catch-up read.

Agentic SaaS and vibe-coding in the wild

Signals

  1. mvschwarz/openrig: Multi-agent harness that runs Claude Code and Codex together as one system
  2. Banger paper from Microsoft Research. (bookmark it) They run 1K+ coding agents at once to test a scalable self-organized multi-agent harnes…
  3. Researchers published the most uncomfortable paper on vibe coding. It’s called "Is Vibe Coding Safe?" They tested 12 of the most advanced A…
  4. I run up to 6 Claude Code sessions on the same codebase. I built a control room so they stop stepping on each other (open source)
  5. 0.158.0
  6. Last few days, I ran 50-100 of my workflows (that earlier used Fable 5.1 as the orchestrator) on Opus 5.5 and found minimal differences. Bu…
  7. How to choose which model for Planning VS Implementing VS Review
  8. Why does Opus 5.5 feel practically unlimited when Fable 5.1 was so heavily limited? It's a combination of two things: Opus's efficiency, an…
  9. Claude Opus 5.5 is out — Anthropic says ~40% cheaper than Opus 5 at similar capability, with external evals before release
  10. Git worktrees solved our parallel-agent file conflicts. The test environment was harder.
  11. I run a solar company and built a Claude-powered operator that works in the field with me. Four tests on a real inverter swap (customer dat…
  12. Claude Opus 5.5 official prompting guide
  13. 2026 in LLMs (so far)
  14. Harmony
  15. Xiaomi just dropped MiMo-V2.6-RL-oss 👀 An open-source RL dataset with 7,780 samples / 12.1 GB covering agentic tasks like software enginee…
  16. we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread…
  17. upcoming playlist youtube:- building a production grade generalized ai agent :- omega tentative stack:- golang, postgres, redis, kafka, AWS…
  18. I can't code, but I vibe-coded a live AI party installation with Claude Code: a local LLM listens to the room and decides what a laser wall…
  19. Coding agents fix bugs where the error shows up, not where it starts
  20. How do you keep large test suites from overwhelming AI coding agents?
  21. "You are a senior engineer" changes how sure the answer sounds, not how right it is
  22. I built an open-source workspace where coding agents can control the app through a CLI
  23. Finally! This is great, now I can stop dropping CLAUDEmd files which just contain "@AGENTSmd"
  24. I let Opus 5.5 run a business on its own for a week: €0.00, 245 dead ideas, and its AI manager ghosted the team on day 1
  25. The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing…

Sources & citations

reddit7512
x2319
rss1403
github351
show_hn400
hackernews300
github_search200
swipe200
yc-rfs130
yc-launch20
seed_urls10
crackscan00
gmail00
ossinsight00
reddit_users00
total60725