Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Agent swarms & scale

  • Ultracode run stats: one build kept 30-45 Claude Opus 5.5 agents running concurrently, burned 1.2B tokens for $415, only 25% of a 20x Max weekly quota — a real cost/throughput baseline for large parallel-agent runs.
  • Navier-Stokes Millennium Prize claim from an agent swarm: OpenAI says a group of agents on a next-gen model produced a proof addressing Navier-Stokes smoothness, a marker that agent swarms are now doing frontier math synthesis, not just code.
  • Shipping 2,500 PRs a month to prod: walkthrough of the pipeline behind that throughput, a concrete data point on how far agent-driven CI/PR flow scales.
  • Drawgent: coding agent on a live Excalidraw canvas: wires your own Claude Code/Codex/opencode into an Excalidraw whiteboard so the agent screenshots the canvas, edits diagrams live, and marks notes DONE; HN pushback says Mermaid+Obsidian still wins for agent architecture collab since Excalidraw lacks the programmatic depth agents need.

Coding agent harness & workflow

LLM mechanics, routing & Jev

Agentic SaaS & agent security

Signals

  1. Some stats creating this with Opus 5.5 set on Ultracode: Most of the time there were between 30-45 agents running at the same time. Total c…
  2. Drawgent: Coding agent on a live Excalidraw canvas
  3. Tons of folks ask me "how do I create a CODING_STANDARDS.md file?" My answer is that if you're using it right, it should only be empty for…
  4. Stop caring so much about model releases. Focus on the harness, and improving the environment your agent operates in. You'll find yourself…
  5. How to give an AI agent safe write access: 1. Read-only by default 2. Writes go through a tool with a strict schema 3. Every write = a dry-…
  6. A lot of people are saying Anthropic has already nerfed Claude Opus 5.5. We're launching NerfBench on BridgeBench tomorrow. We have the day…
  7. If Jev picks the tool, how does the LLM ask for another one?
  8. Turning GLM-5.3-Flash into a Jev-like decision model
  9. RT @theo: Last year, Anthropic was optimizing inference across 3 different types of compute (Nvidia, AWS trainium, Google tpus). There were…
  10. I stayed up til 2am and spent $1,000 benchmarking Jev Router so you don't have to. Performance on DeepSWE was roughly the same as GPT-6 Ast…
  11. here's how i shipped 2,500 PRs last month to production this was originally supposed to be for Cursor Compile in London. i couldn't make it…
  12. v1.3 is cooking /retro, /pr, and /implement-spec https://t.co/wxfS9lC3cF
  13. We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The pro…
  14. Start with the task you repeat most. Install one skill. Select your agent in the installer. Give it a real job and check the result. These…
  15. If you want to learn complete AI Engineering stack, follow these stages: Stage 1: How LLMs actually work Tokens, context window, embeddings…
  16. anthropics/claude-code-action
  17. Hemory
  18. I’ve been experimenting a bit with Jev as a fuzzy linter that runs after edits in your agent harness: looks very promising so far in my eva…
  19. The JEV feature missing from most LLM speed-vs-accuracy comparisons
  20. I built a computer-use AI you can actually install and start using in minutes
  21. OpenAI agents tried to bruteforce a UN website's API fields
  22. With LLMs, we like to think of "smart" and "dumb" as one axis (because we think of humans this way). I'd like to argue against this framing…
  23. This is really smart. There are definitely a set of rules that are too complex to lint deterministically that feel wasted on full-scale int…
  24. Thinking about making a skill called /fix-one-thing: "Read CODING_STANDARDS.md. Find a violation in the codebase, and fix it. Make the PR s…
  25. we’re thinking of killing plan mode and using the shift+tab hotkey to adjust effort levels I don’t think the models need plan mode anymore,…

Sources & citations

x23917
reddit503
hackernews303
rss1401
github421
show_hn400
github_search200
swipe200
yc-rfs130
yc-launch20
gmail10
seed_urls10
crackscan00
ossinsight00
reddit_users00
total59825