Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents

  • v2.1.276 fixes a 2.1.275 regression in Claude Code where every request failed with 400 … Input tag 'advisor_20260301' when ANTHROPIC_BASE_URL points at a proxy or gateway, so upgrade if you route through one.
  • stablyai/orca is an ADE for desktop, mobile and remote that runs Codex, Claude Code, OpenCode or Pi side by side, each in its own worktree, with phone notifications and follow-ups when an agent finishes.
  • Theo on Opus 5.5 filing a PR reports that a vague prompt produced a ready-to-merge T3 Code pull request with a video demo attached, an anecdote about agents owning the full PR loop.
  • Theo on Claude Code vs Codex claims the $200 Claude Code plan is now far ahead of Codex, reversing the picture from a few weeks ago; sentiment only, no data.
  • When chat is the wrong UI is GitHub's pitch for Copilot "canvases" as a more tangible interface than a chat box.
  • AI-powered fuzzing with the GitHub Security Lab Taskflow Agent walks through a new fuzzing taskflow built on the Taskflow Agent framework, a concrete example of agent-driven security tooling.
  • Default Enablement of Copilot Features for Copilot Business and Enterprise adds a global default policy enabling GA Copilot features and supported client capabilities in enterprise and org settings, with a 28-day window, so check admin settings now.

Models and benchmarks

Agent memory and context

  • vectorize-io/hindsight is an agent memory system aimed at agents that learn over time rather than just recall history, positioned as better than RAG and knowledge-graph approaches on long-term memory.
  • AI Engineering bonus thread lays out the full production stack (context, agent loop, memory, tools, harness, guardrails, execution, verification, observability, evals, improvement) around the model; a useful checklist, nothing new.
  • Solo founder on project context says proper project context helped more than any model upgrade; title only, no body captured.

Vibe-coding and repos

  • mattpocock/skills is a set of small, composable agent skills that work with any model, pitched against process-owning frameworks like GSD, BMAD and Spec-Kit that take control away and make bugs hard to trace.
  • Raft goes source-available and accepts "prompt requests" instead of pull requests, an experiment in an open-source contribution model for the agent era.

Agentic SaaS

Signals

  1. v2.1.276
  2. stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
  3. Opus 5.5 just filed a pull request that changes the streaming behavior in T3 code. It's so cool that I can send off a prompt with a vague i…
  4. mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.
  5. (AINews) Jev: a “System One Model” that only decides/classifies/routes/scores — >100x faster, >200x cheaper than small frontier LLMs
  6. v2.1.272
  7. anthropics/financial-services
  8. AI Engineering- Bonus We spent 15 parts exploring the pieces. Now here’s how they fit together. → Context → Agent Loop → Memory → Tools → H…
  9. Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max)…
  10. It's insane how much better the $200 Claude Code plan is compared to Codex right now. Just a few weeks ago, it was the other way around. Ex…
  11. Teaching Everyone to Fish for Tokens
  12. Autonomous Product Delivery
  13. superdesigndev/treg: OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
  14. AI/ML Series (Lesson 15/100): Most AI agents fail because of the model. That is the biggest misconception in enterprise AI. The real failur…
  15. Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
  16. When chat is the wrong UI
  17. TL;DR: We just made Raft source-available. We don't accept pull requests. We want prompt requests instead. We are exploring what a new cont…
  18. GLM-5.2 is the step change for open agents
  19. Parallel cut research time and cost in half with GPT‑6 Astra
  20. vectorize-io/hindsight: Hindsight: Agent Memory That Learns
  21. Default Enablement of Copilot Features for Copilot Business and Enterprise
  22. AI-powered fuzzing with the GitHub Security Lab Taskflow Agent
  23. solo founder: setting up proper project context helped more than any model upgrade
  24. 5 useful things you'll learn in my new post-training textbook (shipping now!)
  25. How invideo improves color grading 3x with GPT‑6 Astra

Sources & citations

rss14013
x2366
github435
reddit251
show_hn400
hackernews300
github_search200
swipe200
yc-rfs130
gmail80
yc-launch20
seed_urls10
crackscan00
ossinsight00
reddit_users00
total57825