Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents & CLI orchestration

  • Claude Code next version: subagents run background by default, keep chatting while they work, foreground opt-in via explicit ask (tweet).
  • Orca — ADE running Codex/ClaudeCode/OpenCode/Pi side-by-side, each own worktree, one dashboard, usage/rate-limit tracking across providers.
  • herdr — terminal agent multiplexer, detach/reattach over ssh, sessions survive restarts, socket API lets agents spawn/control panes themselves.
  • jcode — new coding agent harness targeting multi-session workflows and performance ceiling.
  • pi — self-extensible agent harness: unified LLM API, agent loop, TUI, CLI, modular npm packages.
  • MagicPath 2.0 — multiplayer canvas where humans + Codex/Claude Code build prototypes together, live agent view.
  • mattpocock/skills — small composable Claude/agent skills library, model-agnostic, alternative to heavier frameworks (GSD/BMAD/Spec-Kit).
  • Vercel's Andrew Qu on agents as new software category: eve framework, skills/sandboxes/agent-readable sites as the emerging stack (latent.space).
  • Autoresearch: self-improving agent "recipes" and outer feedback loops, humans still central (latent.space).

Agent security & infra

  • agent-security — open-sourced preflight gate for coding agents: catches poisoned repos, hostile content leaking into commits, destructive git commands before trust boundary crossed.
  • cua — open computer-use drivers for cross-OS agent fleets, drives native apps in background without stealing cursor/focus, plus benchmarks (Cua Bench).
  • SnapState — persistent state layer for agent workflows (HN).
  • Axon self-documents its API for agent-to-agent hiring/payment, no human in loop for integration (tweet).

Token/context economics

  • Deep research pipeline burned a whole Claude Max 5x limit in 30 min; lesson: dynamic prompts kill caching, static retrieval beats micro-optimization (quesma.com).
  • Reddit skill routes "thinking" to Claude, grunt coding to cheaper/free models as executor, judge/executor split for cost control (reddit).

Open model landscape

  • GLM-5.2 flagged as step-change for open agents, capability threshold worth tracking (interconnects.ai).
  • Open model bonanza: Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 all shipped same month, CAISI ran independent V4 assessment (interconnects.ai).
  • Claimed roadmap chatter: GLM-5.2 near-Opus, Kimi-K3 near-Fable, Qwen-3.8 (2.4T) and DeepSeek V4 GA ($0.0028/M) expected to beat Opus (tweet, unverified vendor claims).
  • Alibaba Qoder claims #1 China AI coding market, 47.6% share per IDC, 5M+ users (tweet).
  • LoRA Speedrun — public wall-clock leaderboard for fine-tuning Qwen2.5-1.5B to GSM8K threshold, independently re-run 3x on fixed hardware, modded-nanogpt for fine-tuning.

Vibe-coding & agentic SaaS

  • Founder claims vibe-coded custom CRM replaced Salesforce, cut $600k/yr software spend to zero, better AI-agent integration than off-shelf (tweet, anecdotal).
  • Vibe-Trading — personal trading agent toolkit, note: security warning re impersonating X/token accounts.
  • WrenAI — open-source GenBI, agents generate/govern text-to-SQL dashboards across 20+ data sources (BigQuery, Snowflake, Databricks etc) via governed context layer.
  • AstrBot — open-source agent chatbot platform integrating IM platforms + LLMs + plugins, positions as openclaw alternative.

CLI/TUI

  • Homebrew 6.0.0 — new tap-trust security model, faster internal JSON API, Linux sandboxing, macOS 27 support; community noting forced-upgrade friction pushing users toward Mise for version pinning.

Signals

  1. stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
  2. Introducing MagicPath 2.0. MagicPath is now a multiplayer canvas for humans and agents like Codex or Claude Code to design and build with A…
  3. 1jehuang/jcode: Coding Agent Harness
  4. trycua/cua: Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
  5. earendil-works/pi: AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
  6. "We replaced Salesforce with a vibe-coded CRM built for our own workflows. The custom system integrated our AI agents more effectively, wor…
  7. ogulcancelik/herdr: agent multiplexer that lives in your terminal.
  8. In the next version of Claude Code: subagents run in the background by default, so you can keep talking to Claude while your subagents work…
  9. mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.
  10. Vercel's Andrew Qu on why agents are a new kind of software
  11. I open-sourced `agent-security` today. Coding agents can clone a poisoned starter, read hostile content, leak private context into a public…
  12. SnapState - Persistent state for AI agent workflows
  13. AstrBotDevs/AstrBot: AI Agent Assistant & development framework that integrates lots of IM platforms, LLMs, plugins and AI feature, and can…
  14. If agents are going to hire and pay each other, they have to be able to read the network themselves. Not wait for a human to wire them in.…
  15. I burned all my tokens researching how to save tokens
  16. Autoresearch: The feedback loop behind self-improving agents
  17. Canner/WrenAI: GenBI (Generative BI) for AI agents, an open-source, governed text-to-SQL through an open context layer that turns natural-l…
  18. #AlibabaCloud ranks #1 in China’s AI Coding Market with a 47.6% market share, according to #IDC. Powered by #Qoder, it is driving the shift…
  19. A skill that saves Claude usage for thinking (judge) and hands the grunt-work coding to cheaper/free LLM models (executor)
  20. HKUDS/Vibe-Trading: "Vibe-Trading: Your Personal Trading Agent"
  21. Open Source AI updates : > GLM-5.2 ✅ Almost Opus-level > Kimi-K3 ✅ Almost Fable-level > Qwen-3.8 🔜 2.4T, Expected to beat Opus &g…
  22. GLM-5.2 is the step change for open agents
  23. Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.
  24. LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques
  25. Show HN: Homebrew 6.0.0

Sources & citations

github429
x1837
rss1404
hackernews603
show_hn401
reddit251
yc-rfs160
yc-launch40
gmail00
total51025