Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents & harnesses

Multi-agent systems & security

  • Patterns and problems in emerging multi-agent systems: Anthropic's Frontier Red Team study on agent-to-agent interaction risk as agents increasingly share codebases and markets; community notes Opus 5's shift toward agent-legible (less human-readable) design as a deliberate bet on designed coordination over raw intelligence.
  • MCP reaches a security inflection point: 21,000+ internet-facing MCP servers found exposed, 92% without OAuth; OWASP MCP Top 10 forming, community pushing for npm-style signed lockfiles over server code before trusting it.
  • ProofRun: local, cryptographic verification receipts proving which checks actually ran against the current code state, direct answer to agents claiming "tests pass" from stale memory.

Agent memory & context

  • TencentDB Agent Memory: team-level shared memory hub turning chats/docs/code into four reusable asset types (Chat Memory, Skill, LLM-Wiki, Code-Graph) governed across agents and frameworks.
  • SnapState: persistent state layer for AI agent workflows, aimed at surviving context resets across long agent runs.
  • Ontologies Are So Back: argues formal ontologies are being revived to keep probabilistic agents inside deterministic boundaries, a structured alternative to pure prompt-based context management.

Models & training

  • GLM-5.2 is the step change for open agents: flags GLM-5.2 crossing a capability threshold for open-weight agentic models, paired with a new open RL recipe paper for terminal agents.
  • GLM-5.3: follow-on release billed as a coding capability jump from scaled post-training on the same base model.
  • Frontier post-training recipe review: podcast with Finbarr Timbers on what it'd take to push an Olmo-style open recipe to frontier post-training quality.
  • Soup: one-YAML LLM fine-tuning CLI with layer streaming, claims an 8B model trainable on a 4GB laptop GPU.
  • Inferock Bench: independent, verifiable receipt system for every LLM API call, adjacent to the ProofRun trust-verification trend.

Agentic SaaS & automation

  • citrolabs/ego-lite: browser built to share your logged-in session state with coding agents (Codex, Claude Code) without disrupting your own tabs, positioned against browser-use/agent-browser's separate-browser model.
  • stablyai/orca: desktop/mobile ADE for running a fleet of agents (Codex, Claude Code, OpenCode, Pi) side by side, each in its own worktree with usage/rate-limit tracking.
  • BrowserAct Cloud: one-prompt web scraping product, agent-driven data extraction as a hosted service.
  • Automated Outlook/Zoom scheduling from email content: agent pipeline that schedules meetings by parsing who-said-what in email threads.
  • HKUDS/CLI-Anything: generates CLI wrappers for arbitrary software so agents can drive it, plus a CLI-Hub package registry for sharing agent-usable CLIs.
  • Phind: GPT-4-powered developer search; community flags it eliminates SEO junk but hallucinates on sparse results and has an unsustainable unlimited-API-cost model.

Signals

  1. citrolabs/ego-lite: The fastest browser for AI agents to run browser automation, built for sharing your logged-in browser state with your A…
  2. addyosmani/agent-skills: Production-grade engineering skills for AI coding agents.
  3. 21,000 MCP servers exposed: the protocol reaches a security inflection point
  4. ProofRun – a local verification receipt for AI coding agents
  5. stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
  6. TencentCloud/TencentDB-Agent-Memory: TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and cod…
  7. React for Agents: Astro Creator Brings Hooks to his Meta-Harness, Flue
  8. Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
  9. Patterns and problems in emerging multi-agent systems
  10. ClaudeMax - Made a Free and Open-Source Claude Code Cost analyzer and world user leaderboard. Enjoy, all I ask is a Star! brew install clau…
  11. Made a "fleet view" for multiple Claude (and Codex/other) sessions running at once — no tmux required
  12. SnapState - Persistent state for AI agent workflows
  13. MakazhanAlpamys/Soup: Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
  14. HKUDS/CLI-Anything: "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
  15. Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web
  16. BrowserAct Cloud
  17. I made a free Claude Code plugin that turns your agents into pixel creatures you can watch work 🦀
  18. Running three Claude Code sessions in parallel with git worktrees
  19. nopus - deterministically detect and automatically rewrite complex responses
  20. GLM-5.2 is the step change for open agents
  21. Frontier post-training recipe review with Finbarr Timbers
  22. Inferock Bench
  23. GLM-5.3
  24. Automated scheduling Outlook cal Zoom mtgs by who said what in email.
  25. Show HN: GPT-4-powered web searches for developers

Sources & citations

rss1408
github346
reddit256
hackernews604
show_hn401
yc-rfs130
yc-launch70
gmail10
x00
total32025