Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents

  • Claude Code v2.1.222 fixes worktree-isolated sessions being able to run destructive git against the main checkout, and closes a PreToolUse auto-allow bypass in background agent tasks, security-relevant if you run subagents/worktrees.
  • Claude Code v2.1.216 adds sandbox.filesystem.disabled to skip filesystem isolation while keeping network egress control, and fixes a quadratic slowdown in long sessions that was stalling multi-turn resumes.
  • Claude Code v2.1.215: /verify and /code-review no longer auto-run, invoke them explicitly now.
  • Claude Code v2.1.212: /fork now clones your conversation into a background session instead of an in-session subagent (that's now /subtask), plus claude auto-mode reset.
  • GitHub engineering on turning one giant AI-generated PR into a reviewable stack: teach agents to decompose output into ordered stacked PRs instead of one unreviewable diff.
  • alibaba/open-code-review: Alibaba's internal AI code-review CLI (deterministic pipelines + LLM agent, line-level comments, NPE/thread-safety/XSS/SQLi ruleset) now open source, Anthropic/OpenAI compatible.
  • EveryInc/compound-engineering-plugin: skill pack for Claude Code/Codex/Cursor aimed at making each unit of work compound off the last, moved to a root-native plugin layout.
  • uber/ADR: Uber's open-sourced observability/threat-detection layer for enterprise agents (Cursor, Claude Code, Codex, customer-facing bots), paper accepted at MLSys 2026.
  • browser-use/video-use: open-source video editing driven entirely from Claude Code chat (cuts filler, color grades, assembles from raw footage).

LLM mechanics & models

  • Kimi K3 2.8T-A50B: Moonshot's largest open model release yet, pitched as Opus 4.8-class capability at Sonnet 5 pricing, open-weight frontier gap keeps closing.
  • GLM-5.2: Nathan Lambert flags it as a capability step-change for open agents, alongside a new paper on open RL recipes for terminal agents.
  • llm-anthropic 0.26: Simon Willison's llm plugin adds claude-fable-5, claude-sonnet-5, claude-opus-5 support, riding the LLM 0.32 feature set below.
  • Claude Fable 5 and new AI safety fables: Anthropic's silent classifier-based manipulation of AI-research queries now applies uniformly across safety domains, still contentious on transparency.

Agent memory

  • Zero-Mem: proposes agent memory ops that skip extra LLM calls for writing/retrieving memory, cutting recurring token/latency cost; HN pushback is that the real win is preserving original traces for auditability, not the zero-token framing, compression risks silently dropping evidence.
  • Why an LLM's memory gets expensive and how to fix it: explainer on why long-context prompts cost disproportionately more than short ones at the same model/hardware, KV-cache mechanics.
  • TencentDB Agent Memory: team-level memory hub turning chats/docs/code into four shared asset types (Chat Memory, Skill, LLM-Wiki, Code-Graph) reusable across agents and frameworks.

Agentic SaaS

  • Unpacking ChatGPT Work: outside reconstruction of how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills and Tools actually fit together in ChatGPT Work.
  • OpenAI ships education-focused plugins for ChatGPT Work and Codex, guided modes for K-12 teachers and college students/educators.
  • SnapState: persistent state layer for AI agent workflows, hackernews launch.
  • Phind: GPT-4-powered dev search engine; HN take is the SEO-noise filtering is the real value, but hallucination on thin-result queries and unbounded API cost undercut it.
  • Glasp MCP Connector: exposes your Glasp highlights/notes as an MCP source inside Claude and ChatGPT.
  • Finyuus: code-first language for durable, governed AI workflows.

Vibe-coding

  • Maple-Preview: ternary 20B MoE model running at 120 tok/s on-device on iPhone; HN notes confident hallucination is the real blocker for small on-device models, they need reliable tool-calling since they can't memorize facts or flag their own uncertainty.

Signals

  1. Zero-Mem: Zero-Token Memory Operations for LLM Agents
  2. alibaba/open-code-review: Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines…
  3. Why An LLM’s Memory Gets Expensive and How to Fix It
  4. (AINews) Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
  5. llm-anthropic 0.26
  6. Claude Fable 5 and new AI safety fables
  7. Turn one giant AI-generated pull request to a reviewable stack
  8. EveryInc/compound-engineering-plugin: Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more
  9. Unpacking ChatGPT Work: the Agent for a Billion Users
  10. Finyuus
  11. New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging
  12. v2.1.222
  13. llm 0.32
  14. uber/ADR: ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
  15. browser-use/video-use: Edit videos with coding agents
  16. v2.1.215
  17. SnapState - Persistent state for AI agent workflows
  18. Show HN: GPT-4-powered web searches for developers
  19. v2.1.216
  20. v2.1.212
  21. TencentCloud/TencentDB-Agent-Memory: TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and cod…
  22. Glasp MCP Connector
  23. New ways to learn and teach with ChatGPT Work and Codex
  24. GLM-5.2 is the step change for open agents
  25. Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone

Sources & citations

rss14015
github355
hackernews603
show_hn401
gmail71
reddit250
yc-rfs130
yc-launch50
x00
total32525