Steven Gonsalvez

Software Engineer


CEREBRO


machine-read, human-curated

Coding agents

Coding agents & CLI tooling

  • Claude Code v2.1.211 adds --forward-subagent-text for subagent text/thinking in stream-json, plus fixes bidirectional-override/zero-width/look-alike char neutralization in permission previews relayed to chat channels (prompt-injection defense).
  • Claude Code v2.1.221 ships VSCode Focus view (collapses tool activity into per-turn summary) and mode: "mask" for sandboxed credential files on Linux/WSL.
  • Claude Code v2.1.212 reworks /fork to spawn a background session (own row in claude agents) instead of an in-session subagent; that role now lives in new /subtask.
  • Claude Code v2.1.214 fixes real security bugs: single-segment dir/** allow-rules over-matching nested dirs, a Windows PowerShell 5.1 permission-check bypass, and Bash checks failing open instead of closed.
  • Claude Code v2.1.215 stops auto-running /verify and /code-review, now opt-in only.
  • GitHub Copilot cloud agent: reasoning level now configurable per task, and automations can trigger off issue/PR comments (e.g. "generate docs" via comment).
  • GitHub Copilot: several models deprecate Sept 1 2026 across Chat, inline edits, agent mode, completions, plan migrations now.
  • free-claude-code: proxy layer to run Claude Code/Codex/Pi against your own provider/local models instead of paid subscriptions, admin UI to validate providers.
  • mpai: makes existing Codex/Claude Code sessions multiplayer.
  • Reddit thread comparing Sonnet 5 high-effort vs Opus 5 low-effort, relevant to your reasoning-tier routing calls.

Agentic SaaS / orchestration

  • Orca: ADE for running fleets of parallel coding agents (Codex/Claude Code/OpenCode/Pi), each in own worktree, one tracked view, desktop+mobile, uses your own subscriptions not per-seat billing.
  • Hoplite (YC S26, HN launch): one-thread-one-VM cloud agent deploy, HN notes it's functionally identical to Amp's existing model.
  • AgentSky: cloud-hosted agents, any harness + any LLM.
  • SnapState: persistent state layer for AI agent workflows, addresses the "agent forgets everything between runs" problem.

LLM/model releases

  • GLM-5.2: interconnects.ai calls it a genuine capability threshold for open agentic models, worth tracking vs closed frontier for agent workloads.
  • Qwen 3.8 Max (2.4T) + 27B: Qwen back with a monster MoE flagship plus small model, signals renewed open-weights competition for coding/agent use.
  • Cloudflare Workers AI: running Kimi K-series and GLM at scale, HN pushback that quantization isn't disclosed on the product page.
  • Baseten inference engineering masterclass: infra deep-dive off a $13B round, includes a Kimi K3 breakdown, relevant if you care about inference cost/latency tradeoffs behind agent backends.

Vibe-coding / local inference

  • Swiftlet (HN): runs 80B Qwen in 4.3GB RAM on Mac and 35B on iPhone by streaming MoE experts from disk on demand, HN flags prefill as the real bottleneck (~10 tok/hr disk-swap degenerate case).
  • AirLLM: runs 70B+ models (up to Kimi K3 2.8T) on 4-12GB GPUs without quantization by streaming layers, useful if you want to run big local models on modest hardware.
  • ds4 (DwarfStar): narrow, self-contained local inference engine tuned specifically for DeepSeek V4 Flash/GLM 5.2 on Metal/CUDA/ROCm, not a general GGUF runner.
  • livekit/agents: framework for realtime voice AI agents (STT/LLM/TTS pipeline), server-side multimodal agent building block.
  • alibaba/open-code-review: Alibaba's internal AI code review tool open-sourced, hybrid deterministic-pipeline + LLM-agent architecture with line-level comments for NPE/thread-safety/XSS/SQLi.

Signals

  1. stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
  2. Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
  3. v2.1.211
  4. lyogavin/airllm: AirLLM 70B inference with single 4GB GPU
  5. GLM-5.2 is the step change for open agents
  6. Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
  7. Alishahryar1/free-claude-code: Use Claude Code, Codex and Pi for free from your terminal, app, IDE, or phone like OpenClaw (voice supported)
  8. v2.1.221
  9. v2.1.212
  10. mpai
  11. antirez/ds4: DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
  12. Sonnet 5 high effort vs Opus 5 low effort
  13. v2.1.214
  14. livekit/agents: A framework for building realtime voice AI agents 🤖🎙️📹
  15. alibaba/open-code-review: Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines…
  16. v2.1.215
  17. AgentSky
  18. (AINews) Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
  19. Smaller, faster, safer: running Kimi and GLM at scale
  20. Customize the reasoning level for Copilot cloud agent
  21. Trigger Copilot automations with comments
  22. Upcoming August 2026 model deprecations in GitHub Copilot
  23. I built a terminal with Claude to replace Claude Desktop
  24. SnapState - Persistent state for AI agent workflows
  25. The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

Sources & citations

rss14013
github336
hackernews604
reddit252
show_hn400
yc-rfs130
gmail90
yc-launch50
x00
total32525