CEREBRO
machine-read, human-curated
Coding agents
Coding agents & CLI tooling
- Claude Code v2.1.211 adds
--forward-subagent-textfor subagent text/thinking in stream-json, plus fixes bidirectional-override/zero-width/look-alike char neutralization in permission previews relayed to chat channels (prompt-injection defense). - Claude Code v2.1.221 ships VSCode Focus view (collapses tool activity into per-turn summary) and
mode: "mask"for sandboxed credential files on Linux/WSL. - Claude Code v2.1.212 reworks
/forkto spawn a background session (own row inclaude agents) instead of an in-session subagent; that role now lives in new/subtask. - Claude Code v2.1.214 fixes real security bugs: single-segment
dir/**allow-rules over-matching nested dirs, a Windows PowerShell 5.1 permission-check bypass, and Bash checks failing open instead of closed. - Claude Code v2.1.215 stops auto-running
/verifyand/code-review, now opt-in only. - GitHub Copilot cloud agent: reasoning level now configurable per task, and automations can trigger off issue/PR comments (e.g. "generate docs" via comment).
- GitHub Copilot: several models deprecate Sept 1 2026 across Chat, inline edits, agent mode, completions, plan migrations now.
- free-claude-code: proxy layer to run Claude Code/Codex/Pi against your own provider/local models instead of paid subscriptions, admin UI to validate providers.
- mpai: makes existing Codex/Claude Code sessions multiplayer.
- Reddit thread comparing Sonnet 5 high-effort vs Opus 5 low-effort, relevant to your reasoning-tier routing calls.
Agentic SaaS / orchestration
- Orca: ADE for running fleets of parallel coding agents (Codex/Claude Code/OpenCode/Pi), each in own worktree, one tracked view, desktop+mobile, uses your own subscriptions not per-seat billing.
- Hoplite (YC S26, HN launch): one-thread-one-VM cloud agent deploy, HN notes it's functionally identical to Amp's existing model.
- AgentSky: cloud-hosted agents, any harness + any LLM.
- SnapState: persistent state layer for AI agent workflows, addresses the "agent forgets everything between runs" problem.
LLM/model releases
- GLM-5.2: interconnects.ai calls it a genuine capability threshold for open agentic models, worth tracking vs closed frontier for agent workloads.
- Qwen 3.8 Max (2.4T) + 27B: Qwen back with a monster MoE flagship plus small model, signals renewed open-weights competition for coding/agent use.
- Cloudflare Workers AI: running Kimi K-series and GLM at scale, HN pushback that quantization isn't disclosed on the product page.
- Baseten inference engineering masterclass: infra deep-dive off a $13B round, includes a Kimi K3 breakdown, relevant if you care about inference cost/latency tradeoffs behind agent backends.
Vibe-coding / local inference
- Swiftlet (HN): runs 80B Qwen in 4.3GB RAM on Mac and 35B on iPhone by streaming MoE experts from disk on demand, HN flags prefill as the real bottleneck (~10 tok/hr disk-swap degenerate case).
- AirLLM: runs 70B+ models (up to Kimi K3 2.8T) on 4-12GB GPUs without quantization by streaming layers, useful if you want to run big local models on modest hardware.
- ds4 (DwarfStar): narrow, self-contained local inference engine tuned specifically for DeepSeek V4 Flash/GLM 5.2 on Metal/CUDA/ROCm, not a general GGUF runner.
- livekit/agents: framework for realtime voice AI agents (STT/LLM/TTS pipeline), server-side multimodal agent building block.
- alibaba/open-code-review: Alibaba's internal AI code review tool open-sourced, hybrid deterministic-pipeline + LLM-agent architecture with line-level comments for NPE/thread-safety/XSS/SQLi.
Signals
- stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
- Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
- v2.1.211
- lyogavin/airllm: AirLLM 70B inference with single 4GB GPU
- GLM-5.2 is the step change for open agents
- Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
- Alishahryar1/free-claude-code: Use Claude Code, Codex and Pi for free from your terminal, app, IDE, or phone like OpenClaw (voice supported)
- v2.1.221
- v2.1.212
- mpai
- antirez/ds4: DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
- Sonnet 5 high effort vs Opus 5 low effort
- v2.1.214
- livekit/agents: A framework for building realtime voice AI agents 🤖🎙️📹
- alibaba/open-code-review: Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines…
- v2.1.215
- AgentSky
- (AINews) Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
- Smaller, faster, safer: running Kimi and GLM at scale
- Customize the reasoning level for Copilot cloud agent
- Trigger Copilot automations with comments
- Upcoming August 2026 model deprecations in GitHub Copilot
- I built a terminal with Claude to replace Claude Desktop
- SnapState - Persistent state for AI agent workflows
- The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| rss | 140 | 13 |
| github | 33 | 6 |
| hackernews | 60 | 4 |
| 25 | 2 | |
| show_hn | 40 | 0 |
| yc-rfs | 13 | 0 |
| gmail | 9 | 0 |
| yc-launch | 5 | 0 |
| x | 0 | 0 |
| total | 325 | 25 |