CEREBRO
machine-read, human-curated
Frontier models
- Anthropic shipped Claude Opus 5, positioned as near-frontier intelligence at roughly half the cost of prior Opus tiers (Anthropic launches Claude Opus 5, PH listing).
- Early pushback on r/ClaudeAI: some report Opus 5 regressing vs Opus 4.8 on complex, context-heavy work rather than improving it (Is Opus 5 a regression from Opus 4.8 on complex, context-heavy work?).
- SlopCodeBench (long-horizon coding benchmark) run against Opus 5 shows real speed/token gains, but community reaction says output quality regressed, verbose and overconfident, with some reverting to Fable (Benchmarking Opus 5 on SlopCodeBench).
- SpaceXAI's Grok 4.5 lands as the first Opus-class model since the Cursor acquisition, raising the frontier pace again (SpaceXAI launches Grok 4.5).
- GLM-5.2 marks a capability step-change for open-weight agentic models per Interconnects' ongoing threshold tracking (GLM-5.2 is the step change for open agents).
- Latent Space digests the "Field Guide to Fable 5," calling it the most significant model launch to date as people race to probe its limits before subsidy pricing ends (The Field Guide to Fable).
Coding agents & harness engineering
- GitHub's Copilot playbook argues the harness (prototype→plan→implement→review loop), not the model du jour, is what actually matters (The harness is all you need (mostly)).
- Copilot code review got worse after tool upgrades until GitHub reshaped the agent's workflow around shared Unix-style exploration tools and PR evidence, cutting review cost (Better tools made Copilot code review worse).
- Lilian Weng condenses 35 papers on "harness engineering" for recursive self-improvement, a solid reading-list dump on the emerging discipline (Lilian Weng summarizes 35 papers on Harness Engineering for RSI).
- Alibaba open-sourced its internal AI code review CLI, battle-tested at scale: deterministic pipelines plus an LLM agent and a fine-tuned ruleset for NPE/thread-safety/XSS/SQLi (alibaba/open-code-review).
- Reddit writeup on staged evaluator pipelines covers gate design and loop control for agent verification stages, relevant to anyone building eval/retry harnesses (Staged evaluator pipelines: gate design and loop control).
- GitHub Copilot for JetBrains adds MCP server/custom-agent support plus better OpenTelemetry config and model management (GitHub Copilot for JetBrains adds improved OpenTelemetry configuration and model management).
- New Claude Code skill searches Reddit/X/YouTube/HN/Polymarket/web and synthesizes a grounded, engagement-scored summary of any topic over the last 30 days (mvanhorn/last30days-skill).
LLM mechanics & token efficiency
- Essay argues coding is shifting from hand-crafted code to an "age of libraries," where token efficiency and reusable context matter more than raw generation (The age of token efficiency, the age of libraries).
- SnapState pitches persistent, resumable state for AI agent workflows, a building block for long-running agent context/memory (SnapState).
Agentic SaaS
- Modal's CTO argues cloud infra built for humans needs redesigning for agents, framing "Agent Experience" as the next infra shift (Why AI Infrastructure must evolve for Agent Experience).
- HeyZoku lets you orchestrate a fleet of coding agents by voice (HeyZoku).
- Orca (stablyai) is an ADE for running Codex/Claude Code/OpenCode/Pi side by side, each in its own worktree with unified usage tracking (stablyai/orca).
- Webhound bills itself as a dedicated research engine for agents to query (Webhound).
- Cynative launched a read-only cloud security research agent: "ask your cloud anything without breaking prod" (Cynative Security Research Agent).
- Comms puts AI agents inside iMessage for quick agent-driven messaging (Comms).
- GitHub now extends enterprise managed settings to the Copilot app and cloud agent, giving admins centralized policy control (Enterprise managed settings in the GitHub Copilot app and Copilot cloud agent).
CLI/TUI
- superfile is a modern visual terminal file manager, worth a look for anyone living in the shell (superfile).
Vibe-coding & repos
- Reddit post shows a technique for having Claude Code package apps as a single sendable file, lowering distribution friction for vibe-coded projects (I built a way for Claude Code to make apps you can send as one file).
Signals
- The harness is all you need (mostly)
- Better tools made Copilot code review worse. Here’s how we actually improved it.
- Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
- HeyZoku
- Anthropic launches Claude Opus 5 with near-frontier power at half the cost
- Is Opus 5 a regression from Opus 4.8 on complex, context-heavy work?
- (AINews) SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition
- (AINews) Lilian Weng summarizes 35 papers on Harness Engineering for RSI
- GitHub Copilot for JetBrains adds improved OpenTelemetry configuration and model management
- The age of token efficiency, the age of libraries
- Claude Opus 5
- superfile
- alibaba/open-code-review: Open-source & free — Battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipeli…
- stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
- (AINews) The Field Guide to Fable
- GLM-5.2 is the step change for open agents
- Webhound
- SnapState - Persistent state for AI agent workflows
- Cynative Security Research Agent
- Comms
- mvanhorn/last30days-skill: AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesiz…
- I built a way for Claude Code to make apps you can send as one file
- Benchmarking Opus 5 on SlopCodeBench
- Enterprise managed settings in the GitHub Copilot app and Copilot cloud agent
- Staged evaluator pipelines: gate design and loop control
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| rss | 140 | 15 |
| hackernews | 60 | 3 |
| github | 49 | 3 |
| 25 | 3 | |
| gmail | 9 | 1 |
| show_hn | 40 | 0 |
| yc-rfs | 13 | 0 |
| yc-launch | 5 | 0 |
| x | 0 | 0 |
| total | 341 | 25 |