CEREBRO
machine-read, human-curated
Coding agents
- v2.1.276 fixes a 2.1.275 regression in Claude Code where every request failed with
400 … Input tag 'advisor_20260301'whenANTHROPIC_BASE_URLpoints at a proxy or gateway, so upgrade if you route through one. - stablyai/orca is an ADE for desktop, mobile and remote that runs Codex, Claude Code, OpenCode or Pi side by side, each in its own worktree, with phone notifications and follow-ups when an agent finishes.
- Theo on Opus 5.5 filing a PR reports that a vague prompt produced a ready-to-merge T3 Code pull request with a video demo attached, an anecdote about agents owning the full PR loop.
- Theo on Claude Code vs Codex claims the $200 Claude Code plan is now far ahead of Codex, reversing the picture from a few weeks ago; sentiment only, no data.
- When chat is the wrong UI is GitHub's pitch for Copilot "canvases" as a more tangible interface than a chat box.
- AI-powered fuzzing with the GitHub Security Lab Taskflow Agent walks through a new fuzzing taskflow built on the Taskflow Agent framework, a concrete example of agent-driven security tooling.
- Default Enablement of Copilot Features for Copilot Business and Enterprise adds a global default policy enabling GA Copilot features and supported client capabilities in enterprise and org settings, with a 28-day window, so check admin settings now.
Models and benchmarks
- Terminal-Bench-Science 0.1 leaderboard launches an agentic benchmark for scientific research work, with GPT-6 Astra (max) at 63% and Claude Opus 5.5 (xhigh) at 62%.
- GLM-5.2 is the step change for open agents argues GLM-5.2 is a real jump for open-weight agents, and the post also announces a new paper on open RL recipes for terminal agents.
- Jev, a "System One Model" only decides, classifies, routes and scores, claiming >100x faster and >200x cheaper than small frontier LLMs, which matters for triage and routing layers in agent pipelines.
- Teaching Everyone to Fish for Tokens argues the Linux analogy for open models leaves a narrow path to a self-sustaining open-model ecosystem.
- 5 useful things in the post-training textbook announces the finished Manning book Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs.
Agent memory and context
- vectorize-io/hindsight is an agent memory system aimed at agents that learn over time rather than just recall history, positioned as better than RAG and knowledge-graph approaches on long-term memory.
- AI Engineering bonus thread lays out the full production stack (context, agent loop, memory, tools, harness, guardrails, execution, verification, observability, evals, improvement) around the model; a useful checklist, nothing new.
- Solo founder on project context says proper project context helped more than any model upgrade; title only, no body captured.
Vibe-coding and repos
- mattpocock/skills is a set of small, composable agent skills that work with any model, pitched against process-owning frameworks like GSD, BMAD and Spec-Kit that take control away and make bugs hard to trace.
- Raft goes source-available and accepts "prompt requests" instead of pull requests, an experiment in an open-source contribution model for the agent era.
Agentic SaaS
- superdesigndev/treg is OpenRouter for agent tools: one base URL and token reach 3,000+ endpoints across 60+ providers, priced per call from a cent, plus your team's own keys, skills and CLIs.
- anthropics/financial-services provides reference agents, skills and data connectors for investment banking, equity research, private equity and wealth management, deployable as a Claude Cowork plugin or through the Claude Managed Agents API.
- Underwriting Superintelligence: Backing Agents you can Sue covers AIUC's $40M Series A and AIUC-1, an agent standard backed by real insurance, from a founder who was Anthropic's first product hire.
- Parallel cut research time and cost in half with GPT-6 Astra is a vendor case study of agents synthesizing labor-market data in half the time and at half the cost of prior models.
Signals
- v2.1.276
- stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on d…
- Opus 5.5 just filed a pull request that changes the streaming behavior in T3 code. It's so cool that I can send off a prompt with a vague i…
- mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.
- (AINews) Jev: a “System One Model” that only decides/classifies/routes/scores — >100x faster, >200x cheaper than small frontier LLMs
- v2.1.272
- anthropics/financial-services
- AI Engineering- Bonus We spent 15 parts exploring the pieces. Now here’s how they fit together. → Context → Agent Loop → Memory → Tools → H…
- Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max)…
- It's insane how much better the $200 Claude Code plan is compared to Codex right now. Just a few weeks ago, it was the other way around. Ex…
- Teaching Everyone to Fish for Tokens
- Autonomous Product Delivery
- superdesigndev/treg: OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
- AI/ML Series (Lesson 15/100): Most AI agents fail because of the model. That is the biggest misconception in enterprise AI. The real failur…
- Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
- When chat is the wrong UI
- TL;DR: We just made Raft source-available. We don't accept pull requests. We want prompt requests instead. We are exploring what a new cont…
- GLM-5.2 is the step change for open agents
- Parallel cut research time and cost in half with GPT‑6 Astra
- vectorize-io/hindsight: Hindsight: Agent Memory That Learns
- Default Enablement of Copilot Features for Copilot Business and Enterprise
- AI-powered fuzzing with the GitHub Security Lab Taskflow Agent
- solo founder: setting up proper project context helped more than any model upgrade
- 5 useful things you'll learn in my new post-training textbook (shipping now!)
- How invideo improves color grading 3x with GPT‑6 Astra
Sources & citations
| Source | Fetched | In briefing |
|---|---|---|
| rss | 140 | 13 |
| x | 236 | 6 |
| github | 43 | 5 |
| 25 | 1 | |
| show_hn | 40 | 0 |
| hackernews | 30 | 0 |
| github_search | 20 | 0 |
| swipe | 20 | 0 |
| yc-rfs | 13 | 0 |
| gmail | 8 | 0 |
| yc-launch | 2 | 0 |
| seed_urls | 1 | 0 |
| crackscan | 0 | 0 |
| ossinsight | 0 | 0 |
| reddit_users | 0 | 0 |
| total | 578 | 25 |