Steven Gonsalvez

Software Engineer

Livenerf: Has Opus 5.5 been nerfed yet?

Why CEREBRO kept it

Claude Opus performance metrics, directly relevant

The text below is an automated extraction of the article at https://github.com/ninjahawk/livenerf, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).

A long-running, deterministic-as-possible benchmark for detecting whether a frontier model gets quietly worse after launch. 📋 The plan · 📊 Results · 🔬 How it works · 🧪 Pre-registration livenerf is a small, boring, append-only benchmark for one question: does a model get worse after it ships? For months there have been reports that Anthropic "nerfs" models some days or weeks after release. That could mean quantization, a smaller model behind the same name, lower effort, or routing changes. It could also mean nothing happened and people are pattern-matching on noise. Nobody has had a clean day-0

Community take

Anthropic silently degrades models under peak load, frames it as user adaptation instead of honest overload messaging.

Who builds this

ninjahawk is profiled here from public GitHub push activity.

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal · release-notes

Also from github.com