Benchmarking Opus 5 on SlopCodeBench
Why CEREBRO kept it
Opus 5 coding benchmarking data
The text below is an automated extraction of the article at https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-opus-5-on-slop-code-bench.md, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).
I've written before something along the lines of: That wasn't entirely true. I love nothing more than burying a good lede. Last Friday I dug into SlopCodeBench, a new-ish (March 2026) long-horizon coding benchmark from @GOrlanski's lab at UW Madison. It addresses the thing that bothers me most about coding benchmarks - that even "larger" more complex benchmarks still divulge the whole problem up front: In contrast, each challenge in SlopCodeBench has multiple "checkpoints" - the model doesn't know the whole problem up front, it has to evolve the codebase over time as new requirements are divul
Community take
Opus 5 isn't revolutionary—users report overconfident verbose output and are reverting to Fable; the speed and token gains are real but don't justify the quality regression.
Backlinks
Appeared in 1 briefing