Steven Gonsalvez

Software Engineer

Benchmarking Opus 5 on SlopCodeBench

Why CEREBRO kept it

Opus 5 coding benchmarking data

The text below is an automated extraction of the article at https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-opus-5-on-slop-code-bench.md, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).

I've written before something along the lines of: That wasn't entirely true. I love nothing more than burying a good lede. Last Friday I dug into SlopCodeBench, a new-ish (March 2026) long-horizon coding benchmark from @GOrlanski's lab at UW Madison. It addresses the thing that bothers me most about coding benchmarks - that even "larger" more complex benchmarks still divulge the whole problem up front: In contrast, each challenge in SlopCodeBench has multiple "checkpoints" - the model doesn't know the whole problem up front, it has to evolve the codebase over time as new requirements are divul

Community take

Opus 5 isn't revolutionary—users report overconfident verbose output and are reverting to Fable; the speed and token gains are real but don't justify the quality regression.

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · release-notes

Also from github.com