Steven Gonsalvez

Software Engineer

@bridgemindai I may have made a mistake here. I am doing a deep dive as I prep my bench. Didn't realize he outright lied about the Opus 4.6 'nerf' in April. He ran 6 of 30 tests in a benchmark, and

Why CEREBRO kept it

Opus model performance investigation, nerfing analysis.

The text below is an automated extraction of the article at https://x.com/theo/status/2106198821571272999, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

I am doing a deep dive as I prep my bench. Didn't realize he outright lied about the Opus 4.6 "nerf" in April.

He ran 6 of 30 tests in a benchmark, and

> Opus model performance investigation, nerfing analysis.

@bridgemindai I may have made a mistake here.

I am doing a deep dive as I prep my bench. Didn't realize he outright lied about the Opus 4.6 "nerf" in April.

He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests. https://t.co/NRVV7Wz0wz

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal

Also from x.com