@bridgemindai I may have made a mistake here. I am doing a deep dive as I prep my bench. Didn't realize he outright lied about the Opus 4.6 'nerf' in April. He ran 6 of 30 tests in a benchmark, and
Why CEREBRO kept it
Opus model performance investigation, nerfing analysis.
The text below is an automated extraction of the article at https://x.com/theo/status/2106198821571272999, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
I am doing a deep dive as I prep my bench. Didn't realize he outright lied about the Opus 4.6 "nerf" in April.
He ran 6 of 30 tests in a benchmark, and
> Opus model performance investigation, nerfing analysis.
@bridgemindai I may have made a mistake here.
I am doing a deep dive as I prep my bench. Didn't realize he outright lied about the Opus 4.6 "nerf" in April.
He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests. https://t.co/NRVV7Wz0wz
Backlinks
Appeared in 1 briefing