RT @aicodeking: I agree with Theo here. A Claude Code Update, A Provider Change, A cache hit change, A bad request, New seed anything can change a prompt's outcome in just a matter of minutes. These b
Why CEREBRO kept it
Deep LLM reliability mechanics: nerfs, provider changes, cache.
The text below is an automated extraction of the article at https://x.com/theo/status/2106224871558840504, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
RT @aicodeking: I agree with Theo here. A Claude Code Update, A Provider Change, A cache hit change, A bad request, New seed anything can change a prompt's outcome in just a matter of minutes. These bench's don't make any sense unless done extremely carefully with some kind of private API by each model provider or something as most providers now block Top P and stuff.
Backlinks
Appeared in 1 briefing