Steven Gonsalvez

Software Engineer

RT @theo: Last year, Anthropic was optimizing inference across 3 different types of compute (Nvidia, AWS trainium, Google tpus). There were mistakes made during these optimizations that resulted in ac

Why CEREBRO kept it

Anthropic inference optimization across compute types, token efficiency

The text below is an automated extraction of the article at https://x.com/theo/status/2104008628529410152, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

RT @theo: Last year, Anthropic was optimizing inference across 3 different types of compute (Nvidia, AWS trainium, Google tpus). There were mistakes made during these optimizations that resulted in actual degradation. Initially Anthropic denied it, but eventually they found, fixed and explained what happened. There have been no notable instances of degradation since.

Sadly, this one instance has made half of tech twitter’s brains fall out.

Separately, there’s the statistics side. You’re more likely to see the stupid spikes over longer windows.

If a model has a 1/50 chance of doing weird shi

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal · release-notes

Also from x.com