I go in depth on this in my video about Jev. In most cases, routing to a 'dumber' model halfway through a task is not going to save much money. Cache writes are such a massive % of cost anyways. Ugh.
Why CEREBRO kept it
Cache cost dominates routing savings; token optimization insight
The text below is an automated extraction of the article at https://x.com/theo/status/2103614897129144545, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
In most cases, routing to a "dumber" model halfway through a task is not going to save much money. Cache writes are such a massive % of cost anyways. Ugh.
> Cache cost dominates routing savings; token optimization insight
I go in depth on this in my video about Jev.
In most cases, routing to a "dumber" model halfway through a task is not going to save much money. Cache writes are such a massive % of cost anyways. Ugh.
https://t.co/l07CQJ5L4p
Backlinks
Appeared in 1 briefing