Steven Gonsalvez

Software Engineer

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choos

Why CEREBRO kept it

Token cost optimization, routing, caching strategies

The text below is an automated extraction of the article at https://x.com/brian_armstrong/status/2070670644577280109, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

Better Defaults (not Usage Caps) – Engineers can choos

> Token cost optimization, routing, caching strategies

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching.

Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a d

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics

Also from x.com