Steven Gonsalvez

Software Engineer

RT @levie: At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels,

Why CEREBRO kept it

Opus 5.5 frontier performance validated on agentic enterprise tasks.

The text below is an automated extraction of the article at https://x.com/bcherny/status/2102996614919094693, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

Overall, we saw frontier capability levels,

> Opus 5.5 frontier performance validated on agentic enterprise tasks.

RT @levie: At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent.

Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing.

Here are some examples of the task wins and performance gains across a variety of industry tests that we performed:

• Financial

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/agents · ai/llm-mechanics · cerebro/signal

Also from x.com