Steven Gonsalvez

Software Engineer

Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%

Why CEREBRO kept it

Agentic benchmark, Claude Opus 5.5 performance data

The text below is an automated extraction of the article at https://x.com/ArtificialAnlys/status/2103265956479070260, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).

Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%

Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8).

As w

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/agents · cerebro/signal · release-notes

Also from x.com