VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
Why CEREBRO kept it
Novel SFT+GRPO technique beating Opus; LLM mechanics.
The text below is an automated extraction of the article at https://arxiv.org/abs/2606.16140, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (arxiv.org).
Computer Science > Artificial Intelligence [Submitted on 15 Jun 2026] Title:VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models View PDF HTML (experimental)Abstract:This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime. Building upon the Spectrum-to-Signal post-training paradigm, we systematically enhance the model through an optimized pipeline that includes curriculum-based supervised fine-tuning, multi-domain re
Community take
VibeThinker excels at Python reasoning but fails outside its domain: found zero security bugs in benchmark, won't generalize to other languages.
Backlinks
Appeared in 1 briefing