Qwen3.8-Flash-Next
Why CEREBRO kept it
Open-weights multimodal MoE model, directly enables agentic systems.
The text below is an automated extraction of the article at https://simonwillison.net/2026/Aug/26/qwen38-flash-next/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (simonwillison.net).
<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost.</p> <p>I've been trying it out on a DGX Spark using <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">these Unsloth quantized models</a>. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing <a hre
Backlinks
Appeared in 1 briefing