Steven Gonsalvez

Software Engineer

Qwen3.8-Flash-Next

Why CEREBRO kept it

Open-weights multimodal MoE model, directly enables agentic systems.

The text below is an automated extraction of the article at https://simonwillison.net/2026/Aug/26/qwen38-flash-next/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (simonwillison.net).

<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost.</p> <p>I've been trying it out on a DGX Spark using <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">these Unsloth quantized models</a>. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing <a hre

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal · release-notes

Also from simonwillison.net