Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant
Why CEREBRO kept it
Local LLM token/context optimization: 95GB, 41–52 tok/s
The text below is an automated extraction of the article at https://www.reddit.com/r/LocalLLaMA/comments/1wva7l2/running_955_gib_qwen38flashnext_at_4152_toks_on_a/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (reddit.com).
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant
Backlinks
Appeared in 1 briefing