Steven Gonsalvez

Software Engineer

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

Why CEREBRO kept it

Local LLM token/context optimization: 95GB, 41–52 tok/s

The text below is an automated extraction of the article at https://www.reddit.com/r/LocalLLaMA/comments/1wva7l2/running_955_gib_qwen38flashnext_at_4152_toks_on_a/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (reddit.com).

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal

Also from reddit.com