jundot/omlx: LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Why CEREBRO kept it
LLM inference optimization and macOS menu-bar CLI; both on-topic
The text below is an automated extraction of the article at https://github.com/jundot/omlx, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).
LLM inference, optimized for your Mac Continuous batching and tiered KV caching, managed directly from your menu bar. junkim.dot@gmail.com · https://omlx.ai/me Install · Quickstart · Features · Models · CLI Configuration · Benchmarks · oMLX.ai Every LLM server I tried made me choose between convenience and control. I wanted to pin everyday models in memory, auto-swap heavier ones on demand, set context limits - and manage it all from a menu bar. oMLX persists KV cache across a hot in-memory tier and cold SSD tier - even when context changes mid-conversation, all past context stays cached and r
Backlinks
Appeared in 1 briefing