Benchmarking Pocket-Scale Inference
Why CEREBRO kept it
Edge LLM inference benchmarks; token/cost optimization.
The text below is an automated extraction of the article at https://artificialanalysis.ai/hardware-inference-stack/mobile-phones, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (artificialanalysis.ai).
Benchmarking pocket-scale inference We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process. “Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page. Intelligence and Inference Performance Sum
Community take
Benchmark optimizes for flagships; production bottleneck is older phones/iPads capped at 4GB RAM.
Backlinks
Appeared in 1 briefing