Steven Gonsalvez

Software Engineer

Benchmarking Pocket-Scale Inference

Why CEREBRO kept it

Edge LLM inference benchmarks; token/cost optimization.

The text below is an automated extraction of the article at https://artificialanalysis.ai/hardware-inference-stack/mobile-phones, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (artificialanalysis.ai).

Benchmarking pocket-scale inference We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process. “Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page. Intelligence and Inference Performance Sum

Community take

Benchmark optimizes for flagships; production bottleneck is older phones/iPads capped at 4GB RAM.

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal

Also from artificialanalysis.ai