we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 🧵 part 6: Transformer Inference Arithmetic Carol Chen (ki
Why CEREBRO kept it
Comprehensive AI perf engineering series, infrastructure focus
The text below is an automated extraction of the article at https://x.com/wafer_ai/status/2104441237545898082, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (x.com).
follow and save to keep up with the series. links in thread 🧵
part 6: Transformer Inference Arithmetic
Carol Chen (ki
> Comprehensive AI perf engineering series, infrastructure focus
we launched the most comprehensive ai performance engineering repo in the world
follow and save to keep up with the series. links in thread 🧵
part 6: Transformer Inference Arithmetic
Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s.
Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation:
- prefill and decode expose differ
Backlinks
Appeared in 1 briefing