radixark/miles: Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
Why CEREBRO kept it
RL post-training framework, novel agentic pattern
The text below is an automated extraction of the article at https://github.com/radixark/miles, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).
| Website | Documentation | Quick Start | Supported Models | Miles Diffusion | Blog | Slack (#miles-rl) | - [2026/08] 🔥 Miles v0.1 is released! Read the blog post here: Miles v0.1: Production-level Post-training. - [2026/07] Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles (blog). - [2026/07] 🔥 SGLang and Miles add day-0 support for Kimi K3 (blog). - [2026/07] On-policy distillation lands in Miles (blog). - [2026/07] 🔥 SGLang and Miles add day-0 support for Inkling, a frontier multimodal model (blog). - [2026/07] DeepSeek-V4 Flash RL training comes to AMD Ins
Backlinks
Appeared in 1 briefing