Steven Gonsalvez

Software Engineer

radixark/miles: Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Why CEREBRO kept it

RL post-training framework, novel agentic pattern

The text below is an automated extraction of the article at https://github.com/radixark/miles, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).

| Website | Documentation | Quick Start | Supported Models | Miles Diffusion | Blog | Slack (#miles-rl) | - [2026/08] 🔥 Miles v0.1 is released! Read the blog post here: Miles v0.1: Production-level Post-training. - [2026/07] Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles (blog). - [2026/07] 🔥 SGLang and Miles add day-0 support for Kimi K3 (blog). - [2026/07] On-policy distillation lands in Miles (blog). - [2026/07] 🔥 SGLang and Miles add day-0 support for Inkling, a frontier multimodal model (blog). - [2026/07] DeepSeek-V4 Flash RL training comes to AMD Ins

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics · cerebro/signal · repo/trending

Also from github.com