Miles runs RL end to end on NVIDIA Vera Rubin, out of the box: SGLang rollouts, Megatron training, one container image. And agentic RL with 64 concurrent sandboxes on the Vera CPU, right next to the GPUs. Thanks to @NVIDIAAI for the early access. Details in the blog 👇
SGLang@sgl_project · 20hWe brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference.
Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.
Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20% faster KDA verification, with bitwise-identical output
• 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step
SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Full results and engineering details 👉 lmsys.org/blog/2026-10-0…
Open quoted post →Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.
Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20% faster KDA verification, with bitwise-identical output
• 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step
SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Full results and engineering details 👉 lmsys.org/blog/2026-10-0…
1 3 0 20 1.3K 7
Proximal
Novita AI
Cyrus
Liam Fedus
Adarsh
RadixArk
Aravind Srinivas