SGLang
@sgl_project
We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference.
Working closely with @nvidia, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.
Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20%
Working closely with @nvidia, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.
Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20%
lmsys.orgSGLang and Miles on NVIDIA Vera Rubin
16 145