RadixArk
@radixark
SGLang-Diffusion enables fast, scalable inference for multimodal generation. The latest work with VDN-H3 brings @MiniMax_AI H3 to faster-than-playback video generation with strong scaling across GPUs. 👏
SGLang@sgl_project · Sep 14SGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀
On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup.
No measured quality regression versus dense 50-step H3
On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup.
No measured quality regression versus dense 50-step H3
0 18