RadixArk
@radixark
Excited to team up with @Alibaba_Qwen, @NVIDIAAI, and @AIatAMD on Day-0 support for Qwen3.8-Flash-Next! 🚀
The RadixArk team contributed deep kernel and system optimizations to @sgl_project , alongside releasing the SGLang official Day-0 NVFP4 quantized model.
Try the
The RadixArk team contributed deep kernel and system optimizations to @sgl_project , alongside releasing the SGLang official Day-0 NVFP4 quantized model.
Try the
huggingface.coRadixArk/Qwen3.8-Flash-Next-NVFP4 · Hugging Face
SGLang@sgl_project · Aug 26Congrats to @Alibaba_Qwen on launching Qwen3.8-Flash! SGLang is proud to be a day-0 partner supporting the new architecture preview for Qwen4.
It's a 125B main model with 51B of N-gram embeddings and 6B activated per token.
The 51B N-gram embeddings scale model capacity with
It's a 125B main model with 51B of N-gram embeddings and 6B activated per token.
The 51B N-gram embeddings scale model capacity with
Open quoted post →
2 37