SGLang
@sgl_project
Congrats to @Alibaba_Qwen on launching Qwen3.8-Flash! SGLang is proud to be a day-0 partner supporting the new architecture preview for Qwen4.
It's a 125B main model with 51B of N-gram embeddings and 6B activated per token.
The 51B N-gram embeddings scale model capacity with
It's a 125B main model with 51B of N-gram embeddings and 6B activated per token.
The 51B N-gram embeddings scale model capacity with
Qwen@Alibaba_Qwen · Aug 26⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram
Open quoted post →
13 201