Qwen3.8-Flash-Next
Qwen releases 125B MoE model with only 6B active parameters as Qwen4 architecture preview
“a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4”
Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weights multimodal Mixture-of-Experts model with 125B total parameters but only 6B active, delivering strong performance efficiency. The model is notable as an early preview of the Qwen4 architecture, signaling the direction of Alibaba's next major model generation. Simon Willison tested it on an NVIDIA DGX Spark using quantized versions, finding it capable of generating detailed SVG illustrations.