Qwen3.8-Flash-Next 架构设计:评测、效率与训练稳定性
Qwen 团队公开 28 页技术报告,系统拆解 GDN + QSA 混合注意力、四分支 Gated Residual、51B 主机内存 N-gram Embedding 与 Muon 训练配方,并用能力、训练/推理成本及稳定性联合验证设计取舍。
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
A 28-page Qwen Team report on the GDN/QSA hybrid, four-branch Gated Residual, 51B host-memory N-gram embeddings, and Muon optimization, evaluated jointly across capability, phase-specific efficiency, and training stability.
https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf

GitHub
Qwen3.8-Flash-Next/tech_report.pdf at main · QwenLM/Qwen3.8-Flash-Next
Qwen3.8-Flash-Next is the foundation model developed by Qwen Team, Alibaba Group. - QwenLM/Qwen3.8-Flash-Next
September 1, 2026 5