Startups & Ventures: post #3816 — TG.ME

📊 Alibaba previews Qwen4 architecture

Alibaba released Qwen3.8-Flash-Next, an open multimodal model and early preview of the architecture planned for Qwen4. It has 125B parameters, but uses only 6B for each text fragment. Another 51B sit in separate memory for frequent combinations.

The design keeps the model large while reducing compute. Alibaba says training is about 9 times cheaper than Qwen3.7-Plus, while the model already performs better on coding, tool-use tasks, and long office workflows.

Long-context support reaches 262,000 tokens by default and up to 1 million tokens in extended mode. At 1 million tokens, a new retrieval mechanism can significantly speed up processing. The weights are already open on Hugging Face.

📊@tech
❤271🦄266😁257🔥112👍36
August 27, 2026 3.9K 1 16