OmniFit logo OmniFit Blog AI video ideas, tools, and workflow notes Blog home
News + open weights

Qwen 3.8 open weights are out

📅 Published: August 13, 2026

Alibaba has now released the open weights for Qwen/Qwen3.8-2.4T-A95B, the first Qwen Max-class model to ship in open form. The headline is huge: 2.4T total parameters, 95B activated, 92 layers, 512 experts, and context that stretches from 262K native to about 1.01M extended. The bigger signal is that flagship Chinese models are starting to look like one shared recipe: sparse MoE, hybrid attention, ultra-long context, and agent-first benchmarks.

By Maya Chen August 13, 2026 5 min read AI / open models
Quick take

This is a meaningful open release, not just another preview post. The official Qwen blog announced a 2.4T flagship on August 2, promised open weights the following week, and the Hugging Face repo is now live with the full open model card. The catch is also clear: the managed Qwen3.8-Max cloud product still keeps some important extras, including built-in tools, vision input, and a cleaner 1M-context default experience.

2.4T total / 95B active 92 layers / 512 experts 10 routed + 1 shared 262K native / 1.01M extended
Qwen 3.8 2.4T technical poster used as the hero image for this article
User-provided Qwen 3.8 poster used as the hero image for this post.

What actually shipped

The official record now lines up across three live sources. The Qwen launch post announced Qwen 3.8-Max on August 2 and said open weights would follow the next week. The Hugging Face model card is now live for Qwen/Qwen3.8-2.4T-A95B. And the QwenCloud model page shows what still belongs to the managed Max product.

2.4T total parameter count in the official open model card
95B activated parameters per task according to the MoE design
512 experts, with 10 routed and 1 shared expert active
1.01M extended context limit listed in the open model card
Claim What the live sources confirm Why it matters
First Max-class open release The Hugging Face card explicitly says this is the first time Qwen has brought a Qwen-Max-class model to open release. That makes this more important than a normal incremental checkpoint update.
Hybrid architecture The model card lists a 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)) layout, plus multi-token prediction. The frontier recipe is shifting toward hybrid attention plus sparse activation, not plain dense Transformer scaling.
Very long context Hugging Face lists 262,144 native tokens and extension up to 1,010,000. QwenCloud shows 1M context and 262K max reasoning on the managed service. That matters for repo-scale coding, long documents, and tool-heavy agents.
Cloud product still does more The open model card says managed Qwen3.8-Max adds features like vision input, non-thinking support, 1M context by default, and official built-in tools. Open weights are real, but they are not the entire commercial stack.

Conservative read: the release is real and substantial, but the smoothest multimodal + tool-using experience still appears to live on QwenCloud.

The biggest story is not only that Qwen opened a 2.4T flagship. It is that sparse MoE, hybrid attention, million-token context, and long-horizon agents are now starting to look like the default flagship blueprint.

Watch one good technical review

Bijan Bowen’s “Qwen3.8 Max Is HERE – Is THIS the BEST Open Model Yet?” is a strong companion video for this post because it goes beyond hype and spends real time on first impressions, technical details, browser testing, coding tasks, and multimodal checks.

I kept your poster as the hero image, but this video is still worth adding because it shows how people are pressure-testing Qwen 3.8 in real workflows rather than just repeating the launch numbers.

Why creators should care

1. This is not a video model, but it matters to video businesses

If you run an AI-video workflow, you also run prompts, research, scripting, automations, docs, client notes, assets, and tool glue. A stronger open reasoning model helps on the operations side of creative work.

2. Agent workflows are becoming the real benchmark

The official Qwen post does not pitch Qwen 3.8 as a chatbot upgrade only. It pushes long-horizon coding, closed-loop feedback, and multi-step task completion. That is the same direction many creator teams want for internal assistants and production helpers.

3. The architecture race is narrowing

The numbers in the open card and the launch poster point to a broader market shift: sparse expert routing, hybrid attention, longer context, and multi-token prediction are becoming normal. That means future differentiation may come more from tooling, compute, data, and product fit than from one magical new block diagram.

4. Open weights still do not equal full product parity

The open release is important, but the official sources also make clear that the managed Max product keeps built-in tools and richer multimodal behavior. For teams that just want results fast, cloud convenience still matters.

One practical translation of the original Chinese angle

In simple English, the original point is this: Alibaba did not just release a giant model. It also showed that the top Chinese frontier-model playbook is becoming more standardized. Bigger MoE systems, smarter routing, much longer context, and stronger agent behavior are now starting to converge into one mainstream design direction.

That matters because the next fight may be less about who invented the most exotic architecture, and more about who can run the model cheaply, feed it better real-world data, and turn it into useful products people will actually pay for.

What is still unclear

How practical local deployment really is

The open weights are out, but a 2.4T model is still a very different thing from a casual local install. The release is meaningful for advanced teams, infra operators, and researchers first.

How close open inference gets to cloud behavior

The official model card directly says the managed Max service offers more features. So the important question is not just whether the weights exist, but how much product value still sits above the weights.

Benchmark strength is not the same as workflow fit

The Qwen materials lean heavily into coding and professional work. That is promising, but teams still need hands-on testing for their own repos, documents, and toolchains.

The real business moat may move upward

If frontier architectures keep converging, then the bigger moat may come from product surface, orchestration, private data, distribution, and trust — not from headline parameter count alone.

Bottom line

Qwen 3.8 is a real milestone for open frontier models: a Max-class release, a huge MoE system, long context, agent-first positioning, and a live Hugging Face checkpoint people can actually inspect. But the deeper story is market convergence. The flagship recipe is stabilizing, so the next winners may be decided less by architecture novelty and more by execution.

Sources used for this post

Qwen official launch post

Used for the August 2 launch framing, the 2.4T announcement, and Qwen’s official positioning around coding, work, research, and long-horizon tasks.

Read source
Hugging Face model card

Used for the open-weight confirmation, 95B activated parameter count, 92 layers, 512 experts, 10 routed + 1 shared, hybrid layout, MTP, and native / extended context numbers.

Read source
QwenCloud model page

Used for the managed-product view: 1M context, 262K max reasoning, multimodal inputs, and built-in tools that still sit above the open weights.

Read source
YouTube: Bijan Bowen

Used as the companion video because it includes first-look evaluation and practical technical testing rather than only launch hype.

Watch video