
Hi guys, Mr. Technology here.
Alibaba dropped Qwen 3.8-Max on August 3 — a 2.4-trillion-parameter sparse MoE, multimodal text/image/video-in to text-out, 1M-token context, $2 input / $6 output per million tokens on the OpenRouter routing. It hits the top of Artificial Analysis's Agentic Index at 58.4, beating Claude Opus 4.8's 49.4 and Claude Fable 5's 56.6, and sets the global high-water mark on PaperBench (93.0), IFBench (82.8), MRCR v2 256K (92.9), HealthBench (60.2), $OneMillion-Bench (52.5), OSWorld-Verified (86.1), Parametric CAD Bench (91.5), MathVision (95.2/97.7), VideoMME (90.4), MLVU (90.8), EgoLife (80.3) — and a dozen more.
And — here is the part that should make Anthropic and OpenAI nervous — the weights ship next week. First open-weights Max-class Qwen. Read that twice.
Strip the marketing and the agentic numbers are the story:
| Model | Input $/M | Output $/M | Intelligence Index | Coding Index | **Agentic Index** |
|---|---|---|---|---|---|
| Qwen 3.8-Max | $2.00 | $6.00 | 58.1 | 71.8 | 58.4 |
| Claude Fable 5 | $10.00 | $50.00 | 62.1 | 76.5 | 56.6 |
| Claude Opus 4.8 | $5.00 | $25.00 | 57.3 | 74.3 | 49.4 |
| GPT-5.6 Sol (max) | ~$3.00 | ~$15.00 | — | — | — |
| Meta Muse Spark 1.2 | $1.25 | $4.25 | — | — | — |
Three things to notice:
1. Qwen is the Agentic Index leader. It beats Opus 4.8 by 9.0 points and Fable 5 by 1.8 points. Opus 4.8 is the model every Anthropic-flavored coding agent is wired to. Qwen is now the better agentic model, full stop.
2. Fable 5 wins the Intelligence Index by 4 points — at 5–8× the price. If you are paying Fable 5 prices for "better" on a generic benchmark, you are paying a 5× markup for 7% more intelligence. The math has stopped working.
3. The gap between Qwen and the open-source frontier just collapsed. DeepSeek V4-Flash-0731 at $0.14/$0.28 is still the cheap open-weights play. Qwen 3.8-Max is 14× more expensive but is the first open-weights model that can credibly replace Opus 4.8 in production.
The qwen.ai blog calls out four positioning points:
minimal to xhigh. Default is xhigh (full thinking). Same lever GPT-5.6 and Fable 5 shipped; now table stakes for any frontier model.The architecture is built on the Qwen 3.5 base — Gated Delta Networks (a linear-attention hybrid) plus sparse MoE, scaled to 2.4T parameters, multimodal-pretrained on trillions of tokens, then RL-tuned across million-agent environments. Qwen 3.6 was a stability-and-real-world-utility pass; 3.8-Max is the Max-class payoff.
Most of the launch numbers are Max-leads. A few are not, and those matter more.
Terminal-Bench 2.1: 86.6. Behind GPT-5.6 Sol's 88.8 (2.2 points) and barely ahead of Opus 4.8 and Fable 5 (both 84.6). For raw terminal-use, GPT-5.6 Sol is still the top dog. If you are running Codex CLI workflows, do not switch off GPT-5.6 today for Qwen on terminal alone.
SWE-bench Pro: 67.7. Behind Opus 4.8 (69.2), Fable 5 (80.0). The Fable 5 gap is real.
DeepSWE 1.1: 56.6. Behind Fable 5 (70.0), Opus 4.8 (59.0), GPT-5.6 Sol (73.0). This is becoming a pattern: Anthropic's harness-co-tuned models eat lunch on isolated repo edits, Qwen eats lunch on long-horizon orchestration.
OSWorld-Verified: 86.1. Highest on the public leaderboard. Qwen is the model to beat for GUI-agent work.
PaperBench: 93.0. Higher than Opus 4.8 (80.3) and Fable 5 (88.8). Replicating a research paper end-to-end is the single hardest long-horizon coding task we have a public benchmark for. Qwen is winning it by a real margin.
MRCR v2 256K (8-needle): 92.9. Long-context retrieval at 256K tokens. Highest score. This is the benchmark that kills any model that fakes long context with eviction. Qwen actually reads 256K.
Two honest caveats worth printing. Qwen's charts include Fable 5 results with a footnote that "Fable5 results may involve fallbacks" — Anthropic's flagship scoring is partly on retry-with-fallback configurations. The 4-point Intelligence Index gap is narrower than it looks on an apples-to-apples run. And "autonomously delivers projects spanning 10+ days" is a marketing line until somebody outside Alibaba has run it for ten days. The PaperBench, CoWorkBench, and SkillsBench numbers make it plausible. They do not confirm it.
Every other story about this launch misses the structure.
For two years the agentic-model conversation has been: "Claude Opus, GPT-5.x, and Claude Fable 5 cost $5–$25 per million output tokens, are closed-weight, and you cannot inspect, distill, fine-tune, or host them. If you want any of those you must send your data, your prompts, your proprietary code to a Western API."
Qwen 3.8-Max is the first open-weights Max-class model from any Chinese lab, and it lands the same week as the closed-source flagships — at 2.5× lower input cost and 4× lower output cost than Opus 4.8, and 5× lower input / 8× lower output than Fable 5.
The downstream effects stack fast:
Three concrete moves to make in the next two weeks.
1. Run a head-to-head on your hardest agentic task. If you are on Opus 4.8 or Fable 5 today, pick your most expensive workflow — long-horizon coding, multi-tool research, document-heavy legal/finance work — and route 20% of traffic through Qwen 3.8-Max via OpenRouter or QwenCloud. Track cost-per-completed-task and cost-per-defect. If you are not benchmarking this in August 2026, you are overpaying.
2. Wait for the weights, then distill. Once Qwen/Qwen3.8-Max and the smaller Qwen3.8-27B land on Hugging Face, treat the Max as a teacher. A 4B or 8B distillate will be the new "free Opus" by Q4 2026.
3. Audit your Chinese-model legal posture. "We don't use Chinese models" is not a viable policy in 2026. Qwen 3.8-Max on Alibaba Cloud International is a different posture from DeepSeek-hosted endpoints. Pick deliberately.
Qwen 3.8-Max is not the highest-IQ model — Fable 5 is, by 4 points, at 5–8× the price. It is not the highest-coding model — Fable 5 is. It is not the cheapest — DeepSeek V4-Flash is, by an order of magnitude.
What it is, unambiguously, is the best agentic model you can buy as of August 3, 2026, at the lowest price any frontier agentic model has ever been sold at, with open weights arriving next week.
The smart move this week: stop reading launch posts, run the model on your workflow. The smart move this month: distill the weights into something you can self-host. The smart move this year: rebuild your agent-stack procurement policy around the assumption that the frontier is now Chinese, open-weight, and priced in single-digit dollars per million tokens.
Qwen 3.8-Max is not the end of the frontier-model race. It is the end of the frontier-model pricing model.
— Mr. Technology
Model: Qwen 3.8-Max (flagship of Qwen 3.8, successor to Qwen 3.7-Max) · Lab: Alibaba Tongyi Lab · Release: August 3, 2026 · Architecture: 2.4T sparse MoE on Qwen 3.5 base (Gated Delta Networks + sparse MoE) · Modalities: text + image + video → text · Context: 1,048,576 tokens (1M); max input 991K, max output 131K · Reasoning: max thinking 262K tokens · reasoning_effort: minimal/low/medium/high/xhigh (default xhigh) · Pricing: $2 input / $6 output per million tokens · cache read $0.25, write $2.50 (OpenRouter) · Limits: 2M TPM, 15K RPM · Availability: Qwen Cloud, Alibaba Cloud Model Studio, OpenRouter, Hugging Face (weights "next week") · License: Alibaba Tongyi License (open weights, commercial with case-by-case restrictions) · Artificial Analysis (Aug 3, 2026): Intelligence 58.1, Coding 71.8, Agentic 58.4 (leader) · Sources: Qwen blog · Qwen 3.8 docs hub · QwenCloud product page · Artificial Analysis Agentic Index · AA Qwen 3.8-Max page · OpenRouter · Alibaba Qwen on X