← Back to Payloads
LLM Release2026-08-07

Qwen 3.8-Max Just Stole the Agentic Crown at One-Fifth the Price. The Open Weights Are the Punchline.

Alibaba shipped Qwen 3.8-Max on August 3 — a 2.4T-parameter MoE, 1M context, multimodal, $2 input / $6 output per million tokens. On Artificial Analysis it tops the Agentic Index at 58.4 (vs Opus 4.8's 49.4 and Fable 5's 56.6), sets the high-water mark on PaperBench (93.0), IFBench (82.8), MRCR v2 256K (92.9), HealthBench (60.2), OSWorld-Verified (86.1), and 12+ video/multimodal benchmarks. And the weights drop next week — the first open-weights Max-class Qwen. The Western pricing model just broke.
Quick Access
Install command
$ mrt install qwen
Browse related skills
Qwen 3.8-Max Just Stole the Agentic Crown at One-Fifth the Price. The Open Weights Are the Punchline.

Qwen 3.8-Max Just Stole the Agentic Crown at One-Fifth the Price. The Open Weights Are the Punchline.

Hi guys, Mr. Technology here.

Alibaba dropped Qwen 3.8-Max on August 3 — a 2.4-trillion-parameter sparse MoE, multimodal text/image/video-in to text-out, 1M-token context, $2 input / $6 output per million tokens on the OpenRouter routing. It hits the top of Artificial Analysis's Agentic Index at 58.4, beating Claude Opus 4.8's 49.4 and Claude Fable 5's 56.6, and sets the global high-water mark on PaperBench (93.0), IFBench (82.8), MRCR v2 256K (92.9), HealthBench (60.2), $OneMillion-Bench (52.5), OSWorld-Verified (86.1), Parametric CAD Bench (91.5), MathVision (95.2/97.7), VideoMME (90.4), MLVU (90.8), EgoLife (80.3) — and a dozen more.

And — here is the part that should make Anthropic and OpenAI nervous — the weights ship next week. First open-weights Max-class Qwen. Read that twice.

The Numbers That Matter

Strip the marketing and the agentic numbers are the story:

ModelInput $/MOutput $/MIntelligence IndexCoding Index**Agentic Index**
Qwen 3.8-Max$2.00$6.0058.171.858.4
Claude Fable 5$10.00$50.0062.176.556.6
Claude Opus 4.8$5.00$25.0057.374.349.4
GPT-5.6 Sol (max)~$3.00~$15.00
Meta Muse Spark 1.2$1.25$4.25

Three things to notice:

1. Qwen is the Agentic Index leader. It beats Opus 4.8 by 9.0 points and Fable 5 by 1.8 points. Opus 4.8 is the model every Anthropic-flavored coding agent is wired to. Qwen is now the better agentic model, full stop.

2. Fable 5 wins the Intelligence Index by 4 points — at 5–8× the price. If you are paying Fable 5 prices for "better" on a generic benchmark, you are paying a 5× markup for 7% more intelligence. The math has stopped working.

3. The gap between Qwen and the open-source frontier just collapsed. DeepSeek V4-Flash-0731 at $0.14/$0.28 is still the cheap open-weights play. Qwen 3.8-Max is 14× more expensive but is the first open-weights model that can credibly replace Opus 4.8 in production.

What It Actually Does

The qwen.ai blog calls out four positioning points:

  • 10+ day autonomous coding projects, delivered end-to-end in a single conversation. This is a product claim, not a benchmark cherry-pick.
  • Native visual understanding across the full cycle — image and video inputs into planning, execution, and verification. MMMU-Pro 82.3, VideoMME 90.4, VideoMMMU 88.7, MLVU 90.8 — every video benchmark tested, every video benchmark #1.
  • Long-horizon iteration with closed feedback loops. CoWorkBench 74.8, JobBench 53.4, SkillsBench 70.2 — all near or above Opus 4.8.
  • Reasoning-effort dial from minimal to xhigh. Default is xhigh (full thinking). Same lever GPT-5.6 and Fable 5 shipped; now table stakes for any frontier model.

The architecture is built on the Qwen 3.5 base — Gated Delta Networks (a linear-attention hybrid) plus sparse MoE, scaled to 2.4T parameters, multimodal-pretrained on trillions of tokens, then RL-tuned across million-agent environments. Qwen 3.6 was a stability-and-real-world-utility pass; 3.8-Max is the Max-class payoff.

The Benchmarks, Read Honestly

Most of the launch numbers are Max-leads. A few are not, and those matter more.

Terminal-Bench 2.1: 86.6. Behind GPT-5.6 Sol's 88.8 (2.2 points) and barely ahead of Opus 4.8 and Fable 5 (both 84.6). For raw terminal-use, GPT-5.6 Sol is still the top dog. If you are running Codex CLI workflows, do not switch off GPT-5.6 today for Qwen on terminal alone.

SWE-bench Pro: 67.7. Behind Opus 4.8 (69.2), Fable 5 (80.0). The Fable 5 gap is real.

DeepSWE 1.1: 56.6. Behind Fable 5 (70.0), Opus 4.8 (59.0), GPT-5.6 Sol (73.0). This is becoming a pattern: Anthropic's harness-co-tuned models eat lunch on isolated repo edits, Qwen eats lunch on long-horizon orchestration.

OSWorld-Verified: 86.1. Highest on the public leaderboard. Qwen is the model to beat for GUI-agent work.

PaperBench: 93.0. Higher than Opus 4.8 (80.3) and Fable 5 (88.8). Replicating a research paper end-to-end is the single hardest long-horizon coding task we have a public benchmark for. Qwen is winning it by a real margin.

MRCR v2 256K (8-needle): 92.9. Long-context retrieval at 256K tokens. Highest score. This is the benchmark that kills any model that fakes long context with eviction. Qwen actually reads 256K.

Two honest caveats worth printing. Qwen's charts include Fable 5 results with a footnote that "Fable5 results may involve fallbacks" — Anthropic's flagship scoring is partly on retry-with-fallback configurations. The 4-point Intelligence Index gap is narrower than it looks on an apples-to-apples run. And "autonomously delivers projects spanning 10+ days" is a marketing line until somebody outside Alibaba has run it for ten days. The PaperBench, CoWorkBench, and SkillsBench numbers make it plausible. They do not confirm it.

Why Open Weights Is the Actual Headline

Every other story about this launch misses the structure.

For two years the agentic-model conversation has been: "Claude Opus, GPT-5.x, and Claude Fable 5 cost $5–$25 per million output tokens, are closed-weight, and you cannot inspect, distill, fine-tune, or host them. If you want any of those you must send your data, your prompts, your proprietary code to a Western API."

Qwen 3.8-Max is the first open-weights Max-class model from any Chinese lab, and it lands the same week as the closed-source flagships — at 2.5× lower input cost and 4× lower output cost than Opus 4.8, and 5× lower input / 8× lower output than Fable 5.

The downstream effects stack fast:

  • Distillation is back on the table. A 2.4T model is too big to fine-tune directly, but it is the right size to distill from. Expect a wave of Qwen 3.8-Max-distilled open-weight small models in the next 90 days — DeepSeek V4-Flash-0731 was just the opening shot.
  • Hosting economics shift. The $0.14/$0.28 DeepSeek price point was the previous floor. Qwen 3.8-Max at $2/$6 raises the ceiling for what self-hosted inference on a serious frontier model looks like. A 30B distillate of Qwen 3.8-Max on a single H200 will be roughly equivalent to a fully-quantized Opus 4.8 today, in three months.
  • The Western pricing model just lost its moat. Anthropic's price premium was justified by (a) better benchmarks and (b) the absence of open-weight alternatives. Qwen 3.8-Max kills both. Fable 5 wins on the Intelligence Index by 4 points at 5–8× the price. Opus 4.8 loses on agentic by 9 points at 2.5–4× the price. The "you have to pay us because nobody else can do this" pitch does not work when the open-weights download is sitting on Hugging Face the same week.

What This Means For Builders

Three concrete moves to make in the next two weeks.

1. Run a head-to-head on your hardest agentic task. If you are on Opus 4.8 or Fable 5 today, pick your most expensive workflow — long-horizon coding, multi-tool research, document-heavy legal/finance work — and route 20% of traffic through Qwen 3.8-Max via OpenRouter or QwenCloud. Track cost-per-completed-task and cost-per-defect. If you are not benchmarking this in August 2026, you are overpaying.

2. Wait for the weights, then distill. Once Qwen/Qwen3.8-Max and the smaller Qwen3.8-27B land on Hugging Face, treat the Max as a teacher. A 4B or 8B distillate will be the new "free Opus" by Q4 2026.

3. Audit your Chinese-model legal posture. "We don't use Chinese models" is not a viable policy in 2026. Qwen 3.8-Max on Alibaba Cloud International is a different posture from DeepSeek-hosted endpoints. Pick deliberately.

The Take

Qwen 3.8-Max is not the highest-IQ model — Fable 5 is, by 4 points, at 5–8× the price. It is not the highest-coding model — Fable 5 is. It is not the cheapest — DeepSeek V4-Flash is, by an order of magnitude.

What it is, unambiguously, is the best agentic model you can buy as of August 3, 2026, at the lowest price any frontier agentic model has ever been sold at, with open weights arriving next week.

The smart move this week: stop reading launch posts, run the model on your workflow. The smart move this month: distill the weights into something you can self-host. The smart move this year: rebuild your agent-stack procurement policy around the assumption that the frontier is now Chinese, open-weight, and priced in single-digit dollars per million tokens.

Qwen 3.8-Max is not the end of the frontier-model race. It is the end of the frontier-model pricing model.

Mr. Technology


Model: Qwen 3.8-Max (flagship of Qwen 3.8, successor to Qwen 3.7-Max) · Lab: Alibaba Tongyi Lab · Release: August 3, 2026 · Architecture: 2.4T sparse MoE on Qwen 3.5 base (Gated Delta Networks + sparse MoE) · Modalities: text + image + video → text · Context: 1,048,576 tokens (1M); max input 991K, max output 131K · Reasoning: max thinking 262K tokens · reasoning_effort: minimal/low/medium/high/xhigh (default xhigh) · Pricing: $2 input / $6 output per million tokens · cache read $0.25, write $2.50 (OpenRouter) · Limits: 2M TPM, 15K RPM · Availability: Qwen Cloud, Alibaba Cloud Model Studio, OpenRouter, Hugging Face (weights "next week") · License: Alibaba Tongyi License (open weights, commercial with case-by-case restrictions) · Artificial Analysis (Aug 3, 2026): Intelligence 58.1, Coding 71.8, Agentic 58.4 (leader) · Sources: Qwen blog · Qwen 3.8 docs hub · QwenCloud product page · Artificial Analysis Agentic Index · AA Qwen 3.8-Max page · OpenRouter · Alibaba Qwen on X

Related Dispatches