Qwen 3.8 vs Qwen 3.7 Max: What Actually Changed

Qwen 3.8-Max vs 3.7-Max: benchmark deltas, the $2/$6 price vs the 50%-off promo, image input, and open weights. When to upgrade and when to wait.

INEZA Felin-Michel

INEZA Felin-Michel

3 August 2026

Qwen 3.8 vs Qwen 3.7 Max: What Actually Changed

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Alibaba released Qwen 3.8-Max in early August 2026, and if you’re already running qwen3.7-max in production, you have a real decision to make. The new model posts big gains on agentic and research benchmarks, adds image input, launches at a lower list price, and comes with a promise Alibaba never made for its predecessor: open weights.

But the decision isn’t as simple as “new model wins.” Qwen 3.7-Max is currently running a limited-time 50% discount that makes it cheaper than 3.8-Max in absolute terms. Some benchmark deltas are dramatic. Others are effectively zero. This article walks through every change that matters, so you can decide whether to switch now, wait, or drop down a tier entirely.

If you want the full picture of the new model first, start with our Qwen 3.8-Max explainer. For background on the outgoing flagship, see what Qwen 3.7 brought to the table.

One note on methodology before the numbers. Every benchmark figure below comes from Alibaba’s own table in the official Qwen 3.8 release post. That’s actually useful here: both models were evaluated in the same vendor run, so the deltas are internally consistent, even if independent verification is still pending. Most coding rows were run on the Claude Code harness, and several benchmarks are Qwen in-house.

The short version

What changed Qwen 3.7-Max Qwen 3.8-Max
Terminal Bench 2.1 74.5 86.6
SWE-bench Pro 60.6 67.7
PaperBench 64.8 93.0
IFBench 79.1 82.8
GPQA Diamond 92.4 92.6
HLE 41.4 43.6
List price (per 1M tokens) $2.5 in / $7.5 out $2 in / $6 out
Current price (promo) $1.25 in / $3.75 out $2 in / $6 out
Image input No Yes
Disclosed architecture Not disclosed 2.4T total, 95B active MoE
Open weights Never released Promised “next week” (~Aug 10)
Context window 1M tokens 1M tokens

Now the details.

Benchmark deltas: where the jump is real

The headline improvements cluster around agentic work: long-horizon tasks where the model plans, executes, and self-corrects.

Terminal Bench 2.1: 74.5 to 86.6. This is a 12-point jump on terminal-driven agentic tasks, and it moves Qwen from clearly-behind to genuinely competitive. In the same table, Alibaba scores Claude Opus 4.8 and Fable 5 at 84.6 each, which puts 3.8-Max ahead of both on this row. GPT-5.6 Sol still leads at 88.8.

SWE-bench Pro: 60.6 to 67.7. A 7-point gain on real-world software engineering tasks. Meaningful, but keep it honest: Fable 5 sits at 80.0 in the same table, so this is catching up, not overtaking. If SWE-bench Pro performance is your primary buying criterion, the frontier still lives elsewhere.

PaperBench: 64.8 to 93.0. The single largest delta in the table, a 28-point leap on reproducing AI research papers. It’s also 3.8-Max’s best flagship row, ahead of every rival in Alibaba’s comparison. If your workload involves research reproduction, long technical documents, or multi-step scientific reasoning, this is the number that should get your attention.

IFBench: 79.1 to 82.8. Instruction following improves by nearly 4 points. Notably, 3.7-Max was already strong here (79.1 beat Opus 4.8 and Fable 5 in the same run), so this extends an existing lead rather than fixing a weakness.

GPQA Diamond: 92.4 to 92.6. Effectively flat. Graduate-level science QA was already near saturation for 3.7-Max, and the new model doesn’t move it. If your workload looks like GPQA, the upgrade buys you nothing on quality.

HLE: 41.4 to 43.6. Humanity’s Last Exam improves 2 points but still trails the frontier: Fable 5 posts 53.3 and GPT-5.6 Sol posts 47.2 in the same table. Qwen 3.8-Max closes gaps in many places. This isn’t one of them.

The pattern: massive gains on agentic and research benchmarks, incremental gains on instruction following, no movement on saturated knowledge benchmarks, and a persistent gap on the hardest reasoning evals. For a deeper walk through the full table, including the fine print about harnesses and in-house benchmarks, see our Qwen 3.8 benchmarks breakdown.

Architecture: Alibaba finally shows its numbers

Qwen 3.8-Max is a 2.4 trillion parameter mixture-of-experts model with 95B active parameters per forward pass, built on the Qwen 3.5 architectural foundation. That’s roughly a 4% activation ratio, which matters for serving costs: you pay compute for 95B, not 2.4T.

The contrast with 3.7-Max is the disclosure itself. Alibaba never published comparable scale figures for Qwen 3.7-Max. Parameter counts, activation ratios, expert configuration: all undisclosed. With 3.8-Max, the Model Studio model documentation and the release post spell out the architecture, which makes capacity planning and cost modeling far less of a guessing game.

For context in the open-weight race: Kimi K3 runs 2.8T total with 104B active. Qwen 3.8-Max is smaller on both axes.

Multimodal: the capability 3.7-Max never had

Qwen 3.7-Max is a text model. Qwen 3.8-Max accepts image input natively, and Alibaba published a separate multimodal benchmark table where it leads most rows against Gemini 3.1 Pro and GPT-5.6 Sol.

The standout scores from Alibaba’s multimodal table: MathVision at 95.2, LogicVista at 91.9, and OSWorld-Verified at 86.1 for computer-use tasks. There’s no 3.7-Max column in that table to compare against, because there couldn’t be: the older model doesn’t take images.

Practically, this means workloads you previously had to route to a separate vision model (screenshot understanding, document images, UI automation, chart extraction) can now run on the same model ID as your text traffic. The release post also demonstrates long-document and video understanding, though those are blog-demonstrated capabilities rather than distinct API input types, so verify against the API docs for your specific case.

Pricing: the new model is cheaper, except right now

Here’s where the upgrade math gets interesting.

Qwen 3.8-Max launches at $2 input / $6 output per 1M tokens, a single flat tier across the entire 1M context. Qwen 3.7-Max lists at $2.5 / $7.5. So the new, more capable model undercuts its predecessor’s list price by 20%. That almost never happens in a flagship generation change.

But there’s a wrinkle: Alibaba is currently running a limited-time 50% promotion on 3.7-Max, which brings its effective rate to $1.25 / $3.75. At today’s prices, the older model costs 38% less than the new one. All rates are on the official Model Studio pricing page.

Two things to keep in mind:

  1. The promo can end. Alibaba labels it limited-time and has published no end date. If you architect around $1.25 / $3.75, your budget assumption has an expiry you can’t see. At list prices, 3.8-Max is the cheaper model.
  2. Thinking tokens inflate real bills. Qwen 3.8-Max defaults to reasoning_effort: xhigh, and thinking tokens bill as output. Your effective cost per request can land well above what the sticker suggests. Our Qwen 3.8 pricing guide works through the cache economics and cost-per-task math.

And if neither Max model fits your budget, qwen3.7-plus at $0.4 / $1.6 (currently 20% off) remains the value tier. The Plus vs Max comparison covers when Plus is enough.

Open weights: a first for the Max class

No Qwen-Max-class model has ever shipped open weights. Qwen 3.7-Max didn’t, and Alibaba never suggested it would. Qwen 3.8-Max breaks that pattern: the release post promises weights on Hugging Face and ModelScope “next week,” which points to roughly August 10.

As of this writing (August 3, 2026), the weights are not downloadable. Treat the promise as a promise. The precedent is encouraging: Kimi K3’s weights arrived 11 days after launch, and Alibaba has a long open-weight track record with smaller Qwen models. But a 2.4T parameter download is a multi-node self-hosting project even quantized, so for most teams the practical impact will be third-party hosted access and pricing pressure, not a model running in your rack.

If open weights matter to your procurement or compliance story, that alone may settle the 3.7 vs 3.8 question. One model will (probably) have them. The other never will.

What didn’t change

Worth stating plainly, because upgrade posts tend to oversell:

Should you upgrade? Three honest paths

Upgrade now if your workload is agentic or research-heavy. The Terminal Bench (74.5 to 86.6) and PaperBench (64.8 to 93.0) deltas are the kind of jump you can feel in production, and you get image input plus the lower list price as a bonus. If you’re building agents that operate terminals, reproduce technical work, or process screenshots, the case is clear.

Ride the 3.7-Max promo if your workload lives in the flat zones. GPQA-style knowledge tasks moved 0.2 points. If your traffic is Q&A, summarization, or general chat, and 3.7-Max already meets your quality bar, $1.25 / $3.75 is the best per-token deal either Max model has ever offered. Set a calendar reminder to re-check the promo status monthly, and know your fallback price is 3.8-Max at $2 / $6, not 3.7-Max at list.

Drop to qwen3.7-plus if you’re price-driven and your tasks are routine. At $0.4 / $1.6, it’s 5x cheaper than 3.8-Max on input, and for classification, extraction, and templated generation, Max-class capability is often wasted spend.

Test the switch before you commit

Vendor benchmarks, even internally consistent ones, don’t predict how a model behaves on your prompts. The cheap way to settle this is a side-by-side on your actual workload.

In Apidog, you can set this up in a few minutes: point one collection at Model Studio’s OpenAI-compatible endpoint, then create two environments that differ only in the model ID variable, one set to qwen3.7-max and one to qwen3.8-max. Run the same request set against both, flip environments with one click, and compare outputs, latency, and token usage from the response viewer. Because both models share the same API surface, a single collection covers both, and you can save representative production prompts as test scenarios and re-run them whenever Alibaba ships an update or the promo pricing changes.

That turns “should we upgrade?” from a benchmarks debate into an afternoon of evidence. Download Apidog for free and run the comparison against your own traffic before you change a production config.

FAQ

Is Qwen 3.8-Max cheaper than Qwen 3.7-Max?

At list price, yes: $2 / $6 versus $2.5 / $7.5 per 1M tokens. At today’s actual prices, no: 3.7-Max’s limited-time 50% promo brings it to $1.25 / $3.75, which undercuts 3.8-Max by 38%. The promo has no published end date. Full cost breakdowns are in our Qwen 3.8 pricing guide.

Do I need to change my code to switch from qwen3.7-max to qwen3.8-max?

For basic usage, no. Both models serve through the same OpenAI-compatible Model Studio endpoints, so switching is a model ID swap. You may want to set reasoning_effort explicitly, since 3.8-Max defaults to xhigh and thinking tokens bill as output. If you send images, that’s a 3.8-Max-only capability and requires the standard image-input message format.

Does Qwen 3.7-Max have open weights?

No, and it never will based on Alibaba’s history with the line. Qwen 3.8-Max is the first Max-class model with promised open weights, expected on Hugging Face and ModelScope around August 10, 2026. As of August 3, they’re not yet downloadable.

Is Qwen 3.8-Max better than Claude or GPT now?

Depends on the row, and remember these are Alibaba’s own numbers. In their table, 3.8-Max beats Opus 4.8 and Fable 5 on Terminal Bench 2.1 and leads everyone on PaperBench and IFBench. It trails Fable 5 badly on SWE-bench Pro (67.7 versus 80.0) and everyone at the frontier on HLE. No independent benchmark numbers exist yet, so hold final judgment until third-party evals land.

Explore more

How to use GPT-6.1 Sol APl ?

How to use GPT-6.1 Sol APl ?

GPT-6.1 Sol API guide: your first gpt-6.1-sol request, effort levels, Batch/Flex/Fast pricing, and the four changes to migrate from gpt-6-sol.

30 September 2026

What Is GPT-6.1 Sol?

What Is GPT-6.1 Sol?

GPT-6.1 Sol explained: model ID gpt-6.1-sol, $2/$10 pricing with $0.10 cached input, 922K max input, effort levels, and OpenAI's benchmarks vs Astra.

30 September 2026

OpenAI DevDay 2026 for API developers: what shipped, what it costs, and what's still coming soon

OpenAI DevDay 2026 for API developers: what shipped, what it costs, and what's still coming soon

OpenAI DevDay 2026 for API developers: GPT-6.1 Sol at $2/$10, Ultrafast on Astra, Agents API computer use, MCP Events, and what to change this week.

30 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Qwen 3.8 vs Qwen 3.7 Max: What Actually Changed