Qwen 3.8 Benchmarks: What Alibaba's Table Shows, and What It Doesn't

Qwen 3.8-Max benchmarks, read honestly: PaperBench 93.0 and multimodal wins, HLE and SWE-bench Pro losses, and the fine print most coverage skips.

INEZA Felin-Michel

INEZA Felin-Michel

3 August 2026

Qwen 3.8 Benchmarks: What Alibaba's Table Shows, and What It Doesn't

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

For two weeks, the Qwen 3.8 story had a hole in it. The July press previews promised a frontier-class model but shipped no scores, so every early writeup ran on vibes. That changed on August 3, 2026. Alibaba’s official release post for Qwen 3.8-Max includes a full benchmark table against Claude Opus 4.8, Fable 5, GPT-5.6 Sol, and its own predecessor Qwen3.7-Max, plus a separate multimodal table against Gemini 3.1 Pro and GPT-5.6 Sol.

That table is worth reading closely. It contains genuine wins, honest losses, and fine print that most coverage will skip entirely. This post walks through all three, then gives you a practical way to stop trusting vendor tables altogether: run your own evaluation against the API. If you want the full model rundown first (parameters, pricing, access), start with our Qwen 3.8-Max explainer and come back.

One framing note before the numbers. Every score below comes from Alibaba’s own published table, with leaderboard data the footnotes date to August 3, 2026. No independent evaluations existed at the time of writing. Keep that in mind for every row that follows.

The table, at a glance

Here are the headline text-model rows as Alibaba published them. Higher is better on all of these.

Benchmark Qwen 3.8-Max Claude Opus 4.8 Fable 5 GPT-5.6 Sol Qwen3.7-Max
Terminal Bench 2.1 86.6 84.6 84.6 88.8 74.5
SWE-bench Pro 67.7 69.2 80.0 64.6 60.6
PaperBench 93.0 80.3 88.8 90.5 64.8
GPQA Diamond 92.6 92.0 92.6 94.1 92.4
IFBench 82.8 62.2 63.5 72.7 79.1
HLE (Humanity’s Last Exam) 43.6 45.7 53.3 47.2 41.4

Source: Alibaba’s vendor-run table. We’ll get to what “vendor-run” means in practice further down, because it matters more here than usual.

Where Qwen 3.8-Max wins

PaperBench: its best flagship row

PaperBench, which tests whether a model can reproduce research paper results, is Qwen 3.8-Max’s single strongest showing against the frontier: 93.0, ahead of GPT-5.6 Sol’s 90.5, Fable 5’s 88.8, and Opus 4.8’s 80.3. It’s also a staggering jump from Qwen3.7-Max’s 64.8. A 28-point generational leap on any benchmark deserves scrutiny, but if it holds up under independent testing, it says something real about the model’s long-horizon research ability.

IFBench: instruction following by a wide margin

On IFBench, Qwen 3.8-Max posts 82.8. The nearest non-Qwen competitor in the table is GPT-5.6 Sol at 72.7, with Fable 5 at 63.5 and Opus 4.8 at 62.2. That’s a 10-to-20 point gap, which is unusual on a benchmark this saturated. Interestingly, Qwen3.7-Max already led this row at 79.1, so instruction following looks like a durable Qwen strength rather than a one-off.

Terminal Bench 2.1: edging past Claude

On Alibaba’s run of Terminal Bench 2.1, Qwen 3.8-Max scores 86.6 against 84.6 for both Opus 4.8 and Fable 5. GPT-5.6 Sol still leads the row at 88.8, so this is a second-place finish, not a sweep. But beating both Anthropic flagships on an agentic terminal benchmark is exactly the kind of result Alibaba wanted from this launch, and the two-point margin over Claude is the number you’ll see quoted everywhere.

The multimodal sweep

The separate multimodal table is where Qwen 3.8-Max looks strongest overall. Alibaba reports MathVision at 95.2, LogicVista at 91.9, and OSWorld-Verified at 86.1, and the model leads nearly every OCR row in the table. Neither Opus 4.8 nor Fable 5 competes directly here (they’re text-focused in this comparison), which is partly why the multimodal table reads as a sweep. Against Gemini 3.1 Pro and GPT-5.6 Sol, the visual reasoning and document intelligence rows are the model’s clearest differentiator.

If you’re weighing it against the other big open-weight release of the summer, the multimodal gap is also the sharpest contrast: Kimi K3 is text-only. Our Qwen 3.8 vs Kimi K3 comparison covers that matchup in full.

Where it loses, kept honest

Alibaba left the losses in the table, which deserves some credit. Three stand out.

HLE: 43.6, well behind Fable 5. On Humanity’s Last Exam, the broad-knowledge frontier benchmark, Qwen 3.8-Max scores 43.6. Fable 5 sits at 53.3, GPT-5.6 Sol at 47.2, and Opus 4.8 at 45.7. That’s last place among the four flagships, nearly 10 points behind the leader. It beats Qwen3.7-Max’s 41.4, but only just. HLE resists benchmark-specific tuning better than most evals, which makes this row worth weighting heavily.

SWE-bench Pro: 67.7 vs Fable 5’s 80.0. On the harder software-engineering benchmark, Qwen 3.8-Max lands mid-pack: ahead of GPT-5.6 Sol (64.6), just behind Opus 4.8 (69.2), and a full 12 points behind Fable 5. The generational improvement over Qwen3.7-Max (60.6) is real, but “beats Claude on Terminal Bench” and “trails Fable 5 badly on SWE-bench Pro” are both true at once, and the second claim will appear in far fewer headlines.

DeepSWE and the harder agentic rows. DeepSWE is another row where Qwen 3.8-Max trails the frontier in Alibaba’s own table. The pattern across the coding section is consistent: competitive on terminal-driven agentic tasks, behind on the deepest software-engineering evaluations.

The honest summary: Qwen 3.8-Max beats Opus 4.8 on several agentic rows, trails Fable 5 on most core coding work, and wins big on multimodal and document intelligence. That’s a genuinely strong result for a model priced at $2 input and $6 output per million tokens (see the official pricing page), but it’s not the across-the-board frontier win a skim of the launch coverage might suggest.

The fine print most coverage will skip

This is the section that justifies reading a benchmarks post instead of a headline. Four details from Alibaba’s own publication change how you should read the table.

1. Every number is vendor-run. Alibaba evaluated its own model and its competitors’ models. That’s standard practice for launch tables, and it’s also the standard reason launch tables later get revised. Vendors pick the benchmark versions, the prompting strategies, the sampling settings, and the retry policies. None of that is fraud. All of it is a thumb on the scale.

2. Most coding rows ran on the Claude Code harness. This is the genuinely interesting detail. Alibaba ran most of its coding benchmarks with Claude Code, Anthropic’s own agent harness, pointed at Qwen 3.8-Max through its Anthropic-compatible API. On one hand, that’s a fair-play move: it evaluates every model inside the same widely used tool. On the other, it means “Qwen’s coding scores” are really “Qwen inside Anthropic’s harness” scores, and harness choice can swing agentic results by several points. It also tells you something about where Alibaba expects this model to run in practice.

3. Several benchmarks are Qwen in-house. QwenSWEBench, QwenQoderBench, CoWorkBench, and RecreationBench were all built by the Qwen team. In-house benchmarks aren’t worthless; they often probe capabilities public evals miss. But a model’s strong score on its maker’s own benchmark is a different kind of evidence than a strong score on SWE-bench Pro, and the table presents both side by side without visual distinction. When you quote a number from this table, check which category it belongs to first.

4. The footnote about Fable 5. Alibaba’s own table carries a footnote saying “Fable5 results may involve fallbacks.” That’s a quiet admission that at least one competitor’s numbers may not reflect the model running cleanly. Whatever the technical cause, it means the Fable 5 column, the model Qwen 3.8-Max most often trails, comes with an asterisk from the people who published it.

None of this is unique to Alibaba. Anthropic, OpenAI, and Google all publish self-run tables with their own methodological choices. We flagged the same pattern in our Kimi K3 benchmarks breakdown, and it applied there too. The point isn’t that Alibaba cheated. The point is that a vendor table is a claim, not a measurement.

What to watch, and how to read the next table

As of August 3, 2026, no independent evaluations of Qwen 3.8-Max exist. Artificial Analysis hadn’t published numbers at the time of writing, and none of the community leaderboards had scored the model. That will change within days, and the deltas between Alibaba’s table and the independent runs will be the real story. We’ll update this post when they land.

Until then, here’s a short checklist for judging any vendor benchmark table, this one included:

  1. Who ran the eval? Self-reported numbers for competitors deserve more skepticism than self-reported numbers for the vendor’s own model.
  2. Which harness and settings? Agentic benchmarks are harness-sensitive. A table that names its harness (as Alibaba did) is more useful than one that doesn’t, but the choice still shapes results.
  3. Which benchmarks did the vendor build? In-house evals measure something, but not the same something as public ones.
  4. Read the footnotes. “May involve fallbacks” is doing a lot of work in a small font.
  5. Check what’s missing. Rows a vendor didn’t publish are often as informative as rows it did. Past vendor tables have quietly dropped benchmarks where the model underperformed; compare against how the last generation stacked up across vendors to spot which rows vanished.
  6. Wait for the second source. One week of patience usually buys you independent numbers.

Run your own evaluation instead

The checklist helps you read tables. The better move is to need them less. Qwen 3.8-Max is available now on Alibaba Cloud Model Studio (model ID qwen3.8-max, listed on the official models page) with both an OpenAI-compatible and an Anthropic-compatible API, and a benchmark that uses your actual prompts beats any public eval that doesn’t.

A practical setup in Apidog looks like this. Create one request against the Model Studio chat-completions endpoint with qwen3.8-max and a second identical request against your current baseline model, whether that’s Qwen3.7-Max, Kimi K3, or a Claude or GPT endpoint. Save 20 to 30 of your real production prompts as a test scenario, add assertions for the response properties you actually care about (valid JSON, required fields present, latency under your threshold), and run both scenarios back to back. Apidog’s test reports give you pass rates and response times per model, side by side, on your workload instead of Alibaba’s.

Two Qwen-specific things to check while you’re at it: the model defaults to reasoning_effort: xhigh, so measure cost and latency at the effort level you’d really ship, and if you plan to use the Anthropic-compatible endpoint, test that protocol shape separately, because it’s the one most of Alibaba’s own coding benchmarks ran through. Download Apidog for free, wire up both endpoints as environments, and you’ll have first-party numbers before the independent leaderboards publish theirs.

Frequently asked questions

Are the Qwen 3.8 benchmark numbers independently verified?

No. As of August 3, 2026, every published number comes from Alibaba’s own table. Artificial Analysis and the community leaderboards hadn’t yet scored Qwen 3.8-Max at the time of writing. Treat the table as the vendor’s claim until independent runs land.

What is Qwen 3.8-Max’s best benchmark result?

PaperBench, at 93.0, is its best flagship-table row: ahead of GPT-5.6 Sol (90.5), Fable 5 (88.8), and Opus 4.8 (80.3). In the multimodal table, MathVision (95.2) and the OCR rows are its strongest results overall.

Where does Qwen 3.8-Max clearly lose?

On HLE it scores 43.6, last among the four flagships and well behind Fable 5’s 53.3. On SWE-bench Pro it posts 67.7 against Fable 5’s 80.0, and it also trails on DeepSWE. Deep software-engineering evals and broad-knowledge exams are its weakest areas in Alibaba’s own table.

How much better is Qwen 3.8 than Qwen 3.7-Max on benchmarks?

The generational gains in Alibaba’s table are large: Terminal Bench 2.1 went from 74.5 to 86.6, SWE-bench Pro from 60.6 to 67.7, and PaperBench from 64.8 to 93.0. Whether the upgrade is worth it for your workload depends on price and modality too; our Qwen 3.8 vs Qwen 3.7 comparison walks through the full decision.

Explore more

Qwen 3.8 vs Qwen 3.7 Max: What Actually Changed

Qwen 3.8 vs Qwen 3.7 Max: What Actually Changed

Qwen 3.8-Max vs 3.7-Max: benchmark deltas, the $2/$6 price vs the 50%-off promo, image input, and open weights. When to upgrade and when to wait.

3 August 2026

Qwen 3.8 for Coding: 16-Day Autonomous Runs and the Claude Code Connection

Qwen 3.8 for Coding: 16-Day Autonomous Runs and the Claude Code Connection

Qwen 3.8-Max for coding: Alibaba's 16-day autonomous run, benchmark results, and official configs for Claude Code, Codex, Qoder, Qwen Code, and OpenClaw.

3 August 2026

Qwen 3.8 Pricing Explained: $2 Input / $6 Output Across a 1M Context

Qwen 3.8 Pricing Explained: $2 Input / $6 Output Across a 1M Context

Qwen 3.8 pricing at GA: $2 input / $6 output per 1M tokens, flat across the full 1M context. Cache discounts, free quota fine print, and worked cost math.

3 August 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Qwen 3.8 Benchmarks: What Alibaba's Table Shows, and What It Doesn't