July 2026 turned into a two-horse month for open-weight frontier models, and both horses came out of China. Moonshot AI shipped Kimi K3 on July 16. Three days later, Alibaba previewed Qwen 3.8-Max, then made it generally available in early August. Both are trillion-parameter mixture-of-experts models. Both promise (or deliver) open weights. Both plug into Claude Code-style agent harnesses. And both undercut the US frontier labs on price.
The surface similarity hides real differences: in what you can download today, in what the models can see, and in what a month of heavy usage costs. This comparison walks through every axis you can verify as of August 3, 2026. For the full spec breakdown of Alibaba’s new flagship, start with our Qwen 3.8-Max explainer and come back.
One honesty note before we start: there is no independent head-to-head benchmark of these two models yet. Every number in this article comes from one vendor’s own table, and we label whose run each number is. Treat the benchmark section as indicative, not settled.
The timeline: who shipped what, and when
The release order matters more than usual here, because “open weights” means something different for each model right now.
Kimi K3 launched on July 16, 2026. Moonshot promised open weights at launch, and delivered them 11 days later: the weights hit Hugging Face on July 27 as a 594 GB MXFP4 release. As of today, you can download K3, quantize it further, and run it on your own hardware. That is not a promise. It happened.
Qwen 3.8-Max was previewed on July 19 and reached general availability on Alibaba Cloud Model Studio in early August, per the official Qwen release post. The weights are promised for Hugging Face and ModelScope “next week,” which puts them around August 10. As of August 3, 2026, they are not downloadable. K3’s own 11-day gap between launch and weights shows this arc is normal, but today only one of these models is open in practice.
If self-hosting is your deciding factor, that single fact may end the comparison early.
Parameters: two trillion-scale MoE models, one leaner
Both models are sparse mixture-of-experts designs, which means the headline parameter count and the per-token compute cost are very different numbers.
| Spec | Qwen 3.8-Max | Kimi K3 |
|---|---|---|
| Total parameters | 2.4T | 2.8T |
| Active parameters per token | 95B | 104B |
| Activation ratio | ~4% | ~3.7% |
| Architecture base | Qwen 3.5 foundation, MoE | MoE |
| Open-weight format | Promised (~Aug 10) | MXFP4, 594 GB on Hugging Face |
Qwen 3.8-Max is the smaller model on both counts: 400B fewer total parameters and 9B fewer active parameters than K3. Neither gap is dramatic, and active-parameter counts don’t translate linearly into capability. What they do translate into is serving cost. A model activating 95B parameters per token is cheaper to run at scale than one activating 104B, all else equal, and that difference shows up in the API pricing below.
For self-hosters, the more relevant number is K3’s 594 GB download at MXFP4 precision. That is multi-node territory even before you account for KV cache at long context. Expect Qwen 3.8-Max, at 2.4T total, to land in a similar weight class when its files appear. Most teams will use hosted endpoints for both models and reserve self-hosting for compliance or research cases.
Openness today: live weights vs a dated promise
This axis deserves its own section because the marketing on both sides says “open” while the facts differ.
Kimi K3’s weights are public. You can verify the repo, check the license, and pull the files right now. That brings everything real openness enables: third-party hosting, community quantizations, fine-tuning, and the assurance that the model you benchmarked can’t be silently swapped out from under you.
Qwen 3.8-Max would be the first Qwen-Max-class model ever released with open weights, a genuine shift in Alibaba’s strategy after keeping every previous Max-tier model API-only. But a first is only a first once it happens. Until the Hugging Face repo is live, Qwen 3.8-Max is an API model with an announcement attached.
Score this axis for K3 today, with a pencil note to re-check around August 10. If Alibaba delivers, the axis becomes a tie and the comparison shifts to license terms and quantization quality.
Modality: Qwen sees images, K3 reads text
Here the gap runs the other way, and it isn’t close.
Qwen 3.8-Max accepts text and image input, per the official model configs ("input_modalities": ["text", "image"]). Alibaba’s launch post goes further, demonstrating long-document understanding across 200-plus-page PDFs and 100-hour video comprehension via memory graphs; treat those as blog demos, not documented API input types. The image input, though, is real and in the API today. Alibaba’s vendor benchmarks lean into it: the multimodal table (MathVision 95.2, LogicVista 91.9, OSWorld-Verified 86.1, strong OCR rows) is where Qwen 3.8-Max posts its most one-sided numbers.
Kimi K3 is a text model. If your workload involves screenshots, scanned documents, UI understanding, or any vision-adjacent agent task, K3 needs a separate vision model bolted onto the pipeline, with all the glue code and extra latency that implies.
For text-only coding and agent work, this axis is irrelevant. For document intelligence and computer-use agents, it is decisive.
Context, output, and the API surface
Qwen 3.8-Max ships a 1,000,000-token context window as a single flat tier, with maximum output of 65,536 tokens. Reasoning is controlled by a reasoning_effort parameter (xhigh by default, with medium and low), and thinking output is preserved by default. Notably, thinking and non-thinking modes are billed at the same rate.
Access runs through Alibaba Cloud Model Studio with three regional base URLs (Beijing, Singapore, US-Virginia) and, unusually, two protocol shapes on one platform: an OpenAI-compatible endpoint and an Anthropic-compatible endpoint. That dual-protocol surface is why Qwen 3.8-Max drops into Claude Code with nothing more than ANTHROPIC_BASE_URL and ANTHROPIC_MODEL=qwen3.8-max. The full lineup and regional availability are documented on the Model Studio models page; if you work across regions, tools like Apidog let you store each base URL as an environment and switch between them without editing requests.
Kimi K3 also speaks the Anthropic protocol and slots into Claude Code-style harnesses, which is exactly why these two models compete for the same seat: the “point your existing agent harness at a cheaper frontier model” seat. For K3’s endpoint details and current limits, see our Kimi K3 explainer rather than trusting numbers that may have shifted since launch.
Pricing: flat and simple vs cheap-in, expensive-out
Here are the published rates, per million tokens:
| Rate | Qwen 3.8-Max | Kimi K3 |
|---|---|---|
| Input | $2.00 | $3.00 |
| Output | $6.00 | $15.00 |
| Cache-hit input | $0.20 (10% of input) | $0.30 |
| Tiering | Flat across full 1M context | Per Moonshot’s published schedule |
Qwen 3.8-Max costs $2 in and $6 out, flat from token zero to token one million, thinking or not, per the official Model Studio pricing page. Explicit cache creation bills at 125% of input, cache hits at 10%. There’s also a free quota: 1M tokens, Singapore region only, valid 90 days.
K3’s rates are $0.30 per million on cache hits, $3 per million input, and $15 per million output.
The comparison isn’t one number. On input-heavy workloads with good cache discipline (long system prompts, repeated document context), the two models land closer than the sticker suggests: $0.20 vs $0.30 per cached megatoken is a narrow gap. On output-heavy workloads, the gap is wide: K3’s $15 output rate is 2.5x Qwen’s $6. Agent loops that generate a lot of reasoning and code lean output-heavy, which favors Qwen 3.8-Max on paper.
One caution that applies to Qwen specifically: reasoning_effort defaults to xhigh, and thinking tokens bill as output. Real invoices will run above what a naive tokens-in-tokens-out estimate predicts. Our Qwen 3.8 pricing breakdown works through cost-per-task examples if you want to model this before committing.
Benchmarks: two vendor tables, no referee
This is where most comparisons of these two models go wrong, so let’s be precise about what exists.
What exists: Alibaba published a benchmark table for Qwen 3.8-Max against Claude Opus 4.8, Fable 5, GPT-5.6 Sol, and Qwen 3.7-Max. Moonshot published its own launch table for Kimi K3 against a similar frontier lineup. Both are vendor-run.
What does not exist: any independent evaluation that puts Qwen 3.8-Max and Kimi K3 in the same table under the same harness. No Artificial Analysis numbers for Qwen 3.8-Max were available at the time of writing. The two vendors did not benchmark each other’s models.
With that frame, here is what each side claims on its own run.
From Alibaba’s table (Alibaba’s run, largely executed through the Claude Code harness, a detail Alibaba itself discloses):
| Benchmark | Qwen 3.8-Max | Claude Opus 4.8 | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal Bench 2.1 | 86.6 | 84.6 | 84.6 | 88.8 |
| SWE-bench Pro | 67.7 | 69.2 | 80.0 | 64.6 |
| PaperBench | 93.0 | 80.3 | 88.8 | 90.5 |
| GPQA Diamond | 92.6 | 92.0 | 92.6 | 94.1 |
| HLE | 43.6 | 45.7 | 53.3 | 47.2 |
Alibaba’s own numbers include clear losses: Qwen 3.8-Max trails Fable 5 badly on SWE-bench Pro and trails everyone listed on HLE. Its standout rows are PaperBench, instruction following (IFBench 82.8), and the multimodal sweep. Several other headline results use Alibaba’s in-house benchmarks (QwenSWEBench, CoWorkBench, and others), which can’t be cross-referenced against anyone.
From Moonshot’s side, K3’s launch table showed it beating Claude Opus 4.8 across Moonshot’s headline benchmark rows while trailing Fable 5 and GPT-5.6 Sol overall. We covered those claims, and their caveats, in our Kimi K3 vs Claude Opus 4.8 comparison.
Can you line up the two tables where benchmarks overlap? You can, and people will, but different harnesses, scaffolding, and sampling settings make cross-table reads indicative at best. The honest summary: both vendors show their model competitive with, and selectively ahead of, the US frontier on agentic tasks. Neither table settles which of the two is stronger. For the fine print on Alibaba’s numbers, see our Qwen 3.8 benchmarks analysis.
Run your own head-to-head in Apidog
Since no referee exists, the most useful benchmark is the one you run on your own workload. Both models expose OpenAI-compatible chat-completions endpoints, which makes a side-by-side test straightforward in Apidog.
Set up two environments, one holding the DashScope base URL and qwen3.8-max, the other holding Moonshot’s endpoint and the K3 model ID, each with its own API key variable. Save one request with your real prompt (an actual task from your product, not a riddle) and flip between environments to fire the same payload at both models. Apidog’s response view gives you status, latency, and the full body side by side, and its SSE debugging shows how each model streams reasoning content, which matters if your UI renders thinking tokens. Add a couple of assertions (response schema, latency ceiling, required keywords in the output) and you’ve turned a vibe check into a repeatable test you can rerun when either vendor ships an update. Download Apidog for free to run the comparison; ten prompts through both endpoints will tell you more than either vendor’s table.
The scorecard
| Axis | Winner today | Caveat |
|---|---|---|
| Total/active parameters | Qwen 3.8-Max (leaner: 2.4T/95B vs 2.8T/104B) | Smaller isn’t better or worse by itself; it mainly signals serving cost |
| Open weights, today | Kimi K3 (live on Hugging Face, 594 GB MXFP4) | Qwen’s weights are due ~Aug 10; re-check, this axis may tie soon |
| Modality | Qwen 3.8-Max (text + image input) | Video/long-doc claims are blog demos, not documented API input types |
| Context window | Qwen 3.8-Max (1M flat tier, 65,536 max output) | Verify K3’s current limits against Moonshot’s docs for your use case |
| Price | Qwen 3.8-Max ($2/$6 flat vs $3/$15) | xhigh thinking default inflates Qwen’s real output bills |
| Benchmarks | No call possible | Both tables are vendor-run; no independent head-to-head exists |
| Harness ecosystem | Tie | Both drop into Claude Code-style harnesses via Anthropic-protocol endpoints |
| Track record on promises | Kimi K3 | Weights shipped 11 days post-launch; Qwen’s promise is still open |
Which one when
Pick Kimi K3 if you need weights on your own hardware this week, your compliance posture requires a model you can pin and audit, or your workload is pure text and input-heavy enough that cache pricing dominates.
Pick Qwen 3.8-Max if your workload touches images or documents, you generate a lot of output tokens (agent loops, code generation), or you want a 1M-token context without tiered pricing surprises.
Wait a week if open weights are your whole reason for looking at either model. If Qwen’s weights land as promised around August 10, the openness axis resets and the decision comes down to modality, price, and your own eval results.
The real story is that developers now have two credible, cheap, harness-compatible frontier options from China within three weeks of each other. Whichever you lean toward, run your own prompts through both endpoints before you commit a roadmap to either vendor’s table.
To fully evaluate Qwen 3.8 in this matchup, it helps to know what Alibaba actually changed under the hood — the architectural and benchmark differences between Qwen 3.8 and Qwen 3.7 Max clarify which gains are structural versus the result of additional tuning.
FAQ
Is Kimi K3 more open than Qwen 3.8-Max?
Today, yes, as a plain matter of fact: K3’s weights have been on Hugging Face since July 27, 2026 (594 GB, MXFP4), while Qwen 3.8-Max’s weights are promised for around August 10 and aren’t downloadable yet. If Alibaba delivers, both will be open-weight frontier models and the meaningful differences shift to license terms and community support. Until then, only K3 can be self-hosted; our local K3 guide explains what that takes.
Which model is better at coding, Qwen 3.8-Max or Kimi K3?
Nobody can answer that honestly yet. Both vendors publish strong agentic-coding numbers, but each ran its own evals under its own conditions, and no independent head-to-head exists. Alibaba’s own table even shows Qwen 3.8-Max losing to Fable 5 on SWE-bench Pro, so vendor tables do admit losses; they just aren’t comparable across vendors. Run both models on tasks from your own repo before deciding.
Can I use both models in Claude Code?
Yes. Both expose Anthropic-compatible endpoints. For Qwen 3.8-Max, set ANTHROPIC_BASE_URL to the DashScope Anthropic endpoint and ANTHROPIC_MODEL=qwen3.8-max using the config from Alibaba’s release post. K3 supports the same pattern through Moonshot’s endpoint. That shared harness compatibility is what makes A/B testing the two so cheap to set up.
Which is cheaper for a typical agent workload?
On list price, Qwen 3.8-Max: $2/$6 per million tokens against K3’s $3/$15, and agent loops skew toward output tokens, where the gap is 2.5x. Two qualifiers: K3’s $0.30 cache-hit rate keeps input-heavy, well-cached workloads competitive, and Qwen’s default xhigh reasoning effort bills thinking as output, so its real-world costs run above naive estimates. Model your own token mix before assuming the sticker prices tell the story.



