Moonshot’s Kimi K2.7 Code landed with benchmarks against the two models most developers actually pay for: Claude Opus and GPT-5.5. The short version is that the closed frontier still scores higher on most coding tasks, but not by the margin its price suggests. Kimi is an open-weight model you can download and self-host, and it lands within a few points of models you can only rent.
This comparison puts the numbers side by side, weighs them against cost and openness, and tells you which one to pick for which job.
TL;DR
- Raw coding quality: GPT-5.5 and Claude Opus lead on most suites. Kimi K2.7 Code trails by a few points.
- Cost: Kimi is far cheaper ($0.95/$4.00 per million tokens) and free to self-host. The closed models cost more and can’t be hosted on your own hardware.
- Openness: Kimi ships open weights under a modified MIT license. Opus and GPT-5.5 are closed.
- Verdict: pick the frontier for the highest single-shot quality; pick Kimi when cost, volume, or data control matter more than the last few points.
The contenders at a glance
| Kimi K2.7 Code | Claude Opus | GPT-5.5 | |
|---|---|---|---|
| Type | Open weight (modified MIT) | Closed | Closed |
| Architecture | MoE, 1T total / 32B active | Not disclosed | Not disclosed |
| Context window | 256K tokens | Large | Large |
| Self-host | Yes | No | No |
| Pricing (per M tokens) | $0.95 in / $4.00 out | Higher, rented only | Higher, rented only |
The structural difference is the headline. Kimi tells you exactly what it is and lets you run it yourself. The other two are services you call.
Coding benchmarks
These are Moonshot’s reported scores. Treat them as the vendor’s framing, since several suites are Moonshot’s own, but the pattern is consistent and worth reading.
| Benchmark | Kimi K2.7 Code | GPT-5.5 | Claude Opus |
|---|---|---|---|
| Kimi Code Bench v2 | 62.0 | 69.0 | 67.4 |
| Program Bench | 53.6 | 69.1 | 63.8 |
| MLS Bench Lite | 35.1 | 35.5 | 42.8 |
GPT-5.5 takes the top spot on the first two; Claude Opus leads on MLS Bench Lite and stays close everywhere. Kimi sits a notch below on each. The gap is real but small on Kimi Code Bench v2, wider on Program Bench. If your work looks like that second benchmark, the frontier earns its keep.
Agentic and tool-use benchmarks
Coding agents live or die on tool calls, not just code generation. Here the picture tightens.
| Benchmark | Kimi K2.7 Code | GPT-5.5 | Claude Opus |
|---|---|---|---|
| Kimi Claw 24/7 | 46.9 | 52.8 | 50.4 |
| MCP Atlas | 76.0 | 79.4 | 81.3 |
| MCP Mark Verified | 81.1 | 92.9 | 76.4 |
On MCP Mark Verified, Kimi actually edges out Claude Opus, though GPT-5.5 pulls well ahead. On MCP Atlas the three sit within about five points. For agentic workloads, Kimi is closer to the pack than its raw coding scores suggest, which matters because that’s exactly where token costs pile up.
Cost is where Kimi wins
Benchmarks measure quality. They don’t measure the bill. Kimi K2.7 Code runs $0.95 per million input tokens and $4.00 per million output, with cache hits at $0.19 per million. The closed frontier models cost meaningfully more per token, and you can’t escape that by hosting them; there’s no download.

Two factors compound the savings:
- Open weights mean a zero-token option. Self-host on your own hardware and the per-token cost disappears entirely. You trade it for GPU time.
- Leaner reasoning. K2.7 Code uses about 30% fewer thinking tokens than K2.6 for the same work, so each agent step bills less. Our guide on reducing agent token costs goes deeper on why that adds up.
For an agent making hundreds of tool calls per task, a model that’s 90% as good at a fraction of the price often wins on total value.
Context and openness
All three handle long context well; Kimi’s window is 256K tokens, enough for a full service plus its tests in one prompt. The bigger differentiator is the license. Kimi’s modified MIT weights let you fine-tune, audit, and run the model air-gapped. For teams with compliance or data-residency rules, that’s not a nice-to-have; it’s the deciding factor, and the closed models are simply off the table.
When to pick each
Pick GPT-5.5 or Claude Opus if:
- You need the highest single-shot coding quality and a few benchmark points justify the cost.
- You’re fine renting a closed service with a managed SLA.
- Your hardest tasks look like Program Bench, where the gap is widest.
Pick Kimi K2.7 Code if:
- You run high-volume agent workloads where token cost decides viability.
- Data has to stay on your own infrastructure.
- You want to fine-tune or audit the model.
- You’re optimizing total value, not the top of the leaderboard.
For a broader field, our MiniMax M3 vs DeepSeek V4 vs Qwen 3.7 comparison covers the other open-weight challengers, and DeepSeek V4 vs Claude Opus for coding digs into another open-vs-closed matchup. If you’re comparing the agents rather than the models, see Claude Code vs OpenAI Codex.
Try the comparison yourself
Benchmarks are a starting point, not a verdict for your codebase. The honest test is running all three on your own tasks and reading the results. The Kimi Code CLI gives you the agent for free to start, and you can call each model’s API directly to compare raw outputs.
When you do, test the endpoints in Apidog: send the same prompt to each model, save the calls side by side, and compare response quality, latency, and token usage in one place. Download Apidog to run the bake-off.
FAQ
Is Kimi K2.7 Code better than Claude Opus or GPT-5.5? On Moonshot’s reported coding benchmarks, no; the closed models score higher on most suites. Kimi’s advantage is open weights and much lower cost while staying within a few points.
How much cheaper is Kimi? It’s $0.95 per million input tokens and $4.00 per million output, and free to self-host. The closed frontier models cost more per token and can’t be self-hosted.
Can I run Kimi K2.7 Code myself? Yes. The weights are open under a modified MIT license; serve them with vLLM, SGLang, or KTransformers.
Which is best for coding agents? For raw quality, GPT-5.5 leads. For cost-efficient high-volume agents, Kimi is the better value, and it’s close on agentic benchmarks.
Are these benchmarks neutral? Several suites are Moonshot’s own, so read them as the vendor’s framing. The consistent gap to the frontier is the useful signal, not the exact numbers.
Summary
Kimi K2.7 Code doesn’t beat Claude Opus or GPT-5.5 on most coding benchmarks, and Moonshot doesn’t pretend it does. It scores a few points lower while costing far less and shipping open weights you can run yourself. For the highest possible quality, the closed frontier still wins. For high-volume agents, private deployments, or anyone weighing total value over the top of the chart, Kimi is the smarter pick. Run all three on your own code, compare the outputs in Apidog, and let your workload decide.
