Mistral Large 4 went live on the Mistral API on October 6, 2026, three weeks before its open weights. If you want to try the 1-trillion-parameter “Le Chonk” now, the API is the only way in, and right now it is also the cheapest way: Mistral lists it at $0.68 per million input tokens and $2.09 per million output tokens during the public preview, half the $1.36 / $4.18 list price.
This guide gets you from zero to a working first call in about five minutes, then covers the parts that trip people up: reasoning chunks, image input, function calling, JSON output and cost. Every request can be saved and replayed in Apidog so you can compare Large 4 with whatever model you run today.
New to the model itself? Read Mistral Is Back: Le Chonk Beats GPT-6 Astra and Claude at Cyber first for the benchmarks and the catch behind the cyber headline.
What you need
| Item | Value |
|---|---|
| Base URL | https://api.mistral.ai/v1 |
| Auth | Authorization: Bearer $MISTRAL_API_KEY |
| Model ID | mistral-large-4 (alias mistral-large-4-0) |
| Main endpoint | POST /v1/chat/completions |
| Context window | 1M tokens |
| Input types | Text, images |
| Python SDK | pip install mistralai |
| TypeScript SDK | npm install @mistralai/mistralai |
Step 1: Get an API key
- Sign in to Mistral Studio (formerly La Plateforme).
- Open API Keys and create a new key. Give it a name that says where it will live, such as
local-devorci-staging. - Copy it once. Studio will not show it again.
- Export it in your shell:
export MISTRAL_API_KEY="your-key-here"
Keep the key out of source control. If you are wiring it into several tools, our guide to API key management best practices covers rotation and scoping.
Step 2: Make your first call
The fastest check is plain curl:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{"role": "user", "content": "Give me three edge cases to test on a pagination API."}
]
}'
A successful response returns choices[0].message.content with the answer and a usage block with prompt_tokens, completion_tokens and total_tokens. If you get a 401, the key is wrong or not exported. A 404 on the model usually means a typo in the model ID.
The same call in Python
import os
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[
{"role": "user", "content": "Give me three edge cases to test on a pagination API."}
],
)
print(response.choices[0].message.content)
And in TypeScript
import { Mistral } from "@mistralai/mistralai";
const client = new Mistral({ apiKey: process.env.MISTRAL_API_KEY });
const response = await client.chat.complete({
model: "mistral-large-4",
messages: [
{ role: "user", content: "Give me three edge cases to test on a pagination API." },
],
});
console.log(response.choices[0].message.content);
Step 3: Save it in Apidog
Typing curl commands gets old the moment you start comparing models. In Apidog:
- Create a new HTTP request:
POST https://api.mistral.ai/v1/chat/completions. - Add an environment variable
MISTRAL_API_KEYand set the headerAuthorization: Bearer {{MISTRAL_API_KEY}}. - Paste the JSON body from Step 2 and hit Send.
- Duplicate the request, change
modelto the model you use today (for examplemistral-medium-3-5), and run both.
You now have two saved requests with the same prompt. Apidog shows the response body, status, timing and size for each, so you can compare answer quality, latency and usage token counts without writing a script. Add a post-response assertion that choices[0].message.content is not empty and you have a smoke test you can rerun whenever Mistral updates the preview.
Step 4: Turn reasoning on and off
Large 4 is a hybrid model: the same model handles fast answers and step-by-step reasoning. You control it with one parameter, reasoning_effort:
| Value | Behavior | Use it for |
|---|---|---|
"none" |
Minimal thinking, no thinking chunk in the response | Chat, extraction, classification, anything latency-sensitive |
"high" |
Full thinking chunk before the final answer | Debugging, multi-step planning, math, code review |
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{"role": "user", "content": "Our API returns 200 with an empty body under load. List likely causes in order of probability."}
],
"reasoning_effort": "high"
}'
This is the part that breaks parsers. With reasoning_effort: "high", message.content is no longer a string. It becomes a list of chunks:
- a
thinkingchunk holding the reasoning trace, and - a
textchunk holding the final answer.
So response.choices[0].message.content will print a list, not your answer. Pull the text chunk out explicitly:
response = client.chat.complete(
model="mistral-large-4",
messages=[{"role": "user", "content": "Why would a 200 response have an empty body?"}],
reasoning_effort="high",
)
content = response.choices[0].message.content
if isinstance(content, str):
answer = content
else:
answer = "".join(c.text for c in content if c.type == "text")
print(answer)
Thinking tokens are billed as output tokens, so "high" costs more per request. Default to "none" and switch to "high" only on the calls that need it.
Step 5: Send an image
Large 4 is natively multimodal, with a 1.6B-parameter vision encoder. Pass images as content parts next to your text:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "This is a screenshot of our API error dashboard. Which endpoint is failing most and what is the error code?"},
{"type": "image_url", "image_url": "https://example.com/dashboard.png"}
]
}
]
}'
For local files, send a base64 data URL instead: "image_url": "data:image/png;base64,<encoded>". Mistral reports Large 4 scores 42% on the Dense 200 visual-grounding benchmark, just ahead of GPT-6 Astra’s 41%, so screenshots of dashboards, charts and UI states are a reasonable fit.
Step 6: Function calling
Function calling is where Large 4’s agent benchmarks (59.9% on AutomationBench) become useful. You describe tools, the model decides when to call them, and your code runs the call.
tools = [
{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the status of an order by its ID.",
"parameters": {
"type": "object",
"properties": {
"order_id": {"type": "string", "description": "The order ID, e.g. ORD-1042"}
},
"required": ["order_id"],
},
},
}
]
messages = [{"role": "user", "content": "Where is order ORD-1042?"}]
response = client.chat.complete(
model="mistral-large-4",
messages=messages,
tools=tools,
tool_choice="auto",
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
Run the function yourself, then send the result back with the matching tool_call_id:
import json
result = {"order_id": "ORD-1042", "status": "shipped", "eta": "2026-10-09"}
messages.append(response.choices[0].message)
messages.append({
"role": "tool",
"name": "get_order_status",
"content": json.dumps(result),
"tool_call_id": tool_call.id,
})
final = client.chat.complete(model="mistral-large-4", messages=messages, tools=tools)
print(final.choices[0].message.content)
The tool schema is plain JSON Schema. If your API already has an OpenAPI spec, you can lift the request schema for each operation straight into parameters. Designing the spec in Apidog first keeps the tool definitions and the real API in sync.
Step 7: Get JSON back
When you need machine-readable output, set response_format:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{"role": "user", "content": "Extract method, path and status code from: GET /v1/users/42 returned 404. Reply in JSON."}
],
"response_format": {"type": "json_object"}
}'
Mention JSON in the prompt as well as in response_format. For strict shapes, Mistral also supports {"type": "json_schema", "json_schema": {...}} with a full schema. In Apidog, add a JSON Schema assertion on the response so a drifting shape fails loudly instead of breaking a downstream service.
What it costs
| Usage | Preview price | List price |
|---|---|---|
| Input, per 1M tokens | $0.68 | $1.36 |
| Cached input, per 1M tokens | $0.07 | $0.14 |
| Output, per 1M tokens | $2.09 | $4.18 |
A worked example: an agent that makes 10,000 calls a day, each with 3,000 input tokens (mostly a cached system prompt and tools) and 500 output tokens.
- Input: 30M tokens. If 2,500 of each 3,000 are cached, that is 25M cached at $0.07 and 5M fresh at $0.68, about $5.15/day.
- Output: 5M tokens at $2.09, about $10.45/day.
- Total: roughly $15.60/day at preview pricing, or about $31 at list price.
The same workload on GPT-6 Astra ($10 / $50 per million, before caching discounts) would cost several hundred dollars a day. Mistral has not said when preview pricing ends, so budget against the list price.
Common errors
| Error | Likely cause | Fix |
|---|---|---|
401 Unauthorized |
Missing or wrong key | Check echo $MISTRAL_API_KEY and the Bearer prefix |
404 / invalid model |
Typo in the model ID | Use mistral-large-4 exactly |
422 Unprocessable Entity |
Malformed body, often a bad tools schema |
Validate the JSON Schema in each tool’s parameters |
429 Too Many Requests |
Rate limit for your workspace tier | Back off and retry, or raise limits in Studio |
| Answer prints as a list | reasoning_effort: "high" returns chunks |
Extract the text chunk (Step 4) |
FAQ
Is Mistral Large 4 OpenAI-compatible? The request shape is very close: model, messages, tools, tool_choice and response_format all work the way you expect. Use the Mistral SDKs or plain HTTP to be safe. Reasoning output uses Mistral’s own chunk format.
When can I run it locally? Mistral says the weights ship by the end of October 2026. At 1.05T total parameters it needs multi-GPU server hardware. Our run Mistral 3 locally guide covers the tooling for the smaller models in the meantime.
Is the preview stable enough for production? Not yet. The model is labeled public preview and may change before the weights release. Pin your tests, rerun them when Mistral updates the model, and keep a fallback model configured.
Can I use Large 4 with my existing Mistral code? Yes. Same base URL, same auth, same SDK. Change the model string to mistral-large-4. If you are coming from Medium 3.5, see our Mistral Medium 3.5 API guide for the parts that carry over.
Wrap-up
Five minutes gets you a working call. The next hour is better spent running your real prompts against Large 4 and your current model side by side. Save both requests in Apidog, add assertions on status and response shape, and you will know within a day whether Le Chonk earns a place in your stack, while the preview price is still half off.



