Mistral Large 4's official pages disagree. The 2026-10-06 announcement says 1 trillion total parameters and 49 billion active parameters, while the model card says 1.05 trillion total and 52 billion active, plus a 1.6 billion vision encoder. Neither page explains the 3 billion active-parameter difference, so the safe answer is to report both figures.
Key facts
- On 2026-10-06, Mistral announced a public preview of Mistral Large 4 and linked a preview API on Mistral Studio. (announcement)
- The launch post describes a 1 trillion-parameter model with 49 billion active parameters. (announcement)
- The model card describes 1.05 trillion total parameters, 52 billion active parameters, and a 1.6 billion vision encoder. (model card)
- The model card lists a 1M-token context window and a granular Mixture-of-Experts architecture. (model card)
- Mistral says the model was trained on 3,800 NVIDIA Grace Blackwell GPUs and that the preview runs on the same European infrastructure. (announcement)
- Mistral's 2026-10-06 changelog says launch pricing is 50% off for two weeks and that open weights are coming soon. (changelog)
- The launch post says the weights will be released by the end of October 2026. (announcement)
Why do Mistral Large 4's announcement and model card cite different sizes?
The two official pages publish different headline specifications on 2026-10-06. In the announcement blog post, Mistral writes:
“ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters.”
The model card instead lists 52B active parameters, 1.05T total parameters, and a 1.6B vision encoder. It also lists a 1M-token context window. The launch post does not publish the same component breakdown.
Specification | Announcement blog post | Official model card | What is unresolved |
|---|---|---|---|
Total parameters | 1.00T | 1.05T | The pages do not define whether the totals use the same counting convention. |
Active parameters per token | 49B | 52B | Neither page explains the 3B difference. |
Vision encoder | Not specified | 1.6B | The launch post does not say how vision parameters are counted. |
Context window | Not specified in the post | 1M tokens | The changelog also lists a 1M context window. |
Architecture wording | Natively multimodal | Granular MoE | The descriptions are not a parameter-count reconciliation. |
The 1.00T versus 1.05T gap could reflect different rounding or counting conventions, but Mistral does not say that. The active-parameter gap is more important for capacity and serving estimates, and it remains unresolved. Until Mistral publishes a correction or weights, cite the source and date whenever you repeat either number.
What is confirmed about Mistral Large 4's architecture and performance?
The confirmed description is a multimodal granular Mixture-of-Experts model with a 1M-token context window and a 1.6B vision encoder listed in the model card. The launch post adds that Mistral trained it on 3,800 NVIDIA Grace Blackwell GPUs in its European datacenters and serves the preview on that infrastructure.
Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and 49.8% on its Coding Agent Index. It reports 59.9% on AutomationBench and 1,393 Elo on AA-Briefcase. On security evaluations, Mistral reports 82% on reproducing and patching vulnerabilities and 93% on Cybench. These are vendor-reported results, not independent verification.
Like we analyzed in GPT-5.6 Sol, Terra, and Luna: which tier to use for agent work, model selection for agent work depends on the workload, not just a headline parameter count. What changed in Claude Code during September 2026? and What changed in the Codex CLI refresh of September 2026 show why the surrounding tool also matters when a model is used for coding agents.
What does Mistral Large 4 cost during its preview window?
The 2026-10-06 changelog says launch pricing is 50% off for two weeks. Counting two weeks from 2026-10-06 gives 2026-10-20 as the calendar end date; the changelog itself states the duration rather than an end date. The pricing page shows the original and sale prices together.
Token category | Original price (per M tokens) | Launch price (per M tokens) | Effective discount |
|---|---|---|---|
Input tokens | $1.36 | $0.68 | 50% |
Cached input tokens | $0.14 | $0.07 | 50% |
Output tokens | $4.18 | $2.09 | 50% |
At the original rate, Mistral Large 4 costs more than Mistral Large 3 at $0.50 input and $1.50 output per million tokens, and less than Mistral Medium 3.5 at $1.50 input and $7.50 output. During the launch period, cached input is $0.07 per million tokens.
When will Mistral release the weights for Mistral Large 4?
Mistral says it will release the weights by the end of October 2026. Until that happens, the two official parameter counts are both published claims, not a resolved implementation fact. A careful post should preserve the discrepancy instead of silently choosing 49B or 52B.
Sources
- Introducing Mistral Large 4 (read 2026-10-07)
- Mistral Large 4 model card (read 2026-10-07)
- Mistral Docs changelog (read 2026-10-07)
- Mistral Docs inference pricing (read 2026-10-07)
Last verified: 2026-10-07.