Mistral Large 4 (Le Chonk): Specs, Price, Benchmarks & Release Date
Updated October 6, 2026: Mistral AI has launched the public preview of Mistral Large 4, nicknamed “Le Chonk.” The new flagship is a 1.05-trillion-parameter multimodal Mixture-of-Experts model with 49 billion active parameters, a 1.6B vision encoder, and an official context window of up to 1 million tokens. The API preview is available now, while downloadable weights are expected later in October.
This guide brings together the confirmed specifications, pricing, early benchmark evidence, release timing, deployment implications, and the important caveats that are easy to miss in launch-day coverage. For more new-model coverage, visit AVARIXO’s AI & Tech section.
Mistral Large 4 at a glance
- Official name: Mistral Large 4
- Nickname: Le Chonk
- Preview date: October 6, 2026
- Architecture: Sparse Mixture-of-Experts
- Total parameters: 1.05 trillion
- Active parameters: 49 billion
- Vision encoder: 1.6 billion parameters
- Official context window: Up to 1 million tokens
- API model ID:
mistral-large-4 - Preview price shown in Mistral Docs: $0.68/M input, $0.07/M cached input, $2.09/M output
- Weights: Planned for late October; Reuters reported October 27
What is Mistral Large 4?
Mistral Large 4 is the newest flagship model from Mistral AI and the largest model the company has publicly described. It is designed as a general-purpose multimodal model for reasoning, coding, document analysis, visual understanding, and tool-enabled application workflows.
The headline number is 1.05 trillion total parameters. The more important number for day-to-day inference, however, is 49 billion active parameters. Large 4 uses a sparse Mixture-of-Experts architecture, meaning each token is routed through only a subset of the model’s expert networks instead of activating the entire trillion-parameter system every time.
That design lets Mistral add much more total model capacity without making the compute cost scale linearly with the full parameter count. It also explains why the model can be enormous while still being practical enough to offer through a commercial API.
Why the “Le Chonk” nickname matters less than the underlying release
Mistral has embraced the “Le Chonk” nickname as a playful continuation of a fat-cat meme that circulated around its community earlier in 2026. The meme generated attention, but the technical release is the real story: Large 4 is live in public preview, the API is documented, the model has current pricing, and Mistral says open weights are coming later this month.
The timing also puts Mistral directly into a fast-moving open-weight race that includes models from Reflection AI, Z.ai, Moonshot AI, Qwen, and DeepSeek. Mistral’s pitch is not only benchmark performance. It is also offering a European-developed model that organizations may eventually be able to host and customize themselves.

Mistral Large 4 specifications
| Specification | Mistral Large 4 |
|---|---|
| Launch status | Public Preview |
| Preview date | October 6, 2026 |
| Architecture | Sparse Mixture-of-Experts |
| Total parameters | 1.05T |
| Active parameters | 49B |
| Vision encoder | 1.6B |
| Context window | Up to 1M tokens in official Mistral documentation |
| Input | Text and multimodal inputs |
| Output | Text |
| API model ID | mistral-large-4 |
| Open weights | Expected later in October 2026 |
Mistral’s official model page also lists structured outputs, function calling, document Q&A, chat completions, batching, agents, conversations, and built-in tools. Those features make the model suitable for applications that need predictable data formats or multi-step workflows, not only conversational use.
Mistral Large 4 pricing
Mistral’s official documentation currently shows preview pricing of $0.68 per million input tokens, $0.07 per million cached input tokens, and $2.09 per million output tokens. The page also displays higher figures next to those values — $1.36 input, $0.14 cached input, and $4.18 output — suggesting a discounted public-preview rate.
Because Large 4 is still in preview, teams should check the official Mistral model page before calculating a long-term production budget. Preview pricing can change, and third-party inference providers may eventually offer different prices after the weights are released.
Why cached input pricing could matter
The very low cached-input rate is particularly relevant for long-context workloads. If an application repeatedly sends the same repository snapshot, long reference manual, policy collection, or background context, caching can reduce the cost of resending that material. The important metric is still the total cost per completed task, because multi-step applications may generate many requests and long outputs.
What the early benchmarks show
Mistral has published strong launch numbers, but readers should separate company-reported results from independent measurements. The preview is new, and broad third-party testing has only just started.
| Benchmark or workload | Large 4 result | Status |
|---|---|---|
| DeepSWE v1.1 | 62% | Mistral launch material |
| Finch / FinWorkBench | 67% | Mistral launch material |
| Harvey Legal Agent | About 15% in launch material; Vals lists 15.83% | Early independent result also available |
Vals AI has already added Mistral Large 4 to its model evaluations. Its page currently reports a 48.05% overall Vals Index result and ranks the model sixth on the Harvey Legal Agent benchmark. That does not settle the question of how Large 4 compares across all workloads, but it is a useful independent data point on launch day.
The 1M context window needs one important caveat
Mistral’s official documentation lists a context window of 1 million tokens. Vals AI currently shows a 512K context configuration on its evaluation page. That difference can occur when a testing provider uses a specific endpoint configuration or limits the context during evaluation.
For buyers and developers, the safest interpretation is straightforward: the official model specification is up to 1M tokens, but the usable limit can depend on the endpoint or provider. Always verify the actual API configuration you are using.
Mistral Large 4 vs Reflection Beam
Reflection AI introduced Beam at almost the same time, so comparisons are inevitable. Beam is a 501B-total, 23B-active sparse model focused on coding, reasoning and tool-based software work. Reflection says its weights will be released under Apache 2.0 later in October.
Large 4 is larger, activates more parameters, and adds multimodal input. On Mistral’s own DeepSWE comparison, Large 4 scored 62% while Beam scored 44%. That single result should not be turned into a universal ranking. Beam may offer efficiency advantages, and both systems still need wider independent evaluation.
Mistral Large 4 vs Mistral Large 3
The previous Mistral Large 3 used about 675B total parameters with 41B active. Large 4 increases that to 1.05T total and 49B active. The total parameter jump is much larger than the active-parameter increase, which illustrates the advantage Mistral is trying to get from the sparse expert design.
The product positioning has also broadened. Large 4 puts greater emphasis on multimodal understanding, technical visual analysis, long-context work, coding, documents, and enterprise application workflows.
Multimodal capabilities and visual grounding
Large 4 includes a 1.6B vision encoder. Mistral has highlighted visually demanding material such as satellite images, aerial imagery, technical drawings, and other structured visual inputs. That is potentially important because visual grounding is different from simple image captioning: the model must connect language to specific regions, objects, or details in an image.

Possible applications include geospatial analysis, industrial inspection, engineering review, document intelligence, and research workflows that combine text with images. In every specialized domain, organizations should test the model on representative internal examples rather than assuming a public benchmark will transfer perfectly.
Coding and application workflows
Mistral is clearly targeting developers. Large 4 supports function calling, structured outputs, agents, conversations, document Q&A, and batching according to the official documentation. The combination of long context and tool-oriented features makes it a candidate for repository analysis, software maintenance, data extraction, and business-process automation.
For coding tasks, benchmark scores are only part of the evaluation. Practical tests should measure whether the model follows project conventions, preserves dependencies, makes focused changes, understands large codebases, and completes tasks consistently across repeated runs.
The 1M context window may help with large repositories, but context length is not the same as context quality. A model can accept a huge prompt and still miss the most relevant detail. Retrieval quality, instruction following, and consistency across long tasks remain important.
Can you run Mistral Large 4 locally?
Not yet as a downloadable model. The public preview is currently accessed through Mistral’s API. The weights are expected later in October. Reuters reported that the full public release is planned for October 27.
Even after the weights arrive, this will not be a typical desktop model. A 1.05T-parameter network has a very large storage and memory footprint. Sparse activation reduces the amount of compute used for each token, but the broader expert set still has to be stored and made available to the serving system.
That makes Large 4’s open-weight release most immediately relevant to inference providers, research groups, larger enterprises, and organizations with multi-GPU or multi-node infrastructure. Smaller teams may find the API easier and cheaper than self-hosting.
Why 49B active parameters is an important number
The 49B-active figure explains why Large 4 can combine very high total capacity with manageable inference. In a sparse expert model, a routing system sends each token to a subset of experts. Only those selected experts perform the main work for that token.
However, “49B active” does not mean Large 4 is simply a 49B model in storage. The full expert weights still matter for hosting. Real deployment efficiency will depend on how experts are distributed across hardware, how much communication is needed between devices, quantization, batching, cache behavior, and the serving framework.
Training infrastructure and European positioning
Reporting around the launch says Large 4 was trained from scratch over roughly two months using around 4,000 NVIDIA Grace Blackwell GPUs in European data centers. Mistral is using that detail to reinforce its positioning as a European lab capable of training a frontier-scale model rather than relying only on outside systems.
Hardware count alone does not prove efficiency. Training data, utilization, networking, numerical precision, optimization methods, and post-training all affect the final result. Still, the infrastructure story is part of why this release matters strategically as well as technically.
What to check when the open weights arrive
The late-October weight release will answer several questions that the API preview cannot. The most important items to inspect are:
- License terms: commercial use, redistribution, modification, and any special conditions.
- Weight formats: supported precision levels and sharding.
- Official inference recipes: which serving frameworks are recommended.
- Quantized versions: whether practical lower-precision builds are provided.
- Fine-tuning options: support for adapters or other customization methods.
- Model card: known limitations and evaluation details.
- Reproducibility: whether third parties can reproduce the most important benchmark claims.
Who should consider Mistral Large 4?
Developers building long-context applications
The combination of a large context window, function calling, structured outputs and multimodal input is relevant to complex applications that need to work with large collections of source material.
Teams working with large document collections
Document Q&A and long context can support technical manuals, internal knowledge bases, research collections and structured information extraction.
Organizations that want deployment choice
The promised open weights could let organizations choose between the Mistral API, third-party hosted inference, or their own infrastructure. That flexibility is a major part of the model’s appeal.
Multimodal research and visual analysis teams
Large 4’s vision encoder and Mistral’s visual-grounding claims make it worth watching for workloads that combine complex images with text instructions.
What is still unknown?
Despite the strong launch, several questions remain unresolved. Large 4 is still in preview, independent benchmark coverage is limited, the final downloadable package is not available yet, and production economics at scale have not been widely tested.
- Will the final checkpoint differ materially from the public preview?
- What exact license will ship with the weights?
- How efficiently will third-party providers serve the model?
- How well will the full 1M context perform on difficult retrieval tasks?
- Which quantizations will preserve the best quality?
- How will Large 4 rank once broader independent evaluations are available?
How to access Mistral Large 4 now
The public preview is available through Mistral’s API under the model ID mistral-large-4. The official documentation is the best source for current endpoint support, pricing, context limits and feature availability.
For production evaluation, start with a representative test set instead of moving all traffic immediately. Compare Large 4 with the model you already use on the same prompts and measure quality, latency, consistency and total cost per completed task.
Sources
- Mistral AI Docs — Mistral Large 4: official specifications, features, context and pricing.
- Mistral AI launch page — Mistral Large 4: official announcement.
- Vals AI — Mistral Large 4: early independent evaluation data.
- Reuters coverage via Euronext: launch reporting and October 27 weight-release timing.
Frequently asked questions
When was Mistral Large 4 released?
Mistral Large 4 entered public preview on October 6, 2026.
What is Le Chonk?
Le Chonk is Mistral’s nickname for Mistral Large 4, its new trillion-parameter flagship model.
How many parameters does Mistral Large 4 have?
Mistral lists 1.05 trillion total parameters, 49 billion active parameters, and a 1.6B vision encoder.
What is the Mistral Large 4 context window?
The official Mistral documentation lists up to 1 million tokens. Individual evaluation platforms or providers may use lower configured limits.
How much does Mistral Large 4 cost?
At launch, Mistral’s page shows preview pricing of $0.68 per million input tokens, $0.07 per million cached input tokens, and $2.09 per million output tokens.
Is Mistral Large 4 open source?
Mistral describes it as open-weight. The final downloadable weights and license are expected later in October, so the exact release terms should be reviewed when they arrive.
Can I download Mistral Large 4 now?
Not yet. The API preview is live. Mistral says the weights will be released later in October; Reuters reported October 27.
Can Mistral Large 4 run on a consumer GPU?
The full model is far too large for a typical single consumer GPU. Practical self-hosting is expected to require substantial multi-GPU infrastructure or specialized quantized deployments.
Is Mistral Large 4 better than every closed frontier model?
No universal conclusion is justified yet. Mistral is making strong claims in selected workloads, but broader independent testing is still needed.
Bottom line
Mistral Large 4 is a notable open-weight AI launch because it combines trillion-parameter scale, sparse 49B activation, multimodal input, an official 1M-token context window, competitive preview pricing and a promised downloadable release within weeks. The decisive question is not whether the specifications look impressive on launch day. It is whether the final weights reproduce the early results, run efficiently in real deployments, and give organizations a useful alternative to closed frontier APIs.
For now, the API preview is live, the documentation is public, and independent evaluation has begun. The late-October weight release will be the next major milestone.