Mistral AI's Mistral Large 4 Packs 1 Trillion Parameters Into 4 GPUs
Mistral AI unveils a 1 trillion parameter mixture-of-experts flagship, trained in France on a fraction of competitors' compute, with open weights landing late October.
- Mistral Large 4 is a 1T parameter MoE with 49B active, natively multimodal, nicknamed Le Chonk.
- Available in Mistral API preview today; open weights land October 27 in FP8 and FP4.
- Trained on just 4,000 Grace Blackwell GPUs (~10 MW) in France, 2-3x less compute than Chinese labs.
- State of the art on legal, finance, geospatial and visual grounding; weaker than rivals on agentic coding benchmarks.
- Cybersecurity is a flagship use case, with a private API for vetted defense and state partners.
- Will become the default model powering Vibe and Mistral's product suite after A/B testing.
Mistral AI has announced Mistral Large 4, code-named Le Chonk, a natively multimodal mixture-of-experts model with 1 trillion total parameters. The model is available through a monitored API preview, with quantized open weights scheduled for October 27. Pricing and checkpoint license terms remain undisclosed.
The release targets regulated, sovereign, and on-premises workloads that need control over model deployment. Mistral claims ML4 leads open-weight models developed in the United States and Europe on an aggregate set of benchmarks, with particular strength in document reasoning, visual grounding, cybersecurity, manufacturing, and finance.
One trillion parameters, 49 billion at work
ML4 routes each token through a subset of specialized components, activating 49 billion of its 1 trillion parameters. This sparse architecture reduces computation compared with a dense trillion-parameter model, although deployments must still store the full set of weights across accelerator memory.
The model was designed to reason across text and images and to support agentic workloads, where software gives a model tools and asks it to complete multistep tasks. Mistral plans FP8 and FP4 checkpoints, which use lower numerical precision to reduce memory requirements and increase inference throughput.
The benchmarks favor documents and vision
Mistral emphasizes domain-specific performance instead of uniform leadership across every evaluation. Its reported comparisons show strong results in legal, geospatial, and financial tasks, while agentic software engineering remains a weaker area.
| Domain | Reported result | Practical implication |
|---|---|---|
| Legal | Beats Kimi K3 on Harvey’s Legal Agent Benchmark and scores three times higher than GPT-6 Astra. | Promising for document review and legal-agent workflows. |
| Geospatial | Localizes objects in satellite and aerial imagery more accurately than GPT-6 Astra. | Relevant to mapping, defense, and infrastructure monitoring. Mistral names the French Ministry of the Armed Forces as a partner. |
| Finance | Scores slightly above GLM-5.3 and DeepSeek-V4.1-Flash. | Supports evaluation for financial analysis and document-heavy research. |
| Coding | Trails GPT-6 Astra, DeepSeek-V4.1-Flash, and Kimi K3 on DeepSWE. | Coding teams need workload-specific tests before replacing current models. |
Mistral supplied these comparisons, and independent reproductions are still pending. Exact prompts, model settings, scoring methods, and evaluation harnesses will determine how well the reported rankings transfer to production workloads.
Visual grounding is one of ML4’s more concrete capabilities. It allows the model to connect an answer to a specific image region or document passage, which can improve auditability in satellite analysis, technical diagrams, scanned records, and long PDFs.
Cyber access splits into two tracks
Cybersecurity provides Mistral’s clearest argument for downloadable weights. Cofounder and chief science officer Guillaume Lample says closed model providers often restrict dual-use security requests, which can prevent defensive teams from testing vulnerabilities or analyzing malicious code. A self-hosted checkpoint gives approved operators direct control over inference and safety policies.
Mistral is staging access before publishing the weights. The public API preview monitors requests and automatically blocks malicious cyber queries. A separate private API gives vetted cybersecurity partners and state organizations broader access for offensive and defensive testing.
Because the public preview is monitored, organizations handling regulated or classified material need to review Mistral’s data-retention, logging, and privacy terms before sending sensitive inputs. Once downloaded, the checkpoint can run within infrastructure controlled by the operator, subject to the final license.
The 4,000-GPU training claim
Mistral says it trained ML4 on 4,000 Nvidia Grace Blackwell GPUs using roughly 10 megawatts in its French data center. Lample estimates that the cluster used roughly one-third to one-half as much hardware as many Chinese laboratories and far less than the total accelerator fleets reported by major US companies.
Those comparisons require care because total fleet capacity and model-specific training compute measure different things. A rigorous efficiency comparison would also need token counts, GPU-hours, utilization, failed runs, data composition, and energy consumed throughout pretraining and post-training.
Mistral attributes much of the model’s performance to post-training. Its system generates hundreds of thousands of sandboxed environments for reinforcement learning across mathematics, coding, and vulnerability research. Sandboxes let the model attempt tasks, receive machine-checkable feedback, and improve without affecting production systems.
The company also reports using less distillation than competing laboratories. Distillation trains one model from the outputs of a stronger teacher model, often reducing development costs while inheriting some of the teacher’s behavior. Mistral says its reinforcement-learning methods produced unusually large metric gains, although ablation studies and independent experiments are needed to isolate the contribution of each technique.
From API preview to downloadable weights
| Stage | Access | Details |
|---|---|---|
| Launch preview | Mistral API | Monitored access with automated blocking for malicious cyber requests. |
| Private track | Vetted partners | Broader cybersecurity testing for selected companies and state organizations. |
| October 27 | Open weights | FP8 and FP4 checkpoints intended for deployment across four to eight Nvidia B200 or B300 GPUs. |
| After testing | Mistral products | Planned adoption as the default model in products including Vibe, following A/B tests. |
Mistral has not announced ML4 API pricing. Its previous flagship, Mistral Large 3, costs $0.50 per million input tokens and $1.50 per million output tokens after a 75% reduction from Large 2 pricing. Those figures provide context for Mistral’s pricing strategy but do not establish ML4’s eventual cost.
The downloadable checkpoint is expected to score above the API preview because training remains underway. Developers recording evaluations should treat the preview and October checkpoint as separate model versions, with distinct identifiers, prompts, and results.
For self-hosters, the four-GPU FP4 configuration is the central deployment claim. FP4 compresses each parameter to roughly four bits, allowing a trillion-parameter model to fit within a current-generation multi-GPU server. The sparse architecture lowers per-token computation, while quantization addresses the larger memory burden. Teams still need to measure throughput, latency, routing overhead, and quality loss on their own hardware.
A European sovereignty play
Mistral’s previous public flagship, Mistral Large 3, arrived in December 2025 with 675 billion total parameters and 41 billion active parameters. During that period, DeepSeek, Qwen, and Kimi drove much of the open-weight model competition from China.
ML4 expands Mistral’s total and active parameter counts, adds native multimodality, and concentrates on sectors where data residency and local deployment influence procurement. That positioning aligns with European legal, financial, defense, manufacturing, and critical-infrastructure customers.
A €3 billion Series D announced in September is funding Mistral’s planned expansion toward one gigawatt of computing capacity. That target equals 100 times the power cited for the ML4 training cluster, although the company has not provided a complete construction timetable. Mistral also says another major model update will arrive within two months and form the base for later releases.
Who should test it
ML4’s reported strengths and deployment options define a focused set of early evaluators:
- Regulated enterprises can test legal, financial, and document-processing workflows that require EU data residency or on-premises inference.
- Geospatial and defense teams can evaluate visual grounding against satellite imagery, maps, diagrams, and other domain-specific data.
- Security organizations can assess whether self-hosted access supports defensive research that managed APIs restrict.
- Coding teams should compare ML4 with current models on private repositories, tool use, patch quality, and end-to-end task completion because its DeepSWE result trails several competitors.
- Infrastructure teams can validate the claimed four-to-eight-GPU footprint, including memory use, batch throughput, latency, and quantization quality.
Questions the release still has to answer
Production adoption depends on details that remain absent from the launch information:
- The API price for input, output, image processing, and cached tokens.
- The checkpoint license and its rules for commercial use, modification, and redistribution.
- Context limits, supported image formats, structured output, tool calling, and API compatibility.
- Rate limits, regional availability, service-level commitments, and data-retention policies.
- A complete model card covering training data, safety evaluations, benchmark settings, and quantization effects.
The October 27 checkpoint will determine whether ML4’s API results, hardware requirements, and licensing terms support the self-hosting case Mistral is making. Until then, the preview offers a way to test model quality, while production planning remains contingent on the final release artifacts.