Mistral Drops Large 4, a 1 Trillion Parameter Agent Built for Enterprise

Mistral's new 1T-parameter mixture-of-experts model targets agentic workflows across Gmail, Slack, Salesforce and code, with open weights coming by month-end.

·
·
·
Mistral Drops Large 4, a 1 Trillion Parameter Agent Built for Enterprise
  • Mistral launched Mistral Large 4, a 1T parameter MoE with 49B active parameters, in public preview.
  • Open weights promised by end of October; API live now on Mistral Studio as mistral-large-4.
  • Scores 59.9% on AutomationBench (Gmail, Sheets, Slack, Salesforce workflows), beating DeepSeek V4 Pro and Kimi K3.
  • Pricing in preview is $0.68 input and $2.09 output per million tokens, with 1M token context.
  • Trained on 3,800 Grace Blackwell GPUs in European datacenters, with RL across tens of thousands of parallel rollouts.
  • Independent Artificial Analysis Intelligence Index puts it at 38, behind GLM-5.3 and Kimi K3.

Mistral Large 4 targets tool-using agents

Paris-based Mistral has unveiled Mistral Large 4, nicknamed “le Chonk,” and opened API access through Mistral Studio. The company plans to publish the model weights by the end of the month.

ML4 is a multimodal mixture-of-experts model that accepts text and images, follows instructions, reasons across long inputs, and calls tools. Mistral built it for agents that complete multi-step work across email, spreadsheets, chat, customer relationship management systems, and code repositories. On AutomationBench, which covers 657 workflows across applications including Gmail, Google Sheets, Slack, and Salesforce, Mistral reports a score of 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro.

One trillion parameters, 49 billion active

A router selects a small group of experts for each token, leaving the rest inactive. This sparse design lowers inference compute relative to activating all 1.05 trillion parameters, although serving the complete model still requires storing every expert.

Capability Specification
Architecture Approximately 1.05 trillion total parameters and 49 billion active parameters
Vision 1.6 billion-parameter image encoder
Context window Up to 1 million tokens
Languages More than 160
API features Structured outputs, function calling, document Q&A, batch processing, and agent tools
Model ID mistral-large-4
Availability API preview now; weights promised by month-end

Preview pricing is $0.68 per million input tokens, $0.07 per million cached input tokens, and $2.09 per million output tokens. Mistral lists eventual prices of $1.36 for input and $4.18 for output, making the preview rates roughly half the planned list price.

The one-million-token window can accommodate large repositories, contracts, or ticket histories in a single request. Teams should still test retrieval accuracy, instruction retention, latency, and cost near the limit because context capacity alone does not guarantee reliable use of every token.

Rollouts teach the tool use

Mistral says it trained ML4 from scratch on 3,800 NVIDIA Grace Blackwell GPUs in company-operated European data centers. That provenance may help with European procurement requirements, but customers still need to verify API processing regions, retention policies, contractual residency guarantees, and data-processing terms.

Post-training relies heavily on reinforcement learning in simulated environments. Agents complete full trajectories involving several actions, tool calls, and intermediate decisions; the training system then evaluates the outcome and reinforces successful behavior. These trajectories, known as rollouts, expose the model to the consequences of multi-step decisions.

Mistral reports that its current reinforcement-learning pipeline uses 3,000 GPUs to generate about 33 billion tokens per day. Filtering and masking leave roughly 16 billion trainable completion tokens. Those figures describe pipeline throughput rather than ML4’s total training corpus or overall compute budget.

Tool-heavy work drives the gains

Mistral’s strongest reported results cluster around coding agents and business automation. The company lists the following scores:

  • DeepSWE v1.1: 61.7%
  • SWE-Atlas-QnA: 59.4%
  • Terminal-Bench 4: 28.3%
  • Coding Agent Index: 49.8%
  • AutomationBench: 59.9%

Mistral says the combined Coding Agent Index places ML4 ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. It also claims the leading open-weight result on Harvey’s Legal Agent Benchmark. Benchmark suites use different tools, prompts, execution environments, and scoring rules, so the individual results should be evaluated within their respective harnesses.

Cybersecurity receives dedicated attention. ML4 scored 82% on an Artificial Analysis Cyber Index test covering vulnerability reproduction and patching, along with 93% on Cybench. Mistral reports internal use for malware analysis, vulnerability prioritization, and detection-rule generation. Partners can request a version with reduced moderation for legitimate security research and incident response.

For scientific computing, Mistral highlights data analysis, modeling, simulation, and code generation. The company reports a leading open-weight result on SciCode-Verified and demonstrates the model generating a complete Hartree-Fock simulation in one pass. Reproducibility, numerical correctness, and domain review remain necessary before using generated scientific code.

Independent tests narrow the lead

Artificial Analysis gives the preview a score of 38 on its Intelligence Index, which combines ten evaluations including Terminal-Bench 4, AutomationBench, and Humanity’s Last Exam. That result ranks ML4 64th among 225 models, behind GLM-5.3 at 45 and Kimi K3 at 44.

A separate blind evaluation by Surge AI asked professional annotators to score coding outputs from one to five without seeing model identities. Large 4 Preview placed second among five models with 3.74, while Claude Opus 5 scored 4.22. The result supports ML4’s position among open-weight coding models while showing a measurable gap from the leading closed model in that test.

These evaluations assess different model versions and workloads, and preview behavior may change before a stable release. Teams comparing models should rerun representative tasks with the same prompts, tools, time limits, and scoring criteria.

The weights create a hardware problem

Publishing the weights would give teams more control over deployment, customization, and data handling, subject to the final license. The full parameter set also creates substantial storage requirements. At 16-bit precision, 1.05 trillion parameters require about 2.1 terabytes for weights alone; 8-bit and 4-bit representations would require roughly 1.05 terabytes and 525 gigabytes before runtime overhead.

Actual serving requirements will depend on quantization support, expert parallelism, cache size, batching, and the released inference stack. Mistral has yet to publish the checkpoint format, supported quantizations, recommended GPU topology, throughput figures, or license terms. An open-weight label also does not establish commercial rights or an open-source license.

A production checklist

Teams evaluating ML4 should resolve the following points before committing to production:

  • Confirm the weight license, redistribution terms, and commercial-use rights.
  • Measure task completion, tool-call accuracy, latency, and total token cost on representative workflows.
  • Test long-context retrieval at several input sizes, including near the advertised limit.
  • Verify structured-output schemas, function-calling behavior, retries, and failure recovery.
  • Check image limits, supported formats, batch constraints, and rate limits.
  • Validate API regions, retention settings, logging controls, and data-processing agreements.
  • Estimate self-hosting hardware from the released checkpoint and serving software.
  • Clarify access requirements for the reduced-moderation security version.
  • Plan for the announced list pricing after the preview period.

Who should test it

ML4 is a relevant candidate for agents that coordinate multi-step work across enterprise applications, coding systems that need repository-scale context, and security tools affected by provider refusals. Its API provides an immediate evaluation path, while the planned weights could support private deployment and specialized fine-tuning.

Teams focused on single-turn reasoning or general chat quality should compare ML4 with GLM-5.3, Kimi K3, and leading closed models. The production decision should follow workload-specific tests covering success rate, latency, operating cost, deployment complexity, and data-governance requirements.

Trending
  • No trending articles

Comments

avatar

Next Reads