Mistral's Leanstral 1.5 Solves 587 Putnam Problems at $4 Each
Mistral's Leanstral 1.5 solves 587/672 Putnam problems at ~$4 each, finds 5 unreported open-source bugs, and scales smoothly to 4M tokens
PRO- Leanstral 1.5 is a free, Apache-2.0 Lean 4 formal-reasoning model with 119B total / 6.5B active parameters, available via Hugging Face and a free API.
- It solves 587/672 PutnamBench problems at ~$4 per problem, vs. ~$300 for Seed-Prover 1.5 at comparable performance.
- It achieves 100% on miniF2F and sets new state-of-the-art on FATE-H (87%) and FATE-X (34%) graduate algebra benchmarks.
- An automated Rust-to-Lean pipeline found 5 previously unreported bugs across 57 open-source repositories.
- Test-time scaling is unusually strong: performance climbs smoothly from 44 PutnamBench problems at 50k tokens to 587 at 4M tokens.
- Training used three stages (mid-training, SFT, RL with CISPO) across two environments: a multiturn proof loop and a full code agent filesystem environment.
Leanstral 1.5 is Mistral's latest formal-reasoning model, built specifically for Lean 4, the proof assistant used to mechanically verify everything from graduate-level algebra to software correctness. The model is free, Apache-2.0 licensed, and available right now via a free API endpoint and on Hugging Face. If you have never touched formal verification before, this is the cheapest on-ramp that has ever existed.
What Lean 4 actually is, and why it's hard
In mathematical proofs and software verification, small mistakes can lead to major problems, so formal proof support systems like Lean 4 are used to demonstrate correctness in a way that computers can verify. The catch is that expressions that humans commonly use, such as "obviously true" or "can be proven in the same way," are not understood by computers. Every step needs to be translated into a form Lean 4 can check, which is a time-consuming process, and why dedicated AI models are needed to assist.
The field has been heating up. Lean 4 has been quietly turning into the venue where frontier labs argue about reasoning, with models from DeepSeek, ByteDance (Seed-Prover), and now Mistral all competing on the same benchmarks. The difference with Leanstral 1.5 is the cost-to-performance ratio.
The numbers that matter
Leanstral 1.5 saturates miniF2F completely, reaching 100% on both the validation and test sets. miniF2F covers problems from elementary math up to IMO-level challenges. Beyond that, the headline result is PutnamBench: Pass@8 on PutnamBench climbs from 44 problems solved at 50k tokens to 244 at 200k, 493 at 1M, and 587 at 4M.
Here is how Leanstral 1.5 stacks up across benchmarks:
| Benchmark | Leanstral 1.5 | Notes |
|---|---|---|
| miniF2F (val + test) | 100% | Fully saturated |
| PutnamBench | 587 / 672 | ~$4 per problem |
| FATE-H (grad. algebra) | 87% | New state-of-the-art |
| FATE-X (PhD algebra) | 34% | New state-of-the-art |
| FLTEval pass@1 | 28.9 | Up from 21.9 |
| FLTEval pass@8 |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.