Mistral released Leanstral 1.5 on July 2, 2026, an Apache 2.0 licensed model built specifically for proof engineering in Lean 4. The model keeps a large footprint while activating very little of it, and it pushes formal verification benchmarks to their limits.
A Sparse, Open Model
Leanstral 1.5 has 119 billion total parameters but only 6 billion active per step, making it cheap enough to run repeatedly. It is free to download and use commercially under Apache 2.0.
Benchmark Results
- miniF2F: saturated at 100 percent on validation and test sets
- PutnamBench: 587 of 672 problems solved
- FATE-H: new state of the art at 87 percent
- FATE-X: new state of the art at 34 percent
- FLTEval: pass@1 rose from 21.9 to 28.9
On cost, Mistral reports roughly 4 dollars per problem on PutnamBench, versus an estimated 300 dollars or more for the comparable high setting of Seed-Prover 1.5.
How It Was Trained
Training proceeded through mid-training, supervised fine-tuning, and reinforcement learning with CISPO. Two reinforcement learning environments drove the gains: a multiturn prover that refines proofs against Lean compiler feedback, and a code agent that edits files, runs bash commands and queries the Lean language server across long horizons.
From Mathematics to Real Code
An AVL-tree proof ran for over 2.7 million tokens across 22 compactions to establish O(log n) time complexity. A separate bug-hunting pipeline using Aeneas and SafeVerify flagged 47 violated properties across 57 repositories, yielding 11 genuine bugs, five of them previously unreported.
What This Means
By open sourcing both the weights and FLTEval, Mistral is betting that formal methods become practical tooling rather than academic curiosities. The bug discoveries suggest verified reasoning is already useful on real codebases, not just textbook theorems.