CrystalJev Screens Millions of Crystals 30x Cheaper Without Sacrificing Accuracy

A calibrated decision layer turns machine-learning interatomic potentials into probability engines that match full relaxation accuracy at a thirtieth of the compute.

·
·
·
CrystalJev Screens Millions of Crystals 30x Cheaper Without Sacrificing AccuracyPRO
  • CrystalJev wraps frozen interatomic potentials with a calibrated decision layer outputting probabilities, not thresholded energies.
  • Single unrelaxed forward pass decides nearly as well as full relaxation at one thirtieth of the compute cost.
  • Audit of 65 Matbench Discovery models shows 'stable' calls are hidden probabilities set by model error and base rate.
  • Conformal selection certifies shortlists at a controlled false-discovery rate with finite-sample guarantees.
  • Value-of-information rule escalates candidates to relaxation only inside an indecision band where decisions could flip.
  • Prospective test with 700 new DFT calcs overstated the true 5.8% stable fraction by at most 2.1 points.

CrystalJev turns atomistic potentials into calibrated screeners

Machine-learning interatomic potentials, or MLIPs, commonly triage millions of hypothetical crystals before researchers commit density-functional theory resources. The standard workflow relaxes each candidate, calculates its predicted energy above the convex hull, and labels structures below a fixed threshold as stable. That threshold discards information about model error, candidate prevalence, and the cost of sending uncertain cases to a slower calculation.

The CrystalJev paper adds a calibration and decision layer to frozen atomistic foundation models. A single forward pass on an unrelaxed structure produces calibrated probabilities for specified questions, while a value-of-information policy routes uncertain candidates to full relaxation or DFT. Under the paper’s cost assumptions, the fast path approaches relaxation-level screening performance at about one-thirtieth of the compute.

What the zero cutoff discards

The Matbench Discovery leaderboard evaluates universal interatomic potentials by relaxing candidate crystals and comparing their formation energies with the convex hull of known competing phases. The hull is the lowest-energy combination available at each composition. A candidate at or below zero meV per atom is treated as thermodynamically stable relative to those phases.

Thresholding that estimate hides the model’s error distribution and the population’s base rate of stable structures. Across an audit of 65 leaderboard models, the authors found large differences in the probability that a candidate predicted at the hull would prove stable. Three quantities explained most of the variation: systematic energy bias, the spread of errors near the hull, and the prevalence of stable candidates in the screened pool.

Calibration curves for 65 Matbench Discovery models, comparing true stability probability with predicted hull distance
The audit shows that the same predicted hull distance can imply different stability probabilities across models and training sets.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads