Brown University's Market-1T Stress-Tests AI Stock Models Across 18 Years
A one trillion observation dataset of U.S. equities and a benchmark of 18 encoder strategies set up finance for JEPA-style world models.
- New paper Towards Financial World Modeling benchmarks 18 encoders across 18 years of U.S. equities.
- Market-1T offers nearly one trillion one-second observations from 2008 to 2025, publicly licensed.
- ViT-Tiny beats ViT-Base at 10x the FLOPs on every task tested.
- Month of evaluation explains 97 to 99.9 percent of task variance, exposing single-period benchmarks as noise.
- Forecasting skill and latent organization are nearly uncorrelated (Spearman -0.19) across encoders.
- LeJEPA with time warping emerges as one of the strongest self-supervised recipes for markets.
Market-1T tests financial world models across 18 years of US equities
Researchers at Brown University and Chicago Booth have introduced a large-scale dataset and benchmark for testing whether Joint Embedding Predictive Architectures can learn durable representations of financial markets. JEPAs train an encoder to predict the representation of hidden or future observations, an approach suited to markets where prices are noisy, regimes change, and the full market state remains unobservable.
The NeurIPS paper, Towards Financial World Modeling, contributes Market-1T, a regime-aware evaluation protocol, and a comparison of 18 encoder-training strategies plus a randomly initialized Vision Transformer baseline. Its central finding is practical: the evaluation month explains far more performance variation than the random seed, while forecasting accuracy reveals little about how an encoder organizes market structure.
Regime shifts expose fragile benchmarks
Financial representation-learning studies often evaluate models on a narrow dataset or a single period. A model tuned to the market conditions of 2019 can behave differently during the volatility shock of 2020, even when its architecture and downstream task remain unchanged. To measure that instability, the authors evaluate models across 32 randomly selected months spanning different market regimes.
A useful financial world model must encode several interacting signals, including market-wide conditions, asset-level expected returns, liquidity, volatility, and relationships among securities. The benchmark therefore measures both forecasting utility and the structure of the learned embedding space. This broader evaluation supports future work on planning and decision-making systems that reuse one encoder across multiple tasks.
Inside the trillion-observation dataset
Market-1T contains one-second quote and trade aggregates for all US-traded equities from 2008 through 2025. Its long history and fine time resolution allow researchers to retrain and evaluate encoders across heterogeneous conditions, including the global financial crisis, the pandemic shock, low-volatility periods, and later rate-driven markets.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.