Brown University's Market-1T Stress-Tests AI Stock Models Across 18 Years

A one trillion observation dataset of U.S. equities and a benchmark of 18 encoder strategies set up finance for JEPA-style world models.

·
·
·
Brown University's Market-1T Stress-Tests AI Stock Models Across 18 YearsPRO
Read2 min
TypePaper
TopicData · Llms
SubtopicDatasets
  • New paper Towards Financial World Modeling benchmarks 18 encoders across 18 years of U.S. equities.
  • Market-1T offers nearly one trillion one-second observations from 2008 to 2025, publicly licensed.
  • ViT-Tiny beats ViT-Base at 10x the FLOPs on every task tested.
  • Month of evaluation explains 97 to 99.9 percent of task variance, exposing single-period benchmarks as noise.
  • Forecasting skill and latent organization are nearly uncorrelated (Spearman -0.19) across encoders.
  • LeJEPA with time warping emerges as one of the strongest self-supervised recipes for markets.

Market-1T tests financial world models across 18 years of US equities

Researchers at Brown University and Chicago Booth have introduced a large-scale dataset and benchmark for testing whether Joint Embedding Predictive Architectures can learn durable representations of financial markets. JEPAs train an encoder to predict the representation of hidden or future observations, an approach suited to markets where prices are noisy, regimes change, and the full market state remains unobservable.

The NeurIPS paper, Towards Financial World Modeling, contributes Market-1T, a regime-aware evaluation protocol, and a comparison of 18 encoder-training strategies plus a randomly initialized Vision Transformer baseline. Its central finding is practical: the evaluation month explains far more performance variation than the random seed, while forecasting accuracy reveals little about how an encoder organizes market structure.

Regime shifts expose fragile benchmarks

Financial representation-learning studies often evaluate models on a narrow dataset or a single period. A model tuned to the market conditions of 2019 can behave differently during the volatility shock of 2020, even when its architecture and downstream task remain unchanged. To measure that instability, the authors evaluate models across 32 randomly selected months spanning different market regimes.

Rank information coefficient across market regimes from 2008 to 2025
Predictive performance changes substantially across evaluation months and market regimes.

A useful financial world model must encode several interacting signals, including market-wide conditions, asset-level expected returns, liquidity, volatility, and relationships among securities. The benchmark therefore measures both forecasting utility and the structure of the learned embedding space. This broader evaluation supports future work on planning and decision-making systems that reuse one encoder across multiple tasks.

Inside the trillion-observation dataset

Market-1T contains one-second quote and trade aggregates for all US-traded equities from 2008 through 2025. Its long history and fine time resolution allow researchers to retrain and evaluate encoders across heterogeneous conditions, including the global financial crisis, the pandemic shock, low-volatility periods, and later rate-driven markets.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads