TabFM Beats AutoML on 51 Datasets Without a Single Training Update

A 400M-parameter transformer trained only on synthetic tables predicts classification and regression in one forward pass, topping the TabArena leaderboard without any tuning.

·
·
TabFM Beats AutoML on 51 Datasets Without a Single Training UpdatePRO
Read2 min
TypePaper
  • TabFM is a 400M-parameter transformer that predicts on unseen tables in a single forward pass with frozen weights.
  • Trained entirely on synthetic tables from structural causal models, no real-world pretraining data required.
  • Ranks first among zero-shot tabular foundation models on all 51 TabArena datasets.
  • Beats tuned AutoML pipelines despite doing zero task-specific optimization.
  • TabFM-Auto adds an LLM agent that writes per-dataset feature engineering code, reaching 2013 Elo overall.
  • Technical report and project page now available.

TabFM brings frozen, in-context inference to tabular data

Researchers behind TabFM have published the TabFM technical report, which describes a 400-million-parameter transformer for supervised learning on tables. At inference, the model receives labeled rows and rows requiring predictions in the same context, then produces predictions with frozen weights. The base workflow requires no gradient updates or per-dataset hyperparameter search.

The report also introduces TabFM+, which adds ensembling and calibration, and TabFM-Auto, an LLM agent that writes dataset-specific feature-engineering code. The two extensions occupy the top two positions on the TabArena benchmark while keeping the TabFM backbone frozen.

Frozen weights, three operating modes

Method Dataset adaptation TabFM updates Additional cost
TabFM Uses labeled rows as in-context examples None One forward pass per context or batch
TabFM+ Expands feature views, ensembles predictions, and calibrates outputs None Multiple inference passes and calibration
TabFM-Auto Uses an LLM to generate dataset-specific Python transforms None Code generation, execution, and pipeline validation

Why every CSV resets the clock

Conventional tabular ML starts a new fitting process for each dataset. Teams choose preprocessing rules, train tree ensembles, run hyperparameter searches, compare validation scores, and package the winning pipeline. AutoML automates much of that work, but it still spends compute fitting and evaluating models for every task.

TabFM follows the line established by tabular foundation models such as TabPFN and TabICL. TabPFN pretrained a transformer on synthetic tables to approximate Bayesian inference over a learned tabular prior. TabICL extended in-context learning to larger datasets with separate column and row attention. TabFM combines those ideas with a larger synthetic curriculum, a dedicated table encoder, and a 24-layer in-context predictor.

How TabFM reads a table

TabFM synthetic pretraining pipeline based on structural causal models
TabFM learns from synthetic tables generated by structural causal models.

Tables present two problems for standard transformers. Row order usually carries no meaning, while columns can contain numerical and categorical values with unrelated scales. TabFM handles those constraints through a staged encoder:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads