IBM's Tiny TSPulse Beats Models 100x Bigger Across 75 Datasets
IBM's 1M-parameter TSPulse foundation model handles classification, anomaly detection, imputation, and search with GPU-free inference and open weights.
- IBM released TSPulse-R1, a 1.08M-parameter time-series foundation model under Apache-2.0.
- Supports classification, anomaly detection, imputation, and similarity search with GPU-free CPU inference.
- Reports +20% on TSB-AD anomaly detection, +50% on imputation, +25% on similarity search versus larger baselines.
- Uses dual-space masked reconstruction over time and frequency domains with disentangled temporal, spectral, and semantic embeddings.
- Ships three specialized variants selected via the Hugging Face
revisionargument; paper accepted at ICLR 2026. - Does not do forecasting; context length 512, with anomaly detection needing 1.5-2K points.
IBM’s 1.08M-parameter TSPulse handles four time-series tasks on CPUs
IBM has released Granite TSPulse R1, a pre-trained time-series model for classification, anomaly detection, imputation, and similarity search. The model contains 1.08 million parameters, supports CPU inference, and ships with Apache 2.0-licensed weights. At the time of writing, its Hugging Face page showed more than 860,000 downloads.
IBM says the accompanying paper was accepted at ICLR 2026. Across more than 75 datasets, the paper reports stronger results than models with 10 to 100 times as many parameters.
| Task | Reported improvement | Benchmark scope |
|---|---|---|
| Anomaly detection | 20% | TSB-AD leaderboard |
| Similarity search | 25% | Retrieval benchmarks |
| Imputation | 50% | Missing-value benchmarks |
| Multivariate classification | 5% to 16% | Multiple classification datasets |
Those percentages summarize results across different metrics, baselines, and datasets. The paper’s per-dataset tables provide the necessary detail for comparing TSPulse with a specific production workload.
The point of 1.08 million parameters
A time-series foundation model learns reusable signal representations during pre-training, then applies them to new datasets with limited or no task-specific training. Many recent models use architectures with tens or hundreds of millions of parameters. TSPulse uses a compact MLP-Mixer that alternates between mixing information across time patches and feature channels.
The parameter count reduces deployment costs. Storing 1.08 million parameters in FP32 requires roughly 4.3 MB before packaging, dependencies, and runtime memory. That footprint allows the model to run inside CPU-based monitoring services, IoT gateways, batch pipelines, and developer laptops without dedicated GPU infrastructure.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.