A Horizon Loss Beats Cross-Entropy on ImageNet Across Three Backbones
A new loss function reframes training as a planning problem, beating cross-entropy on ImageNet with a one-line change that scales with label noise.
- New paper Planning to Learn by Ian Osband introduces the horizon loss for classification.
- Exact policy gradient scores only 4% on ImageNet vs 62% for cross-entropy, despite optimizing accuracy directly.
- Cross-entropy is patient accuracy assuming infinite training; policy gradient is the zero-horizon limit.
- Horizon loss truncates at remaining training budget, interpolating between the two extremes.
- Beats cross-entropy top-1 on ImageNet with ResNet-50, ResNet-101, and ViT-S/16 at flat learning rate.
- Gain grows with label noise; cosine learning-rate decay partially closes the gap.
A Training-Horizon Loss Beats Cross-Entropy in ImageNet Tests
Cross-entropy remains the default loss for classification because it supplies strong gradients even when a model assigns the correct label almost no probability. In Planning to Learn, Ian Osband places cross-entropy and exact policy gradient at opposite ends of a family of objectives. A horizon-aware loss between those endpoints improves reported ImageNet top-1 accuracy across ResNet-50, ResNet-101, and ViT-S/16 experiments.
In multiclass classification, deterministic top-1 accuracy depends on an argmax and has no useful gradient. A policy-gradient formulation instead samples a class from the model and awards one point when it matches the label. The expected reward is the correct class probability, py. Because the label and class probabilities are known, the gradient can be computed exactly without sampling noise, exploration, or temporal credit assignment. Even so, the author’s reported experiment produced about 4% ImageNet accuracy with exact policy gradient, compared with 62% for cross-entropy.
A Clean Gradient Allocates Work Badly
For a correct-class log-odds margin s, let p = sigmoid(s). Optimizing expected classification error gives the loss 1 - p, whose gradient magnitude with respect to s is p(1 - p). An example with p = 0.9 therefore contributes about 0.09, while one with p = 10-6 contributes about 10-6. The easier example receives roughly 90,000 times more weight.
Cross-entropy produces a different allocation. Its correct-class gradient magnitude is 1 - p, which is about 0.1 for the easy example and nearly 1 for the hard one. Exact policy gradient concentrates updates where they can increase expected accuracy immediately. Cross-entropy continues investing in examples that require many updates before their predicted class changes.
| Objective | Correct-class loss | Gradient behavior |
|---|---|---|
| Exact policy gradient | 1 - p |
Downweights examples assigned very low probability |
| Cross-entropy | -log(p) |
Maintains strong pressure on low-probability examples |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.