Artificial Analysis Exposes Hidden Family Ties Between Top Video AI Models
Artificial Analysis now flags video models that were post-trained from another base model, with a toggle to filter them out of the leaderboard.
- Artificial Analysis video leaderboards now label models post-trained from another base model.
- Derived entries show the base model's logo next to their name for quick identification.
- New Hide derived models toggle filters the board down to base models only.
- Utopai X and MiniMax H3 Max, both derived from MiniMax H3, sit in the T2V top six.
- MiniMax H3 Max leads AA-Video-I2V v1.0 with an Elo of 1195, ahead of its base.
- Full boards live at artificialanalysis.ai/video.
Artificial Analysis has added lineage labels to its video model leaderboards. Entries post-trained from another listed model now display the base model’s logo, helping developers identify related systems before comparing quality, speed, and price.
Artificial Analysis ranks text-to-video and image-to-video systems through head-to-head preference tests summarized as Elo ratings. A derived model starts from an existing model and receives additional training or tuning, often alongside provider-specific inference optimization. The update adds provenance and filtering while leaving the benchmark scores and evaluation process unchanged.
Lineage enters the leaderboard
- Derived models display the base model’s logo beside their names.
- A Hide derived models toggle limits the board to base models.
- Derived entries remain visible by default because they can differ from their parents in quality, latency, price, and availability.
At the time of the update, two of the top six entries on AA-Video-T2V v2.0, Utopai X and MiniMax H3 Max, were derived from MiniMax H3. The base MiniMax H3 ranked fourth. On AA-Video-I2V v1.0, MiniMax H3 Max led with an Elo rating of 1195, ahead of its parent model.
| Board | Model | Lineage | Standing |
|---|---|---|---|
| AA-Video-T2V v2.0 | Utopai X | Derived from MiniMax H3 | Top six |
| AA-Video-T2V v2.0 | MiniMax H3 Max | Derived from MiniMax H3 | Top six |
| AA-Video-T2V v2.0 | MiniMax H3 | Base model | No. 4 |
| AA-Video-I2V v1.0 | MiniMax H3 Max | Derived from MiniMax H3 | No. 1, Elo 1195 |
Elo measures relative preference within the tested pool, so a higher score indicates stronger head-to-head results on that board. It does not describe every production workload, and rankings can move as votes and models are added.
One family, different products
MiniMax H3 Max illustrates why derived models deserve separate entries. Fal post-trained MiniMax H3 for prompt adherence, audiovisual quality, and aesthetics, then co-optimized the model with its inference stack. Fal says the resulting endpoint can render a five-second 768p clip in under three seconds.
Fine-tuning can improve prompt adherence while reducing motion coherence, or exchange peak quality for faster generation on a particular runtime. Provider-level optimization can also affect cold starts, throughput, regional availability, quotas, and integration requirements even when two endpoints share model lineage.
Commercial terms require the same scrutiny. The text-to-video board listed MiniMax H3 Max and the 768p version of MiniMax H3 at $4.80 per generated minute, despite different Elo ratings and sample counts. Prices and availability can also vary by host.
Reading the revised board
Developers evaluating a production endpoint can use the lineage label to investigate four practical questions:
- Which base model, if any, supplied the endpoint’s starting point?
- Does the tuned version improve the application’s prompts, styles, motion, or audio requirements?
- Which provider, runtime, region, and pricing plan does deployment require?
- Were latency and cost measured at the needed resolution, duration, and concurrency?
The parent logo makes direct family comparisons easier, while the filter reveals how many base-model families sit behind the ranked endpoints. Sample counts, output examples, price units, and provider documentation still matter because an average Elo lead may not transfer to a specific workload.
Flat rankings can make several tuned endpoints from one base appear to represent several independently trained systems. The new labels preserve each deployable endpoint’s result while exposing shared lineage, giving developers a clearer view of product choice and model diversity.