Utopai Studios' Utopai X Beats the MiniMax H3 Model It Was Built On
Utopai Studios post-trained MiniMax H3 into a video model that debuts second on Artificial Analysis' text-to-video board, trailing only Alibaba's Wan 3.0.
- Utopai X debuts at #2 on the Artificial Analysis text-to-video leaderboard, behind Wan 3.0.
- The model is post-trained on MiniMax H3 by Utopai Studios, an AI-native film and TV studio.
- It ranks first in Audio Synchronization and Physics, second in Lighting and Camera Control.
- Available only inside the PAI platform at 93 credits per second of video, no public API.
- PAI plans start at $15 per month for 1,000 credits, billed monthly.
- Utopai X beats its own MiniMax H3 base plus fal's H3 Max variant on the same board.
Utopai X outranks the MiniMax H3 model it was built on
Utopai Studios has released Utopai X, a text-to-video model with audio that uses MiniMax H3 as its base. At publication, Artificial Analysis ranks it second on the AA-Video-T2V v2.0 leaderboard, behind Alibaba’s Wan 3.0 and ahead of Dreamina Seedance 2.5.
Utopai X also leads the three H3-family entries on the board, ranking above the original MiniMax H3 and fal’s H3 Max. The result indicates that specialized post-training can move a general video model closer to the preferences of film and advertising users without replacing the underlying foundation model.
Developers can access Utopai X only through PAI, Utopai’s browser-based production platform. The company has announced no public API, and its launch materials do not disclose the post-training data, optimization method, compute budget, safety evaluation, or inference configuration.
H3, tuned for shot craft
Utopai X applies additional training to MiniMax H3, shaping the base model’s outputs around cinematic motion, camera work, lighting, materials, and synchronized sound. Post-training typically uses curated examples and preference signals to specialize an existing model after its broad initial training.
MiniMax describes H3 as a multimodal model that understands text, images, video, and audio. It can generate clips up to 15 seconds long at resolutions up to 2K, with native stereo audio. Utopai has not published separate maximum duration or resolution specifications for individual Utopai X generations.
PAI exposes the model through a structured film workflow that covers character design, shot generation, editing, and sequence assembly. The platform also tracks characters, environments, and narrative state across multiple shots, addressing continuity problems that isolated prompt-to-clip generation leaves to the user.
Access runs through PAI
Utopai charges 93 PAI credits for each second of Utopai X video. The entry subscription costs $15 per month and includes 1,000 credits, enough for about 10.8 seconds of output if every credit goes toward this model.
| Model | Access | Listed price |
|---|---|---|
| Utopai X | PAI web platform | 93 credits per second |
| MiniMax H3 | API | $4.80 per minute at 768p |
| Wan 3.0 | API | $12.00 per minute |
| Dreamina Seedance 2.5 | API | $34.12 per minute |
The entry plan gives Utopai X an implied cost of about $83.70 per minute of generated footage. That conversion applies only to the included credit allocation and excludes platform features, top-ups, unused credits, and possible differences across subscription tiers. The API prices also cover different resolutions and serving configurations, so they are useful reference points rather than direct cost comparisons.
Where Utopai X gains ground
Artificial Analysis evaluates text-to-video systems across ten technical capabilities and ten use cases. Its AA-Video-T2V v2.0 results place Utopai X near the top in sound, physical behavior, lighting, materials, and camera control.
- Audio Synchronization: first, narrowly ahead of Gemini Omni Flash 1.1 and Dreamina Seedance 2.5
- Physics: first
- Lighting and Materials: second, behind Wan 3.0
- Camera Control: second, behind Wan 3.0
- Social Media and Creator Content: first
- Frontier: first
- Architecture and Real Estate: first
- Consumer: second, behind Wan 3.0
- Marketing and Advertising: second, behind Wan 3.0
Across the detailed results, Utopai X sits closer to the category leader than MiniMax H3 in seven of ten capabilities and seven of ten use cases. Its largest advantages over the base model appear in Audio Synchronization, Camera Control, and Physics, along with the Animation and Gaming, Social Media and Creator Content, and Frontier use cases.
The benchmark defines Audio Synchronization broadly, covering diegetic effects, music, score, and on-screen or off-screen sound. Physics includes mechanics, thermal effects, material interaction, fracture, deformation, optics, cause and effect, and counterfactual behavior. Lighting and Materials combines exposure, shadows, color temperature, textures, cloth, hair, and fur.
What the ranking leaves out
AA-Video-T2V v2.0 aggregates preference evaluations into an Elo-style ranking, so positions can change as models and votes are added. The scores reflect how evaluators prefer generated outputs under the benchmark’s prompts and category definitions.
Benchmark scope excludes several production concerns: API reliability, latency, moderation behavior, licensing terms, revision controls, project-level costs, and continuity across a full scene. PAI’s multi-shot workflow therefore requires separate testing with real scripts, recurring characters, dialogue, and editorial changes.
The missing post-training details also limit reproducibility. Outside researchers cannot determine how much of Utopai X’s gain comes from training data, preference optimization, prompt processing, inference settings, or other serving-time changes.
Wan keeps the broader toolkit
Wan 3.0 retains the overall lead and supports a wider set of creative inputs. Alibaba says the model can generate up to 30 seconds of 1080p video with native audio while accepting text, images, video, audio, documents, and web pages as references.
Utopai centers its product on shot construction and narrative continuity inside a production application. Its strongest benchmark categories align with that focus: synchronized audio, camera movement, physical behavior, lighting, and materials.
The studio strategy
Utopai operates as both a production company and a software vendor. The company says it has three theatrical films and two series scheduled for release in 2027, giving its model team an internal production pipeline for testing generated footage.
In an April 15, 2026 update, Utopai said PAI’s Story Agent could render continuous three-minute sequences at 4K. It also reported $11 million in annual recurring revenue during the technology’s first 60 days, driven by commercial licenses sold to production companies.
The three-minute figure describes PAI’s sequence-level workflow, while MiniMax lists a 15-second limit for individual H3 generations. Utopai has not explained publicly how PAI assembles, extends, or renders those longer sequences, leaving the orchestration layer as a material part of the product.
Who can use it now
- Application developers: Utopai X currently offers no direct integration path. MiniMax H3 provides the closest model lineage and the lowest listed API price among the compared systems, while Wan 3.0 leads the benchmark and accepts more input types.
- Film and advertising teams: PAI provides Utopai X alongside character, shot, and continuity tools. Evaluation should cover complete scenes, revisions, dialogue, and recurring characters, since the leaderboard does not measure the full workflow.
- Model researchers: Utopai X offers evidence that domain-specific post-training can outperform its base model on human-preference evaluations. Reproducing or attributing those gains will require technical details that Utopai has yet to publish.