MiniMax's H3 Tops Video Editing Charts and Goes Open-Source
MiniMax H3 tops video editing benchmarks with native 2K, audio, and instruction-based editing — and open weights are coming
- #1 Video Editing: MiniMax H3 tops Artificial Analysis leaderboard for video editing, ranks top 3 in text-to-video and image-to-video.
- Native 2K + audio: Generates 5-15 second clips at 2560x1440 with synchronized stereo sound in a single pass.
- Omni-reference: Accepts up to 9 images, 3 video clips, and 3 audio clips per request for consistent characters and style across shots.
- Instruction-based editing: Edit finished clips in natural language without regenerating from scratch — a rare capability at this quality level.
- Pricing: $0.13/sec ($7.80/min) for 2K with audio, undercutting Kling 3.0 ($20.16/min) and Seedance 2.0 ($22.45/min).
- Open weights incoming: MiniMax plans to release weights under a community license (free for orgs under $20M revenue), which would make it the strongest open-weights video model, ahead of LTX-2.3.
MiniMax H3 just landed at the top of the Artificial Analysis video leaderboard for editing, and placed top 3 in both text-to-video and image-to-video. It is the successor to the Hailuo family, and it brings a genuinely different architecture to the table: a single model that reads text, images, video clips, and audio, and outputs video with synchronized sound in one pass. MiniMax is also planning to release the weights publicly, which would make it the strongest open-weights video model by a wide margin.
What H3 actually is
MiniMax H3, also called Hailuo 3.0, is the latest flagship video model from MiniMax, the lab behind Hailuo 02 and Hailuo 2.3. The model can generate videos of up to 15 seconds in 2K resolution with native stereo sound. That is a meaningful jump from Hailuo 2.3, which topped out at roughly 10 seconds at 1080p.
H3 upgrades Hailuo's video line across the board: native 2K vs 1080p, 5-15 second clips extendable to roughly 30 seconds via an Extend tool, plus two brand-new capabilities Hailuo 2.3 lacked: omni-reference (up to 9 images, 3 video clips, 3 audio clips) and instruction-based editing, alongside one-pass synchronized audio.
The two features that actually matter
The headline numbers are nice, but the real story is in two capabilities that most video models still cannot do well.
- Omni-reference: You can supply up to 9 reference images, up to three reference video clips, and up to three reference audio clips in one generation, and the model uses them for style, character, motion, or voice guidance without forcing them as keyframes. This is the fix for character identity drift, one of the most painful problems in multi-shot AI video production.