MiniMax's H3 Tops Video Editing Charts and Goes Open-Source

MiniMax H3 tops video editing benchmarks with native 2K, audio, and instruction-based editing — and open weights are coming

·
·
  • #1 Video Editing: MiniMax H3 tops Artificial Analysis leaderboard for video editing, ranks top 3 in text-to-video and image-to-video.
  • Native 2K + audio: Generates 5-15 second clips at 2560x1440 with synchronized stereo sound in a single pass.
  • Omni-reference: Accepts up to 9 images, 3 video clips, and 3 audio clips per request for consistent characters and style across shots.
  • Instruction-based editing: Edit finished clips in natural language without regenerating from scratch — a rare capability at this quality level.
  • Pricing: $0.13/sec ($7.80/min) for 2K with audio, undercutting Kling 3.0 ($20.16/min) and Seedance 2.0 ($22.45/min).
  • Open weights incoming: MiniMax plans to release weights under a community license (free for orgs under $20M revenue), which would make it the strongest open-weights video model, ahead of LTX-2.3.

MiniMax H3 just landed at the top of the Artificial Analysis video leaderboard for editing and placed top 3 in both text-to-video and image-to-video. The successor to the Hailuo family, it brings a genuinely different architecture: a single model that reads text, images, video clips, and audio, then outputs video with synchronized sound in one pass. MiniMax also plans to release the weights publicly, which would make it the strongest open-weights video model by a wide margin.

What H3 actually is

MiniMax H3, also called Hailuo 3.0, is the latest flagship video model from MiniMax, the lab behind Hailuo 02 and Hailuo 2.3. It generates videos up to 15 seconds long at 2K resolution with native stereo sound. That is a meaningful jump from Hailuo 2.3, which topped out at roughly 10 seconds at 1080p.

Beyond the resolution and length increase, H3 adds two capabilities Hailuo 2.3 lacked entirely: omni-reference (up to 9 images, 3 video clips, and 3 audio clips per generation) and instruction-based editing. Clips can also be extended to roughly 30 seconds via a separate Extend tool.

The two features that change production workflows

The resolution numbers are table stakes. The more consequential additions are the ones that address real friction in multi-shot AI video production.

  • Omni-reference: Supply up to 9 reference images, 3 video clips, and 3 audio clips in a single generation. The model uses them for style, character, motion, or voice guidance without locking them in as keyframes. This directly addresses character identity drift, one of the most persistent problems when producing multi-shot AI video.
  • Instruction-based editing: Describe a change in plain text and H3 applies it to a finished clip, adjusting characters, objects, scenes, sound, or pacing without a full regeneration. Combined with motion transfer between videos, this is what pushed H3 to the top of the editing leaderboard ahead of Runway, Kling, and ByteDance.

Where it ranks and what it costs

On the Artificial Analysis leaderboards, H3 sits at #1 in Video Editing, #2 in Text to Video, and #3 in Image to Video. The evaluation was conducted on the 2K tier. Pricing:

  • 2K tier: $0.13/second ($7.80/min)
  • 768p tier: $0.09/second ($5.40/min), listed as coming soon
  • Reference images: first five free, $0.04 each after that
  • Reference audio: free

At $7.80/min for 2K with audio, H3 undercuts Dreamina Seedance 2.0 1080p at $22.45/min and Kling 3.0 1080p at $20.16/min. Google's Gemini Omni Flash is cheaper at $6.00/min but ranks below H3 on the editing leaderboard.

Open weights: the bigger story

MiniMax said it planned to release H3's weights within days of launch, allowing teams to download, self-host, and adapt the model. For context, open-weights video has already closed significant ground on closed SaaS: LTX-2.3 from Lightricks generates synchronized audio and video in one pass and has accumulated 18M+ Hugging Face downloads, suggesting production teams are evaluating open-source options seriously. H3 would arrive with a quality ceiling well above LTX-2.3, making it the new default for teams that need to self-host.

The MiniMax Community License permits commercial use for organizations under $20M in revenue, subject to attribution requirements. That matters most for advertising, entertainment, and enterprise workflows where sending footage to a third-party API is a non-starter. Open weights mean sensitive assets stay inside the team's own environment, and the model can be fine-tuned or integrated into existing pipelines without API dependencies.

Where it fits and where to be cautious

H3 is a strong fit for:

  • Short-form content production: 15-second clips with built-in audio remove the separate voiceover and sound design step for most social and ad formats
  • Multi-shot narratives: omni-reference handles character consistency across scenes without manual compositing
  • Video editing workflows: instruction-based edits let you refine a clip in natural language rather than regenerating from scratch
  • Private deployment: once weights drop, teams can run H3 on-premise under the community license

The leaderboard rankings are based on blind human votes, which is a reliable signal for creative quality but says nothing about how H3 handles edge cases in a specific domain. Test on your own prompts before committing production pipelines.

The competitive picture

H3 competes with Kling 3.0, Google Veo 3.1, ByteDance Seedance 2.0, Alibaba Wan 2.7, and others across the 2026 frontier. What separates it is the combination of top-tier benchmark performance, a price point below most direct competitors, and an imminent open-weights release. Most frontier video models remain API-only. MiniMax is betting that handing teams the weights alongside a commercial license is the right way to win the developer ecosystem.

H3 is live now via the Hailuo AI app and the MiniMax API, with weights expected imminently. If the weights land at the quality level the leaderboard suggests, the bar for what teams can self-host just moved significantly.

Comments

avatar