FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
A new open-source inference engine runs MiniMax H3 video generation locally on 8GB GPUs by streaming weights and adapting kernels to your hardware.
- FlashML-org released FreeVideo, a local inference engine for MiniMax H3 video generation.
- Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.
- Built on OpenVDN's 8-step VDN-H3 model with Video DeltaNet hybrid attention.
- Ships as a ComfyUI plugin with Windows one-click launcher and Linux CLI support.
- Supports community LoRAs, two-pass sampling, batch generation, and reusable adapter caches.
- Apache 2.0 licensed, 546 GitHub stars, actively patched for low-VRAM edge cases.
FreeVideo runs MiniMax H3 on 8 GB GPUs
FlashML-org has published FreeVideo, an Apache 2.0 local inference engine for running an eight-step derivative of MiniMax H3 on consumer hardware. The project targets systems with 8 GB of VRAM and 16 GB of system RAM by distributing model layers across GPU memory, system memory, and disk.
The lower memory requirement brings heavier disk traffic and longer generation times. Open video models commonly require a 24 GB GPU or a hosted API, while FreeVideo makes the workload accessible to smaller cards through layer offloading, streaming, reduced precision, and adaptive attention kernels.
| Area | Details |
|---|---|
| Model | OpenVDN’s eight-step VDN-H3 checkpoint, based on MiniMax H3 |
| Claimed minimum | 8 GB of VRAM and 16 GB of system RAM |
| Interfaces | ComfyUI plugin, Windows launcher, and Linux command-line support |
| Local output | About one megapixel at 24 fps, with clips lasting 5 to 15 seconds |
| Storage behavior | Streams offloaded weights from disk, making SSD performance a major factor |
| License | Apache 2.0 for the FreeVideo code |
One model, several media types
MiniMax H3 is an open-weights, general-purpose multimodal video model. It accepts text, images, video, and audio in one context, then generates video and native stereo audio through the same pipeline. Voices, sound effects, and music are produced alongside the visuals rather than added during a separate post-production step.
Tagged image, video, and audio references give the model guidance on characters, style, motion, and sound. MiniMax describes H3 as 2K-capable, while the open VDN-H3 checkpoint used by FreeVideo outputs with a 768-pixel short edge, roughly one megapixel depending on aspect ratio. That distinction matters when comparing hosted H3 results with local FreeVideo output.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.