MiniMax H3 VAE Doubles Video Resolution but Needs a Reference Image
An ambitious MiniMax H3 VAE project failed its original goal but shipped a working 2X decoder and reference detail enhancer for ComfyUI.
- MiniMax-H3-X2-Detail-VAE ships as a 2-in-1 ComfyUI tool: 2X video decode plus optional reference detail enhancement.
- Mode 1 works as a drop-in X2 VAE via the MiniMax H3 VAE Decode (fast) node.
- Mode 2 enhances RGB reference images using an early-encoder B32 feature tap from down.1.block.1.
- Full-resolution tests showed wins on all 10 holdout frames across RMSE, HP3, HP5 and gradient metrics.
- The approach fails for pure text-to-video because the required pre-bottleneck RGB features do not exist at generation time.
- Requires ComfyUI-MiniMaxH3_LatentUpscaler and is distributed under the MiniMax H3 Community License.
MiniMax H3 VAE doubles frame dimensions; added detail needs a reference
MiniMax-H3-X2-Detail-VAE packages a working 2× video VAE for MiniMax H3, an optional ComfyUI detail-enhancement node, an example workflow, and a report on an unsuccessful attempt to build a higher-information video VAE. The released checkpoint doubles decoded frame dimensions. Its additional detail path works only when the pipeline has an original RGB reference image.
A variational autoencoder, or VAE, compresses video frames into a smaller latent representation and later decodes that representation back into pixels. This checkpoint can serve as MiniMax H3’s decoder or as a source-conditioned enhancement stage, depending on where it is placed in the ComfyUI graph.
One checkpoint, two paths
Placement determines which of the checkpoint’s two operating modes ComfyUI uses:
| Mode | Purpose | Configuration |
|---|---|---|
| 2× VAE Decode | Decodes MiniMax H3 video latents at twice the frame width and height. | Load the .safetensors checkpoint with ComfyUI’s standard Load VAE node, then decode with MiniMax H3 VAE Decode (fast) from the latent upscaler extension. |
| Reference Detail Enhancement | Transfers early encoder detail from an existing RGB reference into image-to-video or reference-to-video processing. |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.