MiniMax H3 Launches with Day-0 ComfyUI Support: Open Weights, Stereo Audio, 2K Video
Original: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
Why This Matters
Open-weights multimodal video generation with native audio marks a significant capability expansion for locally-run creative AI tools.
MiniMax released its third-generation video model H3 on August 3, 2026, with open weights and immediate ComfyUI support. The omni-modal model generates up to 2K, 15-second video clips with native stereo audio from text, image, video, or audio inputs, and can run locally on an NVIDIA RTX 3060.
MiniMax H3 is the company's third-generation open-weights video generation model, succeeding Hailuo 01 and Hailuo 02 — and the first in the series to be released with open weights. ComfyUI announced day-zero native support on the same morning of launch.
H3 supports text-to-video, image-to-video, first-and-last-frame control, and reference-to-video workflows. Outputs reach up to 2K resolution and 15 seconds per clip. A key differentiator is native stereo audio: sound is generated in the same model pass as video rather than added as a post-process step.
The model accepts multimodal inputs — images, audio, and video simultaneously — and uses cross-modal understanding to resolve relationships between inputs against a text prompt. MiniMax describes this as collapsing five separate tasks into a single model.
For workflow users, motion transfer is highlighted as a practical feature: a reference video can supply camera movement or performance style while subject and visual style are sourced independently. This enables iterative shot editing within ComfyUI graphs.
The model is available to try via Comfy Cloud and can also run locally on consumer hardware as modest as an RTX 3060, lowering the barrier for independent creators and developers.