MiniMax-H3 is one of the most-pulled open video models on Hugging Face — its ComfyUI build saw 19.2M downloads in the 30 days to September 2026 — and it's a genuinely different capability: native synchronized audio+video generation in one pass, not video-then-dub.

What it needs: roughly 15.5 GB for the UNet plus ~14.6 GB for its 32B-parameter text encoder — both well past what a single 12 GB consumer card can hold.

Why we're not running it: nothing's wrong with the model. It just doesn't fit the accessible, single-consumer-GPU target this project is built around. If your setup has the headroom (24 GB+, or split across GPUs), it's worth a look — we just can't reproduce it on our own reference hardware, so we're not shipping a workflow we can't stand behind.

In FlixML: not shipped. This entry exists because "what's hot" and "what you can actually run" are different questions, and this is a live example of the gap.