MiniMax-H3 generates synchronized audio-video clips from text or image prompts, producing output with coherent ambient sound and motion together. It supports multiple input modalities including text-to-video, image-to-video, and video-to-video transformation pipelines. The model ships on Diffusers and uses safetensors checkpoints, making it straightforward to integrate into ComfyUI or custom generation workflows.
5,263,381 ↓ · 4,669 ♡
A mixed-precision quantization of MiniMax-H3 using NVFP4, INT4, and INT8 with convolutional rotation (Convrot), targeting NVIDIA hardware for faster video generation with reduced VRAM. This variant is designed for ComfyUI workflows where full-precision weights are not feasible. The quantization is applied to the base MiniMaxAI/MiniMax-H3 weights.
791,069 ↓ · 215 ♡
MiniMax-H3-GGUF is an open-source image-text-to-video model available on HuggingFace. Details are sourced from the public model registry.
401,305 ↓ · 235 ♡