MiniMax H3 GGUF Reference to Video ComfyUI Workflow

In the previous article, we covered how to set up the MiniMax H3 GGUF workflow in ComfyUI for text-to-video and image-to-video generation. In this guide, we will focus on the MiniMax H3 Reference-to-Video (R2V) workflow, which allows you to use reference images to guide video generation while maintaining the character appearance, style, and overall visual consistency.

If you’re thinking about purchasing a new GPU, we’d greatly appreciate it if you used our Amazon Associate links. The price you pay will be exactly the same, but Amazon provides us with a small commission for each purchase. It’s a simple way to support our site and helps us keep creating useful content for you. Recommended GPUs: RTX 5090, RTX 5080, and RTX 5070. #ad

The MiniMax H3 GGUF Reference-to-Video workflow uses the same model files introduced in the previous article. If you have already followed the previous setup guide, there is no need to download the models again. You can continue using your existing MiniMax H3 GGUF models and simply configure the additional nodes required for the reference-to-video workflow.

Reference-to-Video is especially useful for character animation, virtual influencers, cosplay videos, and cinematic storytelling, where maintaining identity and visual details across frames is critical. By providing reference images, MiniMax H3 can generate videos that better preserve the subject’s appearance while adding motion and dynamic scenes.

In this tutorial, we will walk through the ComfyUI workflow setup, explain the required nodes, and demonstrate how to create videos using MiniMax H3 GGUF with reference images.

MiniMax H3 GGUF Models

If you have followed my previous article to download all the models, you can skip this section.

Installation

  • Update your ComfyUI to the latest version if you haven’t already done so. (Run update\update_comfyui.bat for Windows).
  • Download the json file, and open it using ComfyUI.
  • Use ComfyUI Manager to install missing nodes.
  • Restart ComfyUI.

Workflow Nodes

The workflow has 2 Load Image Nodes by default, and you can add more as needed. The model supports up to 9 reference images.

Specify the size and duration for the video. The duration is in seconds.

This is a look up table to show the size settings reference.

Enter prompt here. Please see the prompting guide for reference.

Pick the GGUF model file you downloaded.

Text encoder.

The 2 VAEs. The top one is video VAE and the bottom one is audio VAE.

This is where you connect all the reference images, videos, and audios. It does not show all the connectors initially. A new connector will appear if you use up all the connectors currently shown.

Example

Input images:

Prompt:

Task:
Generate a 10-second reference-to-video travel film.

Reference Subjects:
<Subject 1> <Picture 1>
Brown-haired girl wearing a polka-dot romper dress.

<Subject 2> <Picture 2>
Blonde-haired girl wearing a white sleeveless bodysuit with a lace-up neckline paired with dusty rose high-waisted shorts tied with a bow at the waist.

Requirements:
– Preserve the exact facial identity, hairstyle, body proportions, clothing, and accessories of both reference subjects.
– Do not swap identities.
– Maintain temporal consistency throughout the entire video.
– Natural human motion.
– Cinematic travel documentary style.

Scene:
A beautiful European city during golden hour with elegant architecture, outdoor cafés, flower-lined streets, and warm evening light.

Timeline:

[0.0s–2.0s] Visual Hook
<Subject 1> walks into frame from the left foreground while <Subject 2> is already standing near a stone railing looking toward the city. <Subject 1> briefly glances toward the camera before continuing forward. <Subject 2> notices <Subject 1> and turns naturally. They begin walking together without stopping.

Camera:
Dynamic forward tracking shot with slight handheld travel-camera energy during the first second, then transitions into smooth stabilized motion.

[2.0s–4.0s]
Both subjects walk side by side along a charming European street. They occasionally look at nearby architecture and storefronts instead of looking directly at the camera. Their conversation appears relaxed through subtle head turns and natural hand gestures.

Camera:
Smooth side tracking shot.

[4.0s–6.0s]
The subjects pause beside an old stone bridge overlooking a river. <Subject 1> rests one hand on the railing while observing the scenery. <Subject 2> adjusts a loose strand of hair as a gentle breeze moves naturally through the scene.

Camera:
Slow cinematic orbit around both subjects.

[6.0s–8.0s]
They continue walking through a narrow flower-lined street. They exchange a brief glance before continuing to admire the surroundings. Pedestrians naturally pass in the background. Clothing and hair respond realistically to the breeze.

Camera:
Backward gimbal tracking while maintaining both subjects in frame.

[8.0s–10.0s]
The subjects arrive at a scenic overlook above the city. They stand beside one another, quietly taking in the sunset before slowly turning toward the camera. Their expressions remain calm and confident with only subtle natural expressions.

Camera:
Slow pull-back combined with a gentle crane upward to reveal the panoramic city skyline.

Motion:
Natural walking cadence.
Subtle eye movement.
Occasional brief eye contact between subjects.
Relaxed posture.
Realistic breathing.
Physically accurate clothing simulation.
Natural hair movement driven by a light breeze.
No exaggerated posing or repeated smiling.

Audio:
Soft lo-fi travel music with light acoustic guitar and piano.
Natural city ambience including footsteps on stone pavement, distant conversations, café ambience, birds, a nearby fountain, and gentle wind.
No dialogue.

Style:
Ultra-realistic cinematic travel film.
Luxury travel commercial aesthetic.
Natural lighting.
Premium color grading.
Physically accurate shadows and reflections.
Professional gimbal cinematography.
Editorial travel photography.

Output:

Conclusion

The MiniMax H3 GGUF Reference-to-Video workflow provides a powerful way to create consistent AI-generated videos using reference images. Compared with traditional text-to-video generation, R2V gives you more control over character identity, clothing details, and visual style, making it a valuable tool for creating character-driven videos and storytelling content.

Since this workflow uses the same MiniMax H3 GGUF models covered in the previous article, users who have already completed the initial setup can start experimenting with Reference-to-Video without downloading additional large model files. Only the workflow configuration and reference image inputs need to be added.

While AI video generation is still evolving, MiniMax H3 shows impressive potential for maintaining consistency and producing more controllable video results. With ComfyUI’s flexible workflow system, you can further customize the process, combine different inputs, and explore new creative possibilities.

In the next steps, experiment with different reference images, motion settings, and prompts to discover what works best for your own AI video projects.

References

Further Reading

MiniMax H3 GGUF in ComfyUI: T2V & I2V Guide

LTX-2.3 GGUF Image-to-Video & Text-to-Video in ComfyUI

Be the first to comment

Leave a Reply