self-hosted/ai
§01·guide · video

MiniMax H3 on 16 GB: a 1080p clip in 15 minutes instead of 74

videointermediate16GB+ VRAMOct 1, 2026

A guide for MiniMax H3 (Hailuo 3), written for the RTX 4060 Ti 16GB, RTX 4070 Ti Super and 5 more.

models
tools

Two changes do most of the work

The official ComfyUI template runs MiniMax H3 at 20 steps. On our RTX 5060 Ti 16 GB, a 5-second clip at 1920×1088 took 74 minutes that way. Two changes brought it down to 15:

  • Turn on the template's turbo switch. 8 steps instead of 20.
  • Start ComfyUI with --use-ck-attention. One launch flag, nothing to install.

The same settings at 1344×768 take 5 minutes 20 seconds.

Three more settings are optional. The int8 video VAE saves a minute and we could not see a difference. Sparse attention saves four and a half minutes and the difference is hard to find. A 4-step fused model saves seven minutes and changes the look.

AI-generated with MiniMax H35.2 s · 1920×1088Report this video

1920×1088, 8 steps, Comfy Kitchen attention: 902 s on an RTX 5060 Ti 16 GB.

Prompt

The scene opens exactly on <Picture 1>. Live-action, a bright classroom with a dark green chalkboard. The teacher gives the class a short smile, turns to the empty chalkboard beside her and, with the white chalk in her right hand, writes the single word "MINIMAX" in large clear capital letters, one letter after another from left to right. She steps aside, turns back to the class and points at the word with the white chalk. The camera is still, eye level. Audio: chalk tapping and scratching on the board, quiet classroom ambience. One continuous shot, no cuts. Live-action texture. No subtitles, logos or watermarks; the only text on screen is the word MINIMAX on the board.

AI-generated with MiniMax H35.2 s · 1920×1088Report this video

The same settings: 903 s.

Prompt

The scene opens exactly on <Picture 1>. Live-action, an indoor tennis hall with a blue court, seen from the side. The right-handed woman completes her serve: the yellow ball rises from her left hand, she arches back, drives up and strikes the ball at full reach with the racket in her right hand, follows through across her body, lands forward past the baseline, then settles into a ready stance, eyes following the ball across the net. Her ponytail swings with the motion. The camera is still, side view, full body in frame. Audio: a sharp racket strike, sneakers squeaking on the court, the echo of a big hall. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

AI-generated with MiniMax H35.2 s · 1920×1088Report this video

The same settings: 907 s.

Prompt

The scene opens exactly on <Picture 1>. Live-action, a modern lecture room with a white whiteboard. The teacher gives the class a short smile, turns to the empty whiteboard beside her and, with the black marker in her right hand, writes the single word "MINIMAX" in large clear capital letters, one letter after another from left to right. She steps aside, turns back to the class and points at the word with the black marker. The camera is still, eye level. Audio: a marker squeaking on the whiteboard, quiet classroom ambience. One continuous shot, no cuts. Live-action texture. No subtitles, logos or watermarks; the only text on screen is the word MINIMAX on the board.

The settings we use

SettingValue
Launch flag--use-ck-attention
Templatethe official MiniMax H3 image-to-video template
Turbo switchon, 8 steps (it loads minimax_h3_fl2v_turbo_8step_v1.0, which the template lists)
Canvas1920×1088: set the template's megapixels to 1.98 at 16:9
Length124 frames, 5.17 s
Video VAEfp16 or int8 (see below)
Modelminimax_h3_fl2va_pruned_int8_convrot, as in the install recipe

After the launch, the ComfyUI console must say Using Comfy Kitchen attention. If it says Using pytorch attention, the flag did not reach ComfyUI. The flag exists in ComfyUI 0.37.0 (cli_args.py); it is off by default.

What each change buys

One shot (a woman crossing a wet street at night), one seed, 5.17 s at 1920×1088:

SettingsWall timeChange
Template, 20 steps4,435 s (74 min)
+ turbo switch, 8 steps1,842 s (31 min)−58%
+ --use-ck-attention908 s (15 min)−51%

Wall time is the whole job: text encoding, sampling, VAE decode and saving the file.

  • The flag matters more as the canvas grows. At 1344×768 it took the 8-step clip from 560 s to 319 s (−43%). At 1920×1088 the same flag gave −51%.
  • It does not replace turbo. With the flag and 20 steps, the 1080p clip still took 2,031 s (34 min).
  • Memory did not change. The whole-card peak stayed between 15,514 and 15,601 MiB in all of these runs, on a desktop with a display attached.

Three optional settings

Each was measured on top of the 15-minute settings, on three shots (the street, the back seat of a taxi, a mirror), one seed per shot. The 15-minute clips of the same three shots took 908–912 s.

SettingWall timeChangePicture
int8 video VAE835–840 s−8%the same clip
Sparse attention645–648 s−29%the same scene, small differences
Fused 4-step model507–510 s−44%a different clip, a different look

The int8 video VAE: take it

Swap minimax_h3_video_vae_fp16 for minimax_h3_video_vae_int8_convrot in the VAE loader. Decode drops from about 105 s to 44 s.

  • It is the same clip. The sampler does the same work, so the audio is identical sample for sample, and the frames match at SSIM 0.989–0.991 with no frame below 0.988.
  • We could not see a difference, watching the pairs side by side and in ×3 crops of the eyes.
  • Newer templates already use it. The template package switched to the int8 VAE in v0.11.69, which ComfyUI ships from v0.37.2. If your template loads it, leave it.
AI-generated with MiniMax H35.2 s · 1920×1088Report this video

Left: fp16 video VAE. Right: int8 video VAE. Same seed and the same sampling; only the decode differs. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size.

Prompt

The scene opens exactly on <Picture 1>. Live-action, night, the back seat of a moving taxi. City neon lights slide across the woman's face and the window as the car drives. Her phone vibrates in her hand; she glances down at the screen, a small smile appears, she locks the phone and turns her head to look out of the window at the passing lights. The sequins on her dress catch the moving light. The camera stays still inside the car. Audio: low engine hum, rain on the roof, one short phone vibration, the muffled city outside. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

Sparse attention: worth trying

Add ComfyUI's own Model Sparse Attention node (BlockSparseAttention) on the model wire between the turbo switch and the sampler, with its defaults. Sampling went from 97 s to 60 s per step. The first steps still run dense and the decode takes as long as before, so the whole clip gains 29%, not 38%.

  • The scene stays the same: the same framing and movement, SSIM 0.84–0.90 against the clip without it.
  • Fine skin texture gets a little smoother in close crops. Watching the clips at normal size, we could hardly tell them apart. Three shots is a small sample.
  • One known problem is not ours to confirm. A user with an RTX 5090 reported ripples with sparse attention and larryvrh's turbo LoRA (#16382). With the template's own turbo LoRA we saw none.
  • We did not measure it together with the int8 VAE.
AI-generated with MiniMax H35.2 s · 1920×1088Report this video

Left: the 15-minute settings. Right: the same with sparse attention. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size. Sound: the left clip.

Prompt

The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

The fused 4-step model: a different picture

MATLOWAI's fused file replaces the model itself: 20.98 GB, with the 8-step turbo LoRA and a style LoRA already folded in. Load it in place of the pruned int8 model, turn the turbo switch off, set 4 steps.

  • The whole gain is the step count. Each step took 94.5 s, about what it took before (97 s).
  • The look changes: glossier skin, brighter light, a different background. The style LoRA is baked in and cannot be turned off.
  • Use it for drafts, or if you like the look.
AI-generated with MiniMax H35.2 s · 1920×1088Report this video

Left: the 15-minute settings, 8 steps. Right: MATLOWAI's fused model, 4 steps. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size. Sound: the left clip.

Prompt

The scene opens exactly on <Picture 1>. Live-action, night, the back seat of a moving taxi. City neon lights slide across the woman's face and the window as the car drives. Her phone vibrates in her hand; she glances down at the screen, a small smile appears, she locks the phone and turns her head to look out of the window at the passing lights. The sequins on her dress catch the moving light. The camera stays still inside the car. Audio: low engine hum, rain on the roof, one short phone vibration, the muffled city outside. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

1344×768 or 1920×1088

At megapixels 0.975 the template renders 1344×768, and the clip takes 317–319 s instead of 908–912 s.

  • 1920×1088 looked a little better to us: more texture in skin and hair, less shine. The 768p clips are softer but hold up.
  • MiniMax describes the open model as a 768p model. The model card says H3-Base produces "results at 768p resolution". 1920×1088 is outside what the card promises. ComfyUI 0.37.0 rendered it, and none of our 1080p clips came out broken.
  • Make the start image at the canvas size. A 1920×1088 start image gave a slightly better clip than a 1344×768 one stretched to the same canvas.
AI-generated with MiniMax H35.2 s · 1920×1088Report this video

Left: rendered at 1344×768 and scaled up. Right: rendered at 1920×1088. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size. Sound: the right clip.

Prompt

The scene opens exactly on <Picture 1>. Live-action, night, the back seat of a moving taxi. City neon lights slide across the woman's face and the window as the car drives. Her phone vibrates in her hand; she glances down at the screen, a small smile appears, she locks the phone and turns her head to look out of the window at the passing lights. The sequins on her dress catch the moving light. The camera stays still inside the car. Audio: low engine hum, rain on the roof, one short phone vibration, the muffled city outside. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

Ten seconds

Set the length to 10 s (243 frames).

CanvasWall timeShared GPU memory
1344×768811 s (13.5 min)no change
1920×10885,091 s (85 min)+3.3 GB
  • At 1344×768 ten seconds is practical. The face and the scene held to the last frame.
  • At 1920×1088 it does not fit. The card spilled into shared memory and each step took 10–11 minutes. The clip came out fine, but it is not something to queue twice.
AI-generated with MiniMax H310.1 s · 1344×768Report this video

Ten seconds at 1344×768: 811 s (13.5 min).

Prompt

The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

AI-generated with MiniMax H310.1 s · 1920×1088Report this video

Ten seconds at 1920×1088: 5,091 s (85 min).

Prompt

The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

What these settings do not change

  • Voices. Lip-sync was fine in our clips, but the voice is generic. For dialogue, expect to dub it with a separate voice model.
  • The LoRA errors in the log. If you load a LoRA built for the full model, see the shape '[96768, 8]' is invalid page.

About these numbers

All runs are ours, on one RTX 5060 Ti 16 GB in a Windows desktop with a display attached, ComfyUI 0.37.0 portable, one seed per configuration. ComfyUI was also started with --disable-pinned-memory, which our measuring needs and yours does not. Files were saved at crf 16. Memory figures are whole-card readings and include the desktop's own use.

Licence. MiniMax H3 is released under the MiniMax H3 Community License, which does not cover use in the European Union, the United Kingdom, South Korea or the United States. The files named here carry their own terms. The clips on this page were made under a separate written permission from MiniMax. It covers this site only; your own use is still governed by the Community License and its territory limits. This is a summary, not legal advice.