Two changes do most of the work
The official ComfyUI template runs MiniMax H3 at 20 steps. On our RTX 5060 Ti 16 GB, a 5-second clip at 1920×1088 took 74 minutes that way. Two changes brought it down to 15:
- Turn on the template's turbo switch. 8 steps instead of 20.
- Start ComfyUI with
--use-ck-attention. One launch flag, nothing to install.
The same settings at 1344×768 take 5 minutes 20 seconds.
Three more settings are optional. The int8 video VAE saves a minute and we could not see a difference. Sparse attention saves four and a half minutes and the difference is hard to find. A 4-step fused model saves seven minutes and changes the look.
1920×1088, 8 steps, Comfy Kitchen attention: 902 s on an RTX 5060 Ti 16 GB.
Prompt
The scene opens exactly on <Picture 1>. Live-action, a bright classroom with a dark green chalkboard. The teacher gives the class a short smile, turns to the empty chalkboard beside her and, with the white chalk in her right hand, writes the single word "MINIMAX" in large clear capital letters, one letter after another from left to right. She steps aside, turns back to the class and points at the word with the white chalk. The camera is still, eye level. Audio: chalk tapping and scratching on the board, quiet classroom ambience. One continuous shot, no cuts. Live-action texture. No subtitles, logos or watermarks; the only text on screen is the word MINIMAX on the board.
The same settings: 903 s.
Prompt
The scene opens exactly on <Picture 1>. Live-action, an indoor tennis hall with a blue court, seen from the side. The right-handed woman completes her serve: the yellow ball rises from her left hand, she arches back, drives up and strikes the ball at full reach with the racket in her right hand, follows through across her body, lands forward past the baseline, then settles into a ready stance, eyes following the ball across the net. Her ponytail swings with the motion. The camera is still, side view, full body in frame. Audio: a sharp racket strike, sneakers squeaking on the court, the echo of a big hall. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
The same settings: 907 s.
Prompt
The scene opens exactly on <Picture 1>. Live-action, a modern lecture room with a white whiteboard. The teacher gives the class a short smile, turns to the empty whiteboard beside her and, with the black marker in her right hand, writes the single word "MINIMAX" in large clear capital letters, one letter after another from left to right. She steps aside, turns back to the class and points at the word with the black marker. The camera is still, eye level. Audio: a marker squeaking on the whiteboard, quiet classroom ambience. One continuous shot, no cuts. Live-action texture. No subtitles, logos or watermarks; the only text on screen is the word MINIMAX on the board.
The settings we use
| Setting | Value |
|---|---|
| Launch flag | --use-ck-attention |
| Template | the official MiniMax H3 image-to-video template |
| Turbo switch | on, 8 steps (it loads minimax_h3_fl2v_turbo_8step_v1.0, which the template lists) |
| Canvas | 1920×1088: set the template's megapixels to 1.98 at 16:9 |
| Length | 124 frames, 5.17 s |
| Video VAE | fp16 or int8 (see below) |
| Model | minimax_h3_fl2va_pruned_int8_convrot, as in the install recipe |
After the launch, the ComfyUI console must say Using Comfy Kitchen attention. If it says Using pytorch attention, the flag did not reach ComfyUI. The flag exists in ComfyUI 0.37.0 (cli_args.py); it is off by default.
What each change buys
One shot (a woman crossing a wet street at night), one seed, 5.17 s at 1920×1088:
| Settings | Wall time | Change |
|---|---|---|
| Template, 20 steps | 4,435 s (74 min) | |
| + turbo switch, 8 steps | 1,842 s (31 min) | −58% |
+ --use-ck-attention | 908 s (15 min) | −51% |
Wall time is the whole job: text encoding, sampling, VAE decode and saving the file.
- The flag matters more as the canvas grows. At 1344×768 it took the 8-step clip from 560 s to 319 s (−43%). At 1920×1088 the same flag gave −51%.
- It does not replace turbo. With the flag and 20 steps, the 1080p clip still took 2,031 s (34 min).
- Memory did not change. The whole-card peak stayed between 15,514 and 15,601 MiB in all of these runs, on a desktop with a display attached.
Three optional settings
Each was measured on top of the 15-minute settings, on three shots (the street, the back seat of a taxi, a mirror), one seed per shot. The 15-minute clips of the same three shots took 908–912 s.
| Setting | Wall time | Change | Picture |
|---|---|---|---|
| int8 video VAE | 835–840 s | −8% | the same clip |
| Sparse attention | 645–648 s | −29% | the same scene, small differences |
| Fused 4-step model | 507–510 s | −44% | a different clip, a different look |
The int8 video VAE: take it
Swap minimax_h3_video_vae_fp16 for minimax_h3_video_vae_int8_convrot in the VAE loader. Decode drops from about 105 s to 44 s.
- It is the same clip. The sampler does the same work, so the audio is identical sample for sample, and the frames match at SSIM 0.989–0.991 with no frame below 0.988.
- We could not see a difference, watching the pairs side by side and in ×3 crops of the eyes.
- Newer templates already use it. The template package switched to the int8 VAE in v0.11.69, which ComfyUI ships from v0.37.2. If your template loads it, leave it.
Left: fp16 video VAE. Right: int8 video VAE. Same seed and the same sampling; only the decode differs. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size.
Prompt
The scene opens exactly on <Picture 1>. Live-action, night, the back seat of a moving taxi. City neon lights slide across the woman's face and the window as the car drives. Her phone vibrates in her hand; she glances down at the screen, a small smile appears, she locks the phone and turns her head to look out of the window at the passing lights. The sequins on her dress catch the moving light. The camera stays still inside the car. Audio: low engine hum, rain on the roof, one short phone vibration, the muffled city outside. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
Sparse attention: worth trying
Add ComfyUI's own Model Sparse Attention node (BlockSparseAttention) on the model wire between the turbo switch and the sampler, with its defaults. Sampling went from 97 s to 60 s per step. The first steps still run dense and the decode takes as long as before, so the whole clip gains 29%, not 38%.
- The scene stays the same: the same framing and movement, SSIM 0.84–0.90 against the clip without it.
- Fine skin texture gets a little smoother in close crops. Watching the clips at normal size, we could hardly tell them apart. Three shots is a small sample.
- One known problem is not ours to confirm. A user with an RTX 5090 reported ripples with sparse attention and larryvrh's turbo LoRA (#16382). With the template's own turbo LoRA we saw none.
- We did not measure it together with the int8 VAE.
Left: the 15-minute settings. Right: the same with sparse attention. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size. Sound: the left clip.
Prompt
The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
The fused 4-step model: a different picture
MATLOWAI's fused file replaces the model itself: 20.98 GB, with the 8-step turbo LoRA and a style LoRA already folded in. Load it in place of the pruned int8 model, turn the turbo switch off, set 4 steps.
- The whole gain is the step count. Each step took 94.5 s, about what it took before (97 s).
- The look changes: glossier skin, brighter light, a different background. The style LoRA is baked in and cannot be turned off.
- Use it for drafts, or if you like the look.
Left: the 15-minute settings, 8 steps. Right: MATLOWAI's fused model, 4 steps. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size. Sound: the left clip.
Prompt
The scene opens exactly on <Picture 1>. Live-action, night, the back seat of a moving taxi. City neon lights slide across the woman's face and the window as the car drives. Her phone vibrates in her hand; she glances down at the screen, a small smile appears, she locks the phone and turns her head to look out of the window at the passing lights. The sequins on her dress catch the moving light. The camera stays still inside the car. Audio: low engine hum, rain on the roof, one short phone vibration, the muffled city outside. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
1344×768 or 1920×1088
At megapixels 0.975 the template renders 1344×768, and the clip takes 317–319 s instead of 908–912 s.
- 1920×1088 looked a little better to us: more texture in skin and hair, less shine. The 768p clips are softer but hold up.
- MiniMax describes the open model as a 768p model. The model card says H3-Base produces "results at 768p resolution". 1920×1088 is outside what the card promises. ComfyUI 0.37.0 rendered it, and none of our 1080p clips came out broken.
- Make the start image at the canvas size. A 1920×1088 start image gave a slightly better clip than a 1344×768 one stretched to the same canvas.
Left: rendered at 1344×768 and scaled up. Right: rendered at 1920×1088. Each half is a 960-pixel-wide strip of the 1920×1088 frame at full size. Sound: the right clip.
Prompt
The scene opens exactly on <Picture 1>. Live-action, night, the back seat of a moving taxi. City neon lights slide across the woman's face and the window as the car drives. Her phone vibrates in her hand; she glances down at the screen, a small smile appears, she locks the phone and turns her head to look out of the window at the passing lights. The sequins on her dress catch the moving light. The camera stays still inside the car. Audio: low engine hum, rain on the roof, one short phone vibration, the muffled city outside. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
Ten seconds
Set the length to 10 s (243 frames).
| Canvas | Wall time | Shared GPU memory |
|---|---|---|
| 1344×768 | 811 s (13.5 min) | no change |
| 1920×1088 | 5,091 s (85 min) | +3.3 GB |
- At 1344×768 ten seconds is practical. The face and the scene held to the last frame.
- At 1920×1088 it does not fit. The card spilled into shared memory and each step took 10–11 minutes. The clip came out fine, but it is not something to queue twice.
Ten seconds at 1344×768: 811 s (13.5 min).
Prompt
The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
Ten seconds at 1920×1088: 5,091 s (85 min).
Prompt
The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.
What these settings do not change
- Voices. Lip-sync was fine in our clips, but the voice is generic. For dialogue, expect to dub it with a separate voice model.
- The LoRA errors in the log. If you load a LoRA built for the full model, see the
shape '[96768, 8]' is invalidpage.
About these numbers
All runs are ours, on one RTX 5060 Ti 16 GB in a Windows desktop with a display attached, ComfyUI 0.37.0 portable, one seed per configuration. ComfyUI was also started with --disable-pinned-memory, which our measuring needs and yours does not. Files were saved at crf 16. Memory figures are whole-card readings and include the desktop's own use.
Licence. MiniMax H3 is released under the MiniMax H3 Community License, which does not cover use in the European Union, the United Kingdom, South Korea or the United States. The files named here carry their own terms. The clips on this page were made under a separate written permission from MiniMax. It covers this site only; your own use is still governed by the Community License and its territory limits. This is a summary, not legal advice.