self-hosted/ai
§01·guide · video

MiniMax H3: "shape '[96768, 8]' is invalid" when a LoRA loads on the pruned model

videointermediate12GB+ VRAMOct 1, 2026

A guide for MiniMax H3 (Hailuo 3), written for the RTX 5090, RTX 3090 and 16 more.

models
tools

The short answer

You loaded a LoRA on minimax_h3_fl2va_pruned_int8_convrot (or the pruned_fp8_scaled twin), and ComfyUI printed dozens of lines like this:

ERROR lora diffusion_model.blocks.0.adaln_proj.linear.weight shape '[96768, 8]' is invalid for input of size 260112384
  • Nothing is broken. Your download is fine, the LoRA file is fine, and the clip still renders.
  • The LoRA was made for the full H3, and part of it targets tensors the pruned model does not have. ComfyUI skips that part and applies the rest.
  • We checked what "skips" means. On our RTX 5060 Ti, the original LoRA with the errors produced the same clip as that LoRA's pruned copy, which throws no errors: the decoded video and audio are identical, frame for frame and sample for sample.
  • What you lose is the part the pruned model cannot hold. For most style and character LoRAs that part is small. For a turbo LoRA it may not be, and a node exists that keeps it (below).

Every MiniMax H3 recipe on this site for an NVIDIA or AMD card installs the pruned model, so this applies whichever card you run.

What the error looks like

In the ComfyUI console, when the first sampling step starts:

ERROR lora diffusion_model.blocks.0.adaln_proj.linear.weight shape '[96768, 8]' is invalid for input of size 260112384
ERROR lora diffusion_model.blocks.1.adaln_proj.linear.weight shape '[96768, 8]' is invalid for input of size 260112384
…
ERROR lora diffusion_model.final_layer.adaln_proj.linear.weight shape '[10752, 8]' is invalid for input of size 28901376
  • There is one line for each of the model's 50 blocks and one for the final layer: 51 tensors.
  • The set repeats on every sampling step. At 6 steps that is 306 lines (51 × 6), which is what our log held. It is one problem, not 306.
  • The run finishes normally. ComfyUI does not stop, and the queue shows the job as done.

The same cause has other wordings. Before ComfyUI PR #15908, alibaba-pai's 8-step acceleration LoRAs failed on the pruned model with The size of tensor a (96) must match the size of tensor b (3072) at non-singleton dimension 0 (Kijai discussion #30). Larryvrh's turbo LoRA drew exactly the [96768, 8] line from other users (discussion #11).

Why it happens

"Pruned" does not mean the model has fewer layers. MiniMax's model card says that of the transformer's 33B parameters, "approximately 13B parameters residing in AdaLN-related branches". That is the part that tells each block how far along the denoising it is. Comfy's pruned files replace those branches with a small table: one [96768, 8] projection per block instead of a full weight matrix. That is how the file drops from 34.04 GB to 20.97 GB at int8.

A LoRA trained against the full model usually also adjusts those AdaLN weights. Its adaln_proj tensors are sized for the full [96768, 2688] matrix, 260,112,384 values per block, and cannot be poured into an 8-column table. That mismatch is the error.

Two things that do not cause it:

  • Quantisation. int8, fp8, nvfp4 and w4a8 change how weights are stored, not which weights exist. A LoRA built for the pruned model loads on any of them.
  • Your ComfyUI version. Our run on ComfyUI 0.37.0 printed the same lines another user reported on 0.35.1 (#16382). Updating does not make them go away.

What we measured

One RTX 5060 Ti 16 GB, ComfyUI 0.37.0, fl2va_pruned_int8_convrot, text-to-video at 864×480, 124 frames (5 s), 6 steps, seed 3001, larryvrh's minimax_h3_turbo_v4_step600_ema LoRA loaded three ways:

How the LoRA was loadedError linesWall timeOutput
Original file, stock LoraLoaderModelOnly306112.8 sreference
drbaph's pruned copy (…_pruned_comfyui), stock loader0112.1 sidentical to the row above
Original file, Larryvrh node, bypass (default)0116.3 sdifferent clip
Original file, Larryvrh node, merge (low_vram on)0112.9 sdifferent clip
  • The stock loader with errors and the pruned copy without them are the same thing. The decoded video and audio of the two clips have the same MD5 (the MP4 files themselves differ, because ComfyUI writes the workflow into each one, and the workflow names the LoRA file). So the errors are noise in the log: the pruned copy only makes them go away.
  • The node changes the clip, at no cost in speed. Whole-card peak memory was 15,528–15,750 MiB for all four, and the times differ by a few seconds.
  • How different is "different". SSIM between the stock clip and the node's clips is 0.54 (bypass) and 0.65 (merge). For scale, the stock clip against the official template's own 8-step turbo LoRA, a different LoRA altogether, is 0.62. The node is not a small correction: it is as far from the stock result as switching LoRAs.

Which one looks better we cannot say from 480p clips, so we repeated it at full size.

The same comparison at 1920×1088

Image-to-video this time, on three start images (a street crossing, the back seat of a taxi, a mirror), 1920×1088, 124 frames, 6 steps, the same LoRA, Comfy Kitchen attention on, one seed per shot:

How the LoRA was loadedError linesWall timeWhole-card peakSSIM against the stock clip
Original file, stock LoraLoaderModelOnly306704–708 s15,523–15,701 MiBreference
Original file, node, merge (low_vram on)0704–708 s15,514–15,573 MiB0.87 / 0.92 / 0.95
Original file, node, bypass (default)0734–739 s15,163–15,197 MiB0.55 / 0.77 / 0.65

For scale: the stock clip against the template's 8-step turbo LoRA on the same three shots is 0.55 / 0.82 / 0.68.

  • Merge gave almost the stock clip. Same framing, same movement, the same face; the two drift apart only in small details (a hand position, the spray behind a car). The softness the node's author warns about did not show next to the stock result.
  • In one shot of three, merge was cleaner. In the street shot the woman turns her head quickly, and in the stock clip her eyes smear a little during the turn. In the merge clip they do not. In the taxi and mirror shots we saw no difference between the two. One seed per shot, so take it as a hint, not a rule.
  • Bypass gave a different clip. A little more contrast, the light sits differently on the face, the background differs. Watching the clips, we would not call it better or worse than the other two. It took about 30 s longer per clip (4%).
  • All three fit in 16 GB without spilling into shared memory.
  • None of the nine clips is broken. The stock runs, 306 error lines each, gave clean clips at this size too.

So on a 16 GB card the stock loader is enough, and the node is a small upgrade rather than a fix: merge costs nothing, clears the log and may clean up fast motion; bypass gives a different take for 30 s more.

A caveat on the two tables: at 480p the merge clip was as far from stock as bypass; here it is nearly the same clip. The start image pins down much of an image-to-video clip, and the two tests differ in both task and size, so we cannot say which of the two explains it.

AI-generated with MiniMax H35.2 s · 1920×1088Report this video

The same shot and seed with larryvrh's turbo LoRA loaded three ways. Left to right: the stock loader (306 error lines), the Larryvrh node with low_vram on (merge), the Larryvrh node with low_vram off (bypass). Each strip is the middle third of the 1920×1088 frame at full size. Sound: the merge clip.

Prompt

The scene opens exactly on <Picture 1>. Live-action, a rainy night city street, neon reflections on the wet asphalt. The woman in the beige trench coat hurries across the crosswalk toward the camera, holding her coat closed at the collar with one hand. Halfway across she turns her head to the right as a taxi passes behind her, then looks ahead again and keeps walking, her hair lifting in the wind. Rain falls steadily and splashes on the asphalt. The camera tracks backward slowly, keeping her the same size in frame. Audio: steady rain, her heels on wet asphalt, a car hissing through a puddle, a distant horn. One continuous shot, no cuts. Live-action texture. No text, subtitles, logos or watermarks.

What to do

1. For a style or character LoRA: nothing, or switch to its pruned build to clean up the log. The result is the same either way, as the table shows. Fizgig's author, whose trainer works on the pruned file, argues the AdaLN branch "only sees the timestep, so nothing a LoRA learns lives there" (Fizgig README). Take that as the author's claim, but it matches what a skipped adaln_proj costs a LoRA trained on appearance.

2. For larryvrh's turbo LoRA: decide between the stock loader and the node. A turbo LoRA is trained to make the model take bigger steps, which is exactly what the time-conditioning branch controls. The Larryvrh node "detects a pruned base automatically and re-injects the LoRA's time-conditioning at run time", so one LoRA file serves every base.

  • low_vram off (default, bypass) applies the LoRA at run time. The author calls it the sharpest.
  • low_vram on (merge) folds the LoRA into the weights for less peak memory. The author warns it comes out "softer on quantized (int8 / fp8 / pruned) bases".
  • In our 1920×1088 runs merge came out almost identical to the stock loader, a little cleaner in one fast head turn, and bypass came out different rather than better (above). If the stock result looks right to you, you do not need the node; if you install it, low_vram on keeps the clip you already know and removes the error lines.
  • The turbo LoRA the official ComfyUI template ships (minimax_h3_fl2v_turbo_8step_v1.0, behind the template's turbo switch) is already built for the pruned model: 0 error lines in our runs. It needs neither.

3. When you pick a new LoRA, check which model it was built for. The names people use for pruned-compatible files:

  • _pruned, _pruned_comfy, _pruned_comfyui, Pruned-ComfyUI, Pruned-Fixed;
  • "(No Tensor Errors)" in a Civitai title;
  • anything trained in Fizgig on the pruned file, which is pruned-native by construction.

Match the transformer too: a LoRA trained on fl2va goes with fl2va, one trained on ref2va with ref2va. Publishers of per-transformer files say so on the card.

The reverse case exists. A LoRA built only for the pruned model, loaded on the full 34 GB model, logs keys that "aren't recognized". Kijai, who made the pruned files, explains they are pruned-only keys and the rest applies (discussion #19).

Errors that look similar and are not this

  • SafetensorError: … incomplete metadata, file not fully covered on a pruned turbo LoRA: some early uploads carried 64 extra bytes at the end of the file. Use the re-saved Pruned-Fixed copies.
  • All-black frames: that is the weights or the VAE, not a LoRA. Plain bf16 files rendered black on ComfyUI 0.32.0 (#15563), and the int8 video VAE needs ComfyUI 0.31.0 or newer.

About these numbers

All runs are ours, on one RTX 5060 Ti 16 GB in a desktop with a display attached, one seed per configuration. Times are wall-clock for the whole job, text encoding and VAE decode included. The whole-card memory readings include the desktop's own use.

Licence. MiniMax H3 is released under the MiniMax H3 Community License, which does not cover use in the European Union, the United Kingdom, South Korea or the United States. LoRAs trained on it inherit that licence; the nodes named here carry their own (Larryvrh's node is Apache-2.0). The clips on this page were made under a separate written permission from MiniMax. It covers this site only; your own use is still governed by the Community License and its territory limits. This is a summary, not legal advice.