MiniMax-H3 prediction-preview modular blocks

Custom Modular Diffusers blocks that show what MiniMax-H3 is aiming at while it denoises. These are the video version of the Krea 2 preview blocks: each step, the x0 prediction (what the model thinks the video will be from the current step) is decoded and handed to a callback. From the first step it's a whole clip that moves, and it sharpens on each step after that.

960x544, 124 frames, 25 steps, seed 0. The first part plays the preview as it would look live: the clip loops while the steps advance. Top row: rgb and taeh3. Bottom row: taeh3_fast (scaled up) and the final VAE decode. The final video then plays with its soundtrack.

preview_callback is called as preview_callback(step, total, frames), once per previewed step. frames is a list of PIL images with one per output frame, so it plays at the video's 24 fps. preview_mode chooses how much work goes into each one:

preview_mode="rgb"         24x3 linear projection of the latent channels. No weights, ~8 ms at 960x544
preview_mode="taeh3"       TAEH3 tiny autoencoder. Real detail, ~280 ms at 960x544
preview_mode="taeh3_fast"  the same decoder at quarter size, ~60 ms
preview_mode="latents"     no decode; the raw (1, 24, T, h, w) prediction, for callers that decode themselves

Times are for the whole 124-frame clip, including the conversion to PIL. The taeh3 modes download the decoder weights, 22 MB, the first time they run. Only the video is previewed; the soundtrack isn't.

The blocksets are the stock MiniMax-H3 ones with a preview step added to the denoise loop, for all three workflows (t2va, fl2va, ref2va). With no preview_callback they generate exactly what the stock pipeline does.

Loading & running

import sdnq  # needed to load the quantized text encoder and transformer
import torch

from diffusers import ComponentsManager, ModularPipeline
from diffusers.utils import encode_video


manager = ComponentsManager()
pipe = ModularPipeline.from_pretrained(
    "OzzyGT/minimax_h3_preview_blocks", trust_remote_code=True, workflow="t2va", components_manager=manager
)
pipe.load_components(dtype=torch.bfloat16, trust_remote_code=True)
manager.enable_auto_cpu_offload(device="cuda")


previews = {}


def on_preview(step, total, frames):
    previews[step] = frames  # a list of PIL images, one per output frame


out = pipe(
    prompt="a red fox trotting through a snowy pine forest at dawn, snow crunching underfoot",
    height=544,
    width=960,
    num_frames=124,
    num_inference_steps=25,
    preview_callback=on_preview,
    preview_mode="rgb",
    output=["videos", "audio", "sampling_rate"],
)
encode_video(
    out["videos"][0],
    fps=pipe.fps,
    output_path="result.mp4",
    audio=out["audio"][0],
    audio_sample_rate=out["sampling_rate"],
)

workflow= picks what gets loaded: t2va and fl2va use the transformer/ partition, ref2va uses transformer_ref/. Without it, both partitions load.

These blocks only work with MiniMax-H3.

For loading a different checkpoint or swapping components, see Modular pipeline in the diffusers docs.

Previews

preview_every=N previews every Nth step; the final step is always previewed. The callback runs inside the denoising loop, so whatever it does is time the loop isn't spending on the model. rgb is effectively free and the one I'd recommend. taeh3 costs real time on every previewed step, but next to an H3 step it's a small share.

Raising from the callback aborts the run: the exception propagates out of pipe(...) and the pipeline is reusable afterwards. That's the cancel button.

class Cancelled(Exception):
    pass


def on_preview(step, total, frames):
    if stop_requested:
        raise Cancelled

The loop logs a traceback before re-raising, so a clean cancel still prints one.

References

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OzzyGT/minimax_h3_preview_blocks

Finetuned
(170)
this model