AI video model guide

MiniMax H3

A general-purpose multimodal video system from MiniMax with text, image, video, and audio context, native stereo sound, and separate base and 2K regeneration workflows.

Published model guide
Model input mapMiniMax
Text
First / last frame
Reference media
Audio
model contextgenerated result
Video output

This diagram explains supported context types. It is not presented as a generated sample from the model.

Capability overview

What defines MiniMax H3

The points below summarize the model family as documented by its developer. Product-specific limits are kept separate.

01

Omni-modal context

H3 can interpret combinations of text, images, video, and audio rather than limiting the workflow to one prompt and one starting frame.

02

Native stereo sound

The model jointly produces video and 32 kHz stereo audio, supporting dialogue, ambience, and scene sound in the generated result.

03

Base and 2K workflows

H3-Base produces 768p output. The official full workflow uses H3-Regenerate-2K to regenerate that result at 2K with the original context.

Specifications

At a glance

Published model facts and current UnblurVideo status, last checked 2026-09-27.

Developer
MiniMax
Output duration
4–15 seconds
Common aspect ratios
21:9, 16:9, 4:3, 1:1, 3:4, and 9:16
Base resolution
768 px on the shorter side2K requires the separate H3-Regenerate-2K workflow.
Frame rate
24 FPS
Audio
32 kHz stereo
UnblurVideo availability
PlannedNo fallback to another model is used.
Facts last checked
September 27, 2026
Best fit

Where MiniMax H3 is most relevant

H3 is designed for briefs that need richer context, flexible framing, or synchronized picture and sound.

01

First and last frame control

Plan a transition between a starting image and an ending image, or use either frame independently.

02

Reference-heavy scenes

Use multiple images, video clips, and supported audio references when consistency depends on several source materials.

03

Sound-led short sequences

Create short scenes in which dialogue, effects, ambience, and visual timing need to be considered together.

Workflow

Understand the H3 pipeline before generating

The official system separates context processing, base generation, and optional 2K regeneration.

Step 1

Prepare multimodal context

H3-Context-IR interprets the relationship between prompts and reference media before the generation stage.

Step 2

Generate the base result

H3-Base creates the video and stereo audio at a base output whose shorter side is 768 pixels by default.

Step 3

Regenerate at 2K when needed

The official 2K workflow feeds the base result and original context into H3-Regenerate-2K; 2K should not be described as native H3-Base output.

Model availability and verified limits

This public guide is the canonical UnblurVideo page for MiniMax H3 and separates official model capabilities from features currently offered by the product.

  • This model guide is publicly indexable and ready to be expanded when the production integration is enabled.
  • MiniMax H3 is not in the production generation allowlist yet.
  • The page does not silently substitute a Kling, Seedance, or another MiniMax model.
  • The base model outputs 768p by default; 2K is a separate regeneration stage in the official full workflow.
  • Provider limits and credit pricing will be rendered from the deployed configuration after integration.
FAQ

Questions about MiniMax H3

Answers distinguish the published model family from the version and controls that may later be deployed here.

Related models

Continue comparing model families

Use the same availability and specification checks before choosing the next model for your workflow.

Production-ready options

Generate with a model available today

Open the shared workspace to compare the current production models, supported settings, duration, and credit cost before submitting a job.

Source for this guide: MiniMax H3 technical announcement.

Open generator
MiniMax H3 AI Video Model: Inputs, Audio and 2K Workflow | Unblur Video