China-based AI company MiniMax has released the H3 video generation model. MiniMax states that the system was developed primarily for commercial use.
H3 can process text, image, video, and audio within the same pipeline to generate short videos. Users can combine these inputs to edit existing content or create new scenes. The model processes text commands along with reference images, videos, and audio files together.
According to MiniMax, H3 can generate video at up to 2K resolution for up to 15 seconds. The system generates natural stereo audio simultaneously with the video. This prepares audio and video in the same production process, offering an alternative to workflows requiring separate voiceover and synchronization after image generation.
Motion transfer and commercial focus
H3 can transfer the motion structure of one scene to another video. It extracts character movements and camera angles from reference content and applies them to different scenes. MiniMax states this capability aims to reduce reshooting needs for companies producing commercials, product promotions, and game content.
The target areas for the model include digital ads, product videos, online store content, game scenes, and design work. MiniMax stated that the cost of producing 2K video with H3 is less than one-third of widely available competing products. The company did not publish a detailed pricing table specifying which models and usage conditions the cost comparison covers.
Open weights and hardware compatibility
MiniMax plans to publish H3's model weights within a few days of the announcement. Publishing the weights would allow developers to download the system and adapt it to their own needs. This contrasts with a significant portion of leading video generation models that operate as closed systems.
The licensing terms, commercial use limits, and required hardware capacity for H3's weights have not yet been detailed. Some of MiniMax's previous open-weight models included commercial use restrictions.
MiniMax stated that H3 was designed to run on multiple China-manufactured processors. The company has not yet announced the full list of supported processors or performance results on different hardware. This compatibility move is part of Chinese AI companies' effort to reduce dependence on US-manufactured advanced semiconductors.
Market competition
Competition in China's AI-powered video market accelerated after ByteDance released its Seedance 2.0 model. That model attracted wide interest for its ability to use text, image, video, and audio inputs together. Short-video platform Kuaishou is also positioned in the same market with its Kling 3.0 model.
Chinese companies are competing not only on image quality but also on production time, price, and developer access options.