MiniMax has released MiniMax H3, a general-purpose multimodal generation model that can generate 15-second 2K video clips with native stereo audio, according to MarkTechPost. This model reads text, images, video, and audio as one unified context and returns video with native stereo sound. The release of MiniMax H3 marks a significant development in the field of multimodal generation models.
The MiniMax H3 model has the capability to produce 2K output videos with durations ranging from 4 to 15 seconds, with integer durations, as reported by MarkTechPost. The model's ability to process multiple types of input and generate high-quality video with native stereo audio is a notable feature. However, limited information is available on the model's technical specifications and potential applications.
The release of MiniMax H3 is part of a broader trend in the development of multimodal generation models, which have the potential to revolutionize the way we create and interact with digital content. Related sources, such as MarkTechPost's reports on DeepSeek and Supabase, highlight the ongoing advancements in AI and machine learning technologies.