Contextual Omni Representation
MiniMax describes H3's instruction-following foundation as Contextual Omni Representation. Language connects the target clip to multiple context items and expresses relationships among those items, allowing open-ended directions such as combining a camera move, a character image, and a voice reference in one request.