Models
Browse available AI models across video, image, chat, and music generation.
All Models
MiniMax H3 SH · Text to Video
minimax-h3-sh/text-to-videoOpen Weights text-to-video with native stereo audio, 3–15 second duration, seven aspect ratios, and 480p/540p/768p/1080p output.
MiniMax H3 SH · Image to Video
minimax-h3-sh/image-to-videoAnimate one first-frame image or two first/last frames into 3–15 second video with native audio. Frame images are free; output follows the image ratio.
MiniMax H3 SH · Reference to Video
minimax-h3-sh/reference-to-videoGuide 3–15 second video with up to 9 images, 3 videos, and 3 audio references. Output and reference video seconds are billed, plus 4 credits per image or audio.
Long Video
long-videoCreate consistent long-form videos from 4 to 180 seconds with HappyHorse or Seedance
Seedance 2.5
doubao-seedance-2.5ByteDance Seedance 2.5 — 4–30s multimodal video at up to 1080p with text, image, video, audio-only, and first/last-frame inputs
Seedance 2.0
doubao-seedance-2.0Multi-modal video generation — text, image, video, audio inputs
Seedance 2.0 Fast
doubao-seedance-2.0-fastFaster Seedance 2.0 with lower pricing — same capabilities, quicker generation
MiniMax H3 Fast
minimax-h3-fastMiniMax H3 Fast — 480p text-to-video, first/last-frame, and multimodal reference generation, with free image, video, and audio input.
MiniMax H3 Max Turbo
minimax-h3-max-turboMiniMax H3 Max Turbo — 5–15s text-to-video and first/last frames at 480p or 768p, with prompt expansion and free image input.
MiniMax H3 Max
minimax-h3-maxMiniMax H3 Max — fast 5–15s text-to-video and first/last-frame generation at 480p or 768p, with free image input.
MiniMax H3
minimax-h3MiniMax H3 — multimodal video generation with text, image, video, and audio context, native stereo sound, and 4–15s output at 768p or 2K.
Veo 3.1
veo-3Google Veo 3.1 — text-to-video, image-to-video, first & last frame, with background music.