MAGI 2 is a unified audio-video generation model. Text, video and audio are processed by one Transformer backbone as a single token sequence, so dialogue, music, motion and emotion stay naturally synchronized.
One model, every medium
Instead of separate towers for visuals and sound, MAGI 2 compresses everything into a shared sequence. That keeps character performance consistent — lip-sync, facial expressions, eye movement and body language all move together with the audio.
Built on MagiMoE
MAGI 2 scales with an ultra-fine-grained mixture of experts. The full model has 114B parameters, but only about 6B are activated per token. Fine-grained routing keeps long-form generation practical while preserving expressive capacity.
What you can generate
- Character performance — dialogue, singing and emotion with coordinated lip-sync and body language.
- Music to video — rhythm-aware visuals generated directly from music, with camera and scene cuts that follow the beat.
- Long-form scenes — a unified sequence keeps story, motion, sound and camera rhythm coherent across longer generations.
A step toward physical AGI
Compressing video into generative, composable world knowledge is the foundation for models that understand how the physical world behaves — not just what it looks like.
Try the preview
MAGI-2 Preview is an intermediate research release from Sand.ai. Read the announcement on the Sand blog for architecture details, training systems, and examples.
