Introduction
What is MAGI 2?
MAGI 2 is a unified audio-video generation model. Text, video and audio are processed by one Transformer backbone as a single token sequence, so dialogue, music, motion and emotion stay naturally synchronized.
Key capabilities
- Unified audio-video sequence - one backbone handles text, visuals and sound together.
- MagiMoE scaling - 114B total parameters, only about 6B activated per token.
- Character performance - precise dialogue, singing and emotion with coordinated lip-sync, expressions and body language.
- Music-to-video - turn music into synchronized, directed visuals with rhythm-aware camera and scene cuts.
- Long-form consistency - story, motion, sound and camera rhythm stay coherent across longer generations.
Built to scale
MAGI 2 is built on three interdependent directions: a scalable model architecture that reaches the 100B scale with a much smaller activated-parameter footprint, a scalable training system co-designed with the architecture, and a data methodology that shifts from filtering-centric pipelines toward high-throughput production and precise multimodal annotation.
MAGI-2 Preview
MAGI-2 Preview is an intermediate research release that validates the architecture, training system and data methodology at the 100B level. For the latest details, see the MAGI-2 Preview announcement.