Introduction

What is MAGI 2?

MAGI 2 is a unified audio-video generation model. Text, video and audio are processed by one Transformer backbone as a single token sequence, so dialogue, music, motion and emotion stay naturally synchronized.

Key capabilities

  • Unified audio-video sequence - one backbone handles text, visuals and sound together.
  • MagiMoE scaling - 114B total parameters, only about 6B activated per token.
  • Character performance - precise dialogue, singing and emotion with coordinated lip-sync, expressions and body language.
  • Music-to-video - turn music into synchronized, directed visuals with rhythm-aware camera and scene cuts.
  • Long-form consistency - story, motion, sound and camera rhythm stay coherent across longer generations.

Built to scale

MAGI 2 is built on three interdependent directions: a scalable model architecture that reaches the 100B scale with a much smaller activated-parameter footprint, a scalable training system co-designed with the architecture, and a data methodology that shifts from filtering-centric pipelines toward high-throughput production and precise multimodal annotation.

MAGI-2 Preview

MAGI-2 Preview is an intermediate research release that validates the architecture, training system and data methodology at the 100B level. For the latest details, see the MAGI-2 Preview announcement.