What is MAGI 2

Oct 24, 2025

MAGI 2 is a unified audio-video generation model. Text, video and audio are processed by one Transformer backbone as a single token sequence, so dialogue, music, motion and emotion stay naturally synchronized.

One model, every medium

Instead of separate towers for visuals and sound, MAGI 2 compresses everything into a shared sequence. That keeps character performance consistent — lip-sync, facial expressions, eye movement and body language all move together with the audio.

Built on MagiMoE

MAGI 2 scales with an ultra-fine-grained mixture of experts. The full model has 114B parameters, but only about 6B are activated per token. Fine-grained routing keeps long-form generation practical while preserving expressive capacity.

What you can generate

  • Character performance — dialogue, singing and emotion with coordinated lip-sync and body language.
  • Music to video — rhythm-aware visuals generated directly from music, with camera and scene cuts that follow the beat.
  • Long-form scenes — a unified sequence keeps story, motion, sound and camera rhythm coherent across longer generations.

A step toward physical AGI

Compressing video into generative, composable world knowledge is the foundation for models that understand how the physical world behaves — not just what it looks like.

Try the preview

MAGI-2 Preview is an intermediate research release from Sand.ai. Read the announcement on the Sand blog for architecture details, training systems, and examples.

Admin

Admin

What is MAGI 2 | Blog