MAGI-2 Preview is here

MAGI 2

A unified audio-video generation model.
114B parameters. Only 6B activated per token.

MAGI 2 cinematic generation

What is MAGI 2

MAGI 2 is a unified model that generates video and audio together. Text, visuals and sound live in a single token sequence, so speech, music, motion and emotion stay in sync.

Unified architecture

Text, video and audio are processed by one backbone instead of separate towers.

One model, many outputs

Dialogue, music, sound effects and visuals are generated together.

Built to scale

MagiMoE keeps capacity at 114B while activating only 6B parameters per token.

Foundation for physical AGI

Compressing video into generative, composable world knowledge.

Why MAGI 2

Efficiency and expressiveness designed together - from architecture to training systems to data.

Dialogue, singing and emotional shifts with natural lip-sync, facial expressions, eye movement and body language.

Key capabilities

What MAGI-2 Preview brings to audio-video generation.

Character performance

Precise dialogue, singing and emotion with coordinated lip-sync, expressions and body language.

Music to video

Visuals generated from music with rhythm-aware scene flow.

Efficient MoE scaling

MagiMoE: 12 heads x 256-dim routed subspaces, top-6 of 256 experts per head.

Unified audio-video sequence

One Transformer backbone processes text, video and audio together.

Systems co-design

Head Parallel over NVLink and InfiniBand keeps 114B training practical.

Creative partners

A growing community of creators making ads, films, MVs and more.

Frequently asked questions

Questions about MAGI 2. Learn more at sand.ai.






Experience MAGI 2

Read the MAGI-2 Preview announcement and explore what unified audio-video generation can do.