MAGI 2 vs MAGI 1: What Changed in Sand.ai's Video Model

Sep 16, 2026

MAGI 2 and MAGI 1 are two research milestones from Sand.ai, and the difference between them is not a bigger checkpoint - it is a different answer to the question "how should a video model be built?". MAGI 1 focused on how to generate video over time. MAGI 2 focuses on how to scale a generation model while keeping audio and video together in one sequence. This guide walks through what changed, what stayed, and what it means if you are choosing a model for production work.

What MAGI 1 established

MAGI 1 studied autoregressive video generation with temporal chunking. Instead of predicting a whole clip at once, the model generates the video in ordered chunks, so each new chunk is conditioned on the chunks that came before it. That framing made long video tractable and produced a clear result: temporal structure matters, and how you cut time into chunks changes the quality of motion.

The tradeoff is architectural complexity. Chunked autoregression means the model has to keep track of an ever-growing history, and the machinery that keeps that history coherent - memory, conditioning, and per-chunk consistency - grows with the job.

What MAGI 2 changes

MAGI 2 asks a scaling question instead. The answer is a single-stream architecture: text, video and audio are compressed into one shared token sequence and processed by one Transformer backbone, rather than by separate towers for visuals and sound.

On top of that backbone sits MagiMoE, an ultra-fine-grained mixture of experts. The full model carries 114B parameters, but only about 6B are activated per token. Fine-grained routing spreads capacity across many small experts, so the model can hold expressive capacity without paying for all of it on every token. In practice that is what makes long-form generation affordable: you scale the knowledge, not the per-token bill.

The two changes reinforce each other. Because everything lives in one sequence, there is no cross-tower synchronization step to keep speech aligned with lip movement. Because routing is fine-grained, the sequence can be long without the cost curve exploding.

Where the difference shows up

  • Character performance. Dialogue, singing and emotional shifts keep natural lip-sync, facial expressions, eye movement and body language, because the audio that drives the performance is generated in the same sequence as the frames.
  • Music to video. Rhythm-aware camera moves and scene cuts follow the track instead of being edited in afterwards.
  • Long-form consistency. Story, motion, sound and camera rhythm stay coherent deeper into a generation, which is the practical limit most teams hit first.
  • Inference cost. Per-token computation stays low relative to model size, so a longer clip does not require a proportional jump in compute.

Training and systems work

Scaling a 114B parameter model is a systems problem as much as an architecture one. MAGI 2 relies on co-designed parallelism - head parallel training across NVLink and InfiniBand - to keep the training run practical. MoE routing also changes the shape of the workload: many small experts means more, smaller exchanges between devices, which is why the interconnect and the routing granularity have to be designed together.

If you want the architecture details, the MAGI-2 Preview announcement from Sand.ai is the primary source: read the MAGI 2 research announcement.

Which model should you care about?

If your question is historical - how autoregressive video generation was made to work over time - MAGI 1 is the reference point. If your question is practical - how to generate synchronized audio and video at scale, with a cost curve that lets you ship - MAGI 2 is the model family to follow, and the current milestone is the preview release.

For the background on the model itself, start with what MAGI 2 is, then read the MAGI-2 Preview release and availability notes. If you want to skip straight to running it, the MAGI 2 documentation covers setup and the pricing page covers plans and credits.

FAQ

Is MAGI 2 just a bigger MAGI 1?

No. The scale is larger - 114B parameters - but the more important change is architectural: a single-stream sequence shared by text, video and audio, with ultra-fine-grained mixture-of-experts routing that activates roughly 6B parameters per token.

Does MAGI 1 have audio?

MAGI 1 was primarily a video-generation study built around autoregressive temporal chunking. MAGI 2 is where audio and video are generated together in one sequence.

What is MAGI v2?

"MAGI v2" and "MAGI 2" refer to the same model family. The naming you see in search results varies: MAGI 2, MAGI-2 and magi2 all point here.

Can I use MAGI 2 today?

MAGI-2 Preview is the current release from Sand.ai. Check the preview availability notes for what is public now and what is still to come.

Admin

Admin

MAGI 2 vs MAGI 1: What Changed in Sand.ai's Video Model | MAGI 2 Blog