MAGI 2 generates video and audio together: the model reads text, video and audio as one token sequence, so dialogue, music, motion and mood are produced in the same pass instead of being edited together afterwards. That changes how you prompt. This guide covers the working loop - what to write, in what order, and how to keep long generations coherent.
Start from the sequence, not the shot
Most video tools want you to describe a shot. MAGI 2 rewards you for describing a scene that evolves. Before you write anything, decide four things:
- Who or what is on screen - one character, two characters in conversation, or a music-driven montage with no dialogue.
- What is heard - spoken dialogue, singing, ambient sound, or a track that the visuals should follow.
- How the camera behaves - locked off, slow push in, handheld follow, or cutting on the beat.
- How long the moment lasts - a single beat, or a scene with a beginning and an end.
Those four decisions map onto the token sequence the model builds, so writing them down first is not busywork; it is the prompt skeleton.
Prompting character performance
For dialogue and singing, describe the performance rather than the render quality. Useful phrases name the emotion, the delivery and the body:
- "she speaks quietly at first, then laughs mid-sentence"
- "he sings the last line, shoulders relaxing on the final note"
- "the camera stays on her face; eyes move to the left before the reply"
Because audio and video share one sequence, the lip movement, the expression and the timing of the line stay tied together. If a line looks out of sync, the cause is usually that the performance description changed mid-scene - keep the emotional arc consistent instead of adding new instructions halfway through.
Music-to-video
When the audio is the anchor, describe the relationship between the track and the image rather than the images alone:
- "cuts land on the downbeat, camera drifts during the bridge"
- "abstract shapes expand with the bass, contract on the vocal line"
- "the scene changes colour when the chorus enters"
Rhythm-aware camera and scene cuts are one of the clearest places where a unified audio-video model beats stitching separate tools together: the model already knows where the beat is, because it generated the audio.
Keeping long-form scenes coherent
Long generations fail in predictable ways: characters drift, the audio signature changes, or the camera rhythm resets. Three habits avoid most of it:
- Anchor the constants. Restate the character, wardrobe and location in the same words each time you extend.
- Keep one arc per generation. A scene with a single emotional direction survives longer than one that changes genre midway.
- Move the camera deliberately. Slow, motivated movement reads as intentional; constant motion reads as drift.
If a scene is drifting, cut earlier and start the next segment from the last good frame rather than asking for a longer generation.
Planning credits and cost
Generation cost scales with how much you ask for, not with how many ideas you have. Practical habits:
- Draft short, then extend the takes that work.
- Reuse a locked prompt template across a series so you are not re-exploring the model's behaviour on every run.
- Reserve long-form generations for the shots that will actually be used.
Credits per plan are listed on the MAGI 2 pricing page; each plan includes a monthly or one-time generation allowance.
A five-minute workflow
- Write the four-line skeleton (who, what is heard, camera, length).
- Add the performance or rhythm notes in the language of the medium - emotion and delivery for character work, beat and section for music.
- Generate short, review, then extend the take that works.
- Keep the prompt constants identical when you extend.
- Export and, if you are assembling a longer piece, cut between generations rather than inside one.
New to the model? Start with what MAGI 2 is for the architecture background, or go straight to the MAGI 2 documentation for setup. If you are weighing MAGI 2 against the previous generation, the MAGI 2 vs MAGI 1 comparison covers what changed.
FAQ
Do I need to prompt audio and video separately?
No. MAGI 2 reads text, video and audio as one token sequence, so a single prompt describes the performance and what is heard at the same time.
How do I get synchronized lip-sync?
Describe the performance continuously - emotion, delivery and body language - and avoid changing the character or tone mid-scene. Synchronization comes from the shared sequence, not from post-production.
How long can one generation be?
Longer than a single shot, but the practical limit depends on how much changes inside the scene. Keep one arc per generation and extend from the last good frame.
Where do I start if I have never used it?
Read the documentation for setup, then check the pricing page to pick a plan size that matches how much you intend to generate.
