Capabilities of the unified audio-video generation model - from lip-synced character performance to rhythm-aware music-to-video.

Dialogue, singing and emotional shifts with natural lip-sync, facial expressions and body language.

Turn a track into synchronized visuals with rhythm-aware camera moves and scene cuts.

One token sequence keeps story, motion, sound and camera rhythm coherent across longer generations.

114B parameters with only about 6B activated per token, so long video stays practical to generate.