Bytedance has released Seedance 2.0 to a limited group of users. The previous model was already one of the most capable AI video generators available. The new version pushes things even further.

The multimodal video generation model handles up to four types of input at once: images, videos, audio, and text. Users can combine up to nine images, three videos, and three audio files, up to a total of twelve files. Generated videos run between 4 and 15 seconds long and automatically come with sound effects or music.

The demo videos come straight from ByteDance and were almost certainly cherry-picked from a larger batch of generated clips. Nobody knows yet how consistently the model hits this quality bar in real-world use, what it costs, or how long generation takes. So what we're seeing is likely a best-case scenario—and even when these capabilities look impressive on paper, there are still significant hurdles to getting them into professional workflows, like consistency. Still, the quality on display is genuinely impressive.

Prompt: The camera follows a man in black clothing who flees quickly. Behind him, a crowd of people pursues him. The camera switches to a sideways chase shot. The figure knocks over a roadside fruit stand in panic, picks himself up and runs on. The excited shouts of the crowd can be heard in the background.