Two of the most interesting audio-native video models right now sit in adjacent slots: ByteDance's Seedance 2.5 and MiniMax's H3, also known as Hailuo 03. Both generate speech, sound effects, and music in the same pass as the picture. Both take text, images, and reference media as input. The marketing copy makes them sound interchangeable. The parameter surfaces say otherwise, and the differences are the kind that decide pipelines, not preferences.
I use both through the same third-party platform, which makes the comparison unusually clean: same harness, same uploaders, same meter, two different sets of constraints. The usual caveat applies: these are the models as exposed on the channels available today, not official documentation, and channel-level caps can differ from whatever first-party APIs eventually expose.
The one-table version
Seedance 2.5
MiniMax H3 (Hailuo 03)












