ByteDance released Seedance 2.5 on July 31, doubling the single-generation video duration of its video-creation model from 15 seconds to 30 seconds2. The model can generate a 30-second high-quality video clip in a single run and supports multi-turn extension for longer sequences, enabling production of several minutes of coherent content with unified audio-visual language1.
Seedance 2.5 builds on the Seedance 2.0 unified multi-modal audio-video joint generation architecture, with the new version focusing on three areas: long-narrative capability, multi-modal reference, and editing ability.
On the input side, users can now provide up to 30 images, 10 videos, and 10 audio clips as reference material in a single generation. The model supports white-model reference, motion reference, and creative reference, enabling complex creations involving multiple subjects, scenes, and camera changes.
For long-form content, Seedance 2.5 optimizes shot transitions and scene changes to maintain coherence across extended sequences. The model also improves image quality, audio quality, and motion quality while reducing the oily artifacts common in AI-generated video.
On the editing front, Seedance 2.5 introduces precise timestamp control for targeted video editing and strengthens green-screen editing, viewpoint editing, and reference-based editing aimed at professional film and advertising production.
The model is rolling out to Jimeng AI and the Pro version of Doubao. An API is expected to arrive on Volcano Engine's Ark platform.
ANALYSIS The doubling of single-generation duration from 15 to 30 seconds, combined with multi-round extension for multi-minute sequences, moves Seedance 2.5 closer to the production requirements of professional video workflows where coherent, longer-form output is essential. The expanded multi-modal input capacity—accepting images, videos, and audio simultaneously—positions the model for use cases that require coordinating multiple reference assets in a single generation pass rather than stitching separate outputs together.