Akapulu Labs logo Akapulu Labs Research

Seedance 1.5 Pro

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Seedance 1.5 Pro — method overview

Seedance 1.5 pro is a native audio-visual generation model producing synchronized video and audio with precise multilingual lip-sync and cinematic control. It combines a dual-branch diffusion transformer with advanced training to deliver professional-grade, coherent content.

  • multimodal
  • talking-head
  • lip-sync
  • audio-driven
  • speech-driven
  • face-animation
  • realtime

Authors: Team Seedance, Heyi Chen, Siyan Chen, Xin Chen, Yanfei Chen, Ying Chen, Zhuo Chen, Feng Cheng, Tianheng Cheng, Xinqi Cheng, Xuyan Chi, Jian Cong, Jing Cui, Qinpeng Cui, Qide Dong, Junliang Fan, Jing Fang, Zetao Fang, Chengjian Feng, Han Feng, Mingyuan Gao, Yu Gao, Dong Guo, Qiushan Guo, Boyang Hao, Qingkai Hao, Bibo He, Qian He, Tuyen Hoang, Ruoqing Hu, Xi Hu, Weilin Huang, Zhaoyang Huang, Zhongyi Huang, Donglei Ji, Siqi Jiang, Wei Jiang, Yunpu Jiang, Zhuo Jiang, Ashley Kim, Jianan Kong, Zhichao Lai, Shanshan Lao, Yichong Leng, Ai Li, Feiya Li, Gen Li, Huixia Li, JiaShi Li, Liang Li, Ming Li, Shanshan Li, Tao Li, Xian Li, Xiaojie Li, Xiaoyang Li, Xingxing Li, Yameng Li, Yifu Li, Yiying Li, Chao Liang, Han Liang, Jianzhong Liang, Ying Liang, Zhiqiang Liang, Wang Liao, Yalin Liao, Heng Lin, Kengyu Lin, Shanchuan Lin, Xi Lin, Zhijie Lin, Feng Ling, Fangfang Liu, Gaohong Liu, Jiawei Liu, Jie Liu, Jihao Liu, Shouda Liu, Shu Liu, Sichao Liu, Songwei Liu, Xin Liu, Xue Liu, Yibo Liu, Zikun Liu, Zuxi Liu, Junlin Lyu, Lecheng Lyu, Qian Lyu, Han Mu, Xiaonan Nie, Jingzhe Ning, Xitong Pan, Yanghua Peng, Lianke Qin, Xueqiong Qu, Yuxi Ren, Kai Shen, Guang Shi, Lei Shi, Yan Song, Yinglong Song, Fan Sun, Li Sun, Renfei Sun, Yan Sun, Zeyu Sun, Wenjing Tang, Yaxue Tang, Zirui Tao, Feng Wang, Furui Wang, Jinran Wang, Junkai Wang, Ke Wang, Kexin Wang, Qingyi Wang, Rui Wang, Sen Wang, Shuai Wang, Tingru Wang, Weichen Wang, Xin Wang, Yanhui Wang, Yue Wang, Yuping Wang, Yuxuan Wang, Ziyu Wang, Guoqiang Wei, Wanru Wei, Di Wu, Guohong Wu, Hanjie Wu, Jian Wu, Jie Wu, Ruolan Wu, Xinglong Wu, Yonghui Wu, Ruiqi Xia, Liang Xiang, Fei Xiao, XueFeng Xiao, Pan Xie, Shuangyi Xie, Shuang Xu, Jinlan Xue, Shen Yan, Bangbang Yang, Ceyuan Yang, Jiaqi Yang, Runkai Yang, Tao Yang, Yang Yang, Yihang Yang, ZhiXian Yang, Ziyan Yang, Songting Yao, Yifan Yao, Zilyu Ye, Bowen Yu, Jian Yu, Chujie Yuan, Linxiao Yuan, Sichun Zeng, Weihong Zeng, Xuejiao Zeng, Yan Zeng, Chuntao Zhang, Heng Zhang, Jingjie Zhang, Kuo Zhang, Liang Zhang, Liying Zhang, Manlin Zhang, Ting Zhang, Weida Zhang, Xiaohe Zhang, Xinyan Zhang, Yan Zhang, Yuan Zhang, Zixiang Zhang, Fengxuan Zhao, Huating Zhao, Yang Zhao, Hao Zheng, Jianbin Zheng, Xiaozheng Zheng, Yangyang Zheng, Yijie Zheng, Jiexin Zhou, Jiahui Zhu, Kuan Zhu, Shenhan Zhu, Wenjia Zhu, Benhui Zou, Feilong Zuo

Categories: cs.CV

Comment: Seedance 1.5 pro Technical Report

Published 2025-12-15 · Updated 2025-12-23

Abstract

Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation. Leveraging a dual-branch Diffusion Transformer architecture, the model integrates a cross-modal joint module with a specialized multi-stage data pipeline, achieving exceptional audio-visual synchronization and superior generation quality. To ensure practical utility, we implement meticulous post-training optimizations, including Supervised Fine-Tuning (SFT) on high-quality datasets and Reinforcement Learning from Human Feedback (RLHF) with multi-dimensional reward models. Furthermore, we introduce an acceleration framework that boosts inference speed by over 10X. Seedance 1.5 pro distinguishes itself through precise multilingual and dialect lip-syncing, dynamic cinematic camera control, and enhanced narrative coherence, positioning it as a robust engine for professional-grade content creation. Seedance 1.5 pro is now accessible on Volcano Engine at https://console.volcengine.com/ark/region:ark+cn-beijing/experience/vision?type=GenVideo.


Introduction

Seedance 1.5 pro is presented as a native audio-visual joint generation foundation model for professional content creation. The paper positions the model as a step beyond video-only generation systems by making synchronized audio generation a first-class capability rather than an add-on. The reported target tasks include text-to-audio-video synthesis, image-guided audio-video generation, and the unimodal video regimes that the team already supported in earlier Seedance systems.

The central claim is that native joint generation requires coordinated progress across four layers: data, architecture, training, and inference efficiency. The paper therefore proposes a multi-stage audio-visual data pipeline, a dual-branch diffusion transformer with explicit cross-modal joint modules, a three-stage training recipe consisting of pre-training, supervised fine-tuning, and human-feedback alignment, and a distillation-plus-systems acceleration stack that reduces the inference cost while preserving output quality.

The model is described as particularly strong in multilingual and dialect lip-sync, cinematic camera control, narrative coherence, and audio-visual synchronization. The evaluation section emphasizes Chinese-language dialogue and dialect delivery, sound-event alignment, and professional production scenarios such as film, short drama, and stage-performance style content.

Overall Evaluation. Left: Video Evaluation; Right: Audio Evaluation.
Overall Evaluation. Left: Video Evaluation; Right: Audio Evaluation.

Problem Setting and High-Level Contributions

The paper frames joint audio-visual generation as the natural next step after video generation: users increasingly expect not only visually plausible motion, but also speech, sound effects, music, and synchronization that are consistent with the generated scene. In that setting, the system must jointly satisfy prompt following, visual aesthetics, motion quality, audio realism, and temporal alignment between modalities.

Seedance 1.5 pro is introduced with four stated technical contributions:

  • A comprehensive audio-visual data framework with multi-stage curation, audio-aware segmentation, multimodal quality filtering, joint distribution rebalancing, and curriculum-based staging.
  • A unified multimodal generation architecture based on MMDiT, with a dual-branch topology and dedicated cross-modal joint modules for video and audio streams.
  • Post-training optimization through high-quality supervised fine-tuning and RLHF with multi-dimensional reward models for audio-video alignment, motion dynamics, and aesthetics.
  • Inference acceleration via multi-stage diffusion distillation and system-level optimizations such as quantization, parallelism, offloading, and communication reduction, yielding more than 10× end-to-end acceleration according to the abstract and optimization section.

The paper also emphasizes practical utility: the model is intended for professional-grade generation workflows rather than isolated academic benchmarks.

Data Curation and Captioning

The data section argues that native audio-visual generation depends on the quality of the training corpus as much as on the model architecture. The team therefore constructs a multi-stage curation pipeline designed to preserve semantic continuity across both modalities and to bias the training distribution toward realistic, expressive audio-visual pairs.

Multi-stage audio-visual curation pipeline

  • Holistic audio-visual sourcing: The dataset is built from diverse scenes, subjects, styles, and motion patterns while also covering a wide range of audio types, including human speech in multiple languages, environmental ambience, mechanical and object sounds, and other audio events correlated with visual action.
  • Audio-aware temporal segmentation: Unlike conventional segmentation that follows only visual shot boundaries, the pipeline keeps boundaries only when visual discontinuities align with audio-level evidence such as pauses or low-activity intervals. This is intended to avoid cutting audio events in the middle and to preserve semantically meaningful clip units.
  • Multimodal quality filtering: Visual clips are filtered by fidelity, aesthetics, color consistency, resolution, composition, and motion richness. Audio is screened for silence, low signal-to-noise ratio, clipping, and other degradations. The system also uses an audio-visual synchronization estimator and removes clips with severe mismatches.
  • Joint distribution rebalancing: The authors identify long-tail imbalance across visual categories and audio categories. They rebalance sampling over the Cartesian product of those categories, upweighting rare combinations such as unusual actions paired with rare environmental acoustics.
  • Curriculum-based staging: The curated dataset is organized into stages that first teach basic cross-modal correspondence and then progressively introduce harder motion, more complex acoustics, multi-speaker scenes, and rapidly changing sound events.
Overview of training and inference pipeline.
Overview of training and inference pipeline.

Video captioning and audio caption augmentation

The captioning system is expanded beyond standard video descriptions. The paper says the captioner was enhanced to better describe camera movement, actions, aesthetics, shot transitions, and text recognition. For audio generation, captions also include audio descriptions covering dialogue, music, and sound effects. The captions are designed to be rich, professional-grade, and differentiated by domain and quality.

Captioning is performed with an in-house Seed Omni Model that has video understanding, text recognition, and audio recognition capabilities. The model is quantized for acceleration so that it can annotate very large-scale datasets efficiently.

Engineering infrastructure for large-scale data processing

The paper notes that the data pipeline must process multi-modal inputs from videos, images, audio, and text at very high throughput, with daily processing on the order of hundreds of millions. To support that scale, the infrastructure is designed for flexible resource usage, including unstable compute availability, preemptive resources, reserved time-window resources, and low-cost CPU resources. Asynchronous pipeline parallelism and automation are used to keep CPU and GPU resources busy while reducing manual intervention and waiting time.

Model Design

The model design is built around latent diffusion and a dual-branch multimodal transformer. The core idea is to keep the audio and video generation pathways specialized while still allowing deep cross-modal interaction through learned joint modules.

The proposed diffusion transformer model of .
The proposed diffusion transformer model of .

Latent representation with two VAEs

The architecture uses separate variational autoencoders for the two modalities. For video, the system reuses the pre-trained 3D video VAE from Seedance 1.0 to compress raw visual sequences into latent representations. For audio, the authors introduce a universal audio VAE that is intended to compress speech, sound effects, and music into a compact latent space. This keeps the diffusion model working in lower-dimensional latent space rather than directly on waveforms or pixel space.

Dual-branch Diffusion Transformer

The main generative backbone is an MMDiT-style diffusion transformer. The paper describes a dual-branch topology: the video branch is initialized from the pre-trained Seedance 1.0 Pro weights, while the audio branch is pre-trained from scratch. Cross-modal joint modules connect the two branches so that video tokens, audio tokens, and text conditioning can influence one another during joint training.

During training, the video DiT processes visual and textual inputs, while the audio DiT concurrently handles auditory and textual inputs. The joint modules act as a bridge that supports cross-modality information exchange. The paper explicitly states that all modules are trained.

Diffusion refiner

For high-resolution audio-video generation, the system uses a cascaded diffusion framework analogous to Seedance 1.0, with the key difference that the refiner also incorporates an audio signal input. This indicates that the model is not purely single-pass: the coarse generation stage is followed by a refinement stage for higher-resolution outputs.

Prompt engineering module

The paper introduces a Prompt Engineering model whose job is to rewrite user prompts into a caption-style format that the DiT generator can consume. It is trained on high-quality data from the continue-training and supervised fine-tuning stages, spanning multiple domains. The prompt model is also trained on manually annotated examples with richer expressiveness and narrative structure so that it can better interpret crude, underspecified prompts at test time. The paper reports that this rewriting capability is used for both text-to-video and image-to-video settings.

Training Procedure

The training pipeline has three stages: pre-training, supervised fine-tuning, and reinforcement learning from human feedback. The paper presents these as sequential stages in a single end-to-end system rather than independent experiments.

Stage Primary goal Data / supervision Notable technique
Pre-training Learn native joint audio-video generation and generalize across tasks Large-scale mixed-modality dataset Multi-task joint training on T2AV, I2AV, T2V, and I2V
SFT Align outputs with human preferences and high-fidelity targets High-quality audio-visual-text triplets across hundreds of categories Stepwise annealing and model merging across sub-domain experts
RLHF Optimize quality using reward models Human preference annotations and generated samples Three reward models: alignment, motion dynamics, aesthetics

Pre-training

Pre-training uses a multi-task joint objective over a large mixed-modality corpus. The model is trained to generalize across text-to-audio-video, image-to-audio-video, text-to-video, and image-to-video tasks. The main architectural point is that the cross-modal joint model is trained from the start to learn temporal synchronization and semantic consistency between audio and visual streams.

Supervised fine-tuning

After pre-training, the model is fine-tuned on carefully curated high-quality audio-visual-text triplets. The objective of SFT is to improve visual aesthetics, motion coherence, and audio fidelity. The dataset is organized across hundreds of categories that cover visual style, motion dynamics, and acoustic attributes.

The paper highlights two techniques in this stage. First, a stepwise annealing strategy gradually adjusts the learning rate and data distribution so the model can stabilize without losing text controllability. Second, separate expert models are trained on distinct sub-domains and then integrated using model merging, which is presented as a way to combine strengths from different specializations.

Human feedback alignment

RLHF is used to further align the generative model with human preference. The reward system contains three specialized reward models: an audio-video alignment model, a motion dynamics model, and an aesthetics model. The audio-video alignment reward model is built on the in-house Seed Vision-Language Omni Model, while the motion and aesthetics rewards follow the Seedance 1.0 setup.

The optimization target is to maximize the aggregate reward across these models over tasks such as T2V, T2VA, and I2VA. The paper also reports that RLHF pipeline infrastructure improvements yielded nearly a 3× increase in training speed.

Inference Acceleration and Systems Optimizations

Inference efficiency is a major design goal of the paper. The authors describe a multi-stage distillation framework that reduces the number of function evaluations required at generation time while trying to preserve output quality.

The distillation strategy builds on Seedance 1.0 and refines the TSCD pipeline from HyperSD by incorporating a meanflow-style distillation method. The paper says that multi-segment average-velocity transformations are applied to the original joint video-audio generation trajectories, which helps the model generate both modalities more efficiently.

A second distillation component is score-based and is inspired by RayFlow. Here the student model is trained to approximate the teacher’s score through an expected noise-consistency objective, which is described as improving low-NFE sampling reliability and reducing artifacts. The authors further add human-preference-guided adversarial training to reduce low-step artifacts. Human evaluation is said to confirm that the accelerated model remains on par with the original in video and audio quality.

Beyond model distillation, the systems stack includes kernel fusion, fine-grained mixed-precision quantization, adaptive sparsity, tailored hybrid parallelization for both the DiT and VAE, topology-aware compute-communication overlap, 8-bit communication co-designed with quantization, and hardware-aware asynchronous offloading for out-of-memory cases on heterogeneous GPUs. The combined effect is reported as a substantial reduction in runtime memory and communication overhead and an end-to-end acceleration of more than 10×.

Evaluation Framework

The evaluation section does not rely only on generic user-preference benchmarks. Instead, the team builds an internal multi-dimensional benchmark called SeedVideoBench 1.5, which expands the earlier SeedVideoBench 1.0 with more industry-specific scenarios and audio-aware evaluation criteria. Professional film directors helped codify the criteria, and experts from film production, cinematography, and design conducted human assessments.

Video Absolute Evaluation for Text-to-Video task.
Video Absolute Evaluation for Text-to-Video task.
Video Absolute Evaluation for Image-to-Video task.
Video Absolute Evaluation for Image-to-Video task.
Audio GSB Evaluation for Text-to-Video task.
Audio GSB Evaluation for Text-to-Video task.
Audio GSB Evaluation for Image-to-Video task.
Audio GSB Evaluation for Image-to-Video task.

SeedVideoBench 1.5: video dimension

The updated benchmark covers subjects, motion dynamics, interactions, camera movements, and application scenarios such as advertising, social media content, and short-form narrative content. The paper’s video metric suite includes motion quality and prompt following.

Motion quality is framed as the most immediate user-facing concern. The benchmark retains stability, physical plausibility, and temporal accuracy, but it places extra emphasis on video vividness. The vividness score is defined as a composite perceptual metric over action, camera movement, atmosphere, and emotion. The paper notes that some competing models improve apparent stability by slowing motion, but that this can reduce vividness and expressive quality.

Prompt following is updated to prioritize consistency with the user’s intent rather than only word-level instruction matching. In narrative scenarios, some creative flexibility is allowed as long as the central intention is preserved.

SeedVideoBench 1.5: audio dimension

The benchmark extends into a dedicated audio evaluation space. The primary categories are human voice types, human voice attributes, and non-speech audio.

  • Human voice types: speech, singing, and non-verbal vocalizations such as laughter.
  • Human voice attributes: timbre, accent, and emotional tone.
  • Non-speech audio: sound effects and music, labeled by source, acoustic properties, musical genre, and technical parameters.

For audio in I2V tasks, the generated sound is explicitly conditioned on the reference image to preserve semantic consistency and cross-modal coherence.

Audio metrics

  • Audio prompt following: fidelity to requested vocals, dialogue, and sound effects without semantic drift.
  • Audio quality: artifact level, spatial soundstage, timbre realism, and signal clarity.
  • Audio-visual synchronization: lip-to-speech alignment and the timing of sound effects with on-screen events.
  • Audio expressiveness: emotional fit, background music appropriateness, and contribution to atmospheric immersion.

Human evaluation protocol

The paper evaluates video with both absolute scores and pairwise Good-Same-Bad comparisons. The absolute protocol uses a 5-point Likert scale from 1, “Extremely Dissatisfied,” to 5, “Extremely Satisfied.” The baselines named for video are Kling 2.5, Kling 2.6, Veo 3.1, and Seedance 1.0 Pro. For audio, the comparative systems include Veo 3.1, Wan 2.5, Kling 2.6, and Sora 2.

Reported Results

The paper reports qualitative comparative gains rather than a detailed numeric table in the provided LaTeX excerpt. The main reported findings are:

  • Seedance 1.5 pro shows a significant improvement over Seedance 1.0 Pro across the internal benchmark.
  • In text-to-video generation, it is said to lead in instruction following, while remaining competitive in visual aesthetics and motion dynamics.
  • In image-to-video generation, it remains strong on the same core axes, suggesting the model preserves conditioning from reference imagery while still producing motion and audio.
  • For audio, the model is described as especially strong in Chinese-language speech generation, including dialogue, dialects, and monologues.
  • Audio-visual synchronization is highlighted as a particular strength, with improved alignment of speech and lip motion and better correspondence between sound effects and visual events than Veo 3.1 and Kling 2.6.
  • Compared with Sora 2, Seedance 1.5 pro is described as more balanced and controlled in emotional expressiveness, which the authors argue is beneficial for professional production settings.

The benchmark discussion also states that Seedance 1.5 pro demonstrates practical advantage in Chinese film production, short dramas, traditional performance content, and other narrative workflows that benefit from synchronized dialogue, camera choreography, and stable tone control.

Qualitative application examples

The paper includes several qualitative figures illustrating application-oriented scenarios. The LaTeX excerpt provides the figure identifiers rather than descriptive captions, so the summary below stays at the level stated in the text: Chinese-language content, stylized visual construction, dialect-rich dialogue, traditional stage performance, and close-up emotional continuity.

Paper figure 'img1'
Paper figure 'img1'
Paper figure 'img2'
Paper figure 'img2'
Paper figure 'img3'
Paper figure 'img3'
Paper figure 'img4'
Paper figure 'img4'

Limitations and Scope of the Report

The provided LaTeX does not include a formal limitations section, so the safest reading is that the report focuses on system design, training, and internal evaluation rather than exhaustive failure analysis. The most explicit caveat in the text is that mastery of specific vocal styles across different opera sub-genres is still evolving, even though the model already captures operatic cadence and stylized performance cues.

Another scope note is methodological: the excerpt emphasizes internal benchmarking and qualitative comparisons. It does not provide a full numeric ablation table in the supplied text, so the summary should be interpreted as a report of the authors’ stated design choices and reported comparative outcomes rather than an independently verifiable benchmark audit.

Takeaway for a Conversational-AI / Talking-Head Team

For a talking-head or conversational-AI workflow, the main technical implication of Seedance 1.5 pro is that the model is explicitly optimized for synchronized multimodal generation: speech content, accent and dialect control, lip motion, and sound cues are all treated jointly. The strongest parts of the report are the audio-aware data pipeline, the dual-branch audio-video transformer, the multi-dimensional reward setup, and the acceleration stack that makes the model more usable in production settings. The paper’s evaluation language suggests that the system is especially relevant for scenarios where faithful speech delivery, visible articulation, and cinematic motion need to be generated together.