Akapulu Labs logo Akapulu Labs Research

Real-Time Avatars, Drift-Free TTS, and Empathetic Spoken Dialogue

Today's digest covers streaming talking-head reenactment, person-specific speaking style imitation, synchronized audio-video generation, flow-matching TTS fixes, expressive voice control, empathetic spoken dialogue reasoning, and LLM-based streaming ASR.

Real-Time Avatars, Drift-Free TTS, and Empathetic Spoken Dialogue

Additional cross-identity reenactment examples. Qualitative comparisons on the public (rows 1-4) and NeRSemble (rows 5-7) benchmarks against all evaluated baselines. PixReenact preserves the reference appearance while transferring the driver's facial motion. From Tel Aviv University.

Today's papers push hard on real-time constraints across the full stack — from pixel-level head reenactment and person-specific talking heads, to drift-free zero-shot TTS and streaming ASR. There's also a compelling entry from today's Hugging Face Daily tab on joint audio-video generation. Let's dig in.

Talking Avatars, Lip-Sync & Dubbing

Streaming, identity-preserving reenactment and authentic dubbing dominate this cluster, with methods tackling latency, style fidelity, and audiovisual alignment.

Real-time head reenactment has long struggled with the tension between low latency and maintaining reference identity over extended streams. PixReenact from Tel Aviv University attacks this directly by conditioning a causal video diffusion model on raw pixel-level representations — no specialized motion extractors required. Cross-identity pseudo-supervision and dual-teacher distillation keep the reference identity stable across long video sequences while faithfully transferring driving motion.

Tel Aviv University

Tel Aviv University · Oct 2026

PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment

PixReenact enables real-time head-avatar reenactment by conditioning a causal video diffusion model directly on pixel-level representations without specialized extractors, allowing continuous low-latency streaming. Unlike offline clip-based methods, it uses cross-identity pseudo-supervision and dual-teacher distillation to maintain reference identity over long videos while faithfully transferring driving motion.

Abstract

Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.

A complementary angle on realism comes from TalkLikeYou (Institute of Automation, Chinese Academy of Sciences), which targets person-specific speaking habits — the subtle, consistent articulation quirks that make each person's face uniquely recognizable in motion. Prior talking-head methods collapse these into uniform facial dynamics. Here, habits are modeled in motion space via Flow Matching with single-step inference for real-time performance, and a two-stage imitation learning strategy lets users dial in style via presets or reference videos.

Institute of Automation, Chinese Academy of Sciences

Institute of Automation, Chinese Academy of Sciences · Oct 2026

Talk Like You: Imitating How You Speak in Real-Time Talking Head Generation

TalkLikeYou captures and imitates person-specific speaking habits—subtle, consistent articulation patterns unique to each speaker—in real-time talking head generation. Unlike prior methods that produce uniform facial motions, this work models habits in motion space using efficient Flow Matching with single-step inference, and introduces a two-stage imitation learning strategy enabling users to control speaking style via presets or reference videos.

Abstract

In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou

Dubbing presents its own alignment nightmare: linguistic accuracy and lip synchronization are fundamentally at odds when the source and target languages have different phoneme timings. UltraDub (USTC) resolves this by using visual lip motion as a simultaneous guide for multimodal speech synthesis and a corrector for audiovisual alignment. The key mechanism is a training-free trajectory guidance technique that anchors flow corrections at visual rhythm cues, avoiding the need to retrain for each new language pair.

University of Science and Technology of China

University of Science and Technology of China · Oct 2026

UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance

UltraDub addresses video dubbing by leveraging visual lip motion to simultaneously guide multimodal speech synthesis and rectify audiovisual alignment. The approach uses a novel training-free trajectory guidance mechanism that anchors corrections at visual rhythm cues, resolving the fundamental tradeoff between linguistic accuracy and lip synchronization that plagues existing methods.

Abstract

Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.

From today's Hugging Face Daily tab, Kandinsky 6.0 Video (Kandinsky Lab) takes on the even more ambitious goal of generating synchronized audio and video jointly from text. The architecture uses dual-stream diffusion with bidirectional cross-attention between video and audio streams, built on top of a pretrained video model with a newly trained audio stream. Reinforcement learning post-training pushes speech quality and lip-sync to competitive levels at Full-HD resolution.

Kandinsky Lab

Kandinsky Lab · Oct 2026↑982 comments★ 1

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Kandinsky 6.0 Video generates synchronized text-to-audio-video through dual-stream diffusion with bidirectional cross-attention between video and audio, enabling coherent lip-sync and speech. The approach combines pretrained video with newly trained audio streams and employs reinforcement learning post-training to achieve competitive speech quality and synchronization at Full-HD resolution.

Abstract

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.

TTS & Voice Synthesis

Two papers refine the inference and training axes of expressive, controllable speech generation.

Zero-shot TTS with flow matching is powerful, but there's a subtle failure mode: the reference audio prompt drifts along the ODE trajectory during inference, degrading speaker similarity. Google's Prompt-Consistency Inference paper offers a clean, training-free fix — it analytically restores the reference audio trajectory before each velocity evaluation, preventing drift without adding computation. The result is measurable gains in speaker similarity and intelligibility across multiple TTS models.

Google

Google · Oct 2026

Prompt-Consistency Inference for Zero-Shot Flow-Matching Text-to-Speech Models

This paper improves zero-shot flow-matching text-to-speech by preventing prompt drift during inference: rather than letting the reference audio trajectory drift over ODE steps, a training-free technique analytically restores it before each velocity evaluation. This simple correction significantly improves speaker similarity and intelligibility across TTS models without added computation.

Abstract

In recent years, flow-matching models have produced significant improvements in zero-shot text-to-speech synthesis. Conditioned on an audio prompt and text, these models learn a velocity field and generate speech by iteratively solving an ODE. During inference, the solver evolves a single state spanning both the prompt and the region to be generated, although only the generated region is ultimately retained. The discarded prompt state, however, still matters; its intermediate values influence generation through the velocity field that couples the two regions. As sampling continues, this state can drift away from the prescribed conditional path, introducing a discrepancy into subsequent generation updates. Unlike the unknown generated trajectory, the prompt path is available in closed form from the reference audio and the initial noise. We exploit this observation with Prompt-Consistency Inference (PCI), a training-free rule that restores the prompt block to its analytic value before each velocity evaluation, while leaving the generated block unchanged. PCI improves speaker similarity and intelligibility across the evaluated flow-matching TTS backbones without additional network evaluations. Our ablation studies further show that PCI keeps post-step prompt discrepancies smaller and that corrections covering the later sampling stages recover much of the observed similarity gain.

On the controllability front, VoiceWeaver (Alibaba DAMO Academy) tackles the problem of adding multiple expressive dimensions — emotion, tone, sound events — to a shared text-audio model without catastrophic forgetting. The solution is staged training with replay-based distillation, combined with a structured label-prefix approach and sequential curriculum learning. Each new control attribute is introduced progressively, keeping prior capabilities intact.

Alibaba DAMO Academy

Alibaba DAMO Academy · Oct 2026

VoiceWeaver: Staged Learning of Structured Controls for Expressive Speech and Sound-Event Generation

VoiceWeaver learns expressive speech generation by progressively introducing emotion, tone, and sound-event controls through staged training, preventing catastrophic forgetting via replay-based distillation. A key innovation is the structured label-prefix approach combined with sequential curriculum learning, enabling a shared text-audio model to maintain multi-attribute control without interference across diverse expressive dimensions.

Abstract

Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a shared text-audio model. Replay-based distillation retains earlier predictions, while embedding decorrelation and attribute dropout regularize conditioning. Evaluation separates single-attribute correctness, joint emotion-event generation, and three-attribute correctness. Rounded to whole percentage points, emotion accuracy is 86% in Chinese and 83% in English. Chinese emotion-event joint accuracy is about 70%, versus 56% for mixed-task training and 48% for Ming-omni-tts; English joint accuracy is 66%. Three-attribute accuracy is 59% in Chinese and 55% in English on 500 samples each (57% pooled). Text error rates exceed those of the external TTS baselines. Audio samples are available at https://anonymous.4open.science/api/repo/VoiceWeaver1-4D88/file/index.html.

SpeechLLMs & Spoken Dialogue

Empathetic reasoning over acoustic signals and efficient streaming recognition round out the digest.

Most spoken dialogue systems map input to response in a single shot, losing the opportunity to explicitly model what the speaker is feeling and why. EchoChat (Microsoft Research Asia) breaks this into a structured cognitive pipeline: perception → mental-state inference → response generation, modeled as interdependent stages. Acoustic-grounded attention anchors each stage to the signal, and stage-aware error localization prevents errors from cascading downstream. The work ships with a 400K-sample training dataset and an expert-annotated evaluation benchmark.

Microsoft Research Asia

Microsoft Research Asia · Oct 2026

EchoChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue

EchoChat reformulates empathetic spoken dialogue as a structured cognitive reasoning process, moving beyond direct input-to-response mapping to model perception, mental-state inference, and response generation as interdependent stages. The framework uses acoustic-grounded attention and stage-aware error localization to mitigate cascading errors across reasoning steps, supported by a 400K-sample dataset and expert-annotated evaluation benchmark.

Abstract

Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to "superficially warm" but emotionally hollow interactions. In addition, since empathy relies on a multi-stage process with strong inter-step dependency, errors at any intermediate step can cascade through subsequent steps and lead to inappropriate responses, while existing training paradigms lack mechanisms to precisely localize and improve such errors. In this work, we propose EchoChat, a unified framework that reformulates empathetic spoken dialogue as a structured cognitive reasoning process integrating perception, mental-state reasoning, and response generation. To support this paradigm, we first construct EchoDialogue-400K, an acoustically rich dataset for multi-stage empathetic supervision. During the SFT stage, we strengthen acoustic grounding through proposed Acoustic-Anchored Attention (AAA). During the RL stage, we further introduce a novel stage-aware optimization objective with Step-Decomposed Credit Assignment (SDCA) to localize reasoning errors and mitigate cascaded error propagation. In addition, we introduce EchoEval, an expert-annotated benchmark for multi-dimensional empathy evaluation. Extensive experiments demonstrate that EchoChat achieves state-of-the-art performance in perception, reasoning, and response alignment. Project page: https://github.com/dingdongwang/EchoChat

Finally, LLM-based streaming ASR typically requires special tokens to signal when the model should wait for more audio — bloating the vocabulary and hurting perplexity. Factorized Delayed Streams Modeling (F-DSM) from Waseda University decouples the waiting probability from the text vocabulary entirely, eliminating ASR-specific tokens and reducing computational overhead. The factorization yields better recognition accuracy, lower memory usage, and less perplexity degradation relative to the baseline DSM approach.

Waseda University

Waseda University · Oct 2026

Factorized Delayed Streams Modeling for LLM-based Streaming ASR

This paper proposes Factorized Delayed Streams Modeling (F-DSM), which improves LLM-based streaming automatic speech recognition by separating the waiting probability from the standard text vocabulary, eliminating ASR-specific tokens and reducing computational costs. The factorization approach achieves better recognition performance, lower memory usage, and reduced perplexity degradation compared to the baseline DSM method.

Abstract

Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token <p> and the word-start token <w> to the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show that <w> can be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability for <p> from the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.