Akapulu Labs logo Akapulu Labs Research

Identity-Aware Avatars, Faster TTS, and Smarter Spoken Dialogue

Today's digest covers identity-preserving lip-sync animation, semantic-aware gesture generation, full-duplex multi-party speech models, emotionally grounded dialogue, and a wave of TTS efficiency advances from reward alignment to on-device speculative decoding.

Identity-Aware Avatars, Faster TTS, and Smarter Spoken Dialogue

Figure from From Bangxun Tang.

A strong end-of-September slate today: eight papers spanning talking-head animation, co-speech gestures, full-duplex spoken dialogue, and a cluster of TTS advances — including two from Meta's Hugging Face Daily picks. The through-line is a push toward specificity: individual identity, emotional nuance, semantic grounding, and deployment efficiency all get serious attention.

---

Talking Avatars & Lip Sync

Making animated faces look and move like the actual person they represent.

Generic mouth shapes are one of the most visible failure modes in audio-driven portrait animation. Rather than blending to an average mouth, this work extracts enrollment reference frames and HD mouth patches from the target subject, then trains a paired discriminator to enforce identity-specific lip shape, teeth arrangement, and texture — all while keeping audio synchronization tight and the surrounding face recognizable.

Bangxun Tang

Bangxun Tang · Sep 2026

Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation

This work tackles identity-preserving lip-sync animation by rendering each person's own mouth characteristics—including specific lip shape, teeth arrangement, and texture—rather than an average generic mouth. The key innovation is using enrollment reference frames and HD mouth patches with a paired discriminator to enforce identity-specific detail while maintaining tight audio synchronization and the recognizability of the rest of the face.

Abstract

We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person's reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person's own lip and dental detail while keeping synchronization and the rest of the face intact.

---

Co-Speech Gesture & Animation

Generating body language that reflects what the speaker is actually saying, not just how they're saying it.

Co-speech gesture systems have long leaned hard on acoustic rhythm, leaving semantic content as an afterthought. This paper from Communication University of China argues that semantics should be weighted proportionally to how much they actually contribute in a given segment. The proposed reliability-aware framework runs a dual-branch distribution comparison to estimate that contribution, then selectively amplifies semantic guidance in high-relevance windows — and degrades gracefully when semantic annotations are noisy or missing.

Communication University of China

Communication University of China · Sep 2026

When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation

This paper addresses a key limitation in co-speech gesture generation: over-reliance on acoustic rhythm while neglecting semantic content. It proposes a reliability-aware framework that estimates the actual contribution of textual semantics via dual-branch distribution comparison and selectively strengthens semantic guidance only in semantically-relevant segments, while robustly handling noisy or incomplete semantic annotations.

Abstract

Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.

---

SpeechLLMs & Spoken Dialogue

Scaling conversational speech models to handle the complexity of real human interaction.

Most full-duplex speech work is benchmarked on short, two-person exchanges. MultiTalk from the University of Hong Kong breaks that mold by tackling long-context, multi-party, bilingual (English + Chinese) conversation. The team built a synthetic data engine yielding 57.6k hours of training audio and introduced MultiTalkBench, a real-world evaluation suite that tests speaker tracking, topic coherence, and addressee selection across extended sessions — exactly the conditions where current models fall apart.

University of Hong Kong

University of Hong Kong · Sep 2026

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

This work extends full-duplex speech models to handle long-context, multi-party conversations in both English and Chinese—scaling beyond short dyadic interactions. The key contributions are a synthetic data engine producing 57.6k hours of diverse training data and MultiTalkBench, a real-world benchmark for evaluating speaker tracking, topic coherence, and addressee selection across extended multi-party sessions.

Abstract

End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{https://huggingface.co/datasets/MultiTalk/MultiTalkPT}{MultiTalkPT}$ and $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkFT}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkBench}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.

Empathetic dialogue requires understanding how something is said, not just what is said. The challenge is doing that reasoning without paying the latency and computational cost of explicit chain-of-thought generation. This Tsinghua work uses recurrent latent reasoning to internalize paralinguistic perception of acoustic cues, with a two-stage training regime that decouples reasoning from response planning. At inference time the model speaks directly, bypassing verbose intermediate steps while matching or beating chain-of-thought baselines on emotional dialogue understanding.

Tsinghua University

Tsinghua University · Sep 2026

Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

This work addresses empathetic spoken dialogue by using recurrent latent reasoning to ground paralinguistic perception in acoustic cues, avoiding the latency and inefficiency of explicit chain-of-thought generation. A two-stage training approach separates reasoning from response planning, enabling direct inference while maintaining or exceeding performance on emotional dialogue understanding.

Abstract

Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.

---

TTS & Voice Synthesis

Faster, cheaper, more expressive speech synthesis — with a strong showing from Meta today.

Alignment & Voice Cloning

Zero-shot voice cloning with discrete-diffusion TTS has a dirty secret: computing reverse-trajectory likelihoods for reward alignment is expensive and numerically brittle. RAWD-TTS from Yandex Research sidesteps the problem entirely, using group-relative advantages derived from recognition and speaker-identity rewards to guide token reconstruction without ever needing those likelihoods. The result is a cleaner alignment pipeline that optimizes content accuracy and speaker identity directly at the waveform level.

Yandex Research

Yandex Research · Sep 2026

RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning

This paper addresses alignment of discrete-diffusion TTS models for zero-shot voice cloning by optimizing both content accuracy and speaker identity at the waveform level. RAWD-TTS uses group-relative advantages from recognition and speaker rewards to guide token reconstruction, eliminating the need for reverse-trajectory likelihoods—a key simplification that sidesteps complexity in discrete diffusion.

Abstract

Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.

Emotional Speech

From Meta AI's Hugging Face Daily pick: EmoRES-TTS tackles emotional speech generation without retraining by decomposing emotion vectors into two components — a shared component (pushing away from neutral) and a residual component (steering toward the specific target emotion). These are applied independently, giving finer-grained control over emotional trajectory and significantly outperforming prior vector steering approaches on both emotion adherence and naturalness.

Meta AI

Meta AI · Sep 2026↑51 comment★ 1

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

EmoRES improves emotional speech synthesis by decomposing emotion vectors into shared (moving away from neutral) and residual (directing toward target emotion) components, then steering them independently without retraining. This training-free approach significantly outperforms prior vector steering methods in emotion adherence and naturalness.

Abstract

Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.

Distillation & Few-Step Synthesis

Flow-matching TTS models produce excellent audio but typically require many neural function evaluations. Local Flow-Map Distillation (LFMD) from Meta AI Research avoids the expensive teacher trajectory integration that plagues prior distillation methods, instead deriving a principled sampling schedule directly from teacher dynamics. The payoff: high-quality synthesis in as few as one step, with no external audio metrics needed during training.

Meta AI Research

Meta AI Research · Sep 2026

Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

This paper proposes Local Flow-Map Distillation (LFMD), a training-efficient approach for distilling flow-matching text-to-speech models into few-step synthesizers. By avoiding costly teacher trajectory integration and deriving a principled sampling schedule from teacher dynamics, it achieves high-quality speech synthesis with minimal neural function evaluations—enabling 1-step TTS without external audio metrics.

Abstract

Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.

On-Device Inference

Even with a fast model, autoregressive RVQ-based TTS can bottleneck on token throughput for streaming on-device use. Meta's RVQ Position Aware Speculative Decoding exploits the structured hierarchy of residual vector quantization codes to predict multiple codes simultaneously and verify them efficiently. The approach delivers a 2–2.2× speedup in token generation with negligible parameter overhead — enough to enable real-time streaming TTS on devices as constrained as an iPhone.

Meta

Meta · Sep 2026

RVQ Position Aware Speculative Decoding for On Device Text to Speech

This work optimizes autoregressive speech synthesis for on-device deployment through RVQ position-aware speculative decoding, which predicts multiple quantization codes simultaneously and verifies them efficiently. By exploiting the structured nature of residual vector quantization, it achieves 2–2.2× speedup in token generation with negligible parameter overhead, enabling real-time streaming TTS on resource-constrained devices like iPhones.

Abstract

Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at $5\times10^{-4}$ percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.