Akapulu Labs logo Akapulu Labs Research

Streaming Avatars, Full-Duplex Dialogue, and Smarter Speech Synthesis

Today's digest spans expressive streaming avatars, real-time speech-to-speech translation, full-duplex dialogue memory, and content representations for voice synthesis — eight papers pushing the state of the art in conversational AI from lab to production.

Streaming Avatars, Full-Duplex Dialogue, and Smarter Speech Synthesis

How do we model physically grounded and articulatory-consistent facial motion from speech? (a) Five visible mesh-surface anchors as motion reference points: upper lip (UL), lower lip (LL), left mouth corner (LMC), right mouth corner (RMC), and Jaw. (b) From these anchors, we describe visible articulation with three directional articulatory motions: horizontal rounding/spreading, vertical opening/closing, and depth-wise protrusion/retraction. (c) For an utterance, the trajectories of these motions form temporally localized articulatory events that align with acoustic/phonemic cues. % (d) Modeling speech-driven motion via these directional articulatory motions yields phoneme-faithful 3D lip shapes with physically plausible dynamics. From Seoul National University.

A strong Monday batch lands across three fronts: making talking avatars more expressive and anatomically grounded, squeezing latency and memory out of real-time speech-to-speech pipelines, and sharpening the building blocks of voice synthesis. Here's everything worth reading today.

---

Talking Avatars & Lip Sync

From distillation diversity collapse to phonetically-structured 3D facial motion — two papers rethink how avatars articulate speech.

Knowledge distillation is essential for streaming avatars, but it notoriously flattens expression diversity. Southeast University diagnoses the problem spatially and temporally: not all regions and not all noise stages benefit from the same distillation objective. Their Routed Forcing method applies data-driven supervision to high-variance body regions at high noise stages — where expressiveness is most at risk — while keeping standard distillation for the mouth and background to protect lip-sync stability. The result is dynamic gestures and facial expression diversity recovered without sacrificing synchrony.

Southeast University

Southeast University · Sep 2026

Where and When to Force: Routed Forcing for Streaming Avatars

This paper addresses diversity collapse in distilled streaming avatars by routing the distillation objective spatially and temporally: applying data-driven supervision to high-variance person regions at high noise stages while retaining standard distillation for the mouth and background to preserve lip-sync and stability. The approach recovers dynamic gestures and expression diversity lost during knowledge distillation.

Abstract

Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes a reverse KL divergence, which is inherently mode-seeking: it causes the student to discard high-dynamic modes and collapse onto static outputs, compressing both dynamics and diversity of generated videos. We find that this collapse is region-heterogeneous: person regions involving pose and gesture variations suffer the largest diversity loss, the audio-driven mouth region shows a small loss, and the background remains nearly stable. Based on this observation, we propose Routed Forcing, which routes the distillation objective by semantic region and noise stage to improve dynamics and diversity while preserving visual quality. Specifically, (1) Where to Force: Semantic-Region Routing applies Data-Forcing Distillation (DFD), which supervises the student with real videos, to the person region where diversity collapse is most severe, while retaining DMD for the mouth and background to preserve lip synchronization and scene stability. (2) When to Force: Noise-Stage Routing activates DFD at high noise stages, where real video serves as effective supervision to inject diverse and dynamic motion patterns. At low noise stages, DMD is used to refine details, avoiding blur and artifacts from spatial differences between real video and student-generated video. Experiments show that Routed Forcing improves dynamics by up to 45% and diversity by 7-25% over Self Forcing, while preserving video quality and lip synchronization.

Where routed forcing operates at the diffusion-process level, Seoul National University's Seeing Speech goes deeper into the phonetics of facial movement. Rather than learning a direct audio-to-vertex mapping, they decompose articulation into three anatomically grounded axes — spreading, opening, and protrusion — each learned via phoneme-conditioned memory retrieval and then composed into surface-consistent 3D motion. The structured decomposition respects both anatomical and phonetic constraints, yielding superior lip-sync quality and articulatory realism over end-to-end baselines.

Seoul National University

Seoul National University · Sep 2026

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

This work proposes an articulation-aware framework that decomposes speech-driven facial animation into three directional articulatory motions (spreading, opening, protrusion) learned through phoneme-conditioned memory retrieval, then composes them into surface-consistent 3D motion. This structured approach respects the anatomical and phonetic constraints of speech, achieving superior lip-sync quality and articulatory realism compared to direct audio-to-vertex methods.

Abstract

Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.

---

Speech-to-Speech & Full-Duplex Dialogue

Four papers from Meta AI, Sony AI Research, Tsinghua, and again Meta tackle latency, memory, and natural turn-taking in real-time spoken dialogue systems.

Simultaneous speech-to-speech translation demands both low latency and high fidelity — a notoriously hard trade-off with fixed policies and scarce aligned data. Meta AI's All In Good Time framework attacks both sides: it synthesizes high-fidelity, causally-aligned training data and trains an adaptive translation policy that dynamically adjusts when to translate. Speaker identity is preserved throughout, and the system achieves state-of-the-art quality with significant latency reduction over prior fixed-policy approaches.

Meta AI

Meta AI · Sep 2026

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

This paper introduces a causality-aware framework for LLM-based simultaneous speech-to-speech translation that generates high-fidelity, causally-aligned training data and uses an adaptive translation policy to optimize quality-latency trade-offs. Unlike prior approaches relying on fixed policies and limited aligned data, this method dynamically adjusts translation timing while preserving speaker identity, achieving state-of-the-art results with significant latency reduction.

Abstract

Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.

Tandem speech-to-speech models — where an LLM backend begins composing a response while the user is still speaking — are powerful but expensive to train because simulating the backend at every training step is costly. Sony AI Research sidesteps this with randomized intermediate guidance: instead of simulating, they derive guidance directly from the conversation corpus using both real and randomly sampled backend responses. This teaches the speech frontend to selectively exploit backend information, leading to natural turn-taking and improved audio quality on real conversational data.

Sony AI Research

Sony AI Research · Sep 2026

Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

This paper proposes randomized intermediate guidance for training tandem speech-to-speech models where an LLM backend supplies responses during user utterances. Instead of simulating the backend during training (expensive overhead), guidance is derived directly from the conversation corpus using real and random responses. This teaches the speech frontend to selectively use backend information while achieving natural turn-taking and improved audio quality on real conversations.

Abstract

Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhead when training on real conversations. We propose randomized intermediate guidance, which derives guidance directly from the conversation corpus rather than simulating backend LLM behavior. During training, target responses provide informative guidance, while randomly sampled responses provide potentially irrelevant updates during the utterance. This combination aims to teach the frontend to use backend information selectively. On synthetic dialogues, KAME trained with this recipe achieves response quality comparable to that of the LLM-generated and similarity-based baselines. Training on 3.8k hours of real conversations improves smooth turn-taking and audio-judge naturalness over synthetic-data KAME while retaining a response-quality advantage over Moshi. These results show that randomized guidance offers a practical route to combining the response-quality benefits of tandem models with natural conversational behavior learned from real speech.

Continuous full-duplex models accumulate KV-cache memory rapidly because they must hold both incoming audio and outgoing speech representations simultaneously. Tsinghua University's Acoustic-to-Text KV Compression exploits a key insight: there are "listening-time slack" windows between audio processing chunks and human speech timing where the model can asynchronously transcribe incoming audio into compact text tokens. Trading acoustic KV entries for textual ones frees substantial memory, enabling longer conversations without degrading speaking or listening quality.

Tsinghua University

Tsinghua University · Sep 2026

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

This paper addresses memory constraints in continuous full-duplex speech models by proposing acoustic-to-text KV compression, which converts incoming speech to compact text tokens during "listening-time slack." The approach uniquely exploits gaps between audio processing and human speech timing to trade acoustic memory for textual representations, enabling longer conversations while maintaining speaking and listening quality.

Abstract

Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.

A different angle on spoken dialogue latency: Meta AI's RePlay avoids generation entirely for predictable responses by retrieving pre-recorded voice lines. The trick is knowing when to retrieve — they train hidden-state probes on a foundation model to detect the earliest point at which response content becomes recoverable from the model's internal state. This hybrid retrieve-then-play strategy achieves 3–7× lower latency than ASR–LLM cascade baselines while maintaining user-preferred dialogue quality, at the cost of exact-line accuracy.

Meta AI

Meta AI · Sep 2026

RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

RePlay retrieves and plays back pre-recorded voice lines for multi-turn dialogue by using hidden-state probing to identify when response content becomes recoverable in a foundation model. This hybrid approach achieves 3–7× lower latency than ASR–LLM cascades while maintaining user-preferred dialogue quality, trading exact-line accuracy for real-time responsiveness.

Abstract

Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and playing pre-recorded lines. Using probing, we identify the layer and frame at which the upcoming response becomes recoverable, and use this hidden state as the retrieval query. RePlay retains only the layers up to that point and replaces text and speech generation with lightweight turn-taking and retrieval heads. In simulated multi-turn interviews, RePlay reaches a median latency of 383 ms, 3 to 7 times lower than ASR-LLM cascades of comparable dialogue quality, at the cost of lower exact-line accuracy. In a user study, participants preferred RePlay in 63% of ratings versus 12% for a fast cascade with a small LLM (p = 0.008), and showed a non-significant preference (46% vs. 21%) over a slower cascade with a stronger LLM.

---

TTS & Voice Synthesis

A unified benchmark for content representations and a training-free pronunciation transcription method round out the day.

The choice of speech content representation — discrete tokens, continuous embeddings, self-supervised features — has an outsized effect on downstream synthesis quality, yet comparisons are rarely apples-to-apples. IRCAM addresses this with a comprehensive, unified evaluation framework: they train matched generative models conditioned on each representation and measure content fidelity, speaker identity, and prosody preservation across voice conversion, speech translation, and synthesis tasks. Their headline finding is that speaker disentanglement is not primarily a function of explicit supervision — it emerges from the interaction between training objectives and information capacity constraints.

IRCAM

IRCAM · Sep 2026

A Comprehensive Study of Content Representations for Speech Synthesis

This work provides a unified evaluation framework for comparing speech content representations across voice conversion, speech translation, and synthesis tasks by training generative models conditioned on each representation and measuring content, speaker identity, and prosody preservation. The key finding is that speaker disentanglement emerges from the interaction between training objectives and information capacity constraints, not from supervision alone.

Abstract

Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.

Pronunciation transcription sits at the intersection of text and acoustics, and the dominant approaches either rely on text-only G2P models or speech-only S2P models, both requiring substantial annotated data. The University of Tokyo proposes a training-free alternative: at inference time, constrained candidate pronunciations are generated from a frozen lexical model and then rescored greedily against a frozen acoustic model. The method combines the strengths of both paradigms without any task-specific training, achieving state-of-the-art character error rates with maintained computational efficiency.

University of Tokyo

University of Tokyo · Sep 2026

Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

This paper proposes a training-free pronunciation transcription method that integrates lexical and acoustic information at inference time using frozen pretrained models, eliminating the need for costly annotated training data. Unlike prior text-only (G2P) or speech-only (S2P) approaches, it achieves state-of-the-art results through constrained candidate generation and greedy acoustic rescoring, significantly reducing character error rates while maintaining computational efficiency.

Abstract

Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.