Akapulu Labs logo Akapulu Labs Research

Real-Time Avatars, Latent Reasoning, and Physically Grounded Speech

Today's digest spans real-time 3D avatar animation, lip-sync evaluation, latent-space reasoning for speech LLMs, voice agent grounding benchmarks, encoder-free speech understanding, articulatory TTS, and training-free voice conversion.

Real-Time Avatars, Latent Reasoning, and Physically Grounded Speech

GALA enables efficient, high-fidelity animation across diverse Gaussian avatar models. By introducing shared blendshapes and a lightweight coefficient predictor, GALA accelerates CPU animation by up to three orders of magnitude and enables real-time deployment on mobile devices. From Inria.

Today's digest is a wide-ranging look at the state of conversational and audiovisual AI — from 1000× faster avatar animation to reasoning that happens entirely in latent space, and from physically interpretable TTS to training-free voice conversion. Eight papers across labs from Inria to CMU to Tencent Hunyuan push the boundaries of what real-time, grounded, and controllable speech systems can do.

Talking Avatars, Lip-Sync & Dubbing

Animating faces at scale, evaluating dubbing quality with LLM reasoning, and resolving the ambiguity of silent lips.

Real-time photorealistic avatar animation has long been bottlenecked by costly neural decoding at inference time. GALA from Inria attacks this directly by distilling 3D Gaussian avatar deformations into a compact set of identity-independent blendshape basis vectors, then training a shallow coefficient predictor to replace the heavy neural decoder. The result is up to 1000× speedup in animation inference — fast enough for mobile deployment — while preserving rendering fidelity across both facial and full-body avatars.

Inria

Inria · Oct 2026

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

This paper introduces GALA, a distillation method that accelerates 3D Gaussian avatar animation by replacing costly neural decoding with lightweight blendshape-based linear approximation. By decomposing avatar deformations into identity-independent basis vectors and training a shallow coefficient predictor, the method achieves up to 1000× speedup in animation inference while maintaining rendering fidelity, enabling real-time performance on mobile devices across facial and full-body avatars.

Abstract

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/

Automated dubbing quality evaluation is notoriously hard: existing visual speech recognizers conflate semantic errors with temporal misalignments. The Align Then Reason paper from University of Maryland, College Park (featured on today's Hugging Face Daily tab) introduces a two-stage lip-sync judge that first establishes monotonic alignment between video frames and phonetic units, then feeds that structured representation to an LLM for joint content-and-timing reasoning. By explicitly decoupling alignment from judgment, it produces calibrated scores for both what was said and when.

University of Maryland, College Park

University of Maryland, College Park · Sep 2026↑41 comment

Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

This paper introduces a two-stage lip-sync judge for dubbing that first establishes monotonic alignment between video frames and phonetic units, then uses an LLM to reason jointly about content and timing. Unlike prior visual speech recognizers that struggle with temporal errors, this approach explicitly separates frame-to-phoneme alignment from final judgment, enabling calibrated scoring of both semantic and timing correctness.

Abstract

Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce $\textit{Align Then Reason}$ (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.

Silent talking-face video can correspond to many plausible utterances, making video-to-speech synthesis inherently underdetermined. WYS (Watch Your Speech) from Seoul National University resolves this ambiguity by conditioning speech generation on textual input alongside visual lip movements, fusing both modalities through an attention-based embedding mechanism. The system achieves state-of-the-art audio-visual synchronization and phonetic accuracy on standard benchmarks, with near-human-quality speech output.

Seoul National University

Seoul National University · Oct 2026

Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning

This paper addresses the ambiguity in video-to-speech synthesis—where silent talking faces could produce multiple plausible utterances—by conditioning speech generation on textual input. WYS fuses visual lip movements and textual context through an attention-based embedding mechanism, achieving state-of-the-art audio-visual synchronization and phonetic accuracy on standard benchmarks while generating near-human-quality speech.

Abstract

Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: https://github.com/gunwoo5034/Watch-your-Speech

SpeechLLMs & Voice Agents

Latent-space chain-of-thought, grounding fidelity under conversational load, and encoder-free speaker-aware pretraining.

Chain-of-thought reasoning dramatically improves LLM quality but adds unacceptable latency in real-time speech interaction. AURAL from Tencent Hunyuan sidesteps this tradeoff by moving the reasoning process entirely into latent space and predicting multiple future states jointly in a single forward pass. The result is CoT-level intelligence with dramatically reduced response latency — and because computation is adaptive, harder queries get more hidden computation without explicit token overhead.

Tencent Hunyuan

Tencent Hunyuan · Oct 2026

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

AURAL enables speech language models to achieve chain-of-thought reasoning quality while drastically reducing response latency by moving reasoning into latent space and predicting multiple future states jointly. This solves the fundamental tradeoff between model intelligence and fast interaction through adaptive, hidden computation that scales with problem difficulty.

Abstract

Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.

How well do current voice agents stay grounded to a reference document across a long conversation? DuplexSpeechBench-Document Grounding from the University of Maryland introduces a benchmark specifically targeting three failure modes: context saturation (degradation as context grows), grounding decay (drift from source facts over turns), and proactive grounding (failing to volunteer relevant document content). The benchmark exposes systematic weaknesses across cascaded, proprietary full-duplex, and open-weight speech systems alike.

University of Maryland

University of Maryland · Sep 2026

DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents

This paper introduces a benchmark for evaluating document grounding in voice agents across multi-turn conversation. The benchmark targets three failure modes—context saturation, grounding decay, and proactive grounding—and reveals how grounding fidelity degrades under increasing context and conversational load across cascaded, proprietary full-duplex, and open-weight speech systems.

Abstract

Voice agents enable low-latency, natural interaction, yet their ability to faithfully ground responses in external documents remains underexplored. We introduce DuplexSpeechBench-Document Grounding (DSB-DG), a benchmark for evaluating document grounding in voice agents across five professional domains. DSB-DG targets three failure modes: Context Saturation, which measures grounding under increasing document length; Grounding Decay, which measures retention of document facts across multi-turn dialogue; and Proactive Grounding, which evaluates whether context re-injection mitigates conversational drift. The benchmark contains 1,636 adversarially verified QA pairs from 50 documents covering five professional domains, and supports fully automatic evaluation of grounding accuracy, hallucination, and response latency. Across systems spanning cascaded, proprietary full-duplex and real-time, and open-weight speech2speech architectures, we find substantial differences in effective grounding capacity. While cascaded pipeline (ASR-LLM-TTS) achieves the highest grounding accuracy, Gemini-Live and GPT-Realtime closely trail behind. Open-weight systems exhibit distinct failure modes, most notably an abrupt context-capacity collapse and multi-turn grounding decay. More broadly, grounding fidelity degrades with context and conversational load, and failures frequently manifest as unsupported generations rather than abstention. We show that contextual grounding as a key unresolved challenge for reliable full-duplex voice agents.

Encoder-based Speech-LLMs rely on large pretrained acoustic encoders that can discard paralinguistic information. Microsoft Research shows that encoder-free Speech-LLMs — which map raw acoustic features directly to LLM embeddings — can match or exceed encoder-based counterparts when trained with metadata-supervised pretraining using speaker identity and emotion labels. Crucially, this approach preserves fine-grained acoustic cues needed for speaker discrimination in multi-speaker conversational settings, without any external encoder dependency.

Microsoft Research

Microsoft Research · Oct 2026

Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs

This work demonstrates that encoder-free Speech-LLMs—which directly map acoustic features to language model embeddings—can match or exceed encoder-based models when trained with metadata-supervised pretraining leveraging speaker identity and emotion labels. By eliminating dependence on large pretrained speech encoders, the approach preserves fine-grained acoustic cues essential for speaker discrimination and paralinguistic understanding in multi-speaker conversational scenarios.

Abstract

Encoder-based speech large language models (Speech-LLMs) commonly employ pretrained speech encoders that prioritize linguistic content but may discard fine-grained acoustic cues essential for speaker discrimination and paralinguistic understanding. Encoder-free Speech-LLMs instead map Mel-spectrogram features directly into the LLM input space through lightweight embedding layers, enabling the LLM to learn from low-level acoustic features. However, systematic pretraining strategies for encoder-free Speech-LLMs remain underexplored, limiting their ability to compensate for the absence of large-scale pretrained speech encoders. We propose metadata-supervised pretraining (MSP), which leverages speech attributes such as speaker identity and emotion to develop speaker-discriminative and paralinguistic capabilities. We further introduce speaker-aware utterance composition (SAUC) to strengthen speaker discrimination and apply random span masking to regularize pretraining. We primarily evaluate our approach on joint ASR and speaker diarization in multi-speaker conversations, complemented by experiments on paralinguistic speech-understanding tasks. Under matched training-data conditions, our encoder-free model outperforms its randomly initialized encoder-based counterpart. With limited metadata-annotated data, it is competitive with models using speech encoders pretrained on substantially larger corpora, outperforming them in several settings. These results demonstrate the potential of encoder-free architectures for building native multimodal LLMs that acquire diverse speech capabilities.

TTS & Voice Conversion

Interpretable articulatory control and training-free one-shot conversion without frame alignment.

Neural TTS systems are powerful but opaque. CMU's Articulatory Source-Filter TTS grounds speech synthesis in vocal tract kinematics by decomposing speech into a glottal source and a vocal-tract filter — the classic source-filter model, now learned end-to-end. This decomposition enables direct, interpretable manipulation of speech characteristics: for instance, modifying accent by swapping vocal-tract parameters while preserving speaker identity and maintaining stability across source/filter recombinations.

Carnegie Mellon University

Carnegie Mellon University · Sep 2026

Articulatory Source-Filter TTS: Physically Grounded Control through Vocal Tract Kinematics

This work introduces a controllable TTS system grounded in articulatory kinematics, decomposing speech into interpretable source (glottal) and filter (vocal-tract) components. Unlike black-box neural systems, it enables direct manipulation of speech characteristics—such as accent modification through vocal-tract control—while maintaining speaker identity and stability across source/filter recombination.

Abstract

Modern neural text-to-speech (TTS) systems achieve remarkable acoustic fidelity but act as black boxes, offering little interpretable control over the vocal tract filter. We propose a controllable source-filter TTS architecture grounded in articulatory kinematics. An Acoustic-to-Articulatory Inversion (AAI) model, enhanced by large-scale pretrained representations, generates kinematic pseudo-trajectories for a large TTS corpus. These trajectories condition the filter response, while predicted pitch and energy contours parameterise the glottal source. Source and filter are predicted independently, and the source is refined by an Optimal Transport Conditional Flow Matching (OT-CFM) module before recombination into the final spectrogram. Our model achieves intelligibility and naturalness competitive with similarly sized baselines, with only a modest spectral fidelity cost in exchange for explicit control. Evaluations reveal clear source-filter disentanglement, enabling stable prosodic scaling and cross-speaker source/filter recombination, where F0 remains tied to the source speaker while vocal-tract characteristics follow the filter speaker. Finally, direct manipulation of articulatory trajectories enables fine-grained control, such as accent modification, offering a new direction for interpretable speech synthesis. Audio samples: https://coding-phoenix-12.github.io/ArticulatorySFTTS/

Most voice conversion systems either require training on target speakers or depend on strict frame-level alignment between source and reference utterances. StateVC from EURECOM sidesteps both requirements with a training-free one-shot approach: it jointly discovers shared acoustic states across source and reference via Gaussian mixture modeling, enabling conversion even when source and reference carry entirely different linguistic content — no parallel data or alignment needed.

EURECOM

EURECOM · Oct 2026

Shared-State Local Translations for Training-Free Voice Conversion

StateVC proposes a training-free one-shot voice conversion method that avoids explicit frame-level alignment by jointly discovering shared acoustic states across source and reference utterances via Gaussian mixture modeling. Unlike prior work requiring strict frame correspondence, this approach enables conversion even when source and reference have different linguistic content.

Abstract

In one-shot training-free voice conversion (VC), the source and reference utterances may contain different linguistic content, so reliable frame-level correspondence between them cannot be assumed. We propose StateVC, which jointly defines a common set of local regions from pooled frame-level WavLM representations of the source and reference utterances; we refer to these regions as states. These shared states are obtained by fitting a pair-specific Gaussian mixture model to the pooled representations, without explicit source--reference frame matching. Within each state, StateVC estimates a source-to-reference mean shift in the original WavLM space. Source-frame posterior probabilities then combine the state-specific shifts so that different frames can receive different local updates. For the LibriSpeech one-shot protocol, StateVC achieves the lowest word error rate (WER) and character error rate (CER) among the evaluated systems, at 8.01% and 3.22%, respectively, with a speaker similarity (SIM) of 0.9512. It also achieves the highest mean perceived speaker similarity among the evaluated systems and the highest mean naturalness among the evaluated training-free systems.