Akapulu Labs logo Akapulu Labs Research

Controllable Avatars, Faster Voice Cloning, and Full-Duplex Benchmarking

Today's digest covers controllable talking avatars with full-body pose guidance, transcript-free voice cloning 10× faster than autoregressive baselines, codebook-aware expressive TTS, one-step voice conversion, unified multilingual dubbing, and a new benchmark exposing timing failures in full-duplex speech systems.

Controllable Avatars, Faster Voice Cloning, and Full-Duplex Benchmarking

Qualitative results of MegaAvatar. Each row shows a reference image followed by two SMPL-X meshes and their corresponding generated frames. From Tongji University.

October 1st brings a strong cross-section of speech and avatar research: talking heads gain body awareness and listener intelligence, TTS systems shed autoregressive bottlenecks in favor of masked prediction and codebook routing, and a new benchmark reveals that even semantically correct full-duplex agents routinely fail at behavioral timing. Here's everything worth reading today.

Talking Avatars, Talking Heads & Listener Behavior

From full-body controllable avatars to single-image Gaussian reconstruction and audio-driven listener reactions — the talking-head space is widening its scope on every front.

Most talking avatar systems condition on audio and a reference image, leaving body pose as an afterthought. MegaAvatar from Tongji University addresses this directly by extending diffusion-based video generation with SMPL-X 3D skeletal guidance alongside audio and facial conditioning — giving explicit, decoupled control over full-body pose, head motion, and speech-synchronized expression while preserving speaker identity.

Tongji University

Tongji University · Sep 2026

MegaAvatar: Controllable Talking Avatar Generation

MegaAvatar extends diffusion-based video generation with 3D skeletal guidance (SMPL-X), audio, and facial conditioning to create controllable talking avatars. Unlike prior methods relying mainly on audio or reference images, it enables explicit control over full-body pose and head motion while preserving identity and maintaining speech-synchronized expressions.

Abstract

This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with previous talking-avatar methods that mainly rely on audio or reference-image conditioning, we introduce additional SMPL-X-derived 3D guidance, enabling global control over body pose and head motion. Specifically, we render the driving SMPL-X sequence into dense mesh frames and encode them with a lightweight 3D convolutional encoder, whose outputs are injected into the latent tokens to provide overall motion control. Furthermore, we extend Wan2.2-TI2V-5B with additional audio and face cross-attention modules to enable fine-grained expression control and preserve the input identity, respectively. In addition, we implement an audio-to-SMPL-X model to predict an SMPL-X sequence conditioned on the reference image and input audio, allowing MegaAvatar to support audio-driven inference without user-provided SMPL-X frames. Experiments show that MegaAvatar achieves high-quality talking avatar generation with controllable body and head motion, speech-synchronized facial expressions, and consistent identity preservation. MegaAvatar also supports inference with flexible resolutions and video lengths. Codes, dataset, models will be avaliable in https://github.com/Jeoyal/MegaAvatar

Reconstructing animatable head avatars from a single image is hard because so much of the head (hair, ears, back of skull) is simply never observed. SInGA from Michigan State University sidesteps this by performing semantic inpainting in UV space, exploiting facial symmetry and structured topology to hallucinate plausible completions for unobserved regions. Crucially, SInGA generalizes across identities without any per-identity optimization — a significant departure from prior multi-view pipelines — while still supporting realistic animation with improved identity preservation.

Michigan State University

Michigan State University · Sep 2026

Learning Semantic Inpainting for Animatable Gaussian Head Avatars

SInGA enables reconstruction of animatable head avatars from a single image through semantic inpainting in UV space, leveraging facial symmetry and structured topology to complete unobserved regions. Unlike prior multi-view methods, it generalizes across identities without per-identity optimization and supports realistic animation with improved identity preservation.

Abstract

We present SInGA, a novel method for learning Semantic Inpainting for animatable Gaussian head Avatars from a single image. Existing avatar approaches often rely on multi-view observations and lack effective handling of unobserved regions in single-view settings, limiting their applicability in such scenarios. To address this, we propose a semantic inpainting framework defined in UV space for completing unobserved facial regions. Our key insight lies in the structured topology of the UV representation, which provides consistent spatial correspondences and enables reliable completion of identity-specific features using the inherent symmetry cues of human faces. We extract features from observed regions and use them to complete unobserved regions. The completed representation is then used to regress Gaussian attributes, effectively performing Gaussian inpainting. In addition, instead of relying on a single Gaussian at each surface or pixel location, we stack multiple Gaussians to enhance detail. The resulting avatar generalizes across identities without requiring per-identity optimization and can be animated with driving inputs. Experimental results show that our method generates high-quality head avatars with improved completeness and identity preservation, while supporting realistic animation and consistent rendering from unobserved views.

The talking-head literature has long focused on the speaker, but real conversations require an attentive listener too. GLARE from UIUC tackles this gap by modeling dyadic listener behavior: given speaker audio and prosody, the system predicts when and how listeners should react — nodding, smiling, laughing, or frowning at contextually appropriate moments. The authors also contribute reaction-specific dataset annotations and evaluation metrics that measure behavioral appropriateness rather than pure image quality.

University of Illinois Urbana-Champaign

University of Illinois Urbana-Champaign · Sep 2026

GLARE: Generating Listening Heads with Appropriate Reactions

This paper generates natural listener behavior in dyadic conversations, predicting when and how listeners should react (nod, smile, laugh, frown) based on speaker audio and prosody. Unlike prior talking-head work that focuses on visual realism alone, GLARE combines audio-driven generation with reaction-specific dataset annotations and evaluation metrics that measure behavioral appropriateness rather than just image quality.

Abstract

While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.

TTS, Voice Cloning & Voice Conversion

Three papers push the efficiency and expressiveness frontier of speech synthesis and conversion, each targeting a different bottleneck.

Zero-shot voice cloning has been dominated by autoregressive token decoders, which are slow and require transcripts of the reference audio. Tacit-TTS replaces autoregressive decoding with masked prediction, delivering a reported 10× speedup and completely eliminating the transcript requirement. The transcript-free design unlocks cross-lingual cloning and conditioning on unconventional references — infant babble, synthetic speech — where transcript-dependent systems break down entirely.

Independent Researchers

Independent Researchers · Sep 2026

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning

Tacit-TTS achieves efficient zero-shot voice cloning by replacing autoregressive decoding with masked prediction, enabling 10x faster generation while eliminating the need for reference audio transcripts. This transcript-free approach uniquely enables cross-lingual voice cloning and robust conditioning on unconventional references like infant babble and synthetic speech, where transcript-dependent systems typically fail.

Abstract

TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.

Long-form expressive TTS suffers when prosodic instructions are applied uniformly across all codebook layers and all text spans. SCIC from Alibaba attacks this with a dual insight: different prosodic features concentrate in specific residual codebook layers, and instructions should be scoped to clause-relative targets rather than global document context. By routing instructions to the appropriate codebook layers at the right granularity, SCIC achieves fine-grained, speaker-adapted expressive synthesis with paragraph-level coherence.

Alibaba

Alibaba · Sep 2026

SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS

This paper enables precise prosodic control in long-form TTS by making instruction conditioning aware of both scope (clause-relative targets) and codebook structure (different prosodic features concentrate in specific residual codebook layers). Unlike prior uniform global conditions, SCIC strategically routes instructions to the appropriate codebook layers, achieving fine-grained speaker-adapted expressive synthesis with paragraph-level coherence.

Abstract

Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline Prosody Control, where each Pitch, Energy, or Speed instruction targets a clause relative to the preceding clause from the same speaker, while Pause uses an absolute duration interval. In codec-based TTS, Speed and Pause affect sequence length, whereas Pitch and Energy rely on residual codebooks. By analyzing Qwen3-TTS RVQ codebooks, we find that Energy concentrates in early residual codebooks, whereas Pitch accumulates across a deeper prefix. We therefore propose Scope- and Codebook-Aware Instruction Conditioning (SCIC), combining a Temporal Instruction Router for frame-level tag activation with Tag-Specific Codebook Weighting over residual codebooks. SCIC improves speaker-relative Pitch and Energy control over standard instruction fine-tuning using text-token tags. We further apply multi-reward GDPO post-training to jointly optimize control and quality, improving control accuracy while preserving CER and speaker similarity. In long-form synthesis, SCIC produces a more distinct paragraph-level expressive hierarchy than speaker-adapted SFT without instructions. Audio demos are available at: https://taoliveaigc.github.io/SCIC/

NTT Communication Science Laboratories' MeanVoiceFlow2 takes on zero-shot voice conversion speed. Rather than treating the flow-matching conversion module and content encoder as independent components, it jointly optimizes both, combining conversion distillation, real-data reconstruction, and diffusion-GAN training into a single pipeline. The result is 9× faster inference than its predecessor with no meaningful loss in speech quality or speaker similarity.

NTT Communication Science Laboratories

NTT Communication Science Laboratories · Sep 2026

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

MeanVoiceFlow2 jointly optimizes a flow-based voice conversion module with a computationally efficient content encoder, enabling fast zero-shot voice conversion. The key innovation is combining conversion distillation, real-data reconstruction, and diffusion-GAN training to achieve 9× faster inference than its predecessor while maintaining speech quality and speaker similarity.

Abstract

Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9\times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.

Speech-to-Speech, Dubbing & Voice Agents

A unified multilingual translation family and a new full-duplex benchmark round out today's digest, with implications for real-world deployed conversational systems.

Building separate models for text translation, speech translation, dubbing, and long-document translation is expensive and hard to maintain. Index-Translate from Bilibili unifies all four modalities under a shared foundation with task-specific fine-tuning paths at 2B–35B parameter scales. Notably, it matches or exceeds frontier models exceeding 100B parameters on diverse translation tasks, and its end-to-end speech-to-speech translation outperforms existing dedicated approaches.

Bilibili

Bilibili · Sep 2026

Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation

Index-Translate is a multilingual translation model family that unifies text, speech, dubbing, and long-document translation through a shared foundation with specialized task-specific training paths. It stands out by achieving performance comparable to frontier 100B+ models across diverse translation modalities while maintaining efficiency at 2B–35B scales, and includes end-to-end speech-to-speech translation capabilities that outperform existing approaches.

Abstract

We introduce Index-Translate, a multilingual translation model family that combines a shared multilingual foundation with specialized training for general translation, instruction following, speech translation, controlled dubbing, and long-document translation. It includes three model sizes, 2B, 9B, and 35B-A3B, and supports translation in 150 languages, with multilingual instruction following. Evaluations on general translation and complex translation instructions show that Index-Translate outperforms translation models of comparable size and achieves performance comparable to 100B-scale translation models and frontier models. Index-Echo provides end-to-end speech-to-text and speech-to-speech translation, outperforming existing end-to-end models and achieving performance comparable to frontier omni models. Index-Homura extends the family to syllable-controlled dubbing. Index-NativeLong introduces native long-document translation with a dedicated task formulation and benchmark. These capabilities support diverse translation tasks, including multilingual content production.

"Current systems often achieve semantic correctness while failing at behavioral timing."

Evaluating full-duplex conversational agents is notoriously difficult because most benchmarks measure what a system says, not when it acts. DuplexAct-Bench from Tsinghua University fills this gap with a bilingual benchmark covering six interaction behaviors — interruption, yielding, proactive initiation, backchanneling, and more — under varied conversational contexts. The benchmark exposes a critical and previously underreported gap: systems that score well semantically routinely fail at the real-time behavioral timing that makes conversations feel natural.

Tsinghua University

Tsinghua University · Sep 2026

DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements

DuplexAct-Bench is a bilingual benchmark for evaluating full-duplex conversational systems across six interaction behaviors (interruption, yielding, proactive initiation, backchanneling) under varied contexts. It reveals that current systems often achieve semantic correctness while failing at behavioral timing—exposing a critical gap in real-time interaction robustness that prior benchmarks miss.

Abstract

Existing full-duplex speech benchmarks cover only subsets of real-time interaction behaviors, often under limited contextual conditions. We introduce DuplexAct-Bench, a bilingual benchmark that systematically covers six complementary behaviors, from interruption and yielding to proactive initiation, active silence, and backchanneling, across Pre-session, In-session, and No-explicit conditions. Across 1,290 English and Chinese streaming trials, we evaluate 12 full-duplex speech systems on both Timing and Content. Results reveal substantial variation across behaviors, conditions, and systems, as well as frequent mismatches between semantic quality and behavioral timing. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds. Project page: https://alitaxky.icu/DuplexAct-Bench/