Akapulu Labs logo Akapulu Labs Research

Adaptive Avatars, Full-Duplex Agents, and Smarter Voice Synthesis

Today's digest spans adaptive 3D talking heads, a wave of full-duplex voice agent architectures, and advances in controllable TTS and speech translation — with highlights from both arXiv and the Hugging Face Daily tab.

Adaptive Avatars, Full-Duplex Agents, and Smarter Voice Synthesis

FlowAct-R2 couples proactive interaction with streaming multimodal reference generation. Our demonstrations cover entertainment streaming, live shopping, live vlogging, and video chatting. From ProAudience.

Today's papers push hard on the edges of real-time interaction: avatars that learn your conversational style on the fly, full-duplex voice agents that balance instant response with deep reasoning, and TTS systems that finally get the expressive and temporal details right. Here's everything worth reading from September 29, 2026.

Talking Avatars & Interactive Heads

From per-user adaptation to hour-scale streaming — avatar generation is moving well beyond static inference.

Most talking-head generators treat every user the same, baking motion priors into fixed weights at training time. EvolvingAvatar from USTC challenges that assumption with test-time training on audiovisual context accumulated during a live conversation — no motion labels required. It maintains persistent fast weights to absorb long-term conversational patterns (how this person moves their head, pauses, backchannel) alongside transient jaw adaptation that snaps to the current phoneme. The result is significantly tighter synchronization between speaking and listening motion as a conversation progresses, from Hugging Face Daily.

University of Science and Technology of China

University of Science and Technology of China · Sep 2026↑21 comment

EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold

This paper introduces a real-time 3D talking head generator that adapts to individual users during conversations through test-time training on audiovisual context, without requiring motion labels. Unlike existing fixed-parameter generators, it employs persistent fast weights to capture conversational patterns and transient jaw adaptation to respond to current speech, significantly improving synchronization between speaking and listening motion.

Abstract

Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.

Going beyond single-speaker head generation, FlowAct-R2 from ProAudience targets the full interactive streaming scenario — think live shopping hosts, vloggers, and real-time chat companions. It routes streaming multimodal references (audio, images, video clips) through a diffusion transformer for video synthesis while offloading behavior planning to a proactive agent that prepares reusable skill primitives offline, then deploys them autonomously in response to live audience signals. The decoupling avoids accumulated motion drift that plagues long autoregressive rollouts, enabling stable hour-scale streams.

ProAudience

ProAudience · Sep 2026

FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning

FlowAct-R2 generates interactive avatar videos in real-time by streaming multimodal references (audio, images, video) through a diffusion transformer while a separate proactive agent plans behaviors, prepares reusable skills offline, and responds autonomously to live audience interaction. The approach avoids accumulated motion drift and supports hour-scale interactive streaming across entertainment, shopping, vlogging, and video chatting.

Abstract

We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini reference-to-video backbone to accept rolling action prompts, streaming audio, and dynamically updated image, audio, and video references. Video-driven rotary positional embeddings align reference chunks with the generation timeline, while reference-plus-image conditioning and partially noised historical motion frames preserve appearance and avoid accumulated drift. Second, a Proactive Interaction Agent separates pre-online planning from online scheduling and response: it prepares a persona, a long-horizon agenda, and reusable multimodal skills in advance, then autonomously schedules behaviors, responds to audience input, and handles interruptions during a live session. FlowAct-R2 supports real-time 720p generation and hour-scale streaming across entertainment streaming, live shopping, video chatting, and live vlogging.

Full-Duplex Voice Agents & SpeechLLMs

Three papers this week tackle the same core tension: how do you keep a voice agent instantly responsive while still letting it think deeply or fetch external knowledge?

ByteDance's SALMONN-duo takes a dual-system approach directly inspired by fast/slow thinking. A fast, always-on full-duplex speech LLM handles the real-time audio stream, while an asynchronous backend LLM handles tool calls and deliberative reasoning. The critical piece is adaptive delegation — the real-time model learns to recognize when a query exceeds its in-context capacity and routes it to the backend, maintaining dialogue coherence even mid-reasoning.

ByteDance

ByteDance · Sep 2026

SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents

SALMONN-duo separates real-time voice interaction from deliberative reasoning by pairing a fast, always-on full-duplex speech LLM with an asynchronous backend LLM for tool use. The key innovation is adaptive delegation: the real-time system learns when to answer directly versus when to invoke the backend, maintaining responsiveness and dialogue coherence even during complex reasoning or tool execution.

Abstract

Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $τ$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $τ$-Voice with an acceptable increase in the delegation rate.

CharDuplex from Nanyang Technological University attacks a different gap: existing full-duplex models are responsive but characterless. CharDuplex introduces a dual-stream architecture for simultaneous listening and speaking, pairs it with automated character-conditioned data construction to inject persona into training, and refines the whole system through FDGym, an interactive reinforcement learning environment that rewards both responsiveness and persona consistency. The upshot is voice assistants that sound like someone, not just something.

Nanyang Technological University

Nanyang Technological University · Sep 2026

CharDuplex: Building Character-Consistent Full-Duplex Spoken Dialogue Models

CharDuplex enables full-duplex spoken dialogue systems that maintain consistent character personality throughout real-time conversations, addressing the gap between interactive speech models and coherent persona expression. The approach combines dual-stream architecture with automated character-conditioned data construction and interactive reinforcement learning (FDGym), enabling voice assistants that are both responsive and characteristically consistent.

Abstract

Full-duplex speech models are moving voice interaction beyond conventional turn-taking, yet natural conversation is shaped not only by when an agent speaks, but also by how it behaves as a conversational character. We present CharDuplex, a character-driven full-duplex speech model that combines real-time spoken interaction with persona-conditioned behavior. We first adapt GLM-4-Voice to an always-on dual-stream architecture and train the model for full-duplex conversation. Then a fully automated pipeline constructs character-conditioned dialogue data from open-source character descriptions for character-conditioned supervised fine-tuning. The model is further refined with the proposed FDGym, where an LLM-simulated user dynamically interacts with the model, enabling reinforcement learning over evolving multi-turn interactions. On SpeechRole-Eval, CharDuplex achieves the highest average score among the evaluated open-source models, while remaining competitive with closed-source systems. It also demonstrates competitive general speech intelligence and strong full-duplex interaction capabilities. CharDuplex demonstrates a practical training recipe for building full-duplex voice assistants that are not only interactive, but also character-consistent.

KAIST's Context Spanning tackles the knowledge boundary of end-to-end duplex models by building a communication framework that injects raw retrieved text from external LLM backends directly into the speech model's context — no compression, no lossy summarization. The framework keeps the retrieval path off the critical real-time loop while still allowing the speech model to reason over live retrieved evidence, a clean solution to grounding duplex conversation in up-to-date knowledge.

KAIST

KAIST · Sep 2026

Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends

This paper introduces a framework for injecting real-time external information directly into full-duplex speech models via integration with LLM backends, enabling the model to reason over raw retrieved text rather than compressed representations. Context Spanning demonstrates that external knowledge can be seamlessly incorporated into duplex dialogue systems while maintaining real-time constraints, enabling more informed and natural conversations.

Abstract

Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving them unable to access real-time information and tool execution. Furthermore, even when Large Language Models (LLM) retrieve information, many duplex speech models process it within a compressed latent space rather than in its raw text form, which can lead to information loss from compression. To address this issue, we propose Context Spanning, a framework for information injection between a full-duplex speech model and an external LLM backend via real-time chunked prefill. The injected frame is encoded in a single forward pass inside the real-time frame budget. It feeds the retrieved information to the speech model as-is, enabling it to reason over the information independently and generate responses. With this approach, our model achieves high performance on Full-Duplex benchmarks and strong results on Question Answering tasks, demonstrating its conversation potential. Context Spanning shows that external information can be injected directly into a duplex speech model, introducing a new simple and powerful mechanism for duplex systems.

TTS, Voice Synthesis & Speech Translation

Expressive control, multi-speaker coherence, temporal alignment for dubbing, and reactive listening — this section covers a lot of ground.

Non-verbal vocalizations like laughter and sighs are notoriously hard to control in TTS because they fall outside the phoneme-level supervision most models rely on. NVAlign attacks this with direct-gradient post-training optimization over a flow-matching TTS backbone: it uses an ASR-based reward model and gradient surrogates to align non-verbal output without breaking speech fidelity or speaker identity — a cleaner alternative to SFT-only fine-tuning that tends to over-smooth these subtle sounds.

Independent Researchers

Independent Researchers · Sep 2026

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech

NVAlign enables precise control of non-verbal vocalizations (laughter, sighs, etc.) in flow-matching text-to-speech through direct-gradient post-training optimization. Unlike prior SFT-only approaches, it leverages an ASR-based reward model and gradient surrogates to jointly adapt the TTS backbone while maintaining speech fidelity and speaker identity.

Abstract

While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-ASR model to serve as the reward model for post-training. A two-step gradient surrogate enables efficient reward backpropagation through the flow-matching sampler to jointly update the autoregressive backbone and acoustic flow head. Fidelity penalties and reference-velocity regularization help preserve speaker similarity and speech quality. Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines. These findings demonstrate that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS. Audio samples are available at https://nvalign.github.io/.

Long-form multi-speaker dialogue synthesis is another notoriously underserved scenario. The Script to Drama framework (featuring ControlEdit-TTS) reformulates the problem as critique-driven iterative refinement: a critic identifies prosody, rate, or loudness mismatches at utterance and scene granularity, then routes each issue to selective editing, targeted resynthesis, or timing adjustment — avoiding the cost and coherence loss of full regeneration. It's a practical agentic architecture that treats dialogue TTS as a feedback loop rather than a one-shot decoding problem.

Independent Researcher

Independent Researcher · Sep 2026

From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS

This work formulates multi-speaker dialogue TTS as critique-driven iterative refinement, where a unified speech model (ControlEdit-TTS) performs targeted attribute editing for emotion, speaking rate, and loudness without full regeneration. By routing detected issues to selective editing, resynthesis, or timing adjustment at utterance and scene levels, it achieves more coherent and controllable long-form dialogue synthesis than one-shot generation or direct dialogue models.

Abstract

Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.

Duration mismatch is the silent killer of dubbed speech — translated audio that runs long or short breaks lip sync and audience immersion. DuraS2ST from Microsoft Research Asia addresses this directly in speech-to-speech translation by first generating chain-of-thought rationales for phonetic length planning, then synthesizing speech tokens conditioned on those plans. A Duration Margin Reward in the RL training stage enforces temporal alignment without sacrificing translation quality — filling a gap no prior S2ST system explicitly optimized for.

Microsoft Research Asia

Microsoft Research Asia · Sep 2026

DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation

DuraS2ST addresses the critical problem of duration mismatch in speech-to-speech translation for dubbing by first generating explicit chain-of-thought rationales for phonetic length planning, then synthesizing speech tokens. The system is optimized through reinforcement learning with a novel Duration Margin Reward that balances translation quality with strict temporal alignment—a capability missing from existing S2ST systems.

Abstract

Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.

Reactive listening behavior — the nodding, humming, and micro-gestures a good listener produces — is the subject of REALM from Macquarie University, featured today on Hugging Face Daily. REALM proposes a coarse-to-fine generative framework for embodied reactive listening, generating temporally coherent listener motion that responds appropriately to the speaker's audio and motion cues. The coarse-to-fine design lets the model capture both global conversational rhythm and fine-grained moment-to-moment reactions.

Macquarie University

Macquarie University · Sep 2026↑161 comment★ 24

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

Abstract

Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ

Finally, from Hugging Face Daily with the most upvotes today, Duplex-MPE from Peking University addresses the conspicuous lack of rigorous benchmarking for multi-party full-duplex dialogue — scenarios with three or more simultaneous participants. Existing benchmarks are almost entirely two-party, leaving multi-party turn-taking, interruption, and floor management largely unmeasured. Duplex-MPE fills that gap with a structured evaluation suite designed to stress-test models on the full complexity of group conversation.

Peking University

Peking University · Sep 2026↑361 comment

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

Abstract

Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.