Akapulu Labs logo Akapulu Labs Research

Real-Time Avatars, Steerable Duplex Dialogue, and Expressive TTS

Today's digest covers real-time Gaussian head animation, single-pass talking-head GANs, fine-grained TTS voice control, and a pair of papers pushing full-duplex spoken dialogue toward controllability and speed — plus new evaluation frameworks for co-speech gestures and 3D head avatars.

Real-Time Avatars, Steerable Duplex Dialogue, and Expressive TTS

Streaming audio-to-expression with latency-free acausal noise shaping. Top: audio streams through our method called FaceGAN, which emits every frame in a single forward pass at 25 Hz with 40 ms of audio lookahead; the filmstrip shows rendered output. Bottom: the same causal generator driven by causal noise (left) produces jittery, muted motion, while our acausally shaped noise (right) yields smooth, expressive motion. Bottom traces are illustrative schematics. From Meta Reality Labs.

A strong theme runs through today's eight papers: the relentless push to make conversational AI real-time and controllable — whether that means animating a photorealistic head at interactive frame rates, steering a full-duplex voice agent mid-conversation, or planning expressive prosody that actually matches what a speech synthesizer will produce. Labs from Apple, Meta Reality Labs, Google Research, Tsinghua, NUS, Amsterdam, and UIUC all weigh in.

---

Talking Avatars, Lip Sync & Audio-Driven Faces

From Gaussian splatting to single-pass GANs, the talking-head field is converging on real-time inference without sacrificing quality.

Head animation has long been bottlenecked by the cost of deformable neural rendering. Apple's GHARP attacks this by splitting the problem into two stages: an offline identity stage that builds a 3D Gaussian representation from just a few input images, and a lightweight online stage that predicts only the expression-dependent deltas at runtime. A novel body alignment network handles the otherwise ill-posed variations in torso pose and clothing, and the whole system runs 13× faster than prior work while hitting state-of-the-art benchmark fidelity.

Apple

Apple · Oct 2026

GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior

GHARP decouples real-time head animation into an offline identity stage that builds a Gaussian representation from few input images, and a lightweight online animation stage that predicts expression-dependent changes. A novel body alignment network resolves the ill-posed problem of unmotivated body pose and clothing variations, achieving state-of-the-art fidelity on benchmark while running 13x faster than prior work.

Abstract

We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.

Meta Reality Labs takes a complementary approach to real-time talking heads, asking whether diffusion models are even necessary. Their answer is no — at least for streaming use cases. Rather than distilling a multi-step diffusion pipeline, they train a single-pass GAN and solve the temporal consistency problem through acausal noise shaping, which lets the model look slightly ahead to stabilize latent noise without breaking causality for streaming. The result matches or exceeds diffusion-based SOTA while supporting indefinite streaming without temporal drift.

Meta Reality Labs

Meta Reality Labs · Oct 2026

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

This work replaces expensive multi-step diffusion models with a single-pass GAN for real-time audio-driven facial animation, solving the temporal stochasticity problem through acausal noise shaping while maintaining causal operation. The approach matches or exceeds state-of-the-art generation quality while enabling indefinite streaming without drift at interactive frame rates.

Abstract

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents. And it must run online: each frame emitted from audio observed up to the current time, at interactive rates. Recent progress is dominated by diffusion models, which need many network evaluations per sample and are therefore a poor fit for streaming. We argue the cost is unnecessary in this domain. Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices. The obstacle is not capacity but stochastic structure. We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation. We proposed FaceGAN, which dissolved the limitation by shaping the noise pathway acausally. Because the driving noise is synthetic, its future can be sampled now, so the audio-to-expression path stays causal, and the model supports fully causal operation. FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality. Being feed-forward with bounded attention windows, it generates indefinitely without drift.

Evaluation is the unglamorous but essential companion to generation. The University of Amsterdam presents a holistic evaluation framework for co-speech gesture generation that combines standardized objective metrics with LLM-augmented semantic annotations and a perceptual study involving 101 participants. The headline finding is sobering: individual objective metrics frequently disagree with human judgments, but composite metrics assembled from complementary signals close much of that gap.

University of Amsterdam

University of Amsterdam · Oct 2026

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

This paper presents a comprehensive evaluation framework for co-speech gesture generation that combines standardized objective metrics with human-centered validation and LLM-augmented semantic annotations. The key innovation is demonstrating through perceptual studies with 101 participants that individual objective metrics often fail to reflect human judgment, while composite metrics built from complementary signals achieve significantly better alignment with subjective assessments.

Abstract

Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify. We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation. We first curate a list of 13 objective metrics covering different aspects, including distributional similarity, geometric fidelity, kinematic quality, cross-modal synchrony, and semantic appropriateness. For the semantic-appropriateness category, we propose a new metric, Semantic Gesture Preservation (SGP), which measures how far semantic gestures in the ground truth are preserved in the generated gestures. For this, we augment the BEAT2 dataset's annotations using a multi-modal LLM. We then conduct a perceptual study where 101 participants score generated gestures among five dimensions, including human-likeness, motion diversity, absence of animation errors, speech timing and content match. We systematically analyze objective metric--subjective score correlations. Unlike Semantic Score (SC), which shows no significant association with the evaluated perceptual dimensions, SGP is selectively aligned with speech-aware human judgments. We construct five target-specific composite metrics aligned with the subjective dimensions. These composites improve perceptual alignment across all five dimensions, with the largest gains for absence of animation errors and content match, indicating that complementary objective signals can better approximate human judgments than individual metrics alone. Overall, our results show that objective metrics require validation against subjective evaluations.

---

Digital Humans & 3D Head Avatar Reconstruction

Fine detail in neural head avatars requires taming the optimization trajectory, not just adding capacity.

Fitting high-fidelity head avatars from multi-view video is notoriously unstable — models tend to over-fit high-frequency texture details before the coarse geometry has converged. Tsinghua's Fresco++ addresses this with a progressive frequency curriculum that gates the spatial frequencies available during optimization, ensuring coarse structure settles first. A Canonical Group Consensus mechanism then enforces cross-view consistency by associating observations through shared canonical surface regions, eliminating the floating artifacts common in prior methods.

Tsinghua University

Tsinghua University · Oct 2026

Fresco++: Frequency-Guided and Canonical-Consistent Optimization for Fine-Grained Head Avatar Modeling

Fresco++ optimizes fine-grained head avatars by addressing premature high-frequency fitting and cross-view inconsistency through a progressive frequency curriculum and Canonical Group Consensus. The method stabilizes coarse structures before refining details, while enforcing multi-view consistency by associating observations through shared canonical surface regions.

Abstract

We propose Fresco++, a unified optimization framework for fine-grained and view-consistent head avatar reconstruction. Head avatar optimization is typically driven by per-view image supervision, which can lead to premature fitting of unstable high-frequency details and inconsistent local appearance across viewpoints. Fresco++ addresses these challenges by regulating both the progression of visual detail and the formation of cross-view supervision during optimization. For frequency-aware optimization, a progressive curriculum first stabilizes low-frequency structures and then introduces high-frequency constraints to recover fine facial and hair details without amplifying spurious responses at early stages. For cross-view optimization, we introduce Canonical Group Consensus, which associates local observations through shared canonical surface regions and establishes correspondence across different viewpoints. Geometric and visibility-aware screening removes unreliable observations, while the remaining multi-view evidence is aggregated in feature space to form a consensus target for supervising the current rendering. This design enforces local consistency without relying on a specific image-space parameterization and avoids additional rendering of the auxiliary view. Together, the frequency curriculum and canonical consensus provide stable optimization from coarse structures to fine details while maintaining coherent appearance across viewpoints. Extensive experiments on NeRSemble demonstrate improved reconstruction quality and cross-view consistency, while evaluations across diverse avatar representations further confirm the generality and transferability of Fresco++.

---

TTS, Voice Synthesis & Expressive Control

Two papers push TTS beyond static voice cloning toward instruction-driven timbre editing and reward-optimized prosody planning.

Most TTS systems treat voice identity and expressive delivery as separate concerns. EDICT from NUS unifies them: users can first modify a reference speaker's global timbre via natural-language instructions, then apply segment-level expressive control to individual utterances — all while an edited acoustic reference anchor preserves speaker consistency across the full output. This is a meaningful step toward conversational agents whose voice can be designed and dynamically inflected without re-training.

National University of Singapore

National University of Singapore · Oct 2026

Edit Who Speaks, Control How They Speak: Global Timbre Editing and Local Instruction Control for TTS

EDICT unifies global voice timbre editing with segment-level expressive control for text-to-speech synthesis. Unlike prior work that handles voice cloning or descriptive voice design separately, it allows users to modify a reference voice's timbre through instructions, then apply fine-grained delivery control to individual speech segments while maintaining speaker consistency through an edited acoustic reference anchor.

Abstract

Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce \textbf{EDICT}, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality. Audio demos are available.

Conversational TTS faces a subtler problem: the text-based style planners used to pick prosody labels are trained with supervision that is fundamentally misaligned with what downstream acoustic models actually respond to. Tsinghua's approach replaces weak descriptive pseudo-labels with a direct speech-rewarded optimization loop — the style planner is trained to maximize the likelihood of target speech under a frozen synthesis model, coupling planning and production in a principled way. The result is stronger emotion and prosody alignment within dialogue context.

Tsinghua University

Tsinghua University · Oct 2026

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

This work trains a text-based style planner for conversational TTS that directly optimizes against frozen downstream speech synthesis, using target speech likelihood as a reward signal rather than relying on weak descriptive pseudo-labels. By decoupling style descriptions from acoustic control, the approach achieves stronger emotion and prosody alignment while maintaining contextual coherence in dialogue.

Abstract

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.

---

Full-Duplex Speech Dialogue & Voice Agents

Full-duplex spoken dialogue is maturing fast; this week's papers tackle its two open wounds — controllability and computational cost.

Open-source full-duplex models can hold a conversation, but they can't reliably follow instructions during one. SteerablePlex from UIUC addresses this head-on with a reward-decoupled policy optimization method that trains the model to respond to textual steering signals injected mid-conversation. Crucially, an asynchronous "director" backend monitor watches the conversation and injects instructions only when needed, enabling reliable multi-stage goal completion — something no existing open-source full-duplex model can currently do.

University of Illinois Urbana-Champaign

University of Illinois Urbana-Champaign · Oct 2026

SteerablePlex: Can We Steer Full-Duplex Models?

This paper tackles controllability in full-duplex speech models by training them to follow textual instructions during ongoing conversations via a novel reward-decoupled policy optimization approach. A key innovation is pairing the trained model with an asynchronous backend monitor (the "director") that injects steering instructions when needed, enabling reliable multi-stage goal completion in dialogue scenarios—a capability that existing open-source full-duplex models lack.

Abstract

Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows. When used as user simulators, this lack of control can cause them to deviate from prescribed scenarios and produce unreliable evaluation outcomes. We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order. The benchmark reveals that current open-source full-duplex models struggle to follow such constraints. We then introduce a Group Reward-Decoupled Normalization Policy Optimization (GDPO)-based training recipe that enables a full-duplex model to follow textual instructions during an ongoing conversation while maintaining its turn-taking ability. By connecting the resulting SteerablePlex to an asynchronous backend language model that monitors the conversation and provides instructions when needed, we build a more controllable full-duplex user simulator that follows multi-stage constraints more reliably than existing open-source models and GPT-Realtime.

On the efficiency side, Google Research's DiffuPlex tackles the frame-by-frame autoregressive bottleneck in full-duplex dialogue by predicting multiple future assistant and user frames in a single backbone pass. The system consumes only high-confidence predictions and dynamically replans when actual user input diverges from prediction — a speculative-execution strategy applied to spoken dialogue. The net effect is substantially reduced sequential computation without sacrificing interaction quality or speech naturalness.

Google Research

Google Research · Oct 2026

DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion

DiffuPlex accelerates full-duplex spoken dialogue by predicting multiple future assistant and user frames in a single backbone pass, then consuming only confident predictions while dynamically replanning when actual user input diverges. This reduces sequential computation overhead compared to conventional frame-by-frame autoregressive generation while maintaining interaction quality and speech naturalness.

Abstract

Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-LISTEN consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-SPEAK can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-LISTEN and DiffuPlex-SPEAK achieve $1.46\times$ and $1.59\times$ deployment-path wall-clock speedups and $1.61\times$ and $1.80\times$ Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that LISTEN preserves speech naturalness and conversational quality, while SPEAK retains conversational quality with some degradation in speech naturalness.