Akapulu Labs logo Akapulu Labs Research

Language-Guided Avatars, Few-Step TTS, and Real-Time Speech Translation

Today's digest covers language-driven editing of 3D Gaussian head avatars, a distillation-free few-step TTS approach, spatially grounded gesture generation for VR dialogue, and end-to-end simultaneous speech translation with confidence-based commit timing.

Language-Guided Avatars, Few-Step TTS, and Real-Time Speech Translation

Qualitative comparison. Results for skin, hair, and eyes edits across three subjects. Our method produces localized, natural color changes that preserve identity and fine details (e.g., specular highlights, skin texture). Both GaussianAvatar-Editor~ and InstructPix2Pix~ exhibit color bleeding into unrelated regions and over-saturated tones. From Independent Researchers.

Today's papers span the full stack of conversational AI embodiment — from editing the appearance and geometry of 3D head avatars with natural language, to synthesizing speech more efficiently, grounding gestures in physical space, and translating spoken language in real time. A common thread: making interactive systems more controllable, precise, and latency-aware.

Digital Humans & 3D Head Avatars

Two complementary papers tackle language-guided editing of 3D Gaussian Splatting head avatars — one targeting color, the other shape.

Color editing of photorealistic avatars has historically required retraining or suffered from color bleeding across regions. ChromaGS attacks this by decomposing appearance into region-level base colors and Gaussian-level residuals, enabling deterministic, precisely localized semantic recoloring via natural language — no retraining, no generative ambiguity. The result is real-time edits that stay exactly where you ask them to.

Independent Researchers

Independent Researchers · Oct 2026

ChromaGS: Text-Driven Semantic Editing of 4D Gaussian Avatars

ChromaGS enables real-time, language-guided color editing of 3D Gaussian head avatars by decomposing appearance into region-level base colors and Gaussian-level residuals. Unlike generative editing methods, it provides deterministic, precisely localized semantic control—allowing intuitive recoloring via natural language without retraining or color bleeding artifacts.

Abstract

We present ChromaGS, a method for real-time, language-guided color editing of animatable 3D Gaussian head avatars. Given a trained animatable avatar, users can instantly modify the color of semantic regions through natural language, with edits applied at render time and no retraining required. Our key insight is to augment each Gaussian primitive with learned soft assignments to semantic regions and decompose colors into region-level base colors and Gaussian-level residuals. This decomposition enables coherent color transfer: modifying a region's base color propagates naturally through all associated Gaussians while preserving fine appearance details encoded in residuals. A two-stage language pipeline translates text instructions into target colors, supporting both absolute specifications and relative adjustments. Unlike generative editing methods that may introduce unintended modifications, our approach provides deterministic, precisely localized semantic control. Experiments demonstrate faithful appearance preservation and intuitive interaction across diverse subjects. Project page and code are available at: https://a-canela.github.io/chromags/

Complementing ChromaGS on the geometry side, ManifoldSplat (University of Barcelona) addresses shape editing — a harder problem because naive optimization of unstructured Gaussians tends to break identity and animation rigging. The key insight is to constrain edits to a structured facial manifold, which enforces geometric consistency and preserves the underlying rig. Edits complete in roughly 90 seconds on a consumer GPU at interactive quality, with fine-grained localization that out-of-manifold approaches struggle to match.

University of Barcelona

University of Barcelona · Oct 2026

ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars

ManifoldSplat enables precise language-guided editing of 3D Gaussian head avatars while preserving identity and animation rigging by performing edits within a structured facial manifold rather than optimizing unstructured Gaussians. This approach achieves fine-grained localized control with geometric consistency and runs at interactive speeds (~90 seconds per edit on consumer GPUs).

Abstract

High-fidelity 3D head avatars have reached near-photorealistic quality. While recent methods enable text-driven manipulation, they struggle to provide fine-grained localized control, often entangling features or lacking geometric consistency. Modifying geometry through natural language currently requires slow per-prompt optimization or compromises identity and rigging. We present ManifoldSplat, the first end-toend framework for language-guided semantic shape editing of animatable 3D Gaussian Splatting avatars reconstructed from monocular videos. By performing edits within the structured FLAME manifold rather than directly optimizing an unstructured Gaussian cloud, we strictly preserve identity and animation. We introduce DeltaRegion, a per-region disentangled Conditional Variational Autoencoder (CVAE) delivering feedforward shape deltas, alongside a refining stage to recover view-consistent details. ManifoldSplat reconstructs and edits an avatar in ~90 seconds on a consumer GPU, rendering at ~800 FPS. Extensive evaluations demonstrate our approach sets a new state-of-the-art in localized prompt alignment, geometric coherence, and identity preservation. Project page and code: https://a-canela.github.io/manifoldsplat/

TTS & Voice Synthesis

Rethinking the training objective for fast, high-quality speech synthesis.

Diffusion-based TTS has largely relied on distillation from pretrained teachers to achieve few-step generation — adding complexity and coupling the student to the teacher's quirks. DriftTTS (University of Massachusetts Amherst) breaks this dependency by introducing a distribution-matching drift objective trained via on-policy rollout on the model's own intermediate states. No teacher, no external generative supervision — just the model bootstrapping itself into competitive few-step mel-spectrogram synthesis.

University of Massachusetts Amherst

University of Massachusetts Amherst · Oct 2026

DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

DriftTTS generates speech mel-spectrograms using a distribution-matching drift objective, achieving competitive few-step synthesis without requiring distillation from a pretrained teacher. The method trains directly on its own intermediate states via on-policy rollout, demonstrating that effective few-step TTS is possible without external generative supervision.

Abstract

Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git

Co-Speech Gesture & Grounded Interaction

Beyond naturalness: measuring whether gestures actually point at the right thing.

Existing gesture generation benchmarks reward natural-looking motion, but in conversational VR the critical question is whether a pointing gesture actually indicates the intended spatial referent. The Spatially Grounded Gesture Generation Benchmark (University of Cambridge) separates this into three orthogonal axes: temporal synchronization, spatial grounding precision, and perceived naturalness. A striking finding is that systems can achieve geometric correctness exceeding human performance while still lagging on naturalness — confirming that current metrics conflate two very different goals.

University of Cambridge

University of Cambridge · Oct 2026

A Benchmark for Spatially Grounded Gesture Generation

This work introduces a benchmark for evaluating pointing gestures in conversational VR dialogue, where systems must generate gestures that correctly indicate their intended spatial referents. Unlike prior metrics that reward natural-looking motion regardless of accuracy, this benchmark separately measures temporal synchronization, spatial grounding precision, and perceived naturalness—revealing that geometric correctness can exceed human performance without improving naturalness.

Abstract

Communication in shared space interweaves verbal and non-verbal signals, and pointing gestures anchor language to the environment: "put the cup on that one" is uninterpretable without the gesture that fixes the referent. Yet no common framework exists for evaluating whether generated gestures indicate their intended referent; distributional metrics reward a gesture aimed at the wrong object as long as it looks natural. We introduce a benchmark for spatially grounded gesture generation, comprising ~2K pointing-annotated clips from naturalistic VR dialogue with ground-truth 3D referents, a task in which systems must decide when, how and where to point within conversational speech, and a protocol that separates temporal alignment, spatial grounding and perceived naturalness. We also provide a flow-matching baseline, MM-Conv-Flow. Evaluating it alongside an independent retrieval-based system and captured human motion, we find that geometric grounding can exceed that of human pointing without any gain in perceived naturalness, showing that referential gesture quality must be measured along separate dimensions.

Speech Translation & Spoken Language Models

Teaching a speech LM to know when it has heard enough to commit.

Simultaneous speech translation requires deciding, moment by moment, whether to wait for more audio or output a translation segment — a commit-timing problem that existing systems handle poorly on partial inputs. Researchers at the University of Edinburgh train a speech language model using prefix predictions derived from its own complete and partial translations, requiring neither transcripts nor human references. Confidence-based decoding and multi-turn append-only generation then let users dial the quality–latency tradeoff at inference time, with the largest gains in calibration appearing exactly where it's hardest: early, highly partial inputs.

University of Edinburgh

University of Edinburgh · Oct 2026

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

This paper enables simultaneous speech translation by training a speech language model on prefix predictions derived from its own complete and partial translations—requiring neither transcripts nor human references. It introduces confidence-based decoding and multi-turn append-only generation to let users control the quality–latency tradeoff in real-time, with significant improvements in commit-timing calibration especially at early partial inputs.

Abstract

Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.