Akapulu Labs logo Akapulu Labs Research

Emotion Steering, Gaussian Avatars, and Smarter TTS — Oct 8

Today's digest spans emotion control via activation steering in TTS and full-duplex speech models, advances in 3D Gaussian head avatar quality and efficiency, a new unified lip-sync metric, and cross-modal audio-visual editing from Meta AI.

Emotion Steering, Gaussian Avatars, and Smarter TTS — Oct 8

Example AV editing instructions, inputs, and outputs for CrossEditBench/CrossEdit. From Meta AI.

Today's digest is packed with work pushing the boundaries of controllable speech, photorealistic head avatars, and multimodal editing. Emotion steering emerges as a recurring theme — addressed from two distinct angles in TTS and full-duplex dialogue — while the avatar space sees both quality and efficiency gains. Rounding things out: a principled investigation into diffusion-based TTS training, a new lip-sync evaluation metric, and a movie-scene-level audio-visual editor from Meta AI.

---

Talking Avatars & Lip Sync

Measuring and editing talking heads more precisely, from phoneme timing to full scene composition.

Lip-sync evaluation has long been hampered by metrics that conflate timing errors with articulation errors. PVSync attacks this directly by decomposing lip-sync quality into two orthogonal axes — temporal alignment and phonetic articulation correctness — via viseme-aware contrastive learning. The result is a unified expert model that correlates more closely with human judgments than prior metrics, and generalizes across diverse video generation architectures.

Independent Research

Independent Research · Oct 2026

PVSync: A Unified Lip-Sync Expert for Timing and Articulation

PVSync is a unified model that separates temporal lip-sync quality from phonetic articulation correctness in talking-head videos through viseme-aware contrastive learning. It outperforms prior metrics by more closely matching human judgments of lip-sync quality across diverse video generation models.

Abstract

Lip movements can match the timing of speech without matching the spoken sounds. We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring. PVSync combines window-level contrastive learning for synchronisation with a phoneme-level articulation objective that aligns audio and video embeddings of the same viseme class across clips. Visemes group phonemes with similar visible articulation. Viseme labels are derived automatically from forced-aligned transcripts, without manual annotations. On offset-corrected videos from 13 talking-head video generation models, PVSync matches human rankings of lip-sync quality more closely than LSE-C, achieving a Spearman correlation of 0.83 versus 0.34. On an automatically constructed benchmark from held-out speech, PVSync distinguishes viseme-matched from mismatched audio-visual pairs with an ROC AUC of 0.91. PVSync also outperforms SyncNet and MTD-VocaLiST in temporal offset recovery on held-out in-the-wild clips. Code and benchmark data will be released upon acceptance.

Jumping from evaluation to editing: Meta AI's CrossEdit tackles the much harder problem of compositional, instruction-following audio-visual editing at the movie-scene level. The key insight is cross-modal transfer — complex instruction semantics learned in one modality (say, audio) generalize zero-shot to unseen combinations of modalities and instructions. Paired with a synthetic data generation pipeline, CrossEdit leapfrogs single-attribute, single-modality edits to handle rich, multi-attribute scene transformations.

Meta AI

Meta AI · Oct 2026

CrossEdit: Cross-Modal Training Enables Rich Audio-Visual Editing

CrossEdit enables unified instruction-following for audio-visual editing by leveraging synthetic data generation and cross-modal transfer—complex instruction semantics learned in one modality generalize zero-shot to unseen combinations of modalities and instructions. This advances beyond prior single-attribute edits within one modality to handle compositional movie-scene-level audio-visual edits.

Abstract

Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks. Demos are available at https://wanchichen.github.io/crossedit/.

---

Digital Humans & 3D Head Avatars

Smarter training dynamics and radical efficiency gains for Gaussian-based head synthesis.

Training 3D Gaussian head avatars involves balancing geometry optimization, appearance fitting, and joint refinement — typically managed with hand-crafted loss schedules. Chongqing University's Counterfactual Route Optimization replaces fixed schedules with a counterfactual lookahead strategy that dynamically reweights losses based on actual training outcomes at each state. By asking "what would have happened if we had optimized differently?", the method achieves superior preservation of fine facial details and cross-view consistency.

Chongqing University

Chongqing University · Oct 2026

Counterfactual Route Optimization for Gaussian Head Avatar Modeling

This paper optimizes head avatar modeling by dynamically adjusting loss weights based on actual training outcomes rather than fixed schedules. Using counterfactual lookahead, it determines whether geometry, appearance, or joint optimization is most beneficial at each state, enabling superior preservation of facial details and cross-view consistency.

Abstract

Head avatar modeling requires jointly optimizing multiple objectives with different dominant effects on geometry, appearance, and cross-view consistency. However, their relative effectiveness varies across training states, while existing pipelines typically rely on fixed loss weights or handcrafted stage-wise schedules. A central challenge is therefore to identify which optimization direction is more beneficial at each training state. We propose a counterfactual route optimization framework for Gaussian head avatar modeling, which characterizes state-dependent optimization preference from the realized effects of alternative updates rather than predefined heuristic weighting. Starting from the same training state, we perform short-horizon route-restricted lookahead over geometry, appearance, and joint update routes and evaluate their outcomes under a unified utility. The resulting counterfactual evidence is factorized into a geometry--appearance preference and a residual joint advantage, separately capturing the relative preference between individual update directions and the additional benefit of coordinated optimization. We further amortize this offline evidence into a lightweight controller that directly estimates the current optimization preference and applies bounded modulation to the training objectives during full avatar optimization. Experiments on the NeRSemble dataset validate the effectiveness of the proposed design, consistently outperforming existing methods while preserving clearer local facial structures and finer details.

On the efficiency front, the University of Surrey tackles a real deployment gap: virtually all Gaussian head avatar methods demand GPU acceleration and are thus impractical on mobile or edge hardware. Their approach introduces a parameter-efficient generator built on depth-wise separable convolutions, achieving a striking 94% FLOPs reduction while enabling real-time avatar synthesis in mobile browsers — no specialized hardware or software required.

University of Surrey

University of Surrey · Oct 2026

Efficient 3D Gaussian Head Avatars for Edge Devices

This work introduces an efficient generator architecture for 3D Gaussian head avatar synthesis, dramatically reducing computational requirements through parameter-efficient design and depth-wise separable convolutions. Unlike prior methods requiring GPU acceleration, it achieves 94% FLOPs reduction and enables real-time avatar generation on edge devices and mobile browsers without specialized hardware or software.

Abstract

Generative 3D Gaussian head avatars provide high-quality, efficient rendering, but synthesising the Gaussian representation remains computationally expensive, limiting deployment on resource-constrained and edge devices. We introduce an efficient generator architecture for unconditional 3D Gaussian head synthesis, based on a parameter-efficient synthesis block and depth-wise separable convolutions while retaining style-based conditioning. Our architecture reduces generator complexity without requiring model compression or quantisation. Compared with the baseline model, our approach reduces FLOPs by 94%, parameter count by 70%, and model size by 81%, while maintaining competitive generation quality. We further demonstrate practical CPU inference and browser-based execution on mobile devices using ONNX Runtime, enabling 3D Gaussian avatar synthesis without dedicated GPU hardware or application-specific software. In addition to conventional image-quality metrics, we evaluate multi-view consistency, training cost, and deployment performance. Code, trained models, and evaluation tools will be released publicly.

---

TTS & Voice Synthesis

From semantic-acoustic decomposition to diffusion training mechanics and inference-time emotion control.

Jingdong's JoyAI-Voice 2.0 is an ambitious end-to-end TTS system that decomposes speech into semantically purified and acoustic representations, then jointly processes them through an autoregressive Transformer (for planning) followed by a diffusion-based renderer. The semantic-acoustic disentanglement yields strong separation of content from style and prosody, with state-of-the-art results on multiple benchmarks — including dramatic WER improvements on Seed-TTS.

Jingdong

Jingdong · Sep 2026

JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

JoyAI-Voice 2.0 is an end-to-end speech synthesis model that decomposes speech into semantically purified and acoustic representations, jointly processed through an autoregressive Transformer for planning and diffusion-based rendering. The semantic-acoustic decomposition enables superior disentanglement of speech content from style and prosody, achieving state-of-the-art results on multiple benchmarks including dramatic WER improvements on Seed-TTS.

Abstract

We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48\,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51\% on Seed-TTS, a 14.9\% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.

A tighter mechanistic study comes from Yokohama: the team behind mask-and-replace diffusion for zero-shot TTS asks a pointed question — does the inference-time token revision process actually drive quality gains, or is it the training-time exposure to noisy contexts? By enforcing token commitment at inference while keeping training corruption intact, they isolate the contributions of noisy-context augmentation and replacement supervision, finding both matter independently.

National University of Yokohama

National University of Yokohama · Oct 2026

Beyond Token Revision: Investigating Mask-and-Replace Diffusion for Zero-Shot Text-to-Speech

This paper investigates whether mask-and-replace diffusion training for zero-shot text-to-speech improves generation through inference-time token revision or through training-side robustness to noisy contexts. By enforcing token commitment at inference while preserving training corruption patterns, the work isolates and demonstrates that both noisy-context augmentation and replacement supervision contribute independently to improved speech quality.

Abstract

Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability to revise previously generated tokens. However, exposure to randomly perturbed context during training may itself improve generation, raising the question of whether these gains require inference-time token revision. To investigate this question, we use DeMaR, which combines mask-and-replace training with confidence-ranked mask-only sampling while preserving the total training corruption probability. Trained from scratch on LibriTTS, DeMaR achieves lower word error rates (WER) than autoregressive and mask-only diffusion baselines using the same speech tokenizer. This advantage persists when each token remains unchanged after first being unmasked. Matched training conditions on two heterogeneous speech tokenizers show that both noisy-context augmentation and replacement supervision improve WER under this restriction. These findings demonstrate training-side benefits of replacement beyond enabling inference-time token revision.

Emotion control in TTS usually requires either fine-tuning or fully separate emotion-conditioned models. SteerSpeech offers a leaner path: learned low-rank steering vectors injected at inference time into a frozen pretrained TTS model. A multi-expert objective balances emotion intensity against speaker drift, and a novel generation-and-replay pipeline handles the discrete token space. The result is strong emotion expressiveness without touching model weights.

Independent Research

Independent Research · Oct 2026

Steerspeech: Activation Steering For Emotion Control In Generated Speech

SteerSpeech introduces a lightweight activation-steering framework for inference-time emotion control in pretrained TTS models, using learned low-rank transforms to inject steering vectors while preserving speaker identity and linguistic content without requiring model retraining. The approach uniquely balances strong emotion intensity with minimal speaker drift through a multi-expert objective and novel generation-and-replay pipeline for discrete speech tokens.

Abstract

Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.

---

SpeechLLMs & Spoken Dialogue

Probing the geometric structure of emotion in real-time full-duplex speech models.

Kyutai Lab's work on Moshi (their full-duplex speech model) takes a geometric lens to emotion. Rather than labeling and training with emotion tags, the team maps emotions to directions in activation space via activation steering — pure vector additions at inference, zero retraining. The key finding is that emotion in the model's latent space follows learned geometry, not labels: some emotions converge onto a shared direction, while others occupy distinctly steerable subspaces. This has meaningful implications for how we think about controllability in always-on spoken dialogue systems.

Kyutai Lab

Kyutai Lab · Oct 2026

Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model

This work investigates emotion steering in real-time full-duplex speech models by mapping emotions to geometric directions in the model's latent space via activation steering, requiring only vector additions without retraining. Unlike prior work on identity control or TTS-based emotion synthesis, it reveals that emotion follows learned geometric structure rather than simple labels, with some emotions converging on a shared direction while others remain distinctly steerable.

Abstract

Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally.