Real-Time Avatars, Ultra-Low Bitrate Video, and Score-Native Singing
Today's digest covers a cluster of advances in real-time talking-head avatars — from Gaussian splatting to generative video coding — plus a novel approach to singing voice synthesis that reads directly from musical scores.
Figure from From KAIST.
Today's papers push hard on two fronts: making photorealistic head avatars faster and more expressive, and closing the remaining gaps in faithful facial reenactment. Rounding things out, a new singing synthesis system takes a refreshingly direct route from musical score to voice — no duration prediction required.
Talking Avatars & Head Animation
From single-image Gaussian avatars to tongue-aware reenactment and ultra-low-bitrate video coding, this cluster tackles the full pipeline of realistic talking-head generation.
Single-Image 3D Head Avatars
Single-image avatar creation has long struggled with the trade-off between 3D consistency and rendering speed. KAIST's S-Avatar attacks this by marrying diffusion-guided generation with 3D Gaussian splatting and parametric model fitting. The diffusion prior supplies strong generalization across novel viewpoints and expressions, while the Gaussian representation keeps rendering firmly in real-time territory — outperforming prior work on both realism and 3D coherence.
KAIST · Jul 2026
S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
S-Avatar generates photorealistic 3D head avatars from a single image by combining diffusion-guided Gaussian splatting with parametric model fitting. This achieves superior 3D consistency across novel viewpoints and expressions, enabling real-time dynamic rendering with enhanced realism over prior work.
Abstract
We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation module and strategies for animating 3D Gaussian Splatting (3DGS). While single-image head avatar reconstruction is crucial for lifelike Virtual Reality (VR) applications, existing approaches often struggle to preserve 3D consistency under unseen viewpoints. S-Avatar addresses this limitation through a three-stage pipeline. First, a high-resolution 3DGS is synthesized directly from a single image using a diffusion-based Gaussian splat generation module. Next, the parametric head model FLAME is aligned with the generated 3DGS by optimizing its parameters and spatial transformations. Finally, to adapt the 3DGS to FLAME variations, we construct a binding template that encodes the spatial relationship between the initial splats and FLAME. The dynamic 3D head avatar can then be rendered in real time by deforming the 3DGS with the binding template. By combining diffusion-guided canonical 3DGS generation with FLAME-based control, our method achieves efficient and accurate reconstruction with enhanced 3D consistency. Evaluations on public datasets demonstrate that S-Avatar outperforms state-of-the-art methods in novel-view and expression generation, achieving superior realism and consistency. Consequently, our approach represents a significant advance in accessible avatar creation, applicable to a wide range of VR/AR applications. The project page is available at https://github.com/hailsong/savatar.
A complementary take comes from the Split and Drive framework, which decomposes facial geometry into dual specialized Gaussian branches — one for coarse structure, one for fine-grained detail — while internalizing the entire driving pipeline end-to-end. The key engineering insight is that prior benchmarks routinely excluded tracker latency from inference timings, making systems look faster than they are; Split and Drive reports honest end-to-end numbers and still leads the pack.
University Research Lab · Jul 2026
Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
A real-time single-image head avatar framework that decomposes facial geometry into specialized Gaussian branches while internalizing the driving pipeline for true end-to-end performance, achieving faster inference than prior methods without excluding tracking latency.
Abstract
Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.
Tongue-Aware Face Reenactment
Even the most photorealistic reenactment systems quietly ignore one conspicuous failure mode: tongue articulation. TongueReenact from the University of Manchester is the first dedicated framework for synthesizing realistic tongue dynamics during identity transfer. It bootstraps a tongue segmentation pipeline from existing data, then applies a spatially-constrained diffusion model to blend synthesized tongue geometry seamlessly into the reenacted face — a detail that proves surprisingly visible in open-mouth speech.
University of Manchester · Jul 2026
TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
The first framework for synthesizing realistic tongue dynamics in face reenactment, addressing a critical gap where prior methods ignore tongue articulation. It combines bootstrapping for tongue segmentation and a spatially-constrained diffusion model for seamless tongue synthesis during identity transfer.
Abstract
Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.
Generative Video Coding for Talking Heads
Bandwidth is often the bottleneck for real-world deployment of talking-head video. ReGenVC — from independent researchers — reimagines the codec itself: a generative decoder reconstructs full-quality video from a heavily compressed signal, achieving one-tenth the bitrate of traditional codecs at ultra-low bitrates without the blocking artifacts that plague conventional approaches. The harder problem was latency — generative decoding is expensive — which they solve through model distillation and system-level optimizations to hit real-time 24 fps reconstruction, a first for any generative video codec.
Independent Researchers · Jul 2026
ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
ReGenVC compresses talking-head video to one-tenth the bitrate of traditional codecs while remaining artifact-free at ultra-low bitrates. It solves the latency challenge of generative decoding through distillation and system optimizations, enabling real-time 24 fps reconstruction—a first for generative video codecs.
Abstract
We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The encoder reduces a source clip to a compact bitstream -- a neurally compressed first frame, per-frame pose keypoints, and metadata -- totaling about 26 kB for a 77-frame sequence. The decoder is a four-step distilled diffusion transformer that reconstructs the video conditioned on the transmitted pose and reference frame. Compared with x264/x265, ReGenVC reduces the bitrate to roughly one tenth of that required by traditional codecs (about 26 kB vs. 250--280 kB for essentially artifact-free reconstruction); at a matched ultra-low bitrate, conventional codecs collapse into blocking artifacts while ReGenVC stays sharp by exploiting a strong generative prior. The central obstacle to deploying such a codec is decoder latency: multi-step sampling with transformer and VAE components is too slow for interactive use. We make the decoder real-time through four-step distillation and three model-preserving system techniques: (i) eight-GPU unified sequence parallelism (Ulysses & Ring), (ii) a spatially-split VAE, and (iii) a three-stage overlapped pipeline; an analytical timing model characterizes the real-time feasibility region. On an 8-GPU node, the system sustains 24 fps output (972 ms per 25-frame window, within the 1000 ms budget), enabling a live browser stream without observed frame underruns. A hybrid CPU-GPU deployment further runs the encoder on the CPU at 24 fps and offloads the decoder-side one-shot conditioning encoders to the CPU, reducing the per-GPU memory peak from 21.1 GB to about 7.7 GB. To our knowledge, ReGenVC is the first end-to-end generative video codec to combine ultra-low-bitrate encoding with real-time decoding on an 8-GPU system.
TTS & Voice Synthesis
Singing synthesis takes a score-first approach, letting the musical representation do the heavy lifting.
Most singing voice synthesis (SVS) systems treat duration modeling as a separate, explicit stage and lean on acoustic features as intermediate targets. VocalRender from Nanyang Technological University sidesteps both. By encoding lyrics and notes in an interleaved representation and running autoregressive diffusion directly over it, the model inherently handles variable-length outputs without a dedicated duration predictor. The result integrates cleanly into real-world composition workflows — feed it a score, get back a voice.
Nanyang Technological University · Jul 2026
VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
VocalRender synthesizes singing directly from musical scores without explicit duration prediction or acoustic guidance. Using an interleaved lyric-note representation and autoregressive diffusion, it naturally handles variable-length outputs and integrates seamlessly into real-world composition workflows.
Abstract
Existing singing voice synthesis systems often require predefined durations, explicit duration prediction, or time-aligned acoustic guidance, which limits their compatibility with practical composition workflows. We propose VocalRender, a score-native system that directly synthesizes singing from lyrics, pitches, symbolic note values, and tempo. It uses an interleaved lyric--note representation and an autoregressive diffusion model to generate continuous acoustic latents while predicting the output length, eliminating the need for explicit duration prediction. Trained on a 2,300-hour singing dataset, VocalRender achieves strong intelligibility, strong melody control, and high speaker similarity across both in-domain and out-of-domain benchmarks. Notably, it outperforms the strongest baseline by $0.42$ points in naturalness CMOS, demonstrating the effectiveness of our proposed score-native architecture.
Trending on Hugging Face
Shanghai Jiao Tong University SAI · Jul 2026↑142 comments★ 51
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
Qwen · Jan 2026↑775 comments★ 12,815
Qwen3-TTS Technical Report
The Qwen3-TTS series presents advanced multilingual text-to-speech models with voice cloning and controllable speech generation capabilities, utilizing dual-track LM architecture and specialized speech tokenizers for efficient streaming synthesis.
Fish Audio · Mar 2026↑392 comments★ 32,027
Fish Audio S2 Technical Report
Fish Audio S2 is an open-source text-to-speech system with multi-speaker capabilities, multi-turn generation, and instruction-following control through natural-language descriptions, utilizing a multi-stage training approach and production-ready inference engine.
Microsoft Research · Aug 2025↑17710 comments★ 51,961
VibeVoice Technical Report
VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.
Oct 2024↑171 comment★ 61,897
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture.