Expressive Avatars, Emotional Voice, and Smarter ASR
Today's digest covers speech-driven 3D facial animation, relightable avatar reconstruction from a single image, expressive human motion disentanglement, hierarchical reward optimization for emotional TTS, and edit-flow refinement for non-autoregressive ASR — a broad sweep of lifelike digital human and voice research.
MindFlow teaser image illustrating harmonized cognitive semantics and acoustic dynamics in facial animation of dyadic conversations. From MindFlow.
Talking Avatars & Facial Animation
KM-Speaker
KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization
KM-Speaker is a speech-driven 3D facial animation system that combines global style from full-face keypoints with frame-level control from upper-face keypoints. It enables high-fidelity motion and precise style control, excelling in dialogue localization with accurate lip-sync and expressive performance.
MindFlow
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
MindFlow generates lifelike facial animations in dyadic conversations by combining evolving emotional state reasoning with precise motion control. It models raw audio as emotion states and adaptively fuses acoustic cues to produce semantically rich and temporally accurate facial animation.
Digital Humans & Avatar Reconstruction
MARCUS-Avatar
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
MARCUS-Avatar reconstructs high-quality, relightable 3D face avatars from a single image via cascaded diffusion priors in UV space. It integrates light normalization and differentiable shading to generate physically plausible PBR assets with detailed geometry and robust relighting, trained with limited real 3D scans.
EMOSH
EMOSH: Expressive Motion and Shape Disentanglement for Human Animation
EMOSH presents a new Expressive Human Model that separates body shape from motion for high-fidelity human animation. It prevents shape leakage common in 2D pose methods while capturing detailed facial and gesture motions, enabling expressive, identity-consistent video generation with stable long-term performance.