Dyadic Talking Heads and AI Interview Coaches
Today's digest covers two papers pushing the boundaries of conversational avatar systems: a new framework for dyadic talking-head generation that modulates social interaction over monologic priors, and an LLM-powered mock interview platform with a lip-synced digital interviewer and multimodal coaching feedback.
Comparison of existing audio-driven head motion generation paradigms. Monologic methods (a) ignore interaction dynamics, while existing dyadic approaches (b) either require auxiliary motion inputs or entangle speech and interaction signals through holistic learning. In contrast, Learn2Chat reformulates dyadic generation as interaction modulation over canonical monologic motion, enabling realistic and temporally coherent synthesis. The framework is plug-and-play and converts pretrained monologic backbones to the dyadic setting. From Tsinghua University.
Today's two papers share a common thread: making digital talking heads feel genuinely conversational. One tackles the fundamental motion modeling challenge of dyadic interaction; the other puts a lip-synced avatar to work as an AI interview coach. Together they highlight how far the field has come in turning monologue-optimized systems into socially-aware conversational agents.
Talking Avatars & Conversational Motion
From interaction-aware motion priors to full-stack interview simulations — digital talking heads are getting a social upgrade.
The dominant paradigm in talking-head synthesis has long been monologic: drive a face with speech audio, optimize for lip sync and naturalness, done. But real conversation involves a listener reacting, backchanneling, and dynamically modulating their own motion in response to a speaking partner — a much harder problem. Learn2Chat from Tsinghua University takes a principled new angle on this: rather than training a bespoke dyadic model end-to-end, it frames dyadic talking-head generation as interaction modulation layered on top of any pretrained monologic backbone. Speech-driven motion and social interaction effects are explicitly separated, which means existing monologic models can be adapted plug-and-play without retraining — and the results outperform end-to-end dyadic approaches.
Tsinghua University · Jul 2026
Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors
Learn2Chat reformulates dyadic talking-head generation as interaction modulation over pretrained monologic priors, separating speech-driven motion from social interaction effects. This enables plug-and-play adaptation of diverse monologic backbones without retraining, outperforming end-to-end dyadic approaches.
Abstract
Dyadic conversational motion generation is essential for realistic interactive digital humans. Existing approaches typically model conversational behaviors within unified dyadic generators. However, such holistic formulations tend to couple self-speech-driven motion with partner-responsive social feedback, leaving the interaction-specific component implicit and underutilizing the speech-motion correspondence already learned by pretrained monologic motion models. We propose Learn2Chat, a unified framework that models dyadic motion as interaction modulation over pretrained monologic motion priors. This design separates intrinsic speech-driven motion from social interaction effects and enables more structured interaction modeling. Specifically, we introduce a Monologic-Anchored Motion Factorization scheme that leverages the semantic motion manifold learned from monologic data to disentangle audio-driven motion dynamics from interaction-induced modulation, yielding clean interaction representations from dyadic sequences. On top of this representation space, a Cross-Attentive Interaction Latent Prediction module maps paired speech signals to interaction latents through cross-branch attention and interaction alignment. During inference, the predicted interaction latents modulate canonical monologic motion to generate coherent and synchronized dyadic behaviors in a data-efficient manner. Extensive experiments on the DualTalk benchmark demonstrate that Learn2Chat achieves state-of-the-art performance across both quantitative metrics and perceptual evaluations. Moreover, the framework is model-agnostic and seamlessly integrates with diverse pretrained monologic motion backbones, highlighting the effectiveness of prior reuse and interaction adaptation for scalable conversational motion generation. More visual results are available on the project page.
On the application side, researchers at The Hong Kong Polytechnic University have built PolyInterview, an LLM-based mock interview platform that puts a lip-synced digital interviewer front and center. What sets it apart from simpler chatbot-style practice tools is its answer-aware follow-up generation and competency-grounded feedback pipeline: the system doesn't just ask preset questions, it adapts its dialogue based on what the candidate actually said, then delivers multimodal assessment covering content quality, vocal delivery, and non-verbal behavior. It's a compelling demonstration of how conversational avatar technology — lip sync, adaptive dialogue, and multimodal analysis — can be integrated into a practical coaching product.
The Hong Kong Polytechnic University · Jul 2026
PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment
PolyInterview combines personalized question generation, adaptive dialogue with a lip-synced digital interviewer, and multimodal assessment of content, vocal delivery, and non-verbal behavior. It uniquely integrates answer-aware follow-ups and competency-grounded feedback for realistic, coached interview practice.
Abstract
Preparing for job interviews is important for securing desired positions, yet realistic practice remains difficult to access: real interviews are infrequent, expert mock coaching is costly, and self-practice offers neither adaptive dialogue nor structured assessment. Existing systems typically address only parts of this need through fixed question sequences, limited communication channels, or feedback with little supporting evidence. We present PolyInterview, an LLM-based platform for immersive mock interview practice with comprehensive multimodal assessment. PolyInterview uses the target job description and CV to generate questions tailored to the role and candidate, conducts multi-turn spoken interviews with a lip-synced digital human interviewer that asks answer-aware follow-up questions, and evaluates response content, vocal delivery, and non-verbal behavior. Four parallel evaluators produce 13 behavior-level features that are aggregated into 10 assessment aspects and two competency tracks. Guided by the KSA and STAR frameworks, the report links each score to behavioral evidence and actionable recommendations. PolyInterview is publicly accessible. Its current all-account snapshot contains 101 accounts, 1,564 interview sessions, 7,665 generated questions, and 1,422 five-stage question sets. Generated questions are more closely aligned with their matched job description than with cross-role job descriptions in 93.7% of sessions. An evaluation by ten experts found strong question plans and actionable feedback.
Trending on Hugging Face
Oct 2024↑161 comment★ 61,264
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture.
Jul 2024↑401 comment★ 22,243
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
FunAudioLLM enhances voice interactions by integrating SenseVoice for multilingual speech recognition, emotion detection, and audio event detection with CosyVoice for natural speech generation across languages, timbres, and styles.
Qwen · Jan 2026↑775 comments★ 12,459
Qwen3-TTS Technical Report
The Qwen3-TTS series presents advanced multilingual text-to-speech models with voice cloning and controllable speech generation capabilities, utilizing dual-track LM architecture and specialized speech tokenizers for efficient streaming synthesis.
Feb 2025↑7★ 21,965
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
IndexTTS, an enhanced text-to-speech system combining XTTS and Tortoise models, offers improved naturalness, enhanced voice cloning, and controllable usage through hybrid character-pinyin modeling and optimized vector quantization.
Sep 2025↑11★ 7,716