Akapulu Labs logo Akapulu Labs Research

Full-Duplex Dialogue, Smarter TTS, and Faster Voice Agents

Today's digest covers eight papers spanning full-duplex spoken dialogue systems, emotion-aware and cross-lingual speech LLMs, intelligibility-boosting TTS, and speculative execution for on-device voice agents.

Full-Duplex Dialogue, Smarter TTS, and Faster Voice Agents

Figure from From KAIST.

Today is a strong day for conversational speech research. From reinforcement-learning frameworks that teach models when and what to say simultaneously, to dyadic evaluation setups that treat dialogue as a genuinely two-sided problem, to TTS tricks that mimic how humans shout over noise — the field is getting both more principled and more practical. Here's everything worth reading on October 7, 2026.

SpeechLLMs & Full-Duplex Spoken Dialogue

Factorizing, evaluating, and grounding spoken dialogue models — with emotion in tow.

Full-duplex dialogue demands that a model simultaneously listen, decide when to speak, and choose what to say. KAIST's HiPLEX tackles this with a hierarchical RL framework that separates the control policy (turn-taking, backchannels) from the content policy (what to actually say), routing timing and semantic reward signals through event-causal masks to their respective components. The result is cleaner turn-taking and backchannel coordination compared to unified RL baselines that try to optimize both objectives at once.

KAIST

KAIST · Oct 2026↑31 comment

HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

HiPLEX is an RL framework that factorizes full-duplex speech policies into separate control and content factors to independently optimize when and what to say in real-time dialogue. By routing timing and semantic feedback through event-causal masks to their respective policy components, it achieves better turn-taking and backchannel coordination than unified RL approaches.

Abstract

As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.

Evaluating full-duplex models is just as hard as building them, because standard benchmarks use pre-recorded or static interlocutors that can't react to what the model actually says. Also from KAIST, DyaFDb fixes this by pairing two live model instances in direct conversation under assigned cooperative or conflicting roles, then measuring turn-taking, overlaps, and interruptions as emergent joint phenomena. The key insight is that each model simultaneously acts as examiner and examinee of its partner — a constraint no static benchmark can capture.

KAIST

KAIST · Oct 2026

Conversation Is a Two-Body Problem: Dyadic Evaluation of Full-Duplex Dialogue Models

This paper proposes DyaFDB, a framework for evaluating full-duplex dialogue models through direct peer-to-peer conversation rather than against pre-recorded or static interlocutors. By recording interactions between model pairs under assigned cooperative or conflicting roles, it captures how turn-taking, overlaps, and interruptions emerge as joint products of two coupled speakers—revealing that each model must serve as both examiner and examinee of its partner.

Abstract

Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot react, or an automated examiner that reacts in real time but only administers a fixed sequence of tests and is never graded. These single-sided frameworks evaluate only half of a two-body problem, where turn-taking, overlap, and interruption are joint products of two coupled speakers. We propose DyaFDB, a framework that evaluates full-duplex models in a dyadic setup: two models converse directly under assigned roles with cooperative or conflicting goals, and both sides are scored offline with an external judge. DyaFDB probes how the two models behave toward each other, such as how they take turns or carry an assigned role under different interests. We instantiate four tasks as 140 scenarios and record 7,560 conversations, covering six self- and cross-play pairings. Throughout the experiments, we observe that how a model behaves continually reshapes its partner. We thus demonstrate that each model must be both the examiner and examinee of the other, and no single fixed interlocutor can play both parts. We will release the scenarios, role prompts, and recording protocols between two full-duplex models, without any pre-recorded audio.

Emotion is another dimension that standard speech LLMs tend to collapse. Tsinghua's EMODE addresses this by explicitly splitting acoustic input into semantic and paralinguistic pathways, then dynamically routing them through expert modules to prevent the model from over-leaning on transcribed text for emotional cues. Combined with curriculum-based training, this structured factorization preserves emotional grounding where entangled acoustic representations typically fail.

Tsinghua University

Tsinghua University · Oct 2026

EMODE: Dynamic Para-Semantic Experts for Emotion-Aware Speech Language Modeling

EMODE is an emotion-aware speech language model that decomposes acoustic input into explicit semantic and paralinguistic pathways, dynamically routing them to prevent over-reliance on recovered text. This structured factorization with curriculum-based training better preserves emotional grounding compared to existing entangled acoustic representations.

Abstract

Large speech language models have demonstrated strong capabilities in unified cross-modal understanding and generation, yet paralinguistic cues, especially emotion, remain difficult to preserve. Existing systems typically rely on entangled acoustic representations, which allow the underlying language model to depend excessively on recovered lexical content instead of grounding its behavior in acoustic-prosodic evidence. We address this limitation with EMODE, an emotion-aware speech language model built around \textbf{Dynamic Para-Semantic Experts (DPSE)}. DPSE decomposes continuous speech features into semantic and paralinguistic pathways, routes them dynamically, and fuses them before integration into the language model. To turn this structural decomposition into functional specialization, EMODE is trained with a three-stage curriculum consisting of semantic warm-up, paralinguistic activation, and joint refinement, guided by Orthogonal Expert Guidance (OEG), Semantic-to-Acoustic Alignment (SAA), and Gating Diversity Regularization (GDR). Experiments on SER test, empathetic response evaluation, and the newly constructed bilingual MEPA benchmark show that EMODE improves the balance between lexical fidelity and emotional sensitivity, strengthens affect-grounded response generation, and exposes the value of explicit para-semantic factorization for robust cross-corpus emotion understanding.

Reasoning depth is usually fixed in audio language models, regardless of whether the task is easy speaker ID or hard acoustic scene analysis. The Audio Language Models Lab's AdaLoop changes this with adaptive-depth latent reasoning: the model learns when to stop iterating over refinement loops based on task complexity. Crucially, it achieves larger accuracy gains on perception-heavy tasks without touching the underlying model components.

Audio Language Models Lab

Audio Language Models Lab · Oct 2026

AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models

AdaLoop introduces adaptive-depth reasoning for audio language models, enabling them to dynamically adjust computational effort based on task complexity rather than applying uniform processing. By learning when to halt iterative refinement loops, the method achieves larger accuracy gains on perception-heavy tasks without modifying underlying model components.

Abstract

Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio--question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3\% of the base model's parameters and plugs into any audio encoder--language model pair without modifying either component. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, AdaLoop raises the average accuracy by 2.9 to 3.8 points, with the largest gains on perception-heavy subtasks where the model learns to apply deeper reasoning.

A persistent headache in speech-conditioned LLMs is prompt overfitting — where the speech encoder learns to rely on task-specific prompt patterns rather than genuine speech understanding, hurting zero-shot generalization. IIIT Delhi's DirectSpeech2LLM sidesteps this by aligning speech embeddings directly into the LLM input space using a distance-based CTC loss, trained only on ASR data. Despite never seeing translation or emotion-recognition prompts during training, the model generalizes to both tasks zero-shot — suggesting that implicit geometric grounding is sufficient without explicit multi-task alignment losses.

IIIT Delhi

IIIT Delhi · Oct 2026

DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs

This paper proposes an end-to-end framework that aligns speech embeddings directly to LLM input space using distance-based CTC loss, enabling speech-conditioned LLMs to preserve instruction-following ability across unseen tasks. Trained solely on ASR data, the approach generalizes zero-shot to speech translation and emotion recognition, demonstrating that implicit geometric grounding via modified CTC loss is sufficient without explicit alignment losses.

Abstract

Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.

TTS & Voice Synthesis

Accent-free cross-lingual cloning and noise-adaptive speech generation — no retraining required.

Zero-shot cross-lingual TTS systems often bleed the reference speaker's native accent into the target language output. Google Research's region-aware masking approach fixes this at the training-inference gap: by placing reconstruction masks relative to language boundaries in concatenated utterance pairs, the model learns to suppress accent transfer while retaining speaker identity. Remarkably, this requires only a change in mask geometry — no architectural modifications and no language labels.

Google Research

Google Research · Oct 2026

Region-Aware Masking for Accent-Robust Cross-Lingual Text-to-Speech

This work addresses accent leakage in zero-shot cross-lingual TTS by aligning training and inference through region-aware masking on concatenated utterance pairs. By strategically placing reconstruction masks relative to language boundaries, the approach suppresses unwanted accent transfer while preserving speaker identity—requiring only mask geometry changes with no architectural modifications or language labels.

Abstract

Accent leakage remains a critical challenge in cross-lingual zero-shot text-to-speech (TTS), where models inadvertently transfer both the speaker's timbre and source-language accent into the target-language output. In mask-reconstruction TTS this is amplified by a train--inference mismatch: training reconstructs from same-language context, while inference conditions on cross-lingual context. We close this gap by concatenating utterances from two languages and placing the reconstruction mask relative to the language boundary, so training presents the cross-lingual prompting pattern the model encounters at inference. This requires no language IDs, accent labels, parallel same-speaker recordings, or architectural changes---only the mask geometry differs. In a listening study with native speakers, region-aware masking reduces perceived accent from 4.3 to 0.5 on a 0-5 scale against the unadapted baseline, and by 18-51\% in seen regimes and 35-55\% on unseen source languages against bilingual adaptation alone, with no measurable loss in intelligibility or naturalness. Comparing mask geometries shows that placement, not the amount masked, governs the balance between accent suppression and speaker preservation: masks confined to the reconstruction region suppress accent most but retain the reference speaker least reliably, whereas two-region masking achieves both. Mask geometry is thus a simple, effective lever for accent robustness.

When speech must be heard over noise, humans naturally raise their voice and change their articulation — the Lombard effect. Karlsruhe Institute of Technology's "Loud and Clear" applies activation steering on pretrained TTS models to dynamically induce this effect at inference time, without retraining. A prompt-relative steering mechanism prevents artifact accumulation while enabling adaptive control, yielding 7–22% WER improvements in noisy conditions across multiple speakers and languages.

Karlsruhe Institute of Technology

Karlsruhe Institute of Technology · Oct 2026

Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

This work uses activation steering on pretrained TTS models to dynamically generate more intelligible speech by mimicking the Lombard effect—speakers' natural vocal adaptation in noise—without requiring model retraining. A prompt-relative steering mechanism prevents artifact accumulation while allowing adaptive control, yielding substantial WER improvements (7–22%) in noisy conditions across multiple speakers and languages.

Abstract

Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.

Voice Agents & On-Device Speech

Hiding latency where users feel it most.

On-device voice agents suffer from compounding latency: ASR must finish before the LLM runs, and the LLM must finish before tool calls execute. KAIST's speculative execution paper breaks this pipeline dependency by predicting tool calls from partial ASR hypotheses and firing them during the user's speech turn. A validation mechanism caches correct speculative results and gracefully falls back to standard tool-calling on mispredictions — making the latency savings safe to deploy.

KAIST

KAIST · Oct 2026

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

This work reduces latency in on-device voice agents by predicting tool calls from partial ASR hypotheses and executing them speculatively during speech input, rather than waiting for full transcription and LLM inference. A validation mechanism ensures safety by caching only correct predictions and allowing fallback to standard tool-calling when needed.

Abstract

Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.