English
Related papers

Related papers: AuDirector: A Self-Reflective Closed-Loop Framewor…

200 papers

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

Computation and Language · Computer Science 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

End-to-end (E2E) spoken dialogue systems are increasingly replacing cascaded pipelines for voice-based human-AI interaction, processing raw audio directly without intermediate transcription. Existing benchmarks primarily evaluate these…

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Full-duplex speech interaction, as the most natural and intuitive mode of human communication, is driving artificial intelligence toward more human-like conversational systems. Traditional cascaded speech processing pipelines suffer from…

Artificial Intelligence · Computer Science 2026-05-01 Yadong Li , Guoxin Wu , Haiping Hou , Biye Li

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound…

Encoder-decoder based neural architectures serve as the basis of state-of-the-art approaches in end-to-end open domain dialog systems. Since most of such systems are trained with a maximum likelihood~(MLE) objective they suffer from issues…

We present three enhancements to existing encoder-decoder models for open-domain conversational agents, aimed at effectively modeling coherence and promoting output diversity: (1) We introduce a measure of coherence as the GloVe embedding…

Computation and Language · Computer Science 2018-11-22 Xinnuo Xu , Ondřej Dušek , Ioannis Konstas , Verena Rieser

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited character diversity,…

Computation and Language · Computer Science 2025-04-22 Xiang Li , Duyi Pan , Hongru Xiao , Jiale Han , Jing Tang , Jiabao Ma , Wei Wang , Bo Cheng

Existing automated research systems operate as stateless, linear pipelines -- generating outputs without maintaining any persistent understanding of the research landscape they navigate. They process papers sequentially, propose ideas…

Artificial Intelligence · Computer Science 2026-03-27 Yunbo Long

dAIrector is an automated director which collaborates with humans storytellers for live improvisational performances and writing assistance. dAIrector can be used to create short narrative arcs through contextual plot generation. In this…

Computers and Society · Computer Science 2018-11-09 Markus Eger , Kory W. Mathewson

Large Language Models have demonstrated remarkable capabilities in open-domain dialogues. However, current methods exhibit suboptimal performance in service dialogues, as they rely on noisy, low-quality human conversation data. This…

Computation and Language · Computer Science 2026-05-06 Yuqin Dai , Ning Gao , Wei Zhang , Jie Wang , Zichen Luo , Jinpeng Wang , Yujie Wang , Ruiyuan Wu , Chaozheng Wang

We propose a unified Implicit Dialog framework for goal-oriented, information seeking tasks of Conversational Search applications. It aims to enable dialog interactions with domain data without replying on explicitly encoded the rules but…

Computation and Language · Computer Science 2018-02-14 Song Feng , R. Chulaka Gunasekara , Sunil Shashidhara , Kshitij P. Fadnis , Lazaros C. Polymenakos

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a…

Despite the promising results achieved, state-of-the-art interactive reinforcement learning schemes rely on passively receiving supervision signals from advisor experts, in the form of either continuous monitoring or pre-defined rules,…

Machine Learning · Computer Science 2024-05-27 Shunyu Liu , Kaixuan Chen , Na Yu , Jie Song , Zunlei Feng , Mingli Song

Generating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for practical dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Shiying Li , Xingqun Qi , Bingkun Yang , Chen Weile , Zezhao Tian , Muyi Sun , Qifeng Liu , Man Zhang , Zhenan Sun

Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient framework to unlock their full potential. Existing audio…

Sound · Computer Science 2026-01-01 Cheng Zhu , Jing Han , Qianshuai Xue , Kehan Wang , Huan Zhao , Zixing Zhang

Automatic dubbing, which generates a corresponding version of the input speech in another language, could be widely utilized in many real-world scenarios such as video and game localization. In addition to synthesizing the translated…

Sound · Computer Science 2024-07-08 Jingbei Li , Sipan Li , Ping Chen , Luwen Zhang , Yi Meng , Zhiyong Wu , Helen Meng , Qiao Tian , Yuping Wang , Yuxuan Wang

Enabling Large Language Models (LLMs) to reliably invoke external tools remains a critical bottleneck for autonomous agents. Existing approaches suffer from three fundamental challenges: expensive human annotation for high-quality…

Computation and Language · Computer Science 2025-12-30 Yuwen Li , Wei Zhang , Zelong Huang , Mason Yang , Jiajun Wu , Shawn Guo , Huahao Hu , Lingyi Sun , Jian Yang , Mingjie Tang , Byran Dai

This study presents a deep-learning framework for controlling multichannel acoustic feedback in audio devices. Traditional digital signal processing methods struggle with convergence when dealing with highly correlated noise such as…

Sound · Computer Science 2025-05-30 Yuan-Kuei Wu , Juan Azcarreta , Kashyap Patel , Buye Xu , Jung-Suk Lee , Sanha Lee , Ashutosh Pandey

As communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack…

Human-Computer Interaction · Computer Science 2023-09-12 Zeyuan Huang , Qiang He , Kevin Maher , Xiaoming Deng , Yu-Kun Lai , Cuixia Ma , Sheng-feng Qin , Yong-Jin Liu , Hongan Wang