English
Related papers

Related papers: PaCE: Unified Multi-modal Dialogue Pre-training wi…

200 papers

Task-oriented dialogues often require agents to enact complex, multi-step procedures in order to meet user requests. While large language models have found success automating these dialogues in constrained environments, their widespread…

Computation and Language · Computer Science 2023-06-08 Julia White , Arushi Raghuvanshi , Yada Pruksachatkun

With the availability of massive general-domain dialogue data, pre-trained dialogue generation appears to be super appealing to transfer knowledge from the general domain to downstream applications. In most existing work, such transferable…

Computation and Language · Computer Science 2022-10-25 Xueliang Zhao , Lemao Liu , Tingchen Fu , Shuming Shi , Dongyan Zhao , Rui Yan

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-specific characteristics…

Computation and Language · Computer Science 2024-12-23 Maximillian Chen , Ruoxi Sun , Sercan Ö. Arık

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Recently, instruction-following audio-language models have received broad attention for audio interaction with humans. However, the absence of pre-trained audio models capable of handling diverse audio types and tasks has hindered progress…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Yunfei Chu , Jin Xu , Xiaohuan Zhou , Qian Yang , Shiliang Zhang , Zhijie Yan , Chang Zhou , Jingren Zhou

Multimodal language models now integrate text, audio, and video for unified reasoning. Yet existing RL post-training pipelines treat all input signals as equally relevant, ignoring which modalities each task actually requires. This…

Artificial Intelligence · Computer Science 2026-02-13 Nikhil Verma , Minjung Kim , JooYoung Yoo , Kyung-Min Jin , Manasa Bharadwaj , Kevin Ferreira , Ko Keun Kim , Youngjoon Kim

Recent work in open-domain conversational agents has demonstrated that significant improvements in model engagingness and humanness metrics can be achieved via massive scaling in both pre-training data and model size (Adiwardana et al.,…

Computation and Language · Computer Science 2020-10-05 Kurt Shuster , Eric Michael Smith , Da Ju , Jason Weston

A well-designed interactive human-like dialogue system is expected to take actions (e.g. smiling) and respond in a pattern similar to humans. However, due to the limitation of single-modality (only speech) or small volume of currently…

Human-Computer Interaction · Computer Science 2022-12-13 Zhiling Luo , Qiankun Shi , Sha Zhao , Wei Zhou , Haiqing Chen , Yuankai Ma , Haitao Leng

Audio is a fundamental modality for analyzing speech, music, and environmental sounds. Although pretrained audio models have significantly advanced audio understanding, they remain fragile in real-world settings where data distributions…

Sound · Computer Science 2026-02-04 Chang Li , Kanglei Zhou , Liyuan Wang

In spoken dialogue systems, we aim to deploy artificial intelligence to build automated dialogue agents that can converse with humans. Dialogue systems are increasingly being designed to move beyond just imitating conversation and also…

Computation and Language · Computer Science 2021-11-03 Atharv Singh Patlan , Shiven Tripathi , Shubham Korde

Human-machine interaction has been around for several decades now, with new applications emerging every day. One of the major goals that remain to be achieved is designing an interaction similar to how a human interacts with another human.…

Human-Computer Interaction · Computer Science 2022-12-27 Tauheed Khan Mohd , Nicole Nguyen , Ahmad Y Javaid

The ability to sequence unordered events is an essential skill to comprehend and reason about real world task procedures, which often requires thorough understanding of temporal common sense and multimodal information, as these procedures…

Computation and Language · Computer Science 2024-02-22 Te-Lin Wu , Alex Spangher , Pegah Alipoormolabashi , Marjorie Freedman , Ralph Weischedel , Nanyun Peng

Task-oriented dialogue systems are broadly used in virtual assistants and other automated services, providing interfaces between users and machines to facilitate specific tasks. Nowadays, task-oriented dialogue systems have greatly…

Computation and Language · Computer Science 2024-05-17 Ruolin Su , Biing-Hwang Juang

Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Wenming Tu , Guanrou Yang , Ruiqi Yan , Wenxi Chen , Ziyang Ma , Yipeng Kang , Kai Yu , Xie Chen , Zilong Zheng

Building an end-to-end conversational agent for multi-domain task-oriented dialogues has been an open challenge for two main reasons. First, tracking dialogue states of multiple domains is non-trivial as the dialogue agent must obtain…

Computation and Language · Computer Science 2020-11-17 Hung Le , Doyen Sahoo , Chenghao Liu , Nancy F. Chen , Steven C. H. Hoi

We study video-grounded dialogue generation, where a response is generated based on the dialogue context and the associated video. The primary challenges of this task lie in (1) the difficulty of integrating video data into pre-trained…

Computation and Language · Computer Science 2022-10-25 Xueliang Zhao , Yuxuan Wang , Chongyang Tao , Chenshuo Wang , Dongyan Zhao

Learning an efficient manager of dialogue agent from data with little manual intervention is important, especially for goal-oriented dialogues. However, existing methods either take too many manual efforts (e.g. reinforcement learning…

Computation and Language · Computer Science 2019-08-16 Zhuoxuan Jiang , Xian-Ling Mao , Ziming Huang , Jie Ma , Shaochun Li

Persona-based dialogue generation is an important milestone towards building conversational artificial intelligence. Despite the ever-improving capabilities of large language models (LLMs), effectively integrating persona fidelity in…

Computation and Language · Computer Science 2025-08-12 Arpita Saggar , Jonathan C. Darling , Vania Dimitrova , Duygu Sarikaya , David C. Hogg

We present MoST (Mixture of Speech and Text), a novel multimodal large language model that seamlessly integrates speech and text processing through our proposed Modality-Aware Mixture of Experts (MAMoE) architecture. While current…

Computation and Language · Computer Science 2026-01-16 Yuxuan Lou , Kai Yang , Yang You

End-to-end Task-oriented Dialogue Systems (TDSs) have attracted a lot of attention for their superiority (e.g., in terms of global optimization) over pipeline modularized TDSs. Previous studies on end-to-end TDSs use a single-module model…

Computation and Language · Computer Science 2019-07-12 Jiahuan Pei , Pengjie Ren , Maarten de Rijke