English
Related papers

Related papers: OmniResponse: Online Multimodal Conversational Res…

200 papers

People often capture memories through photos, screenshots, and videos. While existing AI-based tools enable querying this data using natural language, they only support retrieving individual pieces of information like certain objects in…

Human-Computer Interaction · Computer Science 2025-02-24 Jiahao Nick Li , Zhuohao Jerry Zhang , Jiaju Ma

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and…

Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Jin Liu , Xi Wang , Xiaomeng Fu , Yesheng Chai , Cai Yu , Jiao Dai , Jizhong Han

Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data. Retrieval-Augmented Generation (RAG) mitigates these issues by integrating external dynamic information for…

Human-human communication is like a delicate dance where listeners and speakers concurrently interact to maintain conversational dynamics. Hence, an effective model for generating listener nonverbal behaviors requires understanding the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Minh Tran , Di Chang , Maksim Siniukov , Mohammad Soleymani

Pre-trained conversation models (PCMs) have demonstrated remarkable results in task-oriented dialogue (TOD) systems. Many PCMs focus predominantly on dialogue management tasks like dialogue state tracking, dialogue generation tasks like…

Computation and Language · Computer Science 2023-12-29 Mingtao Yang , See-Kiong Ng , Jinlan Fu

Recent advances in video multimodal large language models (Video MLLMs) have significantly enhanced video understanding and multi-modal interaction capabilities. While most existing systems operate in a turn-based manner where the model can…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yueqian Wang , Songxiang Liu , Disong Wang , Nuo Xu , Guanglu Wan , Huishuai Zhang , Dongyan Zhao

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ziqiao Peng , Yanbo Fan , Haoyu Wu , Xuan Wang , Hongyan Liu , Jun He , Zhaoxin Fan

Previous research on multi-party dialogue generation has predominantly leveraged structural information inherent in dialogues to directly inform the generation process. However, the prevalence of colloquial expressions and incomplete…

Computation and Language · Computer Science 2026-04-14 Zhiyu Cao , Peifeng Li , Qiaoming Zhu

Natural Language Understanding (NLU) and Natural Language Generation (NLG) are the two critical components of every conversational system that handles the task of understanding the user by capturing the necessary information in the form of…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Mauajama Firdaus , Avinash Madasu , Asif Ekbal

Real-world live retrieval-augmented generation (RAG) systems face significant challenges when processing user queries that are often noisy, ambiguous, and contain multiple intents. While RAG enhances large language models (LLMs) with…

Computation and Language · Computer Science 2025-06-27 Guanting Dong , Xiaoxi Li , Yuyao Zhang , Mengjie Deng

Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods…

Information Retrieval · Computer Science 2025-11-10 Kuicai Dong , Yujing Chang , Shijie Huang , Yasheng Wang , Ruiming Tang , Yong Liu

According to the Stimulus Organism Response (SOR) theory, all human behavioral reactions are stimulated by context, where people will process the received stimulus and produce an appropriate reaction. This implies that in a specific context…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Siyang Song , Micol Spitale , Yiming Luo , Batuhan Bal , Hatice Gunes

Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among…

Computation and Language · Computer Science 2025-06-06 Haonan Zhang , Run Luo , Xiong Liu , Yuchuan Wu , Ting-En Lin , Pengpeng Zeng , Qiang Qu , Feiteng Fang , Min Yang , Lianli Gao , Jingkuan Song , Fei Huang , Yongbin Li

As digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning. This paper explores using Multimodal Large Language Models (MLLMs) to automatically respond to student questions from online…

Computation and Language · Computer Science 2025-09-30 Sourjyadip Ray , Shubham Sharma , Somak Aditya , Pawan Goyal

Dialogue response generation (DRG) is a critical component of task-oriented dialogue systems (TDSs). Its purpose is to generate proper natural language responses given some context, e.g., historical utterances, system states, etc.…

Computation and Language · Computer Science 2020-02-20 Jiahuan Pei , Pengjie Ren , Christof Monz , Maarten de Rijke

Medical report generation from imaging data remains a challenging task in clinical practice. While large language models (LLMs) show great promise in addressing this challenge, their effective integration with medical imaging data still…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Chunlei Li , Jingyang Hou , Yilei Shi , Jingliang Hu , Xiao Xiang Zhu , Lichao Mou

Understanding emotions accurately is essential for fields like human-computer interaction. Due to the complexity of emotions and their multi-modal nature (e.g., emotions are influenced by facial expressions and audio), researchers have…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Qize Yang , Detao Bai , Yi-Xing Peng , Xihan Wei

Generating personalized responses is one of the major challenges in natural human-robot interaction. Current researches in this field mainly focus on generating responses consistent with the robot's pre-assigned persona, while ignoring the…

Computation and Language · Computer Science 2025-03-26 Bin Li , Hanjun Deng

Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key…