English
Related papers

Related papers: OmniResponse: Online Multimodal Conversational Res…

200 papers

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core breakthrough, it…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Dongyang Yu , Shihao Wang , Yuan Fang , Wangpeng An

This paper explores the potential of constructing an AI spoken dialogue system that "thinks how to respond" and "thinks how to speak" simultaneously, which more closely aligns with the human speech production process compared to the current…

Computation and Language · Computer Science 2023-09-21 Xinyu Zhou , Delong Chen , Yudong Chen

In this paper, we investigate the use of large language models (LLMs) like ChatGPT for document-grounded response generation in the context of information-seeking dialogues. For evaluation, we use the MultiDoc2Dial corpus of task-oriented…

Computation and Language · Computer Science 2023-09-22 Norbert Braunschweiler , Rama Doddipatla , Simon Keizer , Svetlana Stoyanchev

This article introduces Bio-Eng-LMM AI chatbot, a versatile platform designed to enhance user interaction for educational and research purposes. Leveraging cutting-edge open-source Large Language Models (LLMs), Bio-Eng-LMM operates as a…

Systems and Control · Electrical Eng. & Systems 2025-04-22 Ali Forootani , Danial Esmaeili Aliabadi , Daniela Thraen

Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Trong-Thuan Nguyen , Pha Nguyen , Jackson Cothren , Alper Yilmaz , Khoa Luu

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and…

Computation and Language · Computer Science 2025-01-06 Qinglin Zhang , Luyao Cheng , Chong Deng , Qian Chen , Wen Wang , Siqi Zheng , Jiaqing Liu , Hai Yu , Chaohong Tan , Zhihao Du , Shiliang Zhang

In daily life, we encounter a variety of sounds, both desirable and undesirable, with limited control over their presence and volume. Our work introduces "Listen, Chat, and Remix" (LCR), a novel multimodal sound remixer that controls each…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-12 Xilin Jiang , Cong Han , Yinghao Aaron Li , Nima Mesgarani

Evaluating Retrieval-Augmented Generation (RAG) systems using static multi-turn datasets fails to capture the dynamic nature of real-world dialogues. Existing evaluation methods rely on predefined datasets, which restrict them to static,…

Information Retrieval · Computer Science 2026-04-21 Lorenz Brehme , Benedikt Dornauer , Jan-Henrik Böttcher , Klaus Schmid , Mircea-Cristian Racasan , Ruth Breu

The ability to generate sentiment-controlled feedback in response to multimodal inputs comprising text and images addresses a critical gap in human-computer interaction. This capability allows systems to provide empathetic, accurate, and…

Multimedia · Computer Science 2025-10-07 Puneet Kumar , Sarthak Malik , Balasubramanian Raman , Xiaobai Li

Multimodal Retrieval-Augmented Generation (MRAG) enhances reasoning capabilities by integrating external knowledge. However, existing benchmarks primarily focus on simple image-text interactions, overlooking complex visual formats like…

Artificial Intelligence · Computer Science 2025-02-21 Yuming Yang , Jiang Zhong , Li Jin , Jingwang Huang , Jingpeng Gao , Qing Liu , Yang Bai , Jingyuan Zhang , Rui Jiang , Kaiwen Wei

End-to-end neural networks have achieved promising performances in natural language generation (NLG). However, they are treated as black boxes and lack interpretability. To address this problem, we propose a novel framework, heterogeneous…

Computation and Language · Computer Science 2021-02-09 Yangming Li , Kaisheng Yao

Recent Multimodal Large Language Models (MLLMs) achieve promising performance on visual and audio benchmarks independently. However, the ability of these models to process cross-modal information synchronously remains largely unexplored. We…

Artificial Intelligence · Computer Science 2026-03-12 Ziwei Zhou , Rui Wang , Zuxuan Wu , Yu-Gang Jiang

Current general-purpose large language models (LLMs) commonly exhibit knowledge hallucination and insufficient domain-specific adaptability in domain-specific tasks, limiting their effectiveness in specialized question answering scenarios.…

Information Retrieval · Computer Science 2025-09-16 Mengzheng Yang , Yanfei Ren , David Osei Opoku , Ruochang Li , Peng Ren , Chunxiao Xing

Radiology report generation (RRG) aims to automatically produce diagnostic reports from medical images, with the potential to enhance clinical workflows and reduce radiologists' workload. While recent approaches leveraging multimodal large…

Artificial Intelligence · Computer Science 2025-05-16 Ziruo Yi , Ting Xiao , Mark V. Albert

With the advancement of remote sensing satellite technology and the rapid progress of deep learning, remote sensing change detection (RSCD) has become a key technique for regional monitoring. Traditional change detection (CD) methods and…

Image and Video Processing · Electrical Eng. & Systems 2026-03-11 Chengming Wang , Guodong Fan , Jinjiang Li , Min Gan , C. L. Philip Chen

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Hao Zhu , Huaibo Huang , Yi Li , Aihua Zheng , Ran He

Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control capabilities. To…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Xiaoda Yang , Jiayang Xu , Kaixuan Luan , Xinyu Zhan , Hongshun Qiu , Shijun Shi , Hao Li , Shuai Yang , Li Zhang , Checheng Yu , Cewu Lu , Lixin Yang

Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time interactivity, including handling of pauses, interruptions, and…

Computation and Language · Computer Science 2026-05-13 Chung-Ming Chien , Manu Orsini , Eugene Kharitonov , Neil Zeghidour , Karen Livescu , Alexandre Défossez

We propose an online, end-to-end, neural generative conversational model for open-domain dialogue. It is trained using a unique combination of offline two-phase supervised learning and online human-in-the-loop active learning. While most…

Computation and Language · Computer Science 2017-06-19 Nabiha Asghar , Pascal Poupart , Xin Jiang , Hang Li

Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content. In the era of rapidly growing…

Computation and Language · Computer Science 2025-05-20 Xingyu Li , Chen Gong , Guohong Fu