English
Related papers

Related papers: Enabling Chatbots with Eyes and Ears: An Immersive…

200 papers

Perceiving multi-modal information and fulfilling dialogues with humans is a long-term goal of artificial intelligence. Pre-training is commonly regarded as an effective approach for multi-modal dialogue. However, due to the limited…

Computation and Language · Computer Science 2023-06-14 Yunshui Li , Binyuan Hui , ZhiChao Yin , Min Yang , Fei Huang , Yongbin Li

In this work, we present a conceptually simple yet powerful baseline for the multimodal dialog task, an S3 model, that achieves near state-of-the-art results on two compelling leaderboards: MMMU and AI Journey Contest 2023. The system is…

Computation and Language · Computer Science 2024-06-27 Elisei Rykov , Egor Malkershin , Alexander Panchenko

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

The research community has traditionally shown a keen interest in emotion modeling, with a notable emphasis on the detection aspect. In contrast, the exploration of emotion generation has received less attention.This study delves into an…

Human-Computer Interaction · Computer Science 2024-11-06 Taseen Mubassira , Mehedi Hasan , A. B. M. Alim Al Iislam

Emotion Recognition in Conversations (ERC) is an important and active research area. Recent work has shown the benefits of using multiple modalities (e.g., text, audio, and video) for the ERC task. In a conversation, participants tend to…

Computation and Language · Computer Science 2022-11-08 Harsh Agarwal , Keshav Bansal , Abhinav Joshi , Ashutosh Modi

We introduce the concept of "empathic grounding" in conversational agents as an extension of Clark's conceptualization of grounding in conversation in which the grounding criterion includes listener empathy for the speaker's affective…

Human-Computer Interaction · Computer Science 2024-07-03 Mehdi Arjmand , Farnaz Nouraei , Ian Steenstra , Timothy Bickmore

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update episodic and semantic memories, gradually accumulating…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Lin Long , Yichen He , Wentao Ye , Yiyuan Pan , Yuan Lin , Hang Li , Junbo Zhao , Wei Li

In this paper, we present a novel multi-modal attention guidance method designed to address the challenges of turn-taking dynamics in meetings and enhance group conversations within virtual reality (VR) environments. Recognizing the…

Human-Computer Interaction · Computer Science 2024-06-21 Geonsun Lee , Dae Yeol Lee , Guan-Ming Su , Dinesh Manocha

Most chatbot literature that focuses on improving the fluency and coherence of a chatbot, is dedicated to making chatbots more human-like. However, very little work delves into what really separates humans from chatbots -- humans…

Computation and Language · Computer Science 2021-04-26 Hsuan Su , Jiun-Hao Jhan , Fan-yun Sun , Saurav Sahay , Hung-yi Lee

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models. However, training a well-generalized multi-modal dialogue model remains challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Young-Jun Lee , Byungsoo Ko , Han-Gyu Kim , Jonghwan Hyeon , Ho-Jin Choi

Systems powered by artificial intelligence are being developed to be more user-friendly by communicating with users in a progressively human-like conversational way. Chatbots, also known as dialogue systems, interactive conversational…

Human-Computer Interaction · Computer Science 2020-01-03 Kennedy Ralston , Yuhao Chen , Haruna Isah , Farhana Zulkernine

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-specific characteristics…

Computation and Language · Computer Science 2024-12-23 Maximillian Chen , Ruoxi Sun , Sercan Ö. Arık

Dialogue systems, also called chatbots, are now used in a wide range of applications. However, they still have some major weaknesses. One key weakness is that they are typically trained from manually-labeled data and/or written with…

Computation and Language · Computer Science 2021-02-25 Bing Liu , Sahisnu Mazumder

How human-like do conversational robots need to look to enable long-term human-robot conversation? One essential aspect of long-term interaction is a human's ability to adapt to the varying degrees of a conversational partner's engagement…

Human-Computer Interaction · Computer Science 2022-01-25 Maria Tsfasman , Avinash Saravanan , Dekel Viner , Daan Goslinga , Sarah de Wolf , Chirag Raman , Catholijn M. Jonker , Catharine Oertel

As we build towards developing interactive systems that can recognize human emotional states and respond to individual needs more intuitively and empathetically in more personalized and context-aware computing time. This is especially…

Human-Computer Interaction · Computer Science 2024-06-25 Rahul Islam , Sang Won Bae

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio,…

We present the InterviewBot that dynamically integrates conversation history and customized topics into a coherent embedding space to conduct 10 mins hybrid-domain (open and closed) conversations with foreign students applying to U.S.…

Computation and Language · Computer Science 2023-09-06 Zihao Wang , Nathan Keyes , Terry Crawford , Jinho D. Choi

In natural human-to-human communication, multimodal user input is typically used to supplement explicit and complement implicit voice commands, with casualness allowing for flexible input modality combinations and tolerance for imprecise…

Human-Computer Interaction · Computer Science 2026-05-07 Yen-Ting Liu , Chiu-Hsuan Wang , TzuLing Chen , Ting-Ying Lee , Tzu-Hua Wang , Chien-Ming Lin , Bing-Yu Chen , Hsin-Ruey Tsai

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a…

Sound · Computer Science 2025-04-02 Jiachen Luo , Huy Phan , Lin Wang , Joshua Reiss

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While Multimodal Large Language Models (MLLMs) are natural…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Liangyang Ouyang , Yifei Huang , Mingfang Zhang , Caixin Kang , Ryosuke Furuta , Yoichi Sato