中文
相关论文

相关论文: Enabling Chatbots with Eyes and Ears: An Immersive…

200 篇论文

Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves…

机器学习 · 计算机科学 2022-01-19 Anil Rahate , Rahee Walambe , Sheela Ramanna , Ketan Kotecha

We introduce ChatPose, a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Yao Feng , Jing Lin , Sai Kumar Dwivedi , Yu Sun , Priyanka Patel , Michael J. Black

Multimodal Large Language Models (MLLMs) offer an opportunity to support multimedia learning through conversational systems grounded in educational content. However, while conversational AI is known to boost engagement, its impact on…

人机交互 · 计算机科学 2026-04-03 Karan Taneja , Anjali Singh , Ashok K. Goel

Deep learning methods have revolutionized speech recognition, image recognition, and natural language processing since 2010. Each of these tasks involves a single modality in their input signals. However, many applications in the artificial…

人工智能 · 计算机科学 2020-07-15 Chao Zhang , Zichao Yang , Xiaodong He , Li Deng

Affective tactile interaction constitutes a fundamental component of human communication. In natural human-human encounters, touch is seldom experienced in isolation; rather, it is inherently multisensory. Individuals not only perceive the…

机器人学 · 计算机科学 2025-10-09 Qiaoqiao Ren , Tony Belpaeme

A conversational agent (chatbot) is a piece of software that is able to communicate with humans using natural language. Modeling conversation is an important task in natural language processing and artificial intelligence. While chatbots…

计算与语言 · 计算机科学 2019-08-26 Richard Csaky

It may be difficult for some individuals to open up and share their thoughts and feelings in front of a mental health expert. For those who are more at ease with a virtual agent, conversational agents can serve as an intermediate step in…

计算与语言 · 计算机科学 2021-11-17 Harsh Sakhrani , Saloni Parekh , Shubham Mahajan

As chatbots are becoming increasingly popular, we often wonder what users perceive as natural and socially accepted manners of interacting with them. Some researchers maintain that humans should avoid engaging in emotional conversations…

人机交互 · 计算机科学 2020-09-18 Ekaterina Svikhnushina , Pearl Pu

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

机器学习 · 计算机科学 2024-05-01 Paul Pu Liang

This paper presents a research platform that supports spoken dialogue interaction with multiple robots. The demonstration showcases our crafted MultiBot testing scenario in which users can verbally issue search, navigate, and follow…

机器人学 · 计算机科学 2019-10-15 Matthew Marge , Stephen Nogar , Cory J. Hayes , Stephanie M. Lukin , Jesse Bloecker , Eric Holder , Clare Voss

The lack of time-efficient and reliable evaluation methods hamper the development of conversational dialogue systems (chatbots). Evaluations requiring humans to converse with chatbots are time and cost-intensive, put high cognitive demands…

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating…

人机交互 · 计算机科学 2025-09-19 Taesoo Kim , Yongsik Jo , Hyunmin Song , Taehwan Kim

Full-duplex dialog models aim to listen and speak simultaneously, delivering rapid responses to dynamic user input. Among different solutions to full-duplexity, a native solution merges multiple channels in each time step, achieving the…

声音 · 计算机科学 2026-02-02 Yiqun Yao , Xiang Li , Xin Jiang , Xuezhi Fang , Naitong Yu , Wenjia Ma , Aixin Sun , Yequan Wang

In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Xinpeng Li , Shijian Deng , Bolin Lai , Weiguo Pian , James M. Rehg , Yapeng Tian

In this paper, we present a methodology for the development of embodied conversational agents for social virtual worlds. The agents provide multimodal communication with their users in which speech interaction is included. Our proposal…

音频与语音处理 · 电气工程与系统科学 2025-01-29 D. Griol , A. Sanchis , J. M. Molina , Z. Callejas

Many promising applications of multimodal wearables require continuous sensing and heavy computation, yet users reject such devices due to privacy concerns. This paper shares our experiences building an ear-mounted voice-and-vision wearable…

人机交互 · 计算机科学 2025-11-26 Yonatan Tussa , Andy Heredia , Nirupam Roy

In this paper, we introduce EmpBot: an end-to-end empathetic chatbot. Empathetic conversational agents should not only understand what is being discussed, but also acknowledge the implied feelings of the conversation partner and respond…

计算与语言 · 计算机科学 2021-11-02 Emmanouil Zaranis , Georgios Paraskevopoulos , Athanasios Katsamanis , Alexandros Potamianos

The goal of building intelligent dialogue systems has largely been separately pursued under two motives: task-oriented dialogue (TOD) systems, and open-domain systems for chit-chat (CC). Although previous TOD dialogue systems work well in…

计算与语言 · 计算机科学 2022-05-13 Changhong Yu , Chunhong Zhang , Qi Sun

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

人工智能 · 计算机科学 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda