English
Related papers

Related papers: Describing Unseen Videos via Multi-Modal Cooperati…

200 papers

A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through natural language. Here we study how to design artificial…

Recent applications of autonomous agents and robots, such as self-driving cars, scenario-based trainers, exploration robots, and service robots have brought attention to crucial trust-related challenges associated with the current…

Robotics · Computer Science 2022-09-26 Fatai Sado , Chu Kiong Loo , Wei Shiung Liew , Matthias Kerzel , Stefan Wermter

We present dialogue management routines for a system to engage in multiparty agent-infant interaction. The ultimate purpose of this research is to help infants learn a visual sign language by engaging them in naturalistic and socially…

Human-Computer Interaction · Computer Science 2018-09-06 Setareh Nasihati Gilani , David Traum , Arcangelo Merla , Eugenia Hee , Zoey Walker , Barbara Manini , Grady Gallagher , Laura-Ann Petitto

Standard video and movie description tasks abstract away from person identities, thus failing to link identities across sentences. We propose a multi-sentence Identity-Aware Video Description task, which overcomes this limitation and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Jae Sung Park , Trevor Darrell , Anna Rohrbach

Understanding uncertainty plays a critical role in achieving common ground (Clark et al.,1983). This is especially important for multimodal AI systems that collaborate with users to solve a problem or guide the user through a challenging…

Computation and Language · Computer Science 2024-10-21 Qi Cheng , Mert İnan , Rahma Mbarki , Grace Grmek , Theresa Choi , Yiming Sun , Kimele Persaud , Jenny Wang , Malihe Alikhani

Visual dialog has witnessed great progress after introducing various vision-oriented goals into the conversation, especially such as GuessWhich and GuessWhat, where the only image is visible by either and both of the questioner and the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Duo Zheng , Fandong Meng , Qingyi Si , Hairun Fan , Zipeng Xu , Jie Zhou , Fangxiang Feng , Xiaojie Wang

In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xinpeng Li , Shijian Deng , Bolin Lai , Weiguo Pian , James M. Rehg , Yapeng Tian

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully,…

Computer Vision and Pattern Recognition · Computer Science 2019-05-10 Huda Alamri , Vincent Cartillier , Abhishek Das , Jue Wang , Anoop Cherian , Irfan Essa , Dhruv Batra , Tim K. Marks , Chiori Hori , Peter Anderson , Stefan Lee , Devi Parikh

Humans engaged in collaborative activities are naturally able to convey their intentions to teammates through multi-modal communication, which is made up of explicit and implicit cues. Similarly, a more natural form of human-robot…

Robotics · Computer Science 2022-07-01 Simone Macciò , Alessandro Carfì , Fulvio Mastrogiovanni

Currently, dialogue systems have achieved high performance in processing text-based communication. However, they have not yet effectively incorporated visual information, which poses a significant challenge. Furthermore, existing models…

Computation and Language · Computer Science 2023-12-19 Viktor Moskvoretskii , Anton Frolov , Denis Kuznetsov

We aim to develop an AI agent that can watch video clips and have a conversation with human about the video story. Developing video understanding intelligence is a significantly challenging task, and evaluation methods for adequately…

Artificial Intelligence · Computer Science 2021-10-19 Yu-Jung Heo , Minsu Lee , Seongho Choi , Woo Suk Choi , Minjung Shin , Minjoon Jung , Jeh-Kwang Ryu , Byoung-Tak Zhang

High-level understanding of stories in video such as movies and TV shows from raw data is extremely challenging. Modern video question answering (VideoQA) systems often use additional human-made sources like plot synopses, scripts, video…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Deniz Engin , François Schnitzler , Ngoc Q. K. Duong , Yannis Avrithis

We will demonstrate a conversational products recommendation agent. This system shows how we combine research in personalized recommendation systems with research in dialogue systems to build a virtual sales agent. Based on new deep…

Computation and Language · Computer Science 2016-10-06 Yueming Sun , Yi Zhang , Yunfei Chen , Roger Jin

This paper introduces QCaption, a novel video captioning and Q&A pipeline that enhances video analytics by fusing three models: key frame extraction, a Large Multimodal Model (LMM) for image-text analysis, and a Large Language Model (LLM)…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiale Wang , Gee Wah Ng , Lee Onn Mak , Randall Cher , Ng Ding Hei Ryan , Davis Wang

When deploying autonomous agents in the real world, we need effective ways of communicating objectives to them. Traditional skill learning has revolved around reinforcement and imitation learning, each with rigid constraints on the format…

Artificial Intelligence · Computer Science 2019-11-21 Mark Woodward , Chelsea Finn , Karol Hausman

Autonomous vehicles (AVs) are poised to redefine transportation by enhancing road safety, minimizing human error, and optimizing traffic efficiency. The success of AVs depends on their ability to interpret complex, dynamic environments…

Multimedia · Computer Science 2025-07-11 Abolfazl Zarghani , Amirhossein Ebrahimi , Amir Malekesfandiari

As reinforcement learning methods increasingly amass accomplishments, the need for comprehending their solutions becomes more crucial. Most explainable reinforcement learning (XRL) methods generate a static explanation depicting their…

Artificial Intelligence · Computer Science 2025-04-09 Yotam Amitai , Ofra Amir , Guy Avni

A core challenge for an agent learning to interact with the world is to predict how its actions affect objects in its environment. Many existing methods for learning the dynamics of physical interactions require labeled object information.…

Machine Learning · Computer Science 2016-10-19 Chelsea Finn , Ian Goodfellow , Sergey Levine

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represent knowledge using…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-22 Xianghu Yue , Xiaohai Tian , Lu Lu , Malu Zhang , Zhizheng Wu , Haizhou Li

Multimodal AI Agents are AI models that have the capability of interactively and cooperatively assisting human users to solve day-to-day tasks. Augmented Reality (AR) head worn devices can uniquely improve the user experience of solving…

Artificial Intelligence · Computer Science 2025-01-17 Saptarashmi Bandyopadhyay , Vikas Bahirwani , Lavisha Aggarwal , Bhanu Guda , Lin Li , Andrea Colaco