中文
相关论文

相关论文: Multimodal Datasets and Benchmarks for Reasoning a…

200 篇论文

Multimodal question answering tasks can be used as proxy tasks to study systems that can perceive and reason about the world. Answering questions about different types of input modalities stresses different aspects of reasoning such as…

计算与语言 · 计算机科学 2019-11-22 Haytham M. Fayek , Justin Johnson

From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot learning datasets and benchmarks mainly focus on stationary…

机器人学 · 计算机科学 2025-10-13 Zhenyu Zhao , Hongyi Jing , Xiawei Liu , Jiageng Mao , Abha Jha , Hanwen Yang , Rong Xue , Sergey Zakharor , Vitor Guizilini , Yue Wang

Animals perceive the world to plan their actions and interact with other agents to accomplish complex tasks, demonstrating capabilities that are still unmatched by AI systems. To advance our understanding and reduce the gap between the…

Recently, the concept of embodied intelligence has been widely accepted and popularized, leading people to naturally consider the potential for commercialization in this field. In this work, we propose a specific commercial scenario…

机器人学 · 计算机科学 2024-06-27 Zhuoqun Xu , Yang Liu , Xiaoqi Li , Jiyao Zhang , Hao Dong

Existing approaches to video understanding, mainly designed for short videos from a third-person perspective, are limited in their applicability in certain fields, such as robotics. In this paper, we delve into open-ended question-answering…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Shangzhe Di , Weidi Xie

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-specific characteristics…

计算与语言 · 计算机科学 2024-12-23 Maximillian Chen , Ruoxi Sun , Sercan Ö. Arık

Time is an important dimension in our physical world. Lots of facts can evolve with respect to time. For example, the U.S. President might change every four years. Therefore, it is important to consider the time dimension and empower the…

计算与语言 · 计算机科学 2021-10-26 Wenhu Chen , Xinyi Wang , William Yang Wang

Humans with an average level of social cognition can infer the beliefs of others based solely on the nonverbal communication signals (e.g. gaze, gesture, pose and contextual information) exhibited during social interactions. This social…

计算机视觉与模式识别 · 计算机科学 2022-06-23 Jiafei Duan , Samson Yu , Nicholas Tan , Li Yi , Cheston Tan

Human-centered artificial intelligence (AI) posits that machine learning and AI should be developed and applied in a socially aware way. In this article, we argue that qualitative analysis (QA) can be a valuable tool in this process,…

Embodied conversational agents (ECAs) are increasingly more realistic and capable of dynamic conversations. In online surveys, anthropomorphic agents could help address issues like careless responding and satisficing, which originate from…

人机交互 · 计算机科学 2025-08-05 Matus Krajcovic , Peter Demcak , Eduard Kuric

While current visual captioning models have achieved impressive performance, they often assume that the image is well-captured and provides a complete view of the scene. In real-world scenarios, however, a single image may not offer a good…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Anwen Hu , Shizhe Chen , Liang Zhang , Qin Jin

Humans and animals excel in combining information from multiple sensory modalities, controlling their complex bodies, adapting to growth, failures, or using tools. These capabilities are also highly desirable in robots. They are displayed…

机器人学 · 计算机科学 2022-11-08 Matej Hoffmann

Learning meaningful and compact representations with disentangled semantic aspects is considered to be of key importance in representation learning. Since real-world data is notoriously costly to collect, many recent state-of-the-art…

Question Answering (QA) is key for making possible a robust communication between human and machine. Modern language models used for QA have surpassed the human-performance in several essential tasks; however, these models require large…

计算与语言 · 计算机科学 2021-09-08 Liubov Nikolenko , Pouya Rezazadeh Kalehbasti

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronized gaze, speech, and…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Anna Deichler , Jonas Beskow

Spatiotemporal relationships are critical in data science, as many prediction and reasoning tasks require analysis across both spatial and temporal dimensions--for instance, navigating an unfamiliar city involves planning itineraries that…

机器学习 · 计算机科学 2025-05-19 Xiao Han , Dayan Pan , Xiangyu Zhao , Xuyuan Hu , Zhaolin Deng , Xiangjie Kong , Guojiang Shen

As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal expressions are…

机器学习 · 计算机科学 2022-05-17 Anirudh Sundar , Larry Heck

Our goal is a teachable reasoning system for question-answering (QA), where a user can interact with faithful answer explanations, and correct its errors so that the system improves over time. Our approach is to augment a QA model with a…

计算与语言 · 计算机科学 2022-10-25 Bhavana Dalvi Mishra , Oyvind Tafjord , Peter Clark

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

计算机视觉与模式识别 · 计算机科学 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

Large vision-language models have recently demonstrated impressive performance in planning and control tasks, driving interest in their application to real-world robotics. However, deploying these models for reasoning in embodied contexts…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Karmesh Yadav , Yusuf Ali , Gunshi Gupta , Yarin Gal , Zsolt Kira