English
Related papers

Related papers: Motion Question Answering via Modular Motion Progr…

200 papers

In order for conversational AI systems to hold more natural and broad-ranging conversations, they will require much more commonsense, including the ability to identify unstated presumptions of their conversational partners. For example, in…

Artificial Intelligence · Computer Science 2021-02-03 Forough Arabshahi , Jennifer Lee , Mikayla Gawarecki , Kathryn Mazaitis , Amos Azaria , Tom Mitchell

Neural Module Networks (NMN) are a compelling method for visual question answering, enabling the translation of a question into a program consisting of a series of reasoning sub-tasks that are sequentially executed on the image to produce…

Computation and Language · Computer Science 2023-10-25 Wafa Aissa , Marin Ferecatu , Michel Crucianu

Learning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate human movement that consistently reflects the temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Quang Nguyen , Tri Le , Baoru Huang , Minh Nhat Vu , Ngan Le , Thieu Vo , Anh Nguyen

Understanding human motion from video is essential for a range of applications, including pose estimation, mesh recovery and action recognition. While state-of-the-art methods predominantly rely on transformer-based architectures, these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Arnab Kumar Mondal , Stefano Alletto , Denis Tome

Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these systems render human…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Yusu Fang , Tiange Xiang , Tian Tan , Narayan Schuetz , Scott Delp , Li Fei-Fei , Ehsan Adeli

Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent motion benchmarks, primarily due to the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Yulu Gan , Ligeng Zhu , Dandan Shan , Baifeng Shi , Hongxu Yin , Boris Ivanovic , Song Han , Trevor Darrell , Jitendra Malik , Marco Pavone , Boyi Li

Visual Question Answering (VQA) is a challenging task that requires cross-modal understanding and reasoning of visual image and natural language question. To inspect the association of VQA models to human cognition, we designed a survey to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Liben Chen , Long Chen , Tian Ellison-Chen , Zhuoyuan Xu

Video Question Answering is a task which requires an AI agent to answer questions grounded in video. This task entails three key challenges: (1) understand the intention of various questions, (2) capturing various elements of the input…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Ahjeong Seo , Gi-Cheon Kang , Joonhan Park , Byoung-Tak Zhang

Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language…

Robotics · Computer Science 2025-11-03 Simindokht Jahangard , Mehrzad Mohammadi , Abhinav Dhall , Hamid Rezatofighi

We marry two powerful ideas: deep representation learning for visual recognition and language understanding, and symbolic program execution for reasoning. Our neural-symbolic visual question answering (NS-VQA) system first recovers a…

Artificial Intelligence · Computer Science 2019-01-16 Kexin Yi , Jiajun Wu , Chuang Gan , Antonio Torralba , Pushmeet Kohli , Joshua B. Tenenbaum

Human motion prediction is essential for tasks such as human motion analysis and human-robot interactions. Most existing approaches have been proposed to realize motion prediction. However, they ignore an important task, the evaluation of…

Computer Vision and Pattern Recognition · Computer Science 2021-07-09 Pengxiang Ding , Jianqin Yin

In this paper, we propose a novel task, Manipulation Question Answering (MQA), where the robot performs manipulation actions to change the environment in order to answer a given question. To solve this problem, a framework consisting of a…

Artificial Intelligence · Computer Science 2023-02-22 Yuhong Deng , Di Guo , Xiaofeng Guo , Naifu Zhang , Huaping Liu , Fuchun Sun

Action quality assessment (AQA) applies computer vision to quantitatively assess the performance or execution of a human action. Current AQA approaches are end-to-end neural models, which lack transparency and tend to be biased because they…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Lauren Okamoto , Paritosh Parmar

Human reasoning can be understood as a cooperation between the intuitive, associative "System-1" and the deliberative, logical "System-2". For existing System-1-like methods in visual activity understanding, it is crucial to integrate…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Xiaoqian Wu , Yong-Lu Li , Jianhua Sun , Cewu Lu

Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce…

Computation and Language · Computer Science 2022-03-29 Yingshan Chang , Mridu Narang , Hisami Suzuki , Guihong Cao , Jianfeng Gao , Yonatan Bisk

Grasping the intricacies of human motion, which involve perceiving spatio-temporal dependence and multi-scale effects, is essential for predicting human motion. While humans inherently possess the requisite skills to navigate this issue, it…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Kedi Lyu , Haipeng Chen , Zhenguang Liu , Yifang Yin , Yukang Lin , Yingying Jiao

Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical…

Artificial Intelligence · Computer Science 2025-02-28 Lei Li , Sen Jia , Jianhao Wang , Zhaochong An , Jiaang Li , Jenq-Neng Hwang , Serge Belongie

Reasoning is a hallmark of human intelligence, enabling adaptive decision-making in complex and unfamiliar scenarios. In contrast, machine intelligence remains bound to training data, lacking the ability to dynamically refine solutions at…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Shaheer U. Saeed , Yipei Wang , Veeru Kasivisvanathan , Brian R. Davidson , Matthew J. Clarkson , Yipeng Hu , Daniel C. Alexander

Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness in human-machine…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Tan-Hanh Pham , Hoang-Nam Le , Phu-Vinh Nguyen , Chris Ngo , Truong-Son Hy

Anticipating human motion depends on two factors: the past motion and the person's intention. While the first factor has been extensively utilized to forecast short sequences of human motion, the second one remains elusive. In this work we…

Computer Vision and Pattern Recognition · Computer Science 2019-12-16 Julian Tanke , Andreas Weber , Juergen Gall