中文
相关论文

相关论文: EgoTaskQA: Understanding Human Tasks in Egocentric…

200 篇论文

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Yifei Huang , Guo Chen , Jilan Xu , Mingfang Zhang , Lijin Yang , Baoqi Pei , Hongjie Zhang , Lu Dong , Yali Wang , Limin Wang , Yu Qiao

We introduce NExT-QA, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions. Based on the dataset, we set up multi-choice and open-ended QA tasks…

计算机视觉与模式识别 · 计算机科学 2021-05-25 Junbin Xiao , Xindi Shang , Angela Yao , Tat-Seng Chua

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Understanding fine-grained temporal dynamics is crucial in egocentric videos, where continuous streams capture frequent, close-up interactions with objects. In this work, we bring to light that current egocentric video question-answering…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Chiara Plizzari , Alessio Tonioni , Yongqin Xian , Achin Kulshrestha , Federico Tombari

In this work we employ multitask learning to capitalize on the structure that exists in related supervised tasks to train complex neural networks. It allows training a network for multiple objectives in parallel, in order to improve…

计算机视觉与模式识别 · 计算机科学 2019-09-17 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas Noldus , Remco Veltkamp

Modern video understanding systems excel at tasks such as scene classification, object detection, and short video retrieval. However, as video analysis becomes increasingly central to real-world applications, there is a growing need for…

人工智能 · 计算机科学 2025-05-21 Sahil Shah , Harsh Goel , Sai Shankar Narasimhan , Minkyu Choi , S P Sharan , Oguzhan Akcin , Sandeep Chinchali

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Tushar Nagarajan , Kristen Grauman

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Junsheng Huang , Shengyu Hao , Bocheng Hu , Hongwei Wang , Gaoang Wang

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

机器人学 · 计算机科学 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess the cognitive capabilities required for high-level tasks,…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Chenglin Li , Qianglong Chen , Zhi Li , Feng Tao , Yin Zhang

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to focus on surface…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Aaron Foss , Chloe Evans , Sasha Mitts , Koustuv Sinha , Ammar Rizvi , Justine T. Kao

Video Question Answering is a challenging task, which requires the model to reason over multiple frames and understand the interaction between different objects to answer questions based on the context provided within the video, especially…

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Vishakha Lall , Yisi Liu

Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learning-based action…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Thanh-Dat Truong , Khoa Luu

Visual-based human action recognition can be found in various application fields, e.g., surveillance systems, sports analytics, medical assistive technologies, or human-robot interaction frameworks, and it concerns the identification and…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Antonios Gasteratos , Stavros N. Moutsis , Konstantinos A. Tsintotas , Yiannis Aloimonos

Deploying humanoid robots in real-world settings is fundamentally challenging, as it demands tight integration of perception, locomotion, and manipulation under partial-information observations and dynamically changing environments. As well…

机器人学 · 计算机科学 2026-02-05 Yu Bai , MingMing Yu , Chaojie Li , Ziyi Bai , Xinlong Wang , Börje F. Karlsson

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury