English
Related papers

Related papers: Egocentric Video Task Translation @ Ego4D Challeng…

200 papers

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera…

Generating long, coherent egocentric videos is difficult, as hand-object interactions and procedural tasks require reliable long-term memory. Existing autoregressive models suffer from content drift, where object identity and scene…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Liuzhou Zhang , Jiarui Ye , Yuanlei Wang , Ming Zhong , Mingju Cao , Wanke Xia , Bowen Zeng , Zeyu Zhang , Hao Tang

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Shervin Ardeshir , Ali Borji

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Karttikeya Mangalam , Raiymbek Akshulakov , Jitendra Malik

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Pengzhan Sun , Junbin Xiao , Tze Ho Elden Tse , Yicong Li , Arjun Akula , Angela Yao

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

In most settings of practical concern, machine learning practitioners know in advance what end-task they wish to boost with auxiliary tasks. However, widely used methods for leveraging auxiliary data like pre-training and its…

Machine Learning · Computer Science 2022-02-08 Lucio M. Dery , Paul Michel , Ameet Talwalkar , Graham Neubig

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid egomotion and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Runjia Li , Moayed Haji-Ali , Ashkan Mirzaei , Chaoyang Wang , Arpit Sahni , Ivan Skorokhodov , Aliaksandr Siarohin , Tomas Jakab , Junlin Han , Sergey Tulyakov , Philip Torr , Willi Menapace

We pose keystep recognition as a node classification task, and propose a flexible graph-learning framework for fine-grained keystep recognition that is able to effectively leverage long-term dependencies in egocentric videos. Our approach,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-language transformer models do not explicitly fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2022-05-19 Alex Jinpeng Wang , Yixiao Ge , Guanyu Cai , Rui Yan , Xudong Lin , Ying Shan , Xiaohu Qie , Mike Zheng Shou

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-body action…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Yuan-Ming Li , Wei-Jin Huang , An-Lan Wang , Ling-An Zeng , Jing-Ke Meng , Wei-Shi Zheng

Multilingual machine translation addresses the task of translating between multiple source and target languages. We propose task-specific attention models, a simple but effective technique for improving the quality of sequence-to-sequence…

Computation and Language · Computer Science 2018-06-11 Graeme Blackwood , Miguel Ballesteros , Todd Ward

Using an ego-centric camera to do localization and tracking is highly needed for urban navigation and indoor assistive system when GPS is not available or not accurate enough. The traditional hand-designed feature tracking and estimation…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Liang Yang , Hao Jiang , Jizhong Xiao , Zhouyuan Huo

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos…

Robotics · Computer Science 2024-11-01 Simar Kareer , Dhruv Patel , Ryan Punamiya , Pranay Mathur , Shuo Cheng , Chen Wang , Judy Hoffman , Danfei Xu

Egocentric 3D human pose estimation (HPE) from images is challenging due to severe self-occlusions and strong distortion introduced by the fish-eye view from the head mounted camera. Although existing works use intermediate heatmap-based…

Computer Vision and Pattern Recognition · Computer Science 2022-06-13 Jinman Park , Kimathi Kaai , Saad Hossain , Norikatsu Sumi , Sirisha Rambhatla , Paul Fieguth

Pretraining and multitask learning are widely used to improve the speech to text translation performance. In this study, we are interested in training a speech to text translation model along with an auxiliary text to text translation task.…

Computation and Language · Computer Science 2021-07-14 Yun Tang , Juan Pino , Xian Li , Changhan Wang , Dmitriy Genzel

We explore multitask models for neural translation of speech, augmenting them in order to reflect two intuitive notions. First, we introduce a model where the second task decoder receives information from the decoder of the first task,…

Computation and Language · Computer Science 2018-04-27 Antonios Anastasopoulos , David Chiang

Inspired by the success of transformer-based pre-training methods on natural language tasks and further computer vision tasks, researchers have begun to apply transformer to video processing. This survey aims to give a comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Ludan Ruan , Qin Jin

Entity-aware machine translation (EAMT) is a complicated task in natural language processing due to not only the shortage of translation data related to the entities needed to translate but also the complexity in the context needed to…

Computation and Language · Computer Science 2025-06-24 An Trieu , Phuong Nguyen , Minh Le Nguyen