中文
相关论文

相关论文: Retrieval-Augmented Egocentric Video Captioning

200 篇论文

With the rapid development of wearable cameras, a massive collection of egocentric video for first-person visual perception becomes available. Using egocentric videos to predict first-person activity faces many challenges, including limited…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Linfeng Xu , Qingbo Wu , Lili Pan , Fanman Meng , Hongliang Li , Chiyuan He , Hanxin Wang , Shaoxu Cheng , Yu Dai

This paper presents a framework for recognition of human activity from egocentric video and eye tracking data obtained from a head-mounted eye tracker. Three channels of information such as eye movement, ego-motion, and visual features are…

计算机视觉与模式识别 · 计算机科学 2018-05-21 Anjith George , Aurobinda Routray

In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoint changes in egocentric videos cause significant object…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Joohyun Chang , Soyeon Hong , Hyogun Lee , Seong Jong Ha , Dongho Lee , Seong Tae Kim , Jinwoo Choi

We pose keystep recognition as a node classification task, and propose a flexible graph-learning framework for fine-grained keystep recognition that is able to effectively leverage long-term dependencies in egocentric videos. Our approach,…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge. Human-generated queries for video datasets `in the wild' vary a lot in terms of degree of specificity,…

计算机视觉与模式识别 · 计算机科学 2020-02-17 Yang Liu , Samuel Albanie , Arsha Nagrani , Andrew Zisserman

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Yanjun Li , Yuqian Fu , Tianwen Qian , Qi'ao Xu , Silong Dai , Danda Pani Paudel , Luc Van Gool , Xiaoling Wang

In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a…

计算机视觉与模式识别 · 计算机科学 2019-06-05 Wei Zhang , Bairui Wang , Lin Ma , Wei Liu

Pre-training on large scale unlabelled datasets has shown impressive performance improvements in the fields of computer vision and natural language processing. Given the advent of large-scale instructional video datasets, a common strategy…

计算机视觉与模式识别 · 计算机科学 2021-11-04 Valentin Gabeur , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Vishakha Lall , Yisi Liu

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

机器人学 · 计算机科学 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first person perspective. In this paper we address the challenge of egocentric human action recognition by…

计算机视觉与模式识别 · 计算机科学 2019-05-03 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas P. J. J. Noldus , Remco C. Veltkamp

Multimodal large language models (MLLMs) act as essential interfaces, connecting humans with AI technologies in multimodal applications. However, current MLLMs face challenges in accurately interpreting object orientation in images due to…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Ji Hyeok Jung , Eun Tae Kim , Seoyeon Kim , Joo Ho Lee , Bumsoo Kim , Buru Chang

Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contrast, when humans are faced with unfamiliar tasks, such as…

机器人学 · 计算机科学 2025-11-10 Yichen Zhu , Feifei Feng

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Wearable cameras offer a hands-free way to record egocentric images of daily experiences, where social events are of special interest. The first step towards detection of social events is to track the appearance of multiple persons involved…

计算机视觉与模式识别 · 计算机科学 2017-01-24 Maedeh Aghaei , Mariella Dimiccoli , Petia Radeva

Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language encoders and learn…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Shraman Pramanick , Yale Song , Sayan Nag , Kevin Qinghong Lin , Hardik Shah , Mike Zheng Shou , Rama Chellappa , Pengchuan Zhang

We present Ego3DPose, a highly accurate binocular egocentric 3D pose reconstruction system. The binocular egocentric setup offers practicality and usefulness in various applications, however, it remains largely under-explored. It has been…

计算机视觉与模式识别 · 计算机科学 2023-09-22 Taeho Kang , Kyungjin Lee , Jinrui Zhang , Youngki Lee

Can Video-LLMs achieve consistent temporal understanding when videos capture the same event from different viewpoints? To study this, we introduce EgoExo-Con (Consistency), a benchmark of comprehensively synchronized egocentric and…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Minjoon Jung , Junbin Xiao , Junghyun Kim , Byoung-Tak Zhang , Angela Yao

Understanding users' activities from head-mounted cameras is a fundamental task for Augmented and Virtual Reality (AR/VR) applications. A typical approach is to train a classifier in a supervised manner using data labeled by humans. This…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Satoshi Tsutsui , Ruta Desai , Karl Ridgeway