中文
相关论文

相关论文: Retrieval-Augmented Egocentric Video Captioning

200 篇论文

Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameterization.…

计算与语言 · 计算机科学 2025-07-29 George Ibrahim , Rita Ramos , Yova Kementchedjhieva

Assessing human skill levels in complex activities is a challenging problem with applications in sports, rehabilitation, and training. In this work, we present SkillFormer, a parameter-efficient architecture for unified multi-view…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Edoardo Bianchi , Antonio Liotta

Neuro-symbolic representations have proved effective in learning structure information in vision and language. In this paper, we propose a new model architecture for learning multi-modal neuro-symbolic representations for video captioning.…

计算机视觉与模式识别 · 计算机科学 2020-11-20 Hassan Akbari , Hamid Palangi , Jianwei Yang , Sudha Rao , Asli Celikyilmaz , Roland Fernandez , Paul Smolensky , Jianfeng Gao , Shih-Fu Chang

Egocentric interaction recognition aims to recognize the camera wearer's interactions with the interactor who faces the camera wearer in egocentric videos. In such a human-human interaction analysis problem, it is crucial to explore the…

计算机视觉与模式识别 · 计算机科学 2019-06-03 Haoxin Li , Yijun Cai , Wei-Shi Zheng

Recent progress in legged locomotion has allowed highly dynamic and parkour-like behaviors for robots, similar to their biological counterparts. Yet, these methods mostly rely on egocentric (first-person) perception, limiting their…

机器人学 · 计算机科学 2025-12-01 Rémy Rahem , Wael Suleiman

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Tushar Nagarajan , Kristen Grauman

The rapid evolution of egocentric video analysis brings new insights into understanding human activities and intentions from a first-person perspective. Despite this progress, the fragmentation in tasks like action recognition, procedure…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Jing Bi , Yunlong Tang , Luchuan Song , Ali Vosoughi , Nguyen Nguyen , Chenliang Xu

Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Jiaxuan Li , Duc Minh Vo , Akihiro Sugimoto , Hideki Nakayama

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

In a wearable camera video, we see what the camera wearer sees. While this makes it easy to know roughly what he chose to look at, it does not immediately reveal when he was engaged with the environment. Specifically, at what moments did…

计算机视觉与模式识别 · 计算机科学 2016-04-05 Yu-Chuan Su , Kristen Grauman

Humans can rearrange objects in cluttered environments using egocentric perception, navigating occlusions without global coordinates. Inspired by this capability, we study long-horizon multi-object non-prehensile rearrangement for mobile…

机器人学 · 计算机科学 2026-02-23 Boyuan An , Zhexiong Wang , Yipeng Wang , Jiaqi Li , Sihang Li , Jing Zhang , Chen Feng

For reliable autonomous robot navigation in urban settings, the robot must have the ability to identify semantically traversable terrains in the image based on the semantic understanding of the scene. This reasoning ability is based on…

机器人学 · 计算机科学 2024-12-30 Yunho Kim , Jeong Hyun Lee , Choongin Lee , Juhyeok Mun , Donghoon Youm , Jeongsoo Park , Jemin Hwangbo

We present a video summarization approach for egocentric or "wearable" camera data. Given hours of video, the proposed method produces a compact storyboard summary of the camera wearer's day. In contrast to traditional keyframe selection…

计算机视觉与模式识别 · 计算机科学 2015-05-20 Yong Jae Lee , Kristen Grauman

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Zihui Xue , Kristen Grauman

Text-video retrieval contains various challenges, including biases coming from diverse sources. We highlight some of them supported by illustrations to open a discussion. Besides, we address one of the biases, frame length bias, with a…

计算机视觉与模式识别 · 计算机科学 2023-06-08 Burak Satar , Hongyuan Zhu , Hanwang Zhang , Joo Hwee Lim

A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain…

多媒体 · 计算机科学 2024-06-21 Yuchen Yang , Yingxuan Duan

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Sagnik Majumder , Hao Jiang , Pierre Moulon , Ethan Henderson , Paul Calamia , Kristen Grauman , Vamsi Krishna Ithapu

People often struggle to remember specific details of past experiences, which can lead to the need to revisit these memories. Consequently, lifelog retrieval has emerged as a crucial application. Various studies have explored methods to…

信息检索 · 计算机科学 2025-10-07 Yu-Fei Shih , An-Zi Yen , Hen-Hsen Huang , Hsin-Hsi Chen

With the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jinjie Mai , Abdullah Hamdi , Silvio Giancola , Chen Zhao , Bernard Ghanem

While significant progress has been made in the image captioning task, video description is still in its infancy due to the complex nature of video data. Generating multi-sentence descriptions for long videos is even more challenging. Among…

计算机视觉与模式识别 · 计算机科学 2019-04-17 Jae Sung Park , Marcus Rohrbach , Trevor Darrell , Anna Rohrbach