English
Related papers

Related papers: Learning State-Aware Visual Representations from A…

200 papers

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Junhyeok Kim , Min Soo Kim , Jiwan Chung , Jungbin Cho , Jisoo Kim , Sungwoong Kim , Gyeongbo Sim , Youngjae Yu

We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create automatic video…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Yue Zhao , Ishan Misra , Philipp Krähenbühl , Rohit Girdhar

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bingwen Zhu , Yuqian Fu , Qiaole Dong , Guolei Sun , Tianwen Qian , Yuzheng Wu , Danda Pani Paudel , Xiangyang Xue , Yanwei Fu

While visual imitation learning offers one of the most effective ways of learning from visual demonstrations, generalizing from them requires either hundreds of diverse demonstrations, task specific priors, or large, hard-to-train…

Robotics · Computer Science 2021-12-07 Jyothish Pari , Nur Muhammad Shafiullah , Sridhar Pandian Arunachalam , Lerrel Pinto

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Tushar Nagarajan , Yanghao Li , Christoph Feichtenhofer , Kristen Grauman

We propose the use of a proportional-derivative (PD) control based policy learned via reinforcement learning (RL) to estimate and forecast 3D human pose from egocentric videos. The method learns directly from unsegmented egocentric videos…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Ye Yuan , Kris Kitani

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person…

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Triantafyllos Afouras , Andrew Owens , Joon Son Chung , Andrew Zisserman

With the rapid development of artificial intelligence technologies and wearable devices, egocentric vision understanding has emerged as a new and challenging research direction, gradually attracting widespread attention from both academia…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Xiang Li , Heqian Qiu , Lanxiao Wang , Hanwen Zhang , Chenghao Qi , Linfeng Han , Huiyu Xiong , Hongliang Li

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

Machine Learning · Computer Science 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

Computer Vision and Pattern Recognition · Computer Science 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Fiona Ryan , Hao Jiang , Abhinav Shukla , James M. Rehg , Vamsi Krishna Ithapu

"Looking for things" is a mundane but critical task we repeatedly carry on in our daily life. We introduce a method to develop a human character capable of searching for a randomly located target object in a detailed 3D scene using its…

Robotics · Computer Science 2021-09-16 Maks Sorokin , Wenhao Yu , Sehoon Ha , C. Karen Liu

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Yifei Huang , Guo Chen , Jilan Xu , Mingfang Zhang , Lijin Yang , Baoqi Pei , Hongjie Zhang , Lu Dong , Yali Wang , Limin Wang , Yu Qiao

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Analysis and interpretation of egocentric video data is becoming more and more important with the increasing availability and use of wearable cameras. Exploring and fully understanding affinities and differences between ego and allo (or…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Gaurvi Goyal , Nicoletta Noceti , Francesca Odone , Alessandra Sciutti

We present a new computational model for gaze prediction in egocentric videos by exploring patterns in temporal shift of gaze fixations (attention transition) that are dependent on egocentric manipulation tasks. Our assumption is that the…

Computer Vision and Pattern Recognition · Computer Science 2018-12-05 Yifei Huang , Minjie Cai , Zhenqiang Li , Yoichi Sato

Generating long, coherent egocentric videos is difficult, as hand-object interactions and procedural tasks require reliable long-term memory. Existing autoregressive models suffer from content drift, where object identity and scene…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Liuzhou Zhang , Jiarui Ye , Yuanlei Wang , Ming Zhong , Mingju Cao , Wanke Xia , Bowen Zeng , Zeyu Zhang , Hao Tang
‹ Prev 1 4 5 6 7 8 10 Next ›