English
Related papers

Related papers: Learning to Localize and Align Fine-Grained Action…

200 papers

Instruction tuning has emerged as the key in aligning large language models (LLMs) with specific task instructions, thereby mitigating the discrepancy between the next-token prediction objective and users' actual goals. To reduce the labor…

Computation and Language · Computer Science 2024-04-10 Zifeng Wang , Chun-Liang Li , Vincent Perot , Long T. Le , Jin Miao , Zizhao Zhang , Chen-Yu Lee , Tomas Pfister

This paper presents a new method to describe spatio-temporal relations between objects and hands, to recognize both interactions and activities within video demonstrations of manual tasks. The approach exploits Scene Graphs to extract key…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Elena Merlo , Marta Lagomarsino , Edoardo Lamon , Arash Ajoudani

Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agents should possess the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Hai Zhang , Siqi Liang , Li Chen , Yuxian Li , Yukuan Xu , Yichao Zhong , Fu Zhang , Hongyang Li

Recent advances in video-large language models (Video-LLMs) have led to significant progress in video understanding. Current preference optimization methods often rely on proprietary APIs or human-annotated captions to generate preference…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Yogesh Kulkarni , Pooyan Fazli

Recent technological advances have made lightweight, head mounted cameras both practical and affordable and products like Google Glass show first approaches to introduce the idea of egocentric (first-person) video to the mainstream.…

Computer Vision and Pattern Recognition · Computer Science 2015-01-14 Sven Bambach

Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domains with minimal supervision remains an open problem. We explore a simple…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Kanchana Ranasinghe , Michael Ryoo

We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentric visual observations of initial and goal states. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Fengyi Wu , Yifei Dong , Yilong Dai , Guangyu Chen , Qifeng Wu , Huiting Huang , Hang Wang , Qi Dai , Alexander G. Hauptmann , Zhi-Qi Cheng

To perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a…

Robotics · Computer Science 2023-03-01 Chao Tang , Dehao Huang , Lingxiao Meng , Weiyu Liu , Hong Zhang

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Khoa Vo , Thinh Phan , Kashu Yamazaki , Minh Tran , Ngan Le

The potential for agents, whether embodied or software, to learn by observing other agents performing procedures involving objects and actions is rich. Current research on automatic procedure learning heavily relies on action labels or…

Computer Vision and Pattern Recognition · Computer Science 2017-11-23 Luowei Zhou , Chenliang Xu , Jason J. Corso

Learning to infer labels in an open world, i.e., in an environment where the target ``labels'' are unknown, is an important characteristic for achieving autonomy. Foundation models, pre-trained on enormous amounts of data, have shown…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Sanjoy Kundu , Shubham Trehan , Sathyanarayanan N. Aakur

Understanding human actions is a key problem in computer vision. However, recognizing actions is only the first step of understanding what a person is doing. In this paper, we introduce the problem of predicting why a person has performed…

Computer Vision and Pattern Recognition · Computer Science 2016-12-01 Carl Vondrick , Deniz Oktay , Hamed Pirsiavash , Antonio Torralba

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Yifan Shen , Jiateng Liu , Xinzhuo Li , Yuanzhe Liu , Bingxuan Li , Houze Yang , Wenqi Jia , Yijiang Li , Tianjiao Yu , James Matthew Rehg , Xu Cao , Ismini Lourentzou

Large language models (LLMs) hold the promise of solving diverse tasks when provided with appropriate natural language prompts. However, prompting often leads models to make predictions with lower accuracy compared to finetuning a model…

Computation and Language · Computer Science 2024-08-13 Chenyang Zhao , Xueying Jia , Vijay Viswanathan , Tongshuang Wu , Graham Neubig

The ability to actively ground task instructions from an egocentric view is crucial for AI agents to accomplish tasks or assist humans virtually. One important step towards this goal is to localize and track key active objects that undergo…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Te-Lin Wu , Yu Zhou , Nanyun Peng

People often give instructions whose meaning is ambiguous without further context, expecting that their actions or goals will disambiguate their intentions. How can we build assistive agents that follow such instructions in a flexible,…

Artificial Intelligence · Computer Science 2024-02-29 Tan Zhi-Xuan , Lance Ying , Vikash Mansinghka , Joshua B. Tenenbaum

This paper focuses on building object-centric representations for long-term action anticipation in videos. Our key motivation is that objects provide important cues to recognize and predict human-object interactions, especially when the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Ce Zhang , Changcheng Fu , Shijie Wang , Nakul Agarwal , Kwonjoon Lee , Chiho Choi , Chen Sun

Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors and inefficiency. While Vision-Language Models (VLMs)…

Robotics · Computer Science 2026-02-12 Songen Gu , Yunuo Cai , Tianyu Wang , Simo Wu , Yanwei Fu

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

Existing dense or paragraph video captioning approaches rely on holistic representations of videos, possibly coupled with learned object/action representations, to condition hierarchical language decoders. However, they fundamentally lack…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Shih-Han Chou , James J. Little , Leonid Sigal