中文
相关论文

相关论文: Learning to Localize and Align Fine-Grained Action…

200 篇论文

Instruction tuning has emerged as the key in aligning large language models (LLMs) with specific task instructions, thereby mitigating the discrepancy between the next-token prediction objective and users' actual goals. To reduce the labor…

计算与语言 · 计算机科学 2024-04-10 Zifeng Wang , Chun-Liang Li , Vincent Perot , Long T. Le , Jin Miao , Zizhao Zhang , Chen-Yu Lee , Tomas Pfister

This paper presents a new method to describe spatio-temporal relations between objects and hands, to recognize both interactions and activities within video demonstrations of manual tasks. The approach exploits Scene Graphs to extract key…

计算机视觉与模式识别 · 计算机科学 2023-07-10 Elena Merlo , Marta Lagomarsino , Edoardo Lamon , Arash Ajoudani

Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agents should possess the…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Hai Zhang , Siqi Liang , Li Chen , Yuxian Li , Yukuan Xu , Yichao Zhong , Fu Zhang , Hongyang Li

Recent advances in video-large language models (Video-LLMs) have led to significant progress in video understanding. Current preference optimization methods often rely on proprietary APIs or human-annotated captions to generate preference…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yogesh Kulkarni , Pooyan Fazli

Recent technological advances have made lightweight, head mounted cameras both practical and affordable and products like Google Glass show first approaches to introduce the idea of egocentric (first-person) video to the mainstream.…

计算机视觉与模式识别 · 计算机科学 2015-01-14 Sven Bambach

Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domains with minimal supervision remains an open problem. We explore a simple…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Kanchana Ranasinghe , Michael Ryoo

We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentric visual observations of initial and goal states. Unlike…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Fengyi Wu , Yifei Dong , Yilong Dai , Guangyu Chen , Qifeng Wu , Huiting Huang , Hang Wang , Qi Dai , Alexander G. Hauptmann , Zhi-Qi Cheng

To perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a…

机器人学 · 计算机科学 2023-03-01 Chao Tang , Dehao Huang , Lingxiao Meng , Weiyu Liu , Hong Zhang

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Khoa Vo , Thinh Phan , Kashu Yamazaki , Minh Tran , Ngan Le

The potential for agents, whether embodied or software, to learn by observing other agents performing procedures involving objects and actions is rich. Current research on automatic procedure learning heavily relies on action labels or…

计算机视觉与模式识别 · 计算机科学 2017-11-23 Luowei Zhou , Chenliang Xu , Jason J. Corso

Learning to infer labels in an open world, i.e., in an environment where the target ``labels'' are unknown, is an important characteristic for achieving autonomy. Foundation models, pre-trained on enormous amounts of data, have shown…

计算机视觉与模式识别 · 计算机科学 2024-05-06 Sanjoy Kundu , Shubham Trehan , Sathyanarayanan N. Aakur

Understanding human actions is a key problem in computer vision. However, recognizing actions is only the first step of understanding what a person is doing. In this paper, we introduce the problem of predicting why a person has performed…

计算机视觉与模式识别 · 计算机科学 2016-12-01 Carl Vondrick , Deniz Oktay , Hamed Pirsiavash , Antonio Torralba

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Yifan Shen , Jiateng Liu , Xinzhuo Li , Yuanzhe Liu , Bingxuan Li , Houze Yang , Wenqi Jia , Yijiang Li , Tianjiao Yu , James Matthew Rehg , Xu Cao , Ismini Lourentzou

Large language models (LLMs) hold the promise of solving diverse tasks when provided with appropriate natural language prompts. However, prompting often leads models to make predictions with lower accuracy compared to finetuning a model…

计算与语言 · 计算机科学 2024-08-13 Chenyang Zhao , Xueying Jia , Vijay Viswanathan , Tongshuang Wu , Graham Neubig

The ability to actively ground task instructions from an egocentric view is crucial for AI agents to accomplish tasks or assist humans virtually. One important step towards this goal is to localize and track key active objects that undergo…

计算机视觉与模式识别 · 计算机科学 2023-10-24 Te-Lin Wu , Yu Zhou , Nanyun Peng

People often give instructions whose meaning is ambiguous without further context, expecting that their actions or goals will disambiguate their intentions. How can we build assistive agents that follow such instructions in a flexible,…

人工智能 · 计算机科学 2024-02-29 Tan Zhi-Xuan , Lance Ying , Vikash Mansinghka , Joshua B. Tenenbaum

This paper focuses on building object-centric representations for long-term action anticipation in videos. Our key motivation is that objects provide important cues to recognize and predict human-object interactions, especially when the…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Ce Zhang , Changcheng Fu , Shijie Wang , Nakul Agarwal , Kwonjoon Lee , Chiho Choi , Chen Sun

Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors and inefficiency. While Vision-Language Models (VLMs)…

机器人学 · 计算机科学 2026-02-12 Songen Gu , Yunuo Cai , Tianyu Wang , Simo Wu , Yanwei Fu

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

Existing dense or paragraph video captioning approaches rely on holistic representations of videos, possibly coupled with learned object/action representations, to condition hierarchical language decoders. However, they fundamentally lack…

计算机视觉与模式识别 · 计算机科学 2024-01-10 Shih-Han Chou , James J. Little , Leonid Sigal