中文
相关论文

相关论文: Bidirectional Action Sequence Learning for Long-te…

200 篇论文

Long video understanding has emerged as an increasingly important yet challenging task in computer vision. Agent-based approaches are gaining popularity for processing long videos, as they can handle extended sequences and integrate various…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Zhuo Zhi , Qiangqiang Wu , Minghe shen , Wenbo Li , Yinchuan Li , Kun Shao , Kaiwen Zhou

Temporal action segmentation (TAS) demands dense temporal supervision, yet most of the annotation cost in untrimmed videos is spent identifying and refining action transitions, where segmentation errors concentrate and small temporal shifts…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Halil Ismail Helvaci , Sen-ching Samson Cheung

Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's future planning and…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Ramanathan Rajendiran , Debaditya Roy , Basura Fernando

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for…

机器人学 · 计算机科学 2025-03-18 Yue Su , Xinyu Zhan , Hongjie Fang , Han Xue , Hao-Shu Fang , Yong-Lu Li , Cewu Lu , Lixin Yang

Visual reinforcement learning has proven effective in solving control tasks with high-dimensional observations. However, extracting reliable and generalizable representations from vision-based observations remains a central challenge.…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Xiaobo Hu , Youfang Lin , Yue Liu , Jinwen Wang , Shuo Wang , Hehe Fan , Kai Lv

Distinguishing if an action is performed as intended or if an intended action fails is an important skill that not only humans have, but that is also important for intelligent systems that operate in human environments. Recognizing if an…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Olga Zatsarynna , Yazan Abu Farha , Juergen Gall

Behaviour prediction function of an autonomous vehicle predicts the future states of the nearby vehicles based on the current and past observations of the surrounding environment. This helps enhance their awareness of the imminent hazards.…

计算机视觉与模式识别 · 计算机科学 2020-08-07 Sajjad Mozaffari , Omar Y. Al-Jarrah , Mehrdad Dianati , Paul Jennings , Alexandros Mouzakitis

Modeling generalized robot control policies poses ongoing challenges for language-guided robot manipulation tasks. Existing methods often struggle to efficiently utilize cross-dataset resources or rely on resource-intensive vision-language…

机器人学 · 计算机科学 2024-11-05 Wenhui Tan , Bei Liu , Junbo Zhang , Ruihua Song , Jianlong Fu

Event cameras deliver visual information characterized by a high dynamic range and high temporal resolution, offering significant advantages in estimating optical flow for complex lighting conditions and fast-moving objects. Current…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Gangwei Xu , Haotong Lin , Zhaoxing Zhang , Hongcheng Luo , Haiyang Sun , Xin Yang

Video-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-scale video-based models incurs considerably higher costs…

计算机视觉与模式识别 · 计算机科学 2023-08-17 Guangyi Chen , Xiao Liu , Guangrun Wang , Kun Zhang , Philip H. S. Torr , Xiao-Ping Zhang , Yansong Tang

We are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Yiming Li , Ziang Cao , Andrew Liang , Benjamin Liang , Luoyao Chen , Hang Zhao , Chen Feng

Can we better anticipate an actor's future actions (e.g. mix eggs) by knowing what commonly happens after his/her current action (e.g. crack eggs)? What if we also know the longer-term goal of the actor (e.g. making egg fried rice)? The…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Qi Zhao , Shijie Wang , Ce Zhang , Changcheng Fu , Minh Quan Do , Nakul Agarwal , Kwonjoon Lee , Chen Sun

Pedestrian behavior prediction is one of the major challenges for intelligent driving systems. Pedestrians often exhibit complex behaviors influenced by various contextual elements. To address this problem, we propose BiPed, a multitask…

计算机视觉与模式识别 · 计算机科学 2021-08-10 Amir Rasouli , Mohsen Rohani , Jun Luo

Our work explores the task of generating future sensor observations conditioned on the past. We are motivated by `predictive coding' concepts from neuroscience as well as robotic applications such as self-driving vehicles. Predictive video…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Tarasha Khurana , Deva Ramanan

Trajectory prediction, as a critical component of autonomous driving systems, has attracted the attention of many researchers. Existing prediction algorithms focus on extracting more detailed scene features or selecting more reasonable…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Wenyi Xiong , Jian Chen , Ziheng Qi

Anticipating future actions in videos is challenging, as the observed frames provide only evidence of past activities, requiring the inference of latent intentions to predict upcoming actions. Existing transformer-based approaches, which…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Tsung-Ming Tai , Sofia Casarin , Andrea Pilzer , Werner Nutt , Oswald Lanz

We introduce a general self-supervised approach to predict the future outputs of a short-range sensor (such as a proximity sensor) given the current outputs of a long-range sensor (such as a camera); we assume that the former is directly…

机器人学 · 计算机科学 2019-01-18 Mirko Nava , Jerome Guzzi , R. Omar Chavez-Garcia , Luca M. Gambardella , Alessandro Giusti

The ability to predict future visual observations conditioned on past observations and motor commands can enable embodied agents to plan solutions to a variety of tasks in complex environments. This work shows that we can create good video…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Agrim Gupta , Stephen Tian , Yunzhi Zhang , Jiajun Wu , Roberto Martín-Martín , Li Fei-Fei

This paper focuses on weakly-supervised action alignment, where only the ordered sequence of video-level actions is available for training. We propose a novel Duration Network, which captures a short temporal window of the video and learns…

计算机视觉与模式识别 · 计算机科学 2020-11-23 Reza Ghoddoosian , Saif Sayed , Vassilis Athitsos

Trajectory prediction is a critical component of autonomous driving, essential for ensuring both safety and efficiency on the road. However, traditional approaches often struggle with the scarcity of labeled data and exhibit suboptimal…

机器人学 · 计算机科学 2025-09-18 Jianxin Shi , Zengqi Peng , Xiaolong Chen , Tianyu Wo , Jun Ma