English
Related papers

Related papers: Unified Video Action Model

200 papers

This paper studies the joint learning of action recognition and temporal localization in long, untrimmed videos. We employ a multi-task learning framework that performs the three highly related steps of action proposal, action recognition,…

Computer Vision and Pattern Recognition · Computer Science 2017-04-05 Yi Zhu , Shawn Newsam

A core challenge for an agent learning to interact with the world is to predict how its actions affect objects in its environment. Many existing methods for learning the dynamics of physical interactions require labeled object information.…

Machine Learning · Computer Science 2016-10-19 Chelsea Finn , Ian Goodfellow , Sergey Levine

The ability to predict future visual observations conditioned on past observations and motor commands can enable embodied agents to plan solutions to a variety of tasks in complex environments. This work shows that we can create good video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Agrim Gupta , Stephen Tian , Yunzhi Zhang , Jiajun Wu , Roberto Martín-Martín , Li Fei-Fei

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Yucheng Hu , Yanjiang Guo , Pengchao Wang , Xiaoyu Chen , Yen-Jen Wang , Jianke Zhang , Koushil Sreenath , Chaochao Lu , Jianyu Chen

Visual imitation learning enables robotic agents to acquire skills by observing expert demonstration videos. In the one-shot setting, the agent generates a policy after observing a single expert demonstration without additional fine-tuning.…

Robotics · Computer Science 2026-01-01 Raktim Gautam Goswami , Prashanth Krishnamurthy , Yann LeCun , Farshad Khorrami

Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a promising way to…

Video analysis tasks such as action recognition have received increasing research interest with growing applications in fields such as smart healthcare, thanks to the introduction of large-scale datasets and deep learning-based…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yuecong Xu , Haozhi Cao , Zhenghua Chen , Xiaoli Li , Lihua Xie , Jianfei Yang

A dominant paradigm for learning-based approaches in computer vision is training generic models, such as ResNet for image recognition, or I3D for video understanding, on large datasets and allowing them to discover the optimal…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Yubo Zhang , Pavel Tokmakov , Martial Hebert , Cordelia Schmid

Anticipating future actions based on spatiotemporal observations is essential in video understanding and predictive computer vision. Moreover, a model capable of anticipating the future has important applications, it can benefit…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Tsung-Ming Tai , Giuseppe Fiameni , Cheng-Kuang Lee , Simon See , Oswald Lanz

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Yuang Wang , Chao Wen , Haoyu Guo , Sida Peng , Minghan Qin , Hujun Bao , Xiaowei Zhou , Ruizhen Hu

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yicheng Liu , Shiduo Zhang , Zibin Dong , Baijun Ye , Tianyuan Yuan , Xiaopeng Yu , Linqi Yin , Chenhao Lu , Junhao Shi , Luca Jiang-Tao Yu , Liangtao Zheng , Tao Jiang , Jingjing Gong , Xipeng Qiu , Hang Zhao

Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may infer objects, relations, or knowledge cues but fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Fuxiang Zhai , Sixiang Chen , Yingjin Li , Shuaibo Li , Jianyu Lai , Tengjun Huang , Lei Zhu

The Meta Video Dataset (MetaVD) provides annotated relations between action classes in major datasets for human action recognition in videos. Although these annotated relations enable dataset augmentation, it is only applicable to those…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Yuya Yoshikawa , Yutaro Shigeto , Masashi Shimbo , Akikazu Takeuchi

Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of massive tokens leading…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Juyi Lin , Amir Taherin , Arash Akbari , Arman Akbari , Lei Lu , Guangyu Chen , Taskin Padir , Xiaomeng Yang , Weiwei Chen , Yiqian Li , Xue Lin , David Kaeli , Pu Zhao , Yanzhi Wang

Recent advances in learned video compression (LVC) have led to significant performance gains, with codecs such as DCVC-RT surpassing the H.266/VVC low-delay mode in compression efficiency. However, existing LVCs still exhibit key…

Image and Video Processing · Electrical Eng. & Systems 2026-03-09 Yichi Zhang , Ruoyu Yang , Fengqing Zhu

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In…

Large language models, trained on extensive corpora, successfully unify diverse linguistic tasks within a single generative framework. Inspired by this, recent works like Large Vision Model (LVM) extend this paradigm to vision by organizing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Lan Chen , Yuchao Gu , Qi Mao

We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the underlying physics of the environment to refine action…

Recognizing group activities is challenging due to the difficulties in isolating individual entities, finding the respective roles played by the individuals and representing the complex interactions among the participants. Individual…

Computer Vision and Pattern Recognition · Computer Science 2015-03-20 Qiang Qiu , Rama Chellappa

We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the representational strengths of model-based approaches, without incurring planning overhead. By…

Machine Learning · Computer Science 2026-02-16 Jashaswimalya Acharjee , Balaraman Ravindran
‹ Prev 1 4 5 6 7 8 10 Next ›