中文
相关论文

相关论文: Olaf-World: Orienting Latent Actions for Video Wor…

200 篇论文

The Visual-Language-Action (VLA) models can follow text instructions according to visual observations of the surrounding environment. This ability to map multimodal inputs to actions is derived from the training of the VLA model on…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Jinyue Bian , Zhaoxing Zhang , Zhengyu Liang , Shiwei Zheng , Shengtao Zhang , Rong Shen , Chen Yang , Anzhou Hou

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Julong Wei , Shanshuai Yuan , Pengfei Li , Qingda Hu , Zhongxue Gan , Wenchao Ding

We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a neural network capable of predicting an agent's next state given past states and actions. By…

We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model and world model in one single framework. The world model…

机器人学 · 计算机科学 2025-06-27 Jun Cen , Chaohui Yu , Hangjie Yuan , Yuming Jiang , Siteng Huang , Jiayan Guo , Xin Li , Yibing Song , Hao Luo , Fan Wang , Deli Zhao , Hao Chen

Although large-scale video-language pre-training models, which usually build a global alignment between the video and the text, have achieved remarkable progress on various downstream tasks, the idea of adopting fine-grained information…

计算机视觉与模式识别 · 计算机科学 2023-11-10 Weihong Zhong , Mao Zheng , Duyu Tang , Xuan Luo , Heng Gong , Xiaocheng Feng , Bing Qin

For training a video-based action recognition model that accepts multi-view video, annotating frame-level labels is tedious and difficult. However, it is relatively easy to annotate sequence-level labels. This kind of coarse annotations are…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Vijay John , Yasutomo Kawanishi

While the shortage of explicit action data limits Vision-Language-Action (VLA) models, human action videos offer a scalable yet unlabeled data source. A critical challenge in utilizing large-scale human video datasets lies in transforming…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Dujun Nie , Fengjiao Chen , Qi Lv , Jun Kuang , Xiaoyu Li , Xuezhi Cao , Xunliang Cai

Reliable perception during fast motion maneuvers or in high dynamic range environments is crucial for robotic systems. Since event cameras are robust to these challenging conditions, they have great potential to increase the reliability of…

计算机视觉与模式识别 · 计算机科学 2022-02-04 Nico Messikommer , Daniel Gehrig , Mathias Gehrig , Davide Scaramuzza

We describe a latent approach that learns to detect actions in long sequences given training videos with only whole-video class labels. Our approach makes use of two innovations to attention-modeling in weakly-supervised learning. First,…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Phuc Xuan Nguyen , Deva Ramanan , Charless C. Fowlkes

Effective leveraging of real-world driving datasets is crucial for enhancing the training of autonomous driving systems. While Offline Reinforcement Learning enables training autonomous vehicles with such data, most available datasets lack…

机器人学 · 计算机科学 2026-01-27 Vinal Asodia , Barkin Dagda , Yinglong He , Zhenhua Feng , Saber Fallah

Stochastic video generation is particularly challenging when the camera is mounted on a moving platform, as camera motion interacts with observed image pixels, creating complex spatio-temporal dynamics and making the problem partially…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Meenakshi Sarkar , Devansh Bhardwaj , Debasish Ghose

Robot manipulation learning from human demonstrations offers a rapid means to acquire skills but often lacks generalization across diverse scenes and object placements. This limitation hinders real-world applications, particularly in…

机器人学 · 计算机科学 2025-05-22 Yihang Li , Tianle Zhang , Xuelong Wei , Jiayi Li , Lin Zhao , Dongchi Huang , Zhirui Fang , Minhua Zheng , Wenjun Dai , Xiaodong He

We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Shuhei M. Yoshida , Takashi Shibata , Makoto Terao , Takayuki Okatani , Masashi Sugiyama

World models have been developed to support sample-efficient deep reinforcement learning agents. However, it remains challenging for world models to accurately replicate environments that are high-dimensional, non-stationary, and composed…

机器学习 · 计算机科学 2026-03-31 Yosuke Nishimoto , Takashi Matsubara

A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Chuhan Li , Ruilin Han , Joy Hsu , Yongyuan Liang , Rajiv Dhawan , Jiajun Wu , Ming-Hsuan Yang , Xin Eric Wang

Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internalize physical dynamics…

Agents that understand objects and their interactions can learn policies that are more robust and transferable. However, most object-centric RL methods factor state by individual objects while leaving interactions implicit. We introduce the…

机器学习 · 计算机科学 2025-11-05 Fan Feng , Phillip Lippe , Sara Magliacane

Learning new skills by observing humans' behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans' decision-making processes, focusing on learning a model to plan goal-directed actions in…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Jing Bi , Jiebo Luo , Chenliang Xu

The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat the VLM primarily as a multimodal encoder, directly mapping…

机器人学 · 计算机科学 2026-04-29 Yi Chen , Yuying Ge , Hui Zhou , Mingyu Ding , Yixiao Ge , Xihui Liu

In this work, we tackle the challenging problem of unsupervised video domain adaptation (UVDA) for action recognition. We specifically focus on scenarios with a substantial domain gap, in contrast to existing works primarily deal with small…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Hyogun Lee , Kyungho Bae , Seong Jong Ha , Yumin Ko , Gyeong-Moon Park , Jinwoo Choi