English
Related papers

Related papers: Unified Video Action Model

200 papers

Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set of weights covers replacement, removal, style transfer, and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yongsheng Yu , Ziyun Zeng , Zhiyuan Xiao , Zhenghong Zhou , Hang Hua , Wei Xiong , Jiebo Luo

Video prediction is a fundamental task for various downstream applications, including robotics and world modeling. Although general video prediction models have achieved remarkable performance in standard scenarios, occlusion is still an…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Eliyas Suleyman , Paul Henderson , Eksan Firkat , Nicolas Pugeault

Action recognition models have achieved promising results in understanding instructional videos. However, they often rely on dominant, dataset-specific action sequences rather than true video comprehension, a problem that we define as…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Joochan Kim , Minjoon Jung , Byoung-Tak Zhang

Video anomaly detection (VAD) is an important but challenging task in computer vision. The main challenge rises due to the rarity of training samples to model all anomaly cases. Hence, semi-supervised anomaly detection methods have gotten…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Mohammad Baradaran , Robert Bergevin

Despite the rapid development of video Large Language Models (LLMs), a comprehensive evaluation is still absent. In this paper, we introduce a unified evaluation that encompasses multiple video tasks, including captioning, question and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Shuailin Li , Yuang Zhang , Yucheng Zhao , Qiuyue Wang , Fan Jia , Yingfei Liu , Tiancai Wang

Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical world. On the other line, diffusion models have also shown…

Robotics · Computer Science 2024-11-28 Yanjiang Guo , Yucheng Hu , Jianke Zhang , Yen-Jen Wang , Xiaoyu Chen , Chaochao Lu , Jianyu Chen

Amid growing efforts to leverage advances in large language models (LLMs) and vision-language models (VLMs) for robotics, Vision-Language-Action (VLA) models have recently gained significant attention. By unifying vision, language, and…

Robotics · Computer Science 2025-10-09 Kento Kawaharazuka , Jihoon Oh , Jun Yamada , Ingmar Posner , Yuke Zhu

We introduce UViM, a unified approach capable of modeling a wide range of computer vision tasks. In contrast to previous models, UViM has the same functional form for all tasks; it requires no task-specific modifications which require…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Alexander Kolesnikov , André Susano Pinto , Lucas Beyer , Xiaohua Zhai , Jeremiah Harmsen , Neil Houlsby

Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Ting Yu , Kunhao Fu , Shuhui Wang , Qingming Huang , Jun Yu

Video segmentation aims at partitioning video sequences into meaningful segments based on objects or regions of interest within frames. Current video segmentation models are often derived from image segmentation techniques, which struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Chen Liang , Qiang Guo , Xiaochao Qu , Luoqi Liu , Ting Liu

Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a novel VLA approach that leverages the competitive…

Robotics · Computer Science 2025-12-23 Max Argus , Jelena Bratulic , Houman Masnavi , Maxim Velikanov , Nick Heppert , Abhinav Valada , Thomas Brox

Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response to actions. Vision-language-action (VLA), which repurpose…

Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extension of this journey beyond Vision-to-Action (VA). In practice, driving VLAs have often…

Video prediction is a challenging task. The quality of video frames from current state-of-the-art (SOTA) generative models tends to be poor and generalization beyond the training data is difficult. Furthermore, existing prediction…

Computer Vision and Pattern Recognition · Computer Science 2022-10-14 Vikram Voleti , Alexia Jolicoeur-Martineau , Christopher Pal

Vision-Language-Action (VLA) models rely on current observations, including images, language instructions, and robot states, to predict actions and complete tasks. While accurate visual perception is crucial for precise action prediction…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Cheng Yang , Jianhao Jiao , Lingyi Huang , Jinqi Xiao , Zhexiang Tang , Yu Gong , Yibiao Ying , Yang Sui , Jintian Lin , Wen Huang , Bo Yuan

General-purpose robots capable of performing diverse tasks require synergistic reasoning and acting capabilities. However, recent dual-system approaches, which separate high-level reasoning from low-level acting, often suffer from…

Robotics · Computer Science 2026-03-03 Fanqi Lin , Ruiqian Nai , Yingdong Hu , Jiacheng You , Junming Zhao , Yang Gao

Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Matteo Tomei , Lorenzo Baraldi , Simone Calderara , Simone Bronzin , Rita Cucchiara

Despite the rapid progress, existing works on action understanding focus strictly on one type of action agent, which we call actor---a human adult, ignoring the diversity of actions performed by other actors. To overcome this narrow…

Computer Vision and Pattern Recognition · Computer Science 2017-05-01 Chenliang Xu , Caiming Xiong , Jason J. Corso

Deep neural networks have achieved great success for video analysis and understanding. However, designing a high-performance neural architecture requires substantial efforts and expertise. In this paper, we make the first attempt to let…

Computer Vision and Pattern Recognition · Computer Science 2019-07-11 Wei Peng , Xiaopeng Hong , Guoying Zhao

Action recognition is a key problem in computer vision that labels videos with a set of predefined actions. Capturing both, semantic content and motion, along the video frames is key to achieve high accuracy performance on this task. Most…

Computer Vision and Pattern Recognition · Computer Science 2019-10-23 Xia Huang , Hossein Mousavi , Gemma Roig
‹ Prev 1 8 9 10 Next ›