English
Related papers

Related papers: FoundationMotion: Auto-Labeling and Reasoning abou…

200 papers

Temporal Action Detection and Moment Retrieval constitute two pivotal tasks in video understanding, focusing on precisely localizing temporal segments corresponding to specific actions or events. Recent advancements introduced Moment…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Weijun Zhuang , Qizhang Li , Xin Li , Ming Liu , Xiaopeng Hong , Feng Gao , Fan Yang , Wangmeng Zuo

Recent advances in 3D human motion and language integration have primarily focused on text-to-motion generation, leaving the task of motion understanding relatively unexplored. We introduce Dense Motion Captioning, a novel task that aims to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Shiyao Xu , Benedetta Liberatori , Gül Varol , Paolo Rota

The success of foundation models in language has inspired a new wave of general-purpose models for human mobility. However, existing approaches struggle to scale effectively due to two fundamental limitations: a failure to use meaningful…

Artificial Intelligence · Computer Science 2025-11-25 Chonghua Han , Yuan Yuan , Jingtao Ding , Jie Feng , Fanjin Meng , Yong Li

Recently, improving the reasoning ability of large multimodal models (LMMs) through reinforcement learning has made great progress. However, most existing works are based on highly reasoning-intensive datasets such as mathematics and code,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Xingjian Zhang , Siwei Wen , Wenjun Wu , Lei Huang

Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedies often rely on external simulators, teacher models, or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Bo Jiang , Depu Meng , Yihan Hu , Yichen Xie , Tianshuo Xu , Wei Zhan

Motion segmentation in dynamic scenes is highly challenging, as conventional methods heavily rely on estimating camera poses and point correspondences from inherently noisy motion cues. Existing statistical inference or iterative…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Xiankang He , Peile Lin , Ying Cui , Dongyan Guo , Chunhua Shen , Xiaoqin Zhang

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wanyue Zhang , Wenxiang Wu , Wang Xu , Jiaxin Luo , Helu Zhi , Yibin Huang , Shuo Ren , Zitao Liu , Jiajun Zhang

As multimodal large language models (MLLMs) frequently exhibit errors in complex video reasoning scenarios, correcting these errors is critical for uncovering their weaknesses and improving performance. However, existing benchmarks lack…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Xusen Hei , Jiali Chen , Jinyu Yang , Mengchen Zhao , Yi Cai

Foundation models have demonstrated a great ability to achieve general human-level intelligence far beyond traditional approaches. As the technique keeps attracting attention from the AI community, an increasing number of foundation models…

Computation and Language · Computer Science 2024-05-07 Shizhe Diao , Rui Pan , Hanze Dong , Ka Shun Shum , Jipeng Zhang , Wei Xiong , Tong Zhang

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a notable deficiency…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Jinglei Zhang , Yuanfan Guo , Rolandos Alexandros Potamias , Jiankang Deng , Hang Xu , Chao Ma

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

High-quality motion data underpins games, film, XR, and robotics. Vision-based motion capture tools have made significant progress, offering accessible and visually convincing results, yet often fall short in the final stretch -- the last…

Graphics · Computer Science 2026-01-28 Tianxin Tao , Han Liu , Hung Yu Ling

Recent advances in large language models (LLMs) have improved reasoning in text and image domains, yet achieving robust video reasoning remains a significant challenge. Existing video benchmarks mainly assess shallow understanding and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Xuchen Li , Xuzhao Li , Shiyu Hu , Kaiqi Huang , Wentao Zhang

Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Long Qian , Juncheng Li , Yu Wu , Yaobo Ye , Hao Fei , Tat-Seng Chua , Yueting Zhuang , Siliang Tang

We propose a deep neural network for the prediction of future frames in natural video sequences. To effectively handle complex evolution of pixels in videos, we propose to decompose the motion and content, two key components generating…

Computer Vision and Pattern Recognition · Computer Science 2018-01-09 Ruben Villegas , Jimei Yang , Seunghoon Hong , Xunyu Lin , Honglak Lee

Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion representations. However, continuous representations entangle…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Dawei Guan , Di Yang , Chengjie Jin , Jiangtao Wang

While deep convolutional neural networks frequently approach or exceed human-level performance at benchmark tasks involving static images, extending this success to moving images is not straightforward. Having models which can learn to…

Computer Vision and Pattern Recognition · Computer Science 2017-02-07 Tegan Maharaj , Nicolas Ballas , Anna Rohrbach , Aaron Courville , Christopher Pal

Learning from (procedural) videos has increasingly served as a pathway for embodied agents to acquire skills from human demonstrations. To do this, video understanding models must be able to obtain structured understandings, such as the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zitian Tang , Rohan Myer Krishnan , Zhiqiu Yu , Chen Sun

Feature matching across video streams remains a cornerstone challenge in computer vision. Increasingly, robust multimodal matching has garnered interest in robotics, surveillance, remote sensing, and medical imaging. While traditional rely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Jie Wang , Chen Ye Gan , Caoqi Wei , Jiangtao Wen , Yuxing Han

Foundation Models (FMs), e.g., large language models, possess attributes of intelligence which offer promise to endow a robot with the contextual understanding necessary to navigate complex, unstructured tasks in the wild. We see three core…