English
Related papers

Related papers: MoST: Multi-modality Scene Tokenization for Motion…

200 papers

Amodal perception terms the ability of humans to imagine the entire shapes of occluded objects. This gives humans an advantage to keep track of everything that is going on, especially in crowded situations. Typical perception functions,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-02 Jasmin Breitenstein , Tim Fingscheidt

Predicting movement of objects while the action of learning agent interacts with the dynamics of the scene still remains a key challenge in robotics. We propose a multi-layer Long Short Term Memory (LSTM) autoendocer network that predicts…

Machine Learning · Computer Science 2018-10-15 Meenakshi Sarkar , Debasish Ghose

In this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering the visual and linguistic information in two separate…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Yuxin Wang , Hongtao Xie , Shancheng Fang , Jing Wang , Shenggao Zhu , Yongdong Zhang

Predicting the behaviors of other agents on the road is critical for autonomous driving to ensure safety and efficiency. However, the challenging part is how to represent the social interactions between agents and output different possible…

Robotics · Computer Science 2021-09-15 Zhiyu Huang , Xiaoyu Mo , Chen Lv

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we extend this paradigm by leveraging motion and audio that…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Rui Qian , Yeqing Li , Zheng Xu , Ming-Hsuan Yang , Serge Belongie , Yin Cui

World models play a crucial role in understanding and predicting the dynamics of the world, which is essential for video generation. However, existing world models are confined to specific scenarios such as gaming or driving, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Xiaofeng Wang , Zheng Zhu , Guan Huang , Boyuan Wang , Xinze Chen , Jiwen Lu

We present TeSMo, a method for text-controlled scene-aware motion generation based on denoising diffusion models. Previous text-to-motion methods focus on characters in isolation without considering scenes due to the limited availability of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Hongwei Yi , Justus Thies , Michael J. Black , Xue Bin Peng , Davis Rempe

Despite encouraging progress in 3D scene understanding, it remains challenging to develop an effective Large Multi-modal Model (LMM) that is capable of understanding and reasoning in complex 3D environments. Most previous methods typically…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Hanxun Yu , Wentong Li , Song Wang , Junbo Chen , Jianke Zhu

Visual localization remains challenging in dynamic environments where fluctuating lighting, adverse weather, and moving objects disrupt appearance cues. Despite advances in feature representation, current absolute pose regression methods…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Zhongtao Tian , Wenhao Huang , Zhidong Chen , Xiao Wei Sun

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

In monocular videos that capture dynamic scenes, estimating the 3D geometry of video contents has been a fundamental challenge in computer vision. Specifically, the task is significantly challenged by the object motion, where existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Seong Hyeon Park , Jinwoo Shin

We present a novel deep learning architecture for probabilistic future prediction from video. We predict the future semantics, geometry and motion of complex real-world urban scenes and use this representation to control an autonomous…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Anthony Hu , Fergal Cotter , Nikhil Mohan , Corina Gurau , Alex Kendall

The challenge of visual grounding and masking in multimodal machine translation (MMT) systems has encouraged varying approaches to the detection and selection of visually-grounded text tokens for masking. We introduce new methods for…

Computation and Language · Computer Science 2024-03-06 Braeden Bowen , Vipin Vijayan , Scott Grigsby , Timothy Anderson , Jeremy Gwinnup

One significant factor we expect the video representation learning to capture, especially in contrast with the image representation learning, is the object motion. However, we found that in the current mainstream video datasets, some action…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Jinpeng Wang , Yuting Gao , Ke Li , Jianguo Hu , Xinyang Jiang , Xiaowei Guo , Rongrong Ji , Xing Sun

Temporal prediction is critical for making intelligent and robust decisions in complex dynamic environments. Motion prediction needs to model the inherently uncertain future which often contains multiple potential outcomes, due to…

Machine Learning · Computer Science 2019-12-10 Yichuan Charlie Tang , Ruslan Salakhutdinov

Stochastic video prediction enables the consideration of uncertainty in future motion, thereby providing a better reflection of the dynamic nature of the environment. Stochastic video prediction methods based on image auto-regressive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Fei Cui , Jiaojiao Fang , Xiaojiang Wu , Zelong Lai , Mengke Yang , Menghan Jia , Guizhong Liu

Visual place recognition is one of the essential and challenging problems in the fields of robotics. In this letter, we for the first time explore the use of multi-modal fusion of semantic and visual modalities in dynamics-invariant space…

Computer Vision and Pattern Recognition · Computer Science 2022-01-04 Lin Wu , Teng Wang , Changyin Sun

Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Pulkit Kumar , Shuaiyi Huang , Matthew Walmer , Sai Saketh Rambhatla , Abhinav Shrivastava

This paper proposes a neural rendering approach that represents a scene as "compressed light-field tokens (CLiFTs)", retaining rich appearance and geometric information of a scene. CLiFT enables compute-efficient rendering by compressed…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhengqing Wang , Yuefan Wu , Jiacheng Chen , Fuyang Zhang , Yasutaka Furukawa

Compactly representing the visual signals is of fundamental importance in various image/video-centered applications. Although numerous approaches were developed for improving the image and video coding performance by removing the…

Image and Video Processing · Electrical Eng. & Systems 2020-08-14 Rongqun Lin , Linwei Zhu , Shiqi Wang , Sam Kwong
‹ Prev 1 4 5 6 7 8 10 Next ›