English
Related papers

Related papers: MoST: Multi-modality Scene Tokenization for Motion…

200 papers

Transformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual self-supervised learning. Meanwhile, previous methods only…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Zhaowen Li , Zhiyang Chen , Fan Yang , Wei Li , Yousong Zhu , Chaoyang Zhao , Rui Deng , Liwei Wu , Rui Zhao , Ming Tang , Jinqiao Wang

In this work we tackle the task of video-based visual emotion recognition in the wild. Standard methodologies that rely solely on the extraction of bodily and facial features often fall short of accurate emotion prediction in cases where…

Computer Vision and Pattern Recognition · Computer Science 2022-02-03 Ioannis Pikoulis , Panagiotis P. Filntisis , Petros Maragos

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Mengjun Cheng , Yipeng Sun , Longchao Wang , Xiongwei Zhu , Kun Yao , Jie Chen , Guoli Song , Junyu Han , Jingtuo Liu , Errui Ding , Jingdong Wang

Multi-modality spatio-temporal (MoST) data extends spatio-temporal (ST) data by incorporating multiple modalities, which is prevalent in monitoring systems, encompassing diverse traffic demands and air quality assessments. Despite…

Machine Learning · Computer Science 2024-05-07 Jiewen Deng , Renhe Jiang , Jiaqi Zhang , Xuan Song

Motion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Zhiqian Lan , Yuxuan Jiang , Yao Mu , Chen Chen , Shengbo Eben Li

Current multimodal recommendation models have extensively explored the effective utilization of multimodal information; however, their reliance on ID embeddings remains a performance bottleneck. Even with the assistance of multimodal…

Information Retrieval · Computer Science 2024-10-28 Kangning Zhang , Jiarui Jin , Yingjie Qin , Ruilong Su , Jianghao Lin , Yong Yu , Weinan Zhang

Deriving compact and temporally aware visual representations from dynamic scenes is essential for successful execution of sequential scene understanding tasks such as visual tracking and robotic manipulation. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Taekyung Kim , Dongyoon Han , Byeongho Heo , Jeongeun Park , Sangdoo Yun

Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Ziyan Guo , Haoxuan Qu , Hossein Rahmani , Dewen Soh , Ping Hu , Qiuhong Ke , Jun Liu

Trajectory prediction in urban mixed-traffic zones (a.k.a. shared spaces) is critical for many intelligent transportation systems, such as intent detection for autonomous driving. However, there are many challenges to predict the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-24 Hao Cheng , Wentong Liao , Michael Ying Yang , Monika Sester , Bodo Rosenhahn

Predicting the motion of other road agents enables autonomous vehicles to perform safe and efficient path planning. This task is very complex, as the behaviour of road agents depends on many factors and the number of possible future…

In order to plan a safe maneuver an autonomous vehicle must accurately perceive its environment, and understand the interactions among traffic participants. In this paper, we aim to learn scene-consistent motion forecasts of complex urban…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Sergio Casas , Cole Gulino , Simon Suo , Katie Luo , Renjie Liao , Raquel Urtasun

Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection. Despite great progress, these methods show deficiencies in…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Xugong Qin , Pengyuan Lyu , Chengquan Zhang , Yu Zhou , Kun Yao , Peng Zhang , Hailun Lin , Weiping Wang

We present a unified, promptable model capable of simultaneously segmenting, recognizing, and captioning anything. Unlike SAM, we aim to build a versatile region representation in the wild via visual prompting. To achieve this, we train a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Ting Pan , Lulu Tang , Xinlong Wang , Shiguang Shan

This paper proposes a novel generative video compression framework that leverages motion pattern priors, derived from subtle dynamics in common scenes (e.g., swaying flowers or a boat drifting on water), rather than relying on video content…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Shanzhi Yin , Zihan Zhang , Bolin Chen , Shiqi Wang , Yan Ye

Explainability and transparent decision-making are essential for the safe deployment of autonomous driving systems. Scene captioning summarizes environmental conditions and risk factors in natural language, improving transparency, safety,…

Robotics · Computer Science 2026-03-03 Zihang Wang , Xu Li , Benwu Wang , Wenkai Zhu , Xieyuanli Chen , Dong Kong , Kailin Lyu , Yinan Du , Yiming Peng , Haoyang Che

Motion forecasting for autonomous driving is a challenging task because complex driving scenarios result in a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Nigamaa Nayakanti , Rami Al-Rfou , Aurick Zhou , Kratarth Goel , Khaled S. Refaat , Benjamin Sapp

A long-standing goal in scene understanding is to obtain interpretable and editable representations that can be directly constructed from a raw monocular RGB-D video, without requiring specialized hardware setup or priors. The problem is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Yu-Shiang Wong , Niloy J. Mitra

Motion estimation is a crucial component in multi-object tracking (MOT). It predicts the trajectory of objects by analyzing the changes in their positions in consecutive frames of images, reducing tracking failures and identity switches.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jian Song , Wei Mei , Yunfeng Xu , Qiang Fu , Renke Kou , Lina Bu , Yucheng Long

Vision transformers have established a precedent of patchifying images into uniformly-sized chunks before processing. We hypothesize that this design choice may limit models in learning comprehensive and compositional representations from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Neha Kalibhat , Priyatham Kattakinda , Sumit Nawathe , Arman Zarei , Nikita Seleznev , Samuel Sharpe , Senthil Kumar , Soheil Feizi

How to learn discriminative video representation from unlabeled videos is challenging but crucial for video analysis. The latest attempts seek to learn a representation model by predicting the appearance contents in the masked regions.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Xinyu Sun , Peihao Chen , Liangwei Chen , Changhao Li , Thomas H. Li , Mingkui Tan , Chuang Gan