English
Related papers

Related papers: Perception Test 2025: Challenge Summary and a Unif…

200 papers

Depth-aware video panoptic segmentation is a promising approach to camera based scene understanding. However, the current state-of-the-art methods require costly video annotations and use a complex training pipeline compared to their…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Kurt Stolle , Gijs Dubbelman

Small Multi-Object Tracking (SMOT) is particularly challenging when targets occupy only a few dozen pixels, rendering detection and appearance-based association unreliable. Building on the success of the MVA2023 SOD4SB challenge, this paper…

This paper introduces MCTrack, a new 3D multi-object tracking method that achieves state-of-the-art (SOTA) performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xiyang Wang , Shouzheng Qi , Jieyou Zhao , Hangning Zhou , Siyu Zhang , Guoan Wang , Kai Tu , Songlin Guo , Jianbo Zhao , Jian Li , Mu Yang

Tables convey factual and quantitative data with implicit conventions created by humans that are often challenging for machines to parse. Prior work on table recognition (TR) has mainly centered around complex task-specific combinations of…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 ShengYun Peng , Aishwarya Chakravarthy , Seongmin Lee , Xiaojing Wang , Rajarajeswari Balasubramaniyan , Duen Horng Chau

Bird's eye view (BEV)-based 3D perception plays a crucial role in autonomous driving applications. The rise of large language models has spurred interest in BEV-based captioning to understand object behavior in the surrounding environment.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Yunsheng Ma , Burhaneddin Yaman , Xin Ye , Jingru Luo , Feng Tao , Abhirup Mallik , Ziran Wang , Liu Ren

TextVQA requires models to read and reason about text in images to answer questions about them. Specifically, models need to incorporate a new modality of text present in the images and reason over it to answer TextVQA questions. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Yixuan Qiao , Hao Chen , Jun Wang , Shanshan Zhao , Yihao Chen , Xianbin Ye , Ziliang Li , Xianbiao Qi , Peng Gao , Guotong Xie

In recent years, 3D object perception has become a crucial component in the development of autonomous driving systems, providing essential environmental awareness. However, as perception tasks in autonomous driving evolve, their variants…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Yu Wang , Shaohua Wang , Yicheng Li , Mingchun Liu

The major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that attention mechanism is beneficial to the fusion process. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Bin Duan , Hao Tang , Wei Wang , Ziliang Zong , Guowei Yang , Yan Yan

Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under sensor noise, environmental variation, and platform shifts.…

Robotics · Computer Science 2026-01-09 Lingdong Kong , Shaoyuan Xie , Zeying Gong , Ye Li , Meng Chu , Ao Liang , Yuhao Dong , Tianshuai Hu , Ronghe Qiu , Rong Li , Hanjiang Hu , Dongyue Lu , Wei Yin , Wenhao Ding , Linfeng Li , Hang Song , Wenwei Zhang , Yuexin Ma , Junwei Liang , Zhedong Zheng , Lai Xing Ng , Benoit R. Cottereau , Wei Tsang Ooi , Ziwei Liu , Zhanpeng Zhang , Weichao Qiu , Wei Zhang , Ji Ao , Jiangpeng Zheng , Siyu Wang , Guang Yang , Zihao Zhang , Yu Zhong , Enzhu Gao , Xinhan Zheng , Xueting Wang , Shouming Li , Yunkai Gao , Siming Lan , Mingfei Han , Xing Hu , Dusan Malic , Christian Fruhwirth-Reisinger , Alexander Prutsch , Wei Lin , Samuel Schulter , Horst Possegger , Linfeng Li , Jian Zhao , Zepeng Yang , Yuhang Song , Bojun Lin , Tianle Zhang , Yuchen Yuan , Chi Zhang , Xuelong Li , Youngseok Kim , Sihwan Hwang , Hyeonjun Jeong , Aodi Wu , Xubo Luo , Erjia Xiao , Lingfeng Zhang , Yingbo Tang , Hao Cheng , Renjing Xu , Wenbo Ding , Lei Zhou , Long Chen , Hangjun Ye , Xiaoshuai Hao , Shuangzhi Li , Junlong Shen , Xingyu Li , Hao Ruan , Jinliang Lin , Zhiming Luo , Yu Zang , Cheng Wang , Hanshi Wang , Xijie Gong , Yixiang Yang , Qianli Ma , Zhipeng Zhang , Wenxiang Shi , Jingmeng Zhou , Weijun Zeng , Kexin Xu , Yuchen Zhang , Haoxiang Fu , Ruibin Hu , Yanbiao Ma , Xiyan Feng , Wenbo Zhang , Lu Zhang , Yunzhi Zhuge , Huchuan Lu , You He , Seungjun Yu , Junsung Park , Youngsun Lim , Hyunjung Shim , Faduo Liang , Zihang Wang , Yiming Peng , Guanyu Zong , Xu Li , Binghao Wang , Hao Wei , Yongxin Ma , Yunke Shi , Shuaipeng Liu , Dong Kong , Yongchun Lin , Huitong Yang , Liang Lei , Haoang Li , Xinliang Zhang , Zhiyong Wang , Xiaofeng Wang , Yuxia Fu , Yadan Luo , Djamahl Etchegaray , Yang Li , Congfei Li , Yuxiang Sun , Wenkai Zhu , Wang Xu , Linru Li , Longjie Liao , Jun Yan , Benwu Wang , Xueliang Ren , Xiaoyu Yue , Jixian Zheng , Jinfeng Wu , Shurui Qin , Wei Cong , Yao He

Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess the cognitive capabilities required for high-level tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Chenglin Li , Qianglong Chen , Zhi Li , Feng Tao , Yin Zhang

Current popular online multi-object tracking (MOT) solutions apply single object trackers (SOTs) to capture object motions, while often requiring an extra affinity network to associate objects, especially for the occluded ones. This brings…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Junbo Yin , Wenguan Wang , Qinghao Meng , Ruigang Yang , Jianbing Shen

In this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Xin Chen , Ben Kang , Wanting Geng , Jiawen Zhu , Yi Liu , Dong Wang , Huchuan Lu

Although recent efforts in image quality assessment (IQA) have achieved promising performance, there still exists a considerable gap compared to the human visual system (HVS). One significant disparity lies in humans' seamless transition…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Yi Ke Yun , Weisi Lin

Open-world video recognition is challenging since traditional networks are not generalized well on complex environment variations. Alternatively, foundation models with rich knowledge have recently shown their generalization power. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Boyu Chen , Siran Chen , Kunchang Li , Qinglin Xu , Yu Qiao , Yali Wang

Despite the promising performance of current video segmentation models on existing benchmarks, these models still struggle with complex scenes. In this paper, we introduce the 6th Large-scale Video Object Segmentation (LSVOS) challenge in…

Human-centric perception (e.g. detection, segmentation, pose estimation, and attribute analysis) is a long-standing problem for computer vision. This paper introduces a unified and versatile framework (HQNet) for single-stage multi-person…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Sheng Jin , Shuhuai Li , Tong Li , Wentao Liu , Chen Qian , Ping Luo

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vision-language models, such as CLIP and large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Sara Ghazanfari , Siddharth Garg , Nicolas Flammarion , Prashanth Krishnamurthy , Farshad Khorrami , Francesco Croce

Traffic monitoring is crucial for urban mobility, road safety, and intelligent transportation systems (ITS). Deep learning has advanced video-based traffic monitoring through video question answering (VideoQA) models, enabling structured…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Joseph Raj Vishal , Divesh Basina , Rutuja Patil , Manas Srinivas Gowda , Katha Naik , Yezhou Yang , Bharatesh Chakravarthi

While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hindered by a fundamental "temporal blindness". Existing architectures lack intrinsic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Xiaohe Li , Jiahao Li , Kaixin Zhang , Yuqiang Fang , Leilei Lin , Hong Wang , Haohua Wu , Zide Fan

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang