English
Related papers

Related papers: Perception Test 2025: Challenge Summary and a Unif…

200 papers

In the realm of autonomous driving, robust perception under out-of-distribution conditions is paramount for the safe deployment of vehicles. Challenges such as adverse weather, sensor malfunctions, and environmental unpredictability can…

Single object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, RGB+Depth, RGB+Thermal or RGB+Event.). Different reference…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Yinchao Ma , Yuyang Tang , Wenfei Yang , Tianzhu Zhang , Xu Zhou , Feng Wu

This report proposes an improved method for the Tracking Any Point (TAP) task, which tracks any physical surface through a video. Several existing approaches have explored the TAP by considering the temporal relationships to obtain smooth…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Hongpeng Pan , Yang Yang , Zhongtian Fu , Yuxuan Zhang , Shian Du , Yi Xu , Xiangyang Ji

This work introduces a dataset, benchmark, and challenge for the problem of video copy detection and localization. The problem comprises two distinct but related tasks: determining whether a query video shares content with a reference video…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Ed Pizzi , Giorgos Kordopatis-Zilos , Hiral Patel , Gheorghe Postelnicu , Sugosh Nagavara Ravindra , Akshay Gupta , Symeon Papadopoulos , Giorgos Tolias , Matthijs Douze

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detects the temporal…

Visual object tracking is an important task in computer vision, which has many real-world applications, e.g., video surveillance, visual navigation. Visual object tracking also has many challenges, e.g., object occlusion and deformation. To…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Ruize Han , Wei Feng , Qing Guo , Qinghua Hu

Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide the most objective measure of performance and are therefore…

Computer Vision and Pattern Recognition · Computer Science 2020-03-23 Patrick Dendorfer , Hamid Rezatofighi , Anton Milan , Javen Shi , Daniel Cremers , Ian Reid , Stefan Roth , Konrad Schindler , Laura Leal-Taixé

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

Spatio-temporal feature learning is of central importance for action recognition in videos. Existing deep neural network models either learn spatial and temporal features independently (C2D) or jointly with unconstrained parameters (C3D).…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Chao Li , Qiaoyong Zhong , Di Xie , Shiliang Pu

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-08 Davide Berghi , Philip J. B. Jackson

Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce…

Computation and Language · Computer Science 2022-03-29 Yingshan Chang , Mridu Narang , Hisami Suzuki , Guihong Cao , Jianfeng Gao , Yonatan Bisk

The SoccerNet 2023 challenges were the third annual video understanding challenges organized by the SoccerNet team. For this third edition, the challenges were composed of seven vision-based tasks split into three main themes. The first…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Anthony Cioppa , Silvio Giancola , Vladimir Somers , Floriane Magera , Xin Zhou , Hassan Mkhallati , Adrien Deliège , Jan Held , Carlos Hinojosa , Amir M. Mansourian , Pierre Miralles , Olivier Barnich , Christophe De Vleeschouwer , Alexandre Alahi , Bernard Ghanem , Marc Van Droogenbroeck , Abdullah Kamal , Adrien Maglo , Albert Clapés , Amr Abdelaziz , Artur Xarles , Astrid Orcesi , Atom Scott , Bin Liu , Byoungkwon Lim , Chen Chen , Fabian Deuser , Feng Yan , Fufu Yu , Gal Shitrit , Guanshuo Wang , Gyusik Choi , Hankyul Kim , Hao Guo , Hasby Fahrudin , Hidenari Koguchi , Håkan Ardö , Ibrahim Salah , Ido Yerushalmy , Iftikar Muhammad , Ikuma Uchida , Ishay Be'ery , Jaonary Rabarisoa , Jeongae Lee , Jiajun Fu , Jianqin Yin , Jinghang Xu , Jongho Nang , Julien Denize , Junjie Li , Junpei Zhang , Juntae Kim , Kamil Synowiec , Kenji Kobayashi , Kexin Zhang , Konrad Habel , Kota Nakajima , Licheng Jiao , Lin Ma , Lizhi Wang , Luping Wang , Menglong Li , Mengying Zhou , Mohamed Nasr , Mohamed Abdelwahed , Mykola Liashuha , Nikolay Falaleev , Norbert Oswald , Qiong Jia , Quoc-Cuong Pham , Ran Song , Romain Hérault , Rui Peng , Ruilong Chen , Ruixuan Liu , Ruslan Baikulov , Ryuto Fukushima , Sergio Escalera , Seungcheon Lee , Shimin Chen , Shouhong Ding , Taiga Someya , Thomas B. Moeslund , Tianjiao Li , Wei Shen , Wei Zhang , Wei Li , Wei Dai , Weixin Luo , Wending Zhao , Wenjie Zhang , Xinquan Yang , Yanbiao Ma , Yeeun Joo , Yingsen Zeng , Yiyang Gan , Yongqiang Zhu , Yujie Zhong , Zheng Ruan , Zhiheng Li , Zhijian Huang , Ziyu Meng

Continual Visual Question Answering (CVQA) based on pre-trained models(PTMs) has achieved promising progress by leveraging prompt tuning to enable continual multi-modal learning. However, most existing methods adopt cross-modal prompt…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Xu Li , Fan Lyu

The objective of this paper is self-supervised representation learning, with the goal of solving semi-supervised video object segmentation (a.k.a. dense tracking). We make the following contributions: (i) we propose to improve the existing…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Fangrui Zhu , Li Zhang , Yanwei Fu , Guodong Guo , Weidi Xie

Visual Question Answering (VQA) presents a unique challenge as it requires the ability to understand and encode the multi-modal inputs - in terms of image processing and natural language processing. The algorithm further needs to learn how…

Computer Vision and Pattern Recognition · Computer Science 2017-09-26 Supriya Pandhre , Shagun Sodhani

Super-Resolution (SR) is a critical task in computer vision, focusing on reconstructing high-resolution (HR) images from low-resolution (LR) inputs. The field has seen significant progress through various challenges, particularly in…

Image and Video Processing · Electrical Eng. & Systems 2025-07-02 Babak Naderi , Ross Cutler , Juhee Cho , Nabakumar Khongbantabam , Dejan Ivkovic

The rapid advancement of Large Multi-modal Foundation Models (LMM) has paved the way for the possible Explainable Image Quality Assessment (EIQA) with instruction tuning from two perspectives: overall quality explanation, and attribute-wise…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Yiting Lu , Xin Li , Haoning Wu , Bingchen Li , Weisi Lin , Zhibo Chen

Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite their shared objective, existing approaches often tackle these…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Tianlu Zhang , Qiang Zhang , Guiguang Ding , Jungong Han

Video segmentation aims at partitioning video sequences into meaningful segments based on objects or regions of interest within frames. Current video segmentation models are often derived from image segmentation techniques, which struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Chen Liang , Qiang Guo , Xiaochao Qu , Luoqi Liu , Ting Liu

Recent years have witnessed an exponential increase in the demand for face video compression, and the success of artificial intelligence has expanded the boundaries beyond traditional hybrid video coding. Generative coding approaches have…

Image and Video Processing · Electrical Eng. & Systems 2023-10-31 Yixuan Li , Bolin Chen , Baoliang Chen , Meng Wang , Shiqi Wang , Weisi Lin
‹ Prev 1 8 9 10 Next ›