中文
相关论文

相关论文: First Place Solution to the ECCV 2024 BRAVO Challe…

200 篇论文

Video instance segmentation is a challenging task that serves as the cornerstone of numerous downstream applications, including video editing and autonomous driving. In this report, we present further improvements to the SOTA VIS method,…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Tao Zhang , Xingye Tian , Yikang Zhou , Yu Wu , Shunping Ji , Cilin Yan , Xuebo Wang , Xin Tao , Yuan Zhang , Pengfei Wan

Recent vision foundation models (VFMs) have demonstrated proficiency in various tasks but require supervised fine-tuning to perform the task of semantic segmentation effectively. Benchmarking their performance is essential for selecting…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Tommie Kerssies , Daan de Geus , Gijs Dubbelman

The third Pixel-level Video Understanding in the Wild (PVUW CVPR 2024) challenge aims to advance the state of art in video understanding through benchmarking Video Panoptic Segmentation (VPS) and Video Semantic Segmentation (VSS) on…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Qingfeng Liu , Mostafa El-Khamy , Kee-Bong Song

In this report, we descibe our approach to the ECCV 2020 VIPriors Object Detection Challenge which took place from March to July in 2020. We show that by using state-of-the-art data augmentation strategies, model designs, and…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Yinzheng Gu , Yihan Pan , Shizhe Chen

Video panoptic segmentation is a challenging task that serves as the cornerstone of numerous downstream applications, including video editing and autonomous driving. We believe that the decoupling strategy proposed by DVIS enables more…

计算机视觉与模式识别 · 计算机科学 2023-06-09 Tao Zhang , Xingye Tian , Haoran Wei , Yu Wu , Shunping Ji , Xuebo Wang , Xin Tao , Yuan Zhang , Pengfei Wan

In this report, we present our solution to the multi-task robustness track of the 1st Visual Continual Learning (VCL) Challenge at ICCV 2023 Workshop. We propose a vanilla framework named UniNet that seamlessly combines various visual…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Zehui Chen , Qiuchen Wang , Zhenyu Li , Jiaming Liu , Shanghang Zhang , Feng Zhao

This technical report presents the implementation details of 2nd winning for CVPR'24 UG2 WeatherProof Dataset Challenge. This challenge aims at semantic segmentation of images degraded by various degrees of weather from all around the…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Guojin Cao , Jiaxu Li , Jia He , Ying Min , Yunhao Zhang

The recent transformer-based models have dominated the Referring Video Object Segmentation (RVOS) task due to the superior performance. Most prior works adopt unified DETR framework to generate segmentation masks in query-to-instance…

计算机视觉与模式识别 · 计算机科学 2024-01-02 Zhuoyan Luo , Yicheng Xiao , Yong Liu , Yitong Wang , Yansong Tang , Xiu Li , Yujiu Yang

Semantic segmentation is an important task in computer vision, from which some important usage scenarios are derived, such as autonomous driving, scene parsing, etc. Due to the emphasis on the task of video semantic segmentation, we…

计算机视觉与模式识别 · 计算机科学 2021-09-06 Zixuan Chen , Junhong Zou , Xiaotao Wang

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Yunze Man , Shuhong Zheng , Zhipeng Bao , Martial Hebert , Liang-Yan Gui , Yu-Xiong Wang

When designing a semantic segmentation module for a practical application, such as autonomous driving, it is crucial to understand the robustness of the module with respect to a wide range of image corruptions. While there are recent…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Christoph Kamann , Carsten Rother

This technical report outlines the methodologies we applied for the PRCV Challenge, focusing on cognition and decision-making in driving scenarios. We employed InternVL-2.0, a pioneering open-source multi-modal model, and enhanced it by…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Bin Huang , Siyu Wang , Yuanpeng Chen , Yidan Wu , Hui Song , Zifan Ding , Jing Leng , Chengpeng Liang , Peng Xue , Junliang Zhang , Tiankun Zhao

The Agriculture-Vision Challenge at CVPR 2024 aims at leveraging semantic segmentation models to produce pixel level semantic segmentation labels within regions of interest for multi-modality satellite images. It is one of the most famous…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Wang Liu , Zhiyu Wang , Puhong Duan , Xudong Kang , Shutao Li

EarthVision Embed2Scale challenge (CVPR 2025) aims to develop foundational geospatial models to embed SSL4EO-S12 hyperspectral geospatial data cubes into embedding vectors that faciliatetes various downstream tasks, e.g., classification,…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Zirui Xu , Raphael Tang , Mike Bianco , Qi Zhang , Rishi Madhok , Nikolaos Karianakis , Fuxun Yu

In this report, we introduce our (pretty straightforard) two-step "detect-then-match" video instance segmentation method. The first step performs instance segmentation for each frame to get a large number of instance mask proposals. The…

计算机视觉与模式识别 · 计算机科学 2021-11-03 Yuming Du , Wen Guo , Yang Xiao , Vincent Lepetit

Video Panoptic Segmentation (VPS) is a challenging task that is extends from image panoptic segmentation.VPS aims to simultaneously classify, track, segment all objects in a video, including both things and stuff. Due to its wide…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Biao Wu , Diankai Zhang , Si Gao , Chengjian Zheng , Shaoli Liu , Ning Wang

Developing Foundation Models for medical image analysis is essential to overcome the unique challenges of radiological tasks. The first challenges of this kind for 3D brain MRI, SSL3D and FOMO25, were held at MICCAI 2025. Our solution…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Pedro M. Gordaliza , Jaume Banus , Benoît Gérin , Maxence Wynen , Nataliia Molchanova , Jonas Richiardi , Meritxell Bach Cuadra

Masked image modeling (MIM) has become a prevalent pre-training setup for vision foundation models and attains promising performance. Despite its success, existing MIM methods discard the decoder network during downstream applications,…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Qi Han , Yuxuan Cai , Xiangyu Zhang

In order to deal with the task of video panoptic segmentation in the wild, we propose a robust integrated video panoptic segmentation solution. In our solution, we regard the video panoptic segmentation task as a segmentation target…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Jinming Su , Wangwang Yang , Junfeng Luo , Xiaolin Wei

This competition focus on Urban-Sense Segmentation based on the vehicle camera view. Class highly unbalanced Urban-Sense images dataset challenge the existing solutions and further studies. Deep Conventional neural network-based semantic…

计算机视觉与模式识别 · 计算机科学 2022-07-05 Akide Liu , Zihan Wang