中文
相关论文

相关论文: Perception Test 2023: A Summary of the First Chall…

200 篇论文

Following the successful 2023 edition, we organised the Second Perception Test challenge as a half-day workshop alongside the IEEE/CVF European Conference on Computer Vision (ECCV) 2024, with the goal of benchmarking state-of-the-art video…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Joseph Heyward , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Joseph Heyward , Nikhil Parthasarathy , Tyler Zhu , Aravindh Mahendran , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

The IEEE Low-Power Computer Vision Challenge (LPCVC) aims to promote the development of efficient vision models for edge devices, balancing accuracy with constraints such as latency, memory capacity, and energy use. The 2025 challenge…

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on…

This article describes the 2023 IEEE Low-Power Computer Vision Challenge (LPCVC). Since 2015, LPCVC has been an international competition devoted to tackling the challenge of computer vision (CV) on edge devices. Most CV researchers focus…

This survey serves as a review for the 2025 Event-Based Eye Tracking Challenge organized as part of the 2025 CVPR event-based vision workshop. This challenge focuses on the task of predicting the pupil center by processing event camera…

In this paper, we introduce a grounded video question-answering solution. Our research reveals that the fixed official baseline method for video question answering involves two main steps: visual grounding and object tracking. However, a…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Hailiang Zhang , Dian Chao , Zhihao Guan , Yang Yang

We present a new video understanding pentathlon challenge, an open competition held in conjunction with the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2020. The objective of the challenge was to explore and evaluate…

In this report, we present our first-place solution to the Multiple-choice Video Question Answering (QA) track of The Second Perception Test Challenge. This competition posed a complex video understanding task, requiring models to…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Yingzhe Peng , Yixiao Yuan , Zitian Ao , Huapeng Zhou , Kangqi Wang , Qipeng Zhu , Xu Yang

The VALUE (Video-And-Language Understanding Evaluation) benchmark is newly introduced to evaluate and analyze multi-modal representation learning algorithms on three video-and-language tasks: Retrieval, QA, and Captioning. The main…

计算机视觉与模式识别 · 计算机科学 2021-10-14 Minchul Shin , Jonghwan Mun , Kyoung-Woon On , Woo-Young Kang , Gunsoo Han , Eun-Sol Kim

In this report, we describe the technical details of our submission to the EPIC-SOUNDS Audio-Based Interaction Recognition Challenge 2023, by Team "AcieLee" (username: Yuqi\_Li). The task is to classify the audio caused by interactions…

声音 · 计算机科学 2023-06-16 Yuqi Li , Yizhi Luo , Xiaoshuai Hao , Chuanguang Yang , Zhulin An , Dantong Song , Wei Yi

This work introduces a dataset, benchmark, and challenge for the problem of video copy detection and localization. The problem comprises two distinct but related tasks: determining whether a query video shares content with a reference video…

The SoccerNet 2023 challenges were the third annual video understanding challenges organized by the SoccerNet team. For this third edition, the challenges were composed of seven vision-based tasks split into three main themes. The first…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Anthony Cioppa , Silvio Giancola , Vladimir Somers , Floriane Magera , Xin Zhou , Hassan Mkhallati , Adrien Deliège , Jan Held , Carlos Hinojosa , Amir M. Mansourian , Pierre Miralles , Olivier Barnich , Christophe De Vleeschouwer , Alexandre Alahi , Bernard Ghanem , Marc Van Droogenbroeck , Abdullah Kamal , Adrien Maglo , Albert Clapés , Amr Abdelaziz , Artur Xarles , Astrid Orcesi , Atom Scott , Bin Liu , Byoungkwon Lim , Chen Chen , Fabian Deuser , Feng Yan , Fufu Yu , Gal Shitrit , Guanshuo Wang , Gyusik Choi , Hankyul Kim , Hao Guo , Hasby Fahrudin , Hidenari Koguchi , Håkan Ardö , Ibrahim Salah , Ido Yerushalmy , Iftikar Muhammad , Ikuma Uchida , Ishay Be'ery , Jaonary Rabarisoa , Jeongae Lee , Jiajun Fu , Jianqin Yin , Jinghang Xu , Jongho Nang , Julien Denize , Junjie Li , Junpei Zhang , Juntae Kim , Kamil Synowiec , Kenji Kobayashi , Kexin Zhang , Konrad Habel , Kota Nakajima , Licheng Jiao , Lin Ma , Lizhi Wang , Luping Wang , Menglong Li , Mengying Zhou , Mohamed Nasr , Mohamed Abdelwahed , Mykola Liashuha , Nikolay Falaleev , Norbert Oswald , Qiong Jia , Quoc-Cuong Pham , Ran Song , Romain Hérault , Rui Peng , Ruilong Chen , Ruixuan Liu , Ruslan Baikulov , Ryuto Fukushima , Sergio Escalera , Seungcheon Lee , Shimin Chen , Shouhong Ding , Taiga Someya , Thomas B. Moeslund , Tianjiao Li , Wei Shen , Wei Zhang , Wei Li , Wei Dai , Weixin Luo , Wending Zhao , Wenjie Zhang , Xinquan Yang , Yanbiao Ma , Yeeun Joo , Yingsen Zeng , Yiyang Gan , Yongqiang Zhu , Yujie Zhong , Zheng Ruan , Zhiheng Li , Zhijian Huang , Ziyu Meng

Video-text retrieval has many real-world applications such as media analytics, surveillance, and robotics. This paper presents the 1st place solution to the video retrieval track of the ICCV VALUE Challenge 2021. We present a simple yet…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Aiden Seungjoon Lee , Hanseok Oh , Minjoon Seo

This report proposes an improved method for the Tracking Any Point (TAP) task, which tracks any physical surface through a video. Several existing approaches have explored the TAP by considering the temporal relationships to obtain smooth…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Hongpeng Pan , Yang Yang , Zhongtian Fu , Yuxuan Zhang , Shian Du , Yi Xu , Xiangyang Ji

This report presents our team's technical solution for participating in Track 3 of the 2024 ECCV ROAD++ Challenge. The task of Track 3 is atomic activity recognition, which aims to identify 64 types of atomic activities in road scenes based…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Ruyang Li , Tengfei Zhang , Heng Zhang , Tiejun Liu , Yanwei Wang , Xuelei Li

Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglected the necessity of…

计算机视觉与模式识别 · 计算机科学 2023-06-26 Tom Tongjia Chen , Hongshan Yu , Zhengeng Yang , Ming Li , Zechuan Li , Jingwen Wang , Wei Miao , Wei Sun , Chen Chen

The budgeted model training challenge aims to train an efficient classification model under resource limitations. To tackle this task in ImageNet-100, we describe a simple yet effective resource-aware backbone search framework composed of…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Youngjun Kwak , Seonghun Jeong , Yunseung Lee , Changick Kim

This paper presents the learned techniques during the Video Analysis Module of the Master in Computer Vision from the Universitat Aut\`onoma de Barcelona, used to solve the third track of the AI-City Challenge. This challenge aims to track…

计算机视觉与模式识别 · 计算机科学 2021-05-12 Pol Albacar , Òscar Lorente , Eduard Mainou , Ian Riera

This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit pre-trained audio encoders that map raw waveforms to…

‹ 上一页 1 2 3 10 下一页 ›