中文
相关论文

相关论文: Playing for Benchmarks

200 篇论文

Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling highly dynamic scenes. Yet, their integration with natural…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Lingdong Kong , Dongyue Lu , Ao Liang , Rong Li , Yuhao Dong , Tianshuai Hu , Lai Xing Ng , Wei Tsang Ooi , Benoit R. Cottereau

Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Anna Deichler , Jim O'Regan , Fethiye Irmak Dogan , Lubos Marcinek , Anna Klezovich , Iolanda Leite , Jonas Beskow

The ability to simulate the world in a spatially consistent manner is a crucial requirement for effective world models. Such a model enables high-quality visual generation, and also ensures the reliability of world models for downstream…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Kewei Lian , Shaofei Cai , Yitao Liang , Anji Liu

We introduce the OxUvA dataset and benchmark for evaluating single-object tracking algorithms. Benchmarks have enabled great strides in the field of object tracking by defining standardized evaluations on large sets of diverse videos.…

计算机视觉与模式识别 · 计算机科学 2018-08-13 Jack Valmadre , Luca Bertinetto , João F. Henriques , Ran Tao , Andrea Vedaldi , Arnold Smeulders , Philip Torr , Efstratios Gavves

Video is a promising source of knowledge for embodied agents to learn models of the world's dynamics. Large deep networks have become increasingly effective at modeling complex video data in a self-supervised manner, as evaluated by metrics…

计算机视觉与模式识别 · 计算机科学 2023-04-27 Stephen Tian , Chelsea Finn , Jiajun Wu

Researchers have used machine learning approaches to identify motion sickness in VR experience. These approaches demand an accurately-labeled, real-world, and diverse dataset for high accuracy and generalizability. As a starting point to…

In the emerging field of video coding for machines, video datasets with pristine video quality and high-quality annotations are required for a comprehensive evaluation. However, existing video datasets with detailed annotations are severely…

图像与视频处理 · 电气工程与系统科学 2022-05-16 Kristian Fischer , Markus Hofbauer , Christopher Kuhn , Eckehard Steinbach , André Kaup

In order to reach human performance on complexvisual tasks, artificial systems need to incorporate a sig-nificant amount of understanding of the world in termsof macroscopic objects, movements, forces, etc. Inspiredby work on intuitive…

This paper presents a new dataset for Novel View Synthesis, generated from a high-quality, animated film with stunning realism and intricate detail. Our dataset captures a variety of dynamic scenes, complete with detailed textures,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Michal Nazarczuk , Thomas Tanay , Arthur Moreau , Zhensong Zhang , Eduardo Pérez-Pellitero

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Yixuan Ye , Xuanyu Lu , Yuxin Jiang , Yuchao Gu , Rui Zhao , Qiwei Liang , Jiachun Pan , Fengda Zhang , Weijia Wu , Alex Jinpeng Wang

In training deep neural networks for semantic segmentation, the main limiting factor is the low amount of ground truth annotation data that is available in currently existing datasets. The limited availability of such data is due to the…

计算机视觉与模式识别 · 计算机科学 2018-07-18 Matt Angus , Mohamed ElBalkini , Samin Khan , Ali Harakeh , Oles Andrienko , Cody Reading , Steven Waslander , Krzysztof Czarnecki

We introduce Princeton365, a large-scale diverse dataset of 365 videos with accurate camera pose. Our dataset bridges the gap between accuracy and data diversity in current SLAM benchmarks by introducing a novel ground truth collection…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Karhan Kayan , Stamatis Alexandropoulos , Rishabh Jain , Yiming Zuo , Erich Liang , Jia Deng

Given the great interest in creating keyframe summaries from video, it is surprising how little has been done to formalise their evaluation and comparison. User studies are often carried out to demonstrate that a proposed method generates a…

计算机视觉与模式识别 · 计算机科学 2017-12-20 Ludmila I. Kuncheva , Paria Yousefi , Iain A. D. Gunn

Recent advances in event-based vision suggest that these systems complement traditional cameras by providing continuous observation without frame rate limitations and a high dynamic range, making them well-suited for correspondence tasks…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Yijin Li , Yichen Shen , Zhaoyang Huang , Shuo Chen , Weikang Bian , Xiaoyu Shi , Fu-Yun Wang , Keqiang Sun , Hujun Bao , Zhaopeng Cui , Guofeng Zhang , Hongsheng Li

This paper presents a dataset, called Reeds, for research on robot perception algorithms. The dataset aims to provide demanding benchmark opportunities for algorithms, rather than providing an environment for testing application-specific…

计算机视觉与模式识别 · 计算机科学 2021-09-20 Ola Benderius , Christian Berger , Krister Blanch

As automated vehicles are getting closer to becoming a reality, it will become mandatory to be able to characterise the performance of their obstacle detection systems. This validation process requires large amounts of ground-truth data,…

机器人学 · 计算机科学 2018-07-17 Hatem Hajri , Emmanuel Doucet , Marc Revilloud , Lynda Halit , Benoît Lusetti , Mohamed-Cherif Rahal

Many different parametric models for video quality assessment have been proposed in the past few years. This paper presents a review of nine recent models which cover a wide range of methodologies and have been validated for estimating…

多媒体 · 计算机科学 2017-07-03 Tiantian He , Yankai Liu , Rong Xie , Xin Tang , Li Song

Visual localization, i.e., camera pose estimation in a known scene, is a core component of technologies such as autonomous driving and augmented reality. State-of-the-art localization approaches often rely on image retrieval techniques for…

计算机视觉与模式识别 · 计算机科学 2020-12-02 Noé Pion , Martin Humenberger , Gabriela Csurka , Yohann Cabon , Torsten Sattler

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

We propose to build realistic virtual worlds, called 360RVW, for large urban environments directly from 360{\deg} videos. We provide an interface for interactive exploration, where users can freely navigate via their own avatars. 360{\deg}…

多媒体 · 计算机科学 2025-10-14 Mizuki Takenawa , Naoki Sugimoto , Leslie Wöhler , Satoshi Ikehata , Kiyoharu Aizawa