English
Related papers

Related papers: Playing for Benchmarks

200 papers

Semantic segmentation is an essential step for many vision applications in order to understand a scene and the objects within. Recent progress in hyperspectral imaging technology enables the application in driving scenarios and the hope is…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Nick Theisen , Robin Bartsch , Dietrich Paulus , Peer Neubert

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human perceptions; 2) the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Xinran Ling , Chen Zhu , Meiqi Wu , Hangyu Li , Xiaokun Feng , Cundian Yang , Aiming Hao , Jiashu Zhu , Jiahong Wu , Xiangxiang Chu

Although there has been significant progress in the past decade,tracking is still a very challenging computer vision task, due to problems such as occlusion and model drift.Recently, the increased popularity of depth sensors e.g. Microsoft…

Computer Vision and Pattern Recognition · Computer Science 2012-12-13 Shuran Song , Jianxiong Xiao

Investigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due to complex authentic distortions localized in space and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Wen Wen , Mu Li , Yiru Yao , Xiangjie Sui , Yabin Zhang , Long Lan , Yuming Fang , Kede Ma

We investigate research challenges and opportunities for visualization in motion during outdoor physical activities via an initial corpus of real-world recordings that pair egocentric video, biometrics, and think-aloud observations. With…

Human-Computer Interaction · Computer Science 2024-09-11 Ahmed Elshabasi , Lijie Yao , Petra Isenberg , Charles Perin , Wesley Willett

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Xiang Li , Jian Ding , Mohamed Elhoseiny

We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Xingcheng Zhou , Konstantinos Larintzakis , Hao Guo , Walter Zimmer , Mingyu Liu , Hu Cao , Jiajie Zhang , Venkatnarayanan Lakshminarasimhan , Leah Strand , Alois C. Knoll

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, we introduce a large-scale, dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Zheng-Hui Huang , Zhixiang Wang , Jiaming Tan , Ruihan Yu , Yidan Zhang , Bo Zheng , Yu-Lun Liu , Yung-Yu Chuang , Kaipeng Zhang

Efficient vision works maximize accuracy under a latency budget. These works evaluate accuracy offline, one image at a time. However, real-time vision applications like autonomous driving operate in streaming settings, where ground truth…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Gur-Eyal Sela , Ionel Gog , Justin Wong , Kumar Krishna Agrawal , Xiangxi Mo , Sukrit Kalra , Peter Schafhalter , Eric Leong , Xin Wang , Bharathan Balaji , Joseph Gonzalez , Ion Stoica

Autonomous space operations such as on-orbit servicing and active debris removal demand robust part-level semantic understanding and precise relative navigation of target spacecraft, yet collecting large-scale real data in orbit remains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Aodi Wu , Jianhong Zuo , Zeyuan Zhao , Xubo Luo , Ruisuo Wang , Xue Wan

This article introduces a benchmark designed to evaluate the capabilities of multimodal models in analyzing and interpreting images. The benchmark focuses on seven key visual aspects: main object, additional objects, background, detail,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Evgenii Evstafev

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yiqing Shen , Chenjia Li , Chenxiao Fan , Mathias Unberath

Existing visual reasoning benchmarks predominantly rely on natural language prompts, evaluate narrow reasoning modalities, or depend on subjective scoring procedures such as LLM-as-judge. We introduce the TACIT Benchmark, a programmatic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Daniel Nobrega Medeiros

Tactile perception has the potential to significantly enhance dexterous robotic manipulation by providing rich local information that can complement or substitute for other sensory modalities such as vision. However, because tactile sensing…

Robotics · Computer Science 2025-06-17 Tim Schneider , Guillaume Duret , Cristiana de Farias , Roberto Calandra , Liming Chen , Jan Peters

We introduce the ObjectFolder Benchmark, a benchmark suite of 10 tasks for multisensory object-centric learning, centered around object recognition, reconstruction, and manipulation with sight, sound, and touch. We also introduce the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Ruohan Gao , Yiming Dou , Hao Li , Tanmay Agarwal , Jeannette Bohg , Yunzhu Li , Li Fei-Fei , Jiajun Wu

Generating long-form storytelling videos with consistent visual narratives remains a significant challenge in video synthesis. We present a novel framework, dataset, and a model that address three critical limitations: background…

Online continual learning from data streams in dynamic environments is a critical direction in the computer vision field. However, realistic benchmarks and fundamental studies in this line are still missing. To bridge the gap, we present a…

Computer Vision and Pattern Recognition · Computer Science 2021-09-09 Jianren Wang , Xin Wang , Yue Shang-Guan , Abhinav Gupta

High-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable efforts invested in developing various capture systems and reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Xiaoyun Zheng , Liwei Liao , Xufeng Li , Jianbo Jiao , Rongjie Wang , Feng Gao , Shiqi Wang , Ronggang Wang

Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happening within the scene for safe interaction with other road…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Salman Khan , Izzeddin Teeti , Reza Javanmard Alitappeh , Mihaela C. Stoian , Eleonora Giunchiglia , Gurkirt Singh , Andrew Bradley , Fabio Cuzzolin