English
Related papers

Related papers: MoonAnything: A Vision Benchmark with Large-Scale …

200 papers

Panoptic scene understanding and tracking of dynamic agents are essential for robots and automated vehicles to navigate in urban environments. As LiDARs provide accurate illumination-independent geometric depictions of the scene, performing…

Computer Vision and Pattern Recognition · Computer Science 2021-12-28 Whye Kit Fong , Rohit Mohan , Juana Valeria Hurtado , Lubing Zhou , Holger Caesar , Oscar Beijbom , Abhinav Valada

This work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step towards panoptic captioning by formulating it as a task of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Kun-Yu Lin , Hongjun Wang , Weining Ren , Kai Han

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Xiang Li , Jian Ding , Mohamed Elhoseiny

We address the problem of estimating realistic, spatially varying reflectance for complex planetary surfaces such as the lunar regolith, which is critical for high-fidelity rendering and vision-based navigation. Existing lunar rendering…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Clementine Grethen , Nicolas Menga , Roland Brochard , Geraldine Morin , Simone Gasparini , Jeremy Lebreton , Manuel Sanchez Gestido

Multimodal learning is an emerging research topic across multiple disciplines but has rarely been applied to planetary science. In this contribution, we propose a single, unified transformer architecture trained to learn shared…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Tom Sander , Moritz Tenthoff , Kay Wohlfarth , Christian Wöhler

Texture mapping as a fundamental task in 3D modeling has been well established for well-acquired aerial assets under consistent illumination, yet it remains a challenge when it is scaled to large datasets with images under varying views and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Xiao ling , Rongjun Qin

There has been a recent surge in methods that aim to decompose and segment scenes into multiple objects in an unsupervised manner, i.e., unsupervised multi-object segmentation. Performing such a task is a long-standing goal of computer…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Laurynas Karazija , Iro Laina , Christian Rupprecht

The use of robotics in humanitarian demining increasingly involves computer vision techniques to improve landmine detection capabilities. However, in the absence of diverse and realistic datasets, the reliable validation of algorithms…

This article introduces a benchmark designed to evaluate the capabilities of multimodal models in analyzing and interpreting images. The benchmark focuses on seven key visual aspects: main object, additional objects, background, detail,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Evgenii Evstafev

We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (using vision language models). Recently, MVGs have been the…

Graphics · Computer Science 2025-07-02 Xianghui Xie , Chuhang Zou , Meher Gitika Karumuri , Jan Eric Lenssen , Gerard Pons-Moll

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying…

We present egenioussBench, a visual localisation benchmark built on geospatial reference data: a city-scale airborne 3D mesh and a CityGML LoD2 model. This pairing reflects deployable mapping assets and supports true scalability beyond…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Phillipp Fanta-Jende , Francesco Vultaggio , Alexander Kern , Yasmin Loeper , Markus Gerke

Image matching approaches have been widely used in computer vision applications in which the image-level matching performance of matchers is critical. However, it has not been well investigated by previous works which place more emphases on…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 JiaWang Bian , Le Zhang , Yun Liu , Wen-Yan Lin , Ming-Ming Cheng , Ian D. Reid

Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result, their adoption has increased in many computer vision…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Elena Izzo , Luca Parolari , Davide Vezzaro , Lamberto Ballan

The global 21 cm signal from the hyperfine transition of cosmic atomic hydrogen is theorised to track the state of the early Universe via the analysis of its absorption and emission with respect to the radio background. Detecting this…

Instrumentation and Methods for Astrophysics · Physics 2026-02-25 Joe H. N. Pattison , Dominic J. Anstey , Eloy de Lera Acedo

Geometrically accurate and semantically expressive map representations have proven invaluable for robot deployment and task planning in unknown environments. Nevertheless, real-time, open-vocabulary semantic understanding of large-scale…

Low-light image enhancement is crucial for a myriad of applications, from night vision and surveillance, to autonomous driving. However, due to the inherent limitations that come in hand with capturing images in low-illumination…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Manjushree Aithal , Rosaura G. VidalMata , Manikandtan Kartha , Gong Chen , Eashan Adhikarla , Lucas N. Kirsten , Zhicheng Fu , Nikhil A. Madhusudhana , Joe Nasti

With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in real-world…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yipo Huang , Quan Yuan , Xiangfei Sheng , Zhichao Yang , Haoning Wu , Pengfei Chen , Yuzhe Yang , Leida Li , Weisi Lin

Due to the nature of enhancement--the absence of paired ground-truth information, high-level vision tasks have been recently employed to evaluate the performance of low-light image enhancement. A widely-used manner is to see how accurately…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Mingjia Li , Hao Zhao , Xiaojie Guo

We developed a Python based framework for astronomical image processing and analysis. Astronomical image loading, normalizing, stacking, and filtering processes represent visible range images from grayscale. Besides, the blending process…

Instrumentation and Methods for Astrophysics · Physics 2024-10-10 Tanmoy Bhowmik , MD Fardin Islam , Kazi Nusrat Tasneem , Rantideb Roy , Rownok Shahariar
‹ Prev 1 3 4 5 6 7 10 Next ›