English
Related papers

Related papers: MoonAnything: A Vision Benchmark with Large-Scale …

200 papers

We introduce the first comprehensive 3D dataset for the task of unsupervised anomaly detection and localization. It is inspired by real-world visual inspection scenarios in which a model has to detect various types of defects on…

Computer Vision and Pattern Recognition · Computer Science 2022-02-25 Paul Bergmann , Xin Jin , David Sattlegger , Carsten Steger

Multiple benchmarks have been developed to assess the alignment between deep neural networks (DNNs) and human vision. In almost all cases these benchmarks are observational in the sense they are composed of behavioural and brain responses…

Image classification with small datasets has been an active research area in the recent past. However, as research in this scope is still in its infancy, two key ingredients are missing for ensuring reliable and truthful progress: a…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 L. Brigato , B. Barz , L. Iocchi , J. Denzler

The design of a plenoptic camera requires the combination of two dissimilar optical systems, namely a main lens and an array of microlenses. And while the construction process of a conventional camera is mainly concerned with focusing the…

Image and Video Processing · Electrical Eng. & Systems 2022-04-12 Tim Michels , Reinhard Koch

Given a long list of anomaly detection algorithms developed in the last few decades, how do they perform with regard to (i) varying levels of supervision, (ii) different types of anomalies, and (iii) noisy and corrupted data? In this work,…

Machine Learning · Computer Science 2022-09-20 Songqiao Han , Xiyang Hu , Hailiang Huang , Mingqi Jiang , Yue Zhao

Surface prediction and completion have been widely studied in various applications. Recently, research in surface completion has evolved from small objects to complex large-scale scenes. As a result, researchers have begun increasing the…

Robotics · Computer Science 2024-03-19 Guiyong Zheng , Jinqi Jiang , Chen Feng , Shaojie Shen , Boyu Zhou

We consider the problem of cross-view geo-localization. The primary challenge of this task is to learn the robust feature against large viewpoint changes. Existing benchmarks can help, but are limited in the number of viewpoints. Image…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Zhedong Zheng , Yunchao Wei , Yi Yang

While recent advancements in vision-language models have had a transformative impact on multi-modal comprehension, the extent to which these models possess the ability to comprehend generated images remains uncertain. Synthetic images, in…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Keqiang Sun , Junting Pan , Yuying Ge , Hao Li , Haodong Duan , Xiaoshi Wu , Renrui Zhang , Aojun Zhou , Zipeng Qin , Yi Wang , Jifeng Dai , Yu Qiao , Limin Wang , Hongsheng Li

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first…

Computation and Language · Computer Science 2025-09-23 Fei Zhao , Chengqiang Lu , Yufan Shen , Qimeng Wang , Yicheng Qian , Haoxin Zhang , Yan Gao , Yi Wu , Yao Hu , Zhen Wu , Shangyu Xing , Xinyu Dai

Egocentric human video data, which captures rich human-environment interactions and can be collected at scale, has become a key driver of embodied intelligence research. However, existing egocentric datasets typically lack tactile sensing,…

With the rapid development of high-speed communication and artificial intelligence technologies, human perception of real-world scenes is no longer limited to the use of small Field of View (FoV) and low-dimensional scene detection devices.…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Shaohua Gao , Kailun Yang , Hao Shi , Kaiwei Wang , Jian Bai

We propose a novel method for combining synthetic and real images when training networks to determine geometric information from a single image. We suggest a method for mapping both image types into a single, shared domain. This is…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Koutilya PNVR , Hao Zhou , David Jacobs

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Shuo Cao , Jiayang Li , Xiaohui Li , Yuandong Pu , Kaiwen Zhu , Yuanting Gao , Siqi Luo , Yi Xin , Qi Qin , Yu Zhou , Xiangyu Chen , Wenlong Zhang , Bin Fu , Yu Qiao , Yihao Liu

Astronaut photography, spanning six decades of human spaceflight, presents a unique Earth observations dataset with immense value for both scientific research and disaster response. Despite its significance, accurately localizing the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Gabriele Berton , Alex Stoken , Barbara Caputo , Carlo Masone

Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pristine visual inputs and overlook the degradations that commonly occur in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xiaolong Zhou , Yifei Liu , Ziyang Gong , Jiarui Li , Qiyue Zhao , Muyao Niu , Yuanyuan Gao , Le Ma , Xue Yang , Hongjie Zhang , Zhihang Zhong

Unraveling the hierarchical structure-property relationships is the central challenge of materials science, necessitating the interpretation of data across vast physical scales from micro to macro. Despite the rapid integration of Large…

Digital Libraries · Computer Science 2026-03-23 Yuting Zheng , Zijian Chen , Qi Jia

Vision-aided wireless sensing is emerging as a cornerstone of 6G mobile computing. While data-driven approaches have advanced rapidly, establishing a precise geometric correspondence between ego-centric visual data and radio propagation…

Signal Processing · Electrical Eng. & Systems 2026-01-28 Yingzhe Mao , Chao Zou , Yanqun Tang

The rapidly developing field of large multimodal models (LMMs) has led to the emergence of diverse models with remarkable capabilities. However, existing benchmarks fail to comprehensively, objectively and accurately evaluate whether LMMs…

We present a rigorous mathematical solution to photometric redshift estimation and the more general inversion problem. The challenge we address is to meaningfully constrain unknown properties of astronomical sources based on given…

Astrophysics · Physics 2011-02-11 Tamas Budavari

Multimodal remote sensing image (MRSI) matching is pivotal for cross-modal fusion, localization, and object detection, but it faces severe challenges due to geometric, radiometric, and viewpoint discrepancies across imaging modalities.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Peihao Wu , Yongxiang Yao , Wenfei Zhang , Dong Wei , Yi Wan , Yansheng Li , Yongjun Zhang