中文
相关论文

相关论文: MoonAnything: A Vision Benchmark with Large-Scale …

200 篇论文

With the rapid advancement of e-commerce, exploring general representations rather than task-specific ones has attracted increasing research attention. For product understanding, although existing discriminative dual-flow architectures…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Daoze Zhang , Chenghan Fu , Zhanheng Nie , Jianyu Liu , Wanxian Guan , Yuan Gao , Jun Song , Pengjie Wang , Jian Xu , Bo Zheng

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

An event-based camera outputs an event whenever a change in scene brightness of a preset magnitude is detected at a particular pixel location in the sensor plane. The resulting sparse and asynchronous output coupled with the high dynamic…

计算机视觉与模式识别 · 计算机科学 2023-08-02 Loïc J. Azzalini , Emmanuel Blazquez , Alexander Hadjiivanov , Gabriele Meoni , Dario Izzo

The visual detection and tracking of surface terrain is required for spacecraft to safely land on or navigate within close proximity to celestial objects. Current approaches rely on template matching with pre-gathered patch-based features,…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Timothy Chase , Karthik Dantu

Modern astronomical observatories generate a massive volume of multimodal data, creating a critical bottleneck for expert human review. While multimodal large language models (LLMs) have shown promise in interpreting complex visual and…

In this paper, we introduce a novel benchmark designed to propel the advancement of general-purpose, large-scale 3D vision models for remote sensing imagery. While several datasets have been proposed within the realm of remote sensing, many…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Jiayu Wang , Ruizhi Wang , Jie Song , Haofei Zhang , Mingli Song , Zunlei Feng , Li Sun

Accurate object geometry estimation is essential for many downstream tasks, including robotic manipulation and physical interaction. Although vision is the dominant modality for shape perception, it becomes unreliable under occlusions or…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Langzhe Gu , Hung-Jui Huang , Mohamad Qadri , Michael Kaess , Wenzhen Yuan

Monocular relative and metric depth estimation has seen a tremendous boost in the last few years due to the sharp advancements in foundation models and in particular transformer based networks. As we start to see applications to the domain…

Recent advances in deep learning for remote sensing rely heavily on large annotated datasets, yet acquiring high-quality ground truth for geometric, radiometric, and multi-domain tasks remains costly and often infeasible. In particular, the…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Safouane El Ghazouali , Nicola Venturi , Michael Rueegsegger , Umberto Michelucci

Utilizing the complementary strengths of wavelength-specific range or depth sensors is crucial for robust computer-assisted tasks such as autonomous driving. Despite this, there is still little research done at the intersection of optical…

图像与视频处理 · 电气工程与系统科学 2025-11-11 Vanessa Wirth , Johanna Bräunig , Nikolai Hofmann , Martin Vossiek , Tim Weyrich , Marc Stamminger

Global localization is necessary for autonomous operations on the lunar surface where traditional Earth-based navigation infrastructure, such as GPS, is unavailable. As NASA advances toward sustained lunar presence under the Artemis…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Annika Thomas , Robaire Galliath , Aleksander Garbuz , Luke Anger , Cormac O'Neill , Trevor Johst , Dami Thomas , George Lordos , Jonathan P. How

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yolo Y. Tang , Pinxin Liu , Zhangyun Tan , Mingqian Feng , Rui Mao , Chao Huang , Jing Bi , Yunzhong Xiao , Susan Liang , Hang Hua , Ali Vosoughi , Luchuan Song , Zeliang Zhang , Chenliang Xu

Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus on general multimodal capabilities but fail to capture the…

天体物理仪器与方法 · 物理学 2025-10-22 Jinghang Shi , Xiaoyu Tang , Yang Huang , Yuyang Li , Xiao Kong , Yanxia Zhang , Caizhan Yue

Multimodal large language models (MLLMs) have made significant progress in integrating visual and linguistic understanding. Existing benchmarks typically focus on high-level semantic capabilities, such as scene understanding and visual…

计算与语言 · 计算机科学 2025-02-18 Shangyu Xing , Changhao Xiang , Yuteng Han , Yifan Yue , Zhen Wu , Xinyu Liu , Zhangtai Wu , Fei Zhao , Xinyu Dai

This paper compares scale-invariant (SIFT) and scale-variant (ORB) feature detection methods, alongside our novel feature detector, IntFeat, specifically applied to lunar imagery. We evaluate these methods using low (128x128) and…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Ashutosh Kumar , Sarthak Kaushal , Shiv Vignesh Murthy

Current Large Multimodal Models (LMMs) in Earth Observation typically neglect the critical "vertical" dimension, limiting their reasoning capabilities in complex remote sensing geometries and disaster scenarios where physical spatial…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Xuran Hu , Zhitong Xiong , Zhongcheng Hong , Yifang Ban , Xiaoxiang Zhu , Wufan Zhao

Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Zhekai Chen , Yuqing Wang , Manyuan Zhang , Xihui Liu

Vision Based Navigation consists in utilizing cameras as precision sensors for GNC after extracting information from images. To enable the adoption of machine learning for space applications, one of obstacles is the demonstration that…

Existing all-in-one image restoration approaches, which aim to handle multiple weather degradations within a single framework, are predominantly trained and evaluated using mixed single-weather synthetic datasets. However, these datasets…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Qiyuan Guan , Qianfeng Yang , Xiang Chen , Tianyu Song , Guiyue Jin , Jiyu Jin

Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Baorui Ma , Jiahui Yang , Donglin Di , Xuancheng Zhang , Jianxun Cui , Hao Li , Yan Xie , Wei Chen