English
Related papers

Related papers: Identification of Stone Deterioration Patterns wit…

200 papers

This technical report aims to fill a deficiency in the assessment of large multimodal models (LMMs) by specifically examining the self-consistency of their outputs when subjected to common corruptions. We investigate the cross-modal…

Machine Learning · Computer Science 2024-01-23 Jiawei Zhang , Tianyu Pang , Chao Du , Yi Ren , Bo Li , Min Lin

The reliability of substation equipment is crucial to the stability of power systems, but traditional fault analysis methods heavily rely on manual expertise, limiting their effectiveness in handling complex and large-scale data. This paper…

Artificial Intelligence · Computer Science 2024-12-24 Jinzhi Wang , Qinfeng Song , Lidong Qian , Haozhou Li , Qinke Peng , Jiangbo Zhang

Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and interacting with the physical world. While traditional methods achieve high fidelity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Weijie Wang , Qihang Cao , Sensen Gao , Donny Y. Chen , Haofei Xu , Wenjing Bian , Songyou Peng , Tat-Jen Cham , Chuanxia Zheng , Andreas Geiger , Jianfei Cai , Jia-Wang Bian , Bohan Zhuang

Multi-modality (or multi-channel) imaging is becoming increasingly important and more widely available, e.g. hyperspectral imaging in remote sensing, spectral CT in material sciences as well as multi-contrast MRI and PET-MR in medicine.…

Image and Video Processing · Electrical Eng. & Systems 2020-12-25 Leon Bungert , Matthias J. Ehrhardt

Vision foundation models (FMs) are accelerating the development of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly…

Many animal species can approximately judge the number of objects in a visual scene at a single glance, and humans can further determine the exact cardinality of a set by deploying systematic counting procedures. In contrast, it has been…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Alberto Testolin , Kuinan Hou , Marco Zorzi

Maintaining software artifacts is among the hardest tasks an engineer faces. Like any other piece of code, model transformations developed by engineers are also subject to maintenance. To facilitate the comprehension of programs, software…

Software Engineering · Computer Science 2020-10-13 Chihab eddine Mokaddem , Houari Sahraoui , Eugene Syriani

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent methods have shown strong performance by identifying degradation characteristics to guide…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Eunho Lee , Rei Kawakami , Youngbae Hwang

Large Language Models (LLMs) are rapidly being adopted in conservation to address the biodiversity crisis, yet their reliability for species evaluation is uncertain. This study systematically validates five leading models on 21,955 species…

Computation and Language · Computer Science 2025-10-06 Shinya Uryu

The landscape of computer graphics has undergone significant transformations with the recent advances of differentiable rendering models. These rendering models often rely on heuristic designs that may not fully align with the final…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Fangneng Zhan , Hanxue Liang , Yifan Wang , Michael Niemeyer , Michael Oechsle , Adam Kortylewski , Cengiz Oztireli , Gordon Wetzstein , Christian Theobalt

Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite these gains, unified models often underperform compared to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Zhiyang Xu , Jiuhai Chen , Zhaojiang Lin , Xichen Pan , Lifu Huang , Tianyi Zhou , Madian Khabsa , Qifan Wang , Di Jin , Michihiro Yasunaga , Lili Yu , Xi Victoria Lin , Shaoliang Nie

We introduce a set of image transformations that can be used as corruptions to evaluate the robustness of models as well as data augmentation mechanisms for training neural networks. The primary distinction of the proposed transformations…

Computer Vision and Pattern Recognition · Computer Science 2022-05-02 Oğuzhan Fatih Kar , Teresa Yeo , Andrei Atanov , Amir Zamir

Nowdays, most datasets used to train and evaluate super-resolution models are single-modal simulation datasets. However, due to the variety of image degradation types in the real world, models trained on single-modal simulation datasets do…

Image and Video Processing · Electrical Eng. & Systems 2020-04-14 Haoran Li , Weihong Quan , Meijun Yan , Jin zhang , Xiaoli Gong , Jin Zhou

Depth estimation and 3D reconstruction have been extensively studied as core topics in computer vision. Starting from rigid objects with relatively simple geometric shapes, such as vehicles, the research has expanded to address general…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Muhammad Aamir , Naoya Muramatsu , Sangyun Shin , Matthew Wijers , Jia-Xing Zhong , Xinyu Hou , Amir Patel , Andrew Loveridge , Andrew Markham

Foundation Models (FMs), such as OpenAI's GPT, are fundamentally transforming the practice of software engineering by enabling the development of \emph{FMware} -- applications and infrastructures built around these models. FMware systems…

Software Engineering · Computer Science 2025-10-14 Zitao Wang , Zhimin Zhao , Michael W. Godfrey

Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of scenarios, including autonomous driving and medical…

Machine Learning · Computer Science 2024-11-04 Qingyang Zhang , Yake Wei , Zongbo Han , Huazhu Fu , Xi Peng , Cheng Deng , Qinghua Hu , Cai Xu , Jie Wen , Di Hu , Changqing Zhang

Photo restoration technology enables preserving visual memories in photographs. However, physical prints are vulnerable to various forms of deterioration, ranging from physical damage to loss of image quality, etc. While restoration by…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Seung-Yeon Back , Geonho Son , Dahye Jeong , Eunil Park , Simon S. Woo

Standard 3D reconstruction pipelines assume stationary world, therefore suffer from `ghost artifacts' whenever dynamic objects are present in the scene. Recent approaches has started tackling this issue, however, they typically either only…

Computer Vision and Pattern Recognition · Computer Science 2019-04-17 Ondrej Miksik , Vibhav Vineet

Multimodal image registration plays a key role in creating digital patient models by combining data from different imaging techniques into a single coordinate system. This process often involves multiple sequential and interconnected…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Agnieszka Anna Tomaka , Dariusz Pojda , Michał Tarnawski , Leszek Luchowski