English
Related papers

Related papers: Identification of Stone Deterioration Patterns wit…

200 papers

While there is much excitement about the potential of large multimodal models (LMM), a comprehensive evaluation is critical to establish their true capabilities and limitations. In support of this aim, we evaluate two state-of-the-art LMMs,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Mengchen Liu , Chongyan Chen , Danna Gurari

Current autonomous driving perception models primarily rely on supervised learning with predefined categories. However, these models struggle to detect general obstacles not included in the fixed category set due to their variability and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Tamás Matuszka , Péter Hajas , Dávid Szeghy

The robustness of object detection models is a major concern when applied to real-world scenarios. The performance of most models tends to degrade when confronted with images affected by corruptions, since they are usually trained and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Haodong He , Jian Ding , Bowen Xu , Gui-Song Xia

With the rapid development of mobile devices, modern widely-used mobile phones typically allow users to capture 4K resolution (i.e., ultra-high-definition) images. However, for image demoireing, a challenging task in low-level vision,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Xin Yu , Peng Dai , Wenbo Li , Lan Ma , Jiajun Shen , Jia Li , Xiaojuan Qi

The steel structure demolition scheme needs to be compiled according to the specific engineering characteristics and the update results of the finite element model. The designers need to refer to the relevant engineering cases according to…

Computation and Language · Computer Science 2025-08-25 Zhifeng Yang , Peizong Wu

3D object detection is fundamentally important for various emerging applications, including autonomous driving and robotics. A key requirement for training an accurate 3D object detector is the availability of a large amount of LiDAR-based…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Ruiyu Mao , Sarthak Kumar Maharana , Rishabh K Iyer , Yunhui Guo

Earth observation data presents a unique challenge: it is spatial like images, sequential like video or text, and highly multimodal. We present OlmoEarth: a multimodal, spatio-temporal foundation model that employs a novel self-supervised…

Foundation models can be disruptive for future AI development by scaling up deep learning in terms of model size and training data's breadth and size. These models achieve state-of-the-art performance (often through further adaptation) on a…

Artificial Intelligence · Computer Science 2022-12-20 Johannes Schneider

Recently, large models, or foundation models, have exhibited remarkable performance, profoundly impacting research paradigms in diverse domains. Foundation models, trained on extensive and diverse datasets, provide exceptional…

Geophysics · Physics 2024-12-30 Qi Liu , Jianwei Ma

This work uniquely identifies and characterizes four prevalent multimodal model architectural patterns in the contemporary multimodal landscape. Systematically categorizing models by architecture type facilitates monitoring of developments…

Artificial Intelligence · Computer Science 2024-05-29 Shakti N. Wadekar , Abhishek Chaurasia , Aman Chadha , Eugenio Culurciello

In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These…

We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 TsaiChing Ni , ZhenQi Chen , YuanFu Yang

Multi-modal foundation models combining vision and language models such as Flamingo or GPT-4 have recently gained enormous interest. Alignment of foundation models is used to prevent models from providing toxic or harmful output. While…

Machine Learning · Computer Science 2023-08-22 Christian Schlarmann , Matthias Hein

Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled…

This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics…

Computation and Language · Computer Science 2025-11-10 Jonghyun Lee , Dojun Park , Jiwoo Lee , Hoekeon Choi , Sung-Eun Lee

Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio…

Sound · Computer Science 2025-06-23 Charilaos Papaioannou , Emmanouil Benetos , Alexandros Potamianos

With the widespread adoption and development of mobile devices, vision-based recognition applications have become a hot topic in research. Jade, as an important cultural heritage and artistic item, has significant applications in fields…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Zhenyu Wang , Wenjia Li , Pengyu Zhu

Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-source models against a very challenging real-world medical form that mixes dates;…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Nicholas Pather , Joshua Fouché , Sitwala Mundia , Karl-Günter Technau , Thokozile Malaba , Alex Welte , Ushma Mehta , Bruce A. Bassett

Due to the increase in computational resources and accessibility of data, an increase in large, deep learning models trained on copious amounts of multi-modal data using self-supervised or semi-supervised learning have emerged. These…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Madeline Chantry Schiappa , Shehreen Azad , Sachidanand VS , Yunhao Ge , Ondrej Miksik , Yogesh S. Rawat , Vibhav Vineet

With the development of large models, watermarks are increasingly employed to assert copyright, verify authenticity, or monitor content distribution. As applications become more multimodal, the utility of watermarking techniques becomes…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Jielin Qiu , William Han , Xuandong Zhao , Shangbang Long , Christos Faloutsos , Lei Li