中文
相关论文

相关论文: Identification of Stone Deterioration Patterns wit…

200 篇论文

While there is much excitement about the potential of large multimodal models (LMM), a comprehensive evaluation is critical to establish their true capabilities and limitations. In support of this aim, we evaluate two state-of-the-art LMMs,…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Mengchen Liu , Chongyan Chen , Danna Gurari

Current autonomous driving perception models primarily rely on supervised learning with predefined categories. However, these models struggle to detect general obstacles not included in the fixed category set due to their variability and…

计算机视觉与模式识别 · 计算机科学 2024-08-23 Tamás Matuszka , Péter Hajas , Dávid Szeghy

The robustness of object detection models is a major concern when applied to real-world scenarios. The performance of most models tends to degrade when confronted with images affected by corruptions, since they are usually trained and…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Haodong He , Jian Ding , Bowen Xu , Gui-Song Xia

With the rapid development of mobile devices, modern widely-used mobile phones typically allow users to capture 4K resolution (i.e., ultra-high-definition) images. However, for image demoireing, a challenging task in low-level vision,…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Xin Yu , Peng Dai , Wenbo Li , Lan Ma , Jiajun Shen , Jia Li , Xiaojuan Qi

The steel structure demolition scheme needs to be compiled according to the specific engineering characteristics and the update results of the finite element model. The designers need to refer to the relevant engineering cases according to…

计算与语言 · 计算机科学 2025-08-25 Zhifeng Yang , Peizong Wu

3D object detection is fundamentally important for various emerging applications, including autonomous driving and robotics. A key requirement for training an accurate 3D object detector is the availability of a large amount of LiDAR-based…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Ruiyu Mao , Sarthak Kumar Maharana , Rishabh K Iyer , Yunhui Guo

Earth observation data presents a unique challenge: it is spatial like images, sequential like video or text, and highly multimodal. We present OlmoEarth: a multimodal, spatio-temporal foundation model that employs a novel self-supervised…

Foundation models can be disruptive for future AI development by scaling up deep learning in terms of model size and training data's breadth and size. These models achieve state-of-the-art performance (often through further adaptation) on a…

人工智能 · 计算机科学 2022-12-20 Johannes Schneider

Recently, large models, or foundation models, have exhibited remarkable performance, profoundly impacting research paradigms in diverse domains. Foundation models, trained on extensive and diverse datasets, provide exceptional…

地球物理 · 物理学 2024-12-30 Qi Liu , Jianwei Ma

This work uniquely identifies and characterizes four prevalent multimodal model architectural patterns in the contemporary multimodal landscape. Systematically categorizing models by architecture type facilitates monitoring of developments…

人工智能 · 计算机科学 2024-05-29 Shakti N. Wadekar , Abhishek Chaurasia , Aman Chadha , Eugenio Culurciello

In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These…

We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution…

计算机视觉与模式识别 · 计算机科学 2026-01-13 TsaiChing Ni , ZhenQi Chen , YuanFu Yang

Multi-modal foundation models combining vision and language models such as Flamingo or GPT-4 have recently gained enormous interest. Alignment of foundation models is used to prevent models from providing toxic or harmful output. While…

机器学习 · 计算机科学 2023-08-22 Christian Schlarmann , Matthias Hein

Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled…

This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics…

计算与语言 · 计算机科学 2025-11-10 Jonghyun Lee , Dojun Park , Jiwoo Lee , Hoekeon Choi , Sung-Eun Lee

Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio…

声音 · 计算机科学 2025-06-23 Charilaos Papaioannou , Emmanouil Benetos , Alexandros Potamianos

With the widespread adoption and development of mobile devices, vision-based recognition applications have become a hot topic in research. Jade, as an important cultural heritage and artistic item, has significant applications in fields…

计算机视觉与模式识别 · 计算机科学 2025-02-21 Zhenyu Wang , Wenjia Li , Pengyu Zhu

Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-source models against a very challenging real-world medical form that mixes dates;…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Nicholas Pather , Joshua Fouché , Sitwala Mundia , Karl-Günter Technau , Thokozile Malaba , Alex Welte , Ushma Mehta , Bruce A. Bassett

Due to the increase in computational resources and accessibility of data, an increase in large, deep learning models trained on copious amounts of multi-modal data using self-supervised or semi-supervised learning have emerged. These…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Madeline Chantry Schiappa , Shehreen Azad , Sachidanand VS , Yunhao Ge , Ondrej Miksik , Yogesh S. Rawat , Vibhav Vineet

With the development of large models, watermarks are increasingly employed to assert copyright, verify authenticity, or monitor content distribution. As applications become more multimodal, the utility of watermarking techniques becomes…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Jielin Qiu , William Han , Xuandong Zhao , Shangbang Long , Christos Faloutsos , Lei Li