中文
相关论文

相关论文: Towards Open-Vocabulary Industrial Defect Understa…

200 篇论文

Steel surface defect analysis is critical for industrial quality control, yet existing benchmarks rely primarily on label-only annotations, limiting fine-grained semantic understanding and systematic evaluation of vision-language models. To…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Shuxian Zhao , Jie Gui , Baosheng Yu , Dacheng Tao

In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite…

人工智能 · 计算机科学 2025-02-24 Xi Jiang , Jian Li , Hanqiu Deng , Yong Liu , Bin-Bin Gao , Yifeng Zhou , Jialin Li , Chengjie Wang , Feng Zheng

In industrial settings, the accurate detection of anomalies is essential for maintaining product quality and ensuring operational safety. Traditional industrial anomaly detection (IAD) models often struggle with flexibility and…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Zhiling Chen , Hanning Chen , Mohsen Imani , Farhad Imani

Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains largely unexplored. We introduce SoM-1K, the first large-scale multimodal benchmark dataset…

计算与语言 · 计算机科学 2025-09-26 Qixin Wan , Zilong Wang , Jingwen Zhou , Wanting Wang , Ziheng Geng , Jiachen Liu , Ran Cao , Minghui Cheng , Lu Cheng

In recent years, the upstream of Large Language Models (LLM) has also encouraged the computer vision community to work on substantial multimodal datasets and train models on a scale in a self-/semi-supervised manner, resulting in Vision…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Keno Moenck , Duc Trung Thieu , Julian Koch , Thorsten Schüppstuhl

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Vaclav Kosar , Antonín Hoskovec , Milan Šulc , Radek Bartyzal

Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Peng Xia , Siwei Han , Shi Qiu , Yiyang Zhou , Zhaoyang Wang , Wenhao Zheng , Zhaorun Chen , Chenhang Cui , Mingyu Ding , Linjie Li , Lijuan Wang , Huaxiu Yao

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image…

Despite progress in vision-based inspection algorithms, real-world industrial challenges -- specifically in data availability, quality, and complex production requirements -- often remain under-addressed. We introduce the VISION Datasets, a…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Haoping Bai , Shancong Mou , Tatiana Likhomanenko , Ramazan Gokberk Cinbis , Oncel Tuzel , Ping Huang , Jiulong Shan , Jianjun Shi , Meng Cao

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhenyue Qin , Yu Yin , Dylan Campbell , Xuansheng Wu , Ke Zou , Yih-Chung Tham , Ninghao Liu , Xiuzhen Zhang , Qingyu Chen

Defect inspection is paramount within the closed-loop manufacturing system. However, existing datasets for defect inspection often lack precision and semantic granularity required for practical applications. In this paper, we introduce the…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Shuai Yang , Zhifei Chen , Pengguang Chen , Xi Fang , Yixun Liang , Shu Liu , Yingcong Chen

Object 6DoF (6D) pose estimation is essential for robotic perception, especially in industrial settings. It enables robots to interact with the environment and manipulate objects. However, existing benchmarks on object 6D pose estimation…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Ruimin Ma , Sebastian Zudaire , Zhen Li , Chi Zhang

Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements,…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Yuan Liu , Saihui Hou , Saijie Hou , Jiabao Du , Shibei Meng , Yongzhen Huang

Large language models (LLMs) are effective at capturing complex, valuable conceptual representations from textual data for a wide range of real-world applications. However, in fields like Intelligent Fault Diagnosis (IFD), incorporating…

人工智能 · 计算机科学 2024-12-03 Hamzah A. A. M. Qaid , Bo Zhang , Dan Li , See-Kiong Ng , Wei Li

The emergence of vision-language models has transformed medical AI, enabling unprecedented advances in diagnostic capability and clinical applications. However, progress in dermatology has lagged behind other medical domains due to the lack…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Siyuan Yan , Ming Hu , Yiwen Jiang , Xieji Li , Hao Fei , Philipp Tschandl , Harald Kittler , Zongyuan Ge

Utility companies increasingly rely on drone imagery for post-event and routine inspection, but training accurate defect-type classifiers remains difficult because defect examples are rare and inspection datasets are often limited or…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Xuesong Wang , Caisheng Wang

In real-world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any interleaved multimedia form. This capability defines the goal…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Yanlin Li , Minghui Guo , Kaiwen Zhang , Shize Zhang , Yiran Zhao , Haodong Li , Congyue Zhou , Weijie Zheng , Yushen Yan , Shengqiong Wu , Wei Ji , Lei Cui , Furu Wei , Hao Fei , Mong-Li Lee , Wynne Hsu

With the rapid advancement of Multimodal Large Language Models (MLLMs), numerous evaluation benchmarks have emerged. However, comprehensive assessments of their performance across diverse industrial applications remain limited. In this…

计算与语言 · 计算机科学 2025-01-29 Dongyi Yi , Guibo Zhu , Chenglin Ding , Zongshu Li , Dong Yi , Jinqiao Wang

In precision agriculture, the detection and recognition of insects play an essential role in the ability of crops to grow healthy and produce a high-quality yield. The current machine vision model requires a large volume of data to achieve…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Hoang-Quan Nguyen , Thanh-Dat Truong , Xuan Bac Nguyen , Ashley Dowling , Xin Li , Khoa Luu

While tobacco advertising innovates at unprecedented speed, traditional surveillance methods remain frozen in time, especially in the context of social media. The lack of large-scale, comprehensive datasets and sophisticated monitoring…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Naga VS Raviteja Chappa , Matthew Shepard , Connor McCurtain , Charlotte McCormick , Page Daniel Dobbs , Khoa Luu
‹ 上一页 1 2 3 10 下一页 ›