中文

面向半导体电子显微镜分析的多模态指令调优小规模视觉-语言助手

计算机视觉与模式识别 2024-09-13 v1 机器学习

摘要

我们提出了一个新框架,用于利用视觉-语言指令调优分析和解释半导体制造中的电子显微镜图像。该框架采用独特的教师-学生方法,利用诸如 GPT-4 等预训练的多模态大型语言模型生成用于零样本视觉问答(VQA)和分类任务的指令-following 数据,定制较小的多模态模型(SMMs)进行显微镜图像分析, resulting in an instruction-tuned language-and-vision assistant。Our framework merges knowledge engineering with machine learning to integrate domain-specific expertise from larger to smaller multimodal models within this specialized field, greatly reducing the need for extensive human labeling. Our study presents a secure, cost-effective, and customizable approach for analyzing microscopy images, addressing the challenges of adopting proprietary models in semiconductor manufacturing.

关键词

引用

@article{arxiv.2409.07463,
  title  = {Multi-Modal Instruction-Tuning Small-Scale Language-and-Vision Assistant for Semiconductor Electron Micrograph Analysis},
  author = {Sakhinana Sagar Srinivas and Geethan Sannidhi and Venkataramana Runkana},
  journal= {arXiv preprint arXiv:2409.07463},
  year   = {2024}
}

备注

Paper published at AAAI 2024 Spring Symposium Series