中文
相关论文

相关论文: Image and Data Mining in Reticular Chemistry Using…

200 篇论文

Multimodal large language models (MLLMs) represent an evolutionary expansion in the capabilities of traditional large language models, enabling them to tackle challenges that surpass the scope of purely text-based applications. It leverages…

计算与语言 · 计算机科学 2025-01-17 Jinlong He , Pengfei Li , Gang Liu , Genrong He , Zhaolin Chen , Shenjun Zhong

We use prompt engineering to guide ChatGPT in the automation of text mining of metal-organic frameworks (MOFs) synthesis conditions from diverse formats and styles of the scientific literature. This effectively mitigates ChatGPT's tendency…

信息检索 · 计算机科学 2023-10-04 Zhiling Zheng , Oufan Zhang , Christian Borgs , Jennifer T. Chayes , Omar M. Yaghi

Physically realistic materials are pivotal in augmenting the realism of 3D assets across various applications and lighting conditions. However, existing 3D assets and generative models often lack authentic material properties. Manual…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Ye Fang , Zeyi Sun , Tong Wu , Jiaqi Wang , Ziwei Liu , Gordon Wetzstein , Dahua Lin

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

物理教育 · 物理学 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic

Multi-modal retrieval-augmented generation (MM-RAG) promises grounded biomedical QA, but it is unclear when to (i) convert figures/tables into text versus (ii) use optical character recognition (OCR)-free visual retrieval that returns page…

计算与语言 · 计算机科学 2025-12-19 Primož Kocbek , Azra Frkatović-Hodžić , Dora Lalić , Vivian Hui , Gordan Lauc , Gregor Štiglic

Traditional deep learning-based methods for classifying cellular features in microscopy images require time- and labor-intensive processes for training models. Among the current limitations are major time commitments from domain experts for…

图像与视频处理 · 电气工程与系统科学 2025-01-22 Abhiram Kandiyana , Peter R. Mouton , Yaroslav Kolinko , Lawrence O. Hall , Dmitry Goldgof

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

This paper describes a rapid feasibility study of using GPT-4, a large language model (LLM), to (semi)automate data extraction in systematic reviews. Despite the recent surge of interest in LLMs there is still a lack of understanding of how…

This paper discusses the effectiveness of leveraging Chatbot: Generative Pre-trained Transformer (ChatGPT) versions 3.5 and 4 for analyzing research papers for effective writing of scientific literature surveys. The study selected the…

Predicting pedestrian behavior is the key to ensure safety and reliability of autonomous vehicles. While deep learning methods have been promising by learning from annotated video frame sequences, they often fail to fully grasp the dynamic…

计算机视觉与模式识别 · 计算机科学 2024-01-29 Jia Huang , Peng Jiang , Alvika Gautam , Srikanth Saripalli

In radiology, Artificial Intelligence (AI) has significantly advanced report generation, but automatic evaluation of these AI-produced reports remains challenging. Current metrics, such as Conventional Natural Language Generation (NLG) and…

Multi-modality promises to unlock further uses for large language models. Recently, the state-of-the-art language model GPT-4 was enhanced with vision capabilities. We carry out a prompting evaluation of GPT-4V and five other baselines on…

计算与语言 · 计算机科学 2023-12-20 Mukul Singh , José Cambronero , Sumit Gulwani , Vu Le , Gust Verbruggen

The rapidly evolving sector of Multi-modal Large Language Models (MLLMs) is at the forefront of integrating linguistic and visual processing in artificial intelligence. This paper presents an in-depth comparative study of two pioneering…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Zhangyang Qi , Ye Fang , Mengchen Zhang , Zeyi Sun , Tong Wu , Ziwei Liu , Dahua Lin , Jiaqi Wang , Hengshuang Zhao

We present a comprehensive benchmark dataset for Knowledge Graph Question Answering in Materials Science (KGQA4MAT), with a focus on metal-organic frameworks (MOFs). A knowledge graph for metal-organic frameworks (MOF-KG) has been…

Recently, GPT-4o has garnered significant attention for its strong performance in image generation, yet open-source models still lag behind. Several studies have explored distilling image data from GPT-4o to enhance open-source models,…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Junyan Ye , Dongzhi Jiang , Zihao Wang , Leqi Zhu , Zhenghao Hu , Zilong Huang , Jun He , Zhiyuan Yan , Jinghua Yu , Hongsheng Li , Conghui He , Weijia Li

The transition from task-specific artificial intelligence toward general-purpose foundation models raises fundamental questions about their capacity to support the integrated reasoning required in clinical medicine, where diagnosis demands…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Alexandru Florea , Shansong Wang , Mingzhe Hu , Qiang Li , Zach Eidex , Luke del Balzo , Mojtaba Safari , Xiaofeng Yang

Medical image analysis is crucial in modern radiological diagnostics, especially given the exponential growth in medical imaging data. The demand for automated report generation systems has become increasingly urgent. While prior research…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Hao Chen , Wei Zhao , Yingli Li , Tianyang Zhong , Yisong Wang , Youlan Shang , Lei Guo , Junwei Han , Tianming Liu , Jun Liu , Tuo Zhang

The automatic analysis of chemical literature has immense potential to accelerate the discovery of new materials and drugs. Much of the critical information in patent documents and scientific articles is contained in figures, depicting the…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Lucas Morin , Martin Danelljan , Maria Isabel Agea , Ahmed Nassar , Valery Weber , Ingmar Meijer , Peter Staar , Fisher Yu

Multimodal scientific reasoning remains a significant challenge for large language models (LLMs), particularly in chemistry, where problem-solving relies on symbolic diagrams, molecular structures, and structured visual data. Here, we…

计算与语言 · 计算机科学 2025-12-18 Yiming Cui , Xin Yao , Yuxuan Qin , Xin Li , Shijin Wang , Guoping Hu

Although gold nanorods have been the subject of much research, the pathways for controlling their shape and thereby their optical properties remain largely heuristically understood. Although it is apparent that the simultaneous presence of…