中文
相关论文

相关论文: A Multimodal Benchmark Dataset and Model for Crop …

200 篇论文

Prenatal diagnosis of Congenital Heart Diseases (CHDs) holds great potential for Artificial Intelligence (AI)-driven solutions. However, collecting high-quality diagnostic data remains difficult due to the rarity of these conditions,…

Automated breast cancer detection via computer vision techniques is challenging due to the complex nature of breast tissue, the subtle appearance of cancerous lesions, and variations in breast density. Mainstream techniques primarily focus…

定量方法 · 定量生物学 2025-12-11 Noor Ul Huda Shah , Tanveer Hussain , Amr Ahmed , Yonghuai Liu , Usman Ali , Ardhendu Behera

Leaf disease identification plays a pivotal role in smart agriculture. However, many existing studies still struggle to integrate image and textual modalities to compensate for each other's limitations. Furthermore, many of these approaches…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Khang Nguyen Quoc , Lan Le Thi Thu , Luyl-Da Quach

Medicine is inherently multimodal and multitask, with diverse data modalities spanning text, imaging. However, most models in medical field are unimodal single tasks and lack good generalizability and explainability. In this study, we…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Lijian Xu , Hao Sun , Ziyu Ni , Hongsheng Li , Shaoting Zhang

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition.…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Abdurrahim Yilmaz , Ozan Erdem , Ece Gokyayla , Ayda Acar , Burc Bugra Dagtas , Dilara Ilhan Erdil , Gulsum Gencoglan , Burak Temelkuran

Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and domain-specific terminology. Such heterogeneous evidence is difficult for laypersons to…

计算与语言 · 计算机科学 2026-05-14 Sike Xiang , Shuang Chen , Kevin Qinghong Lin , Jialin Yu , Yijia Sun , Philip Torr , Amir Atapour-Abarghouei

Recent advancements in large language models (LLMs) have demonstrated extraordinary comprehension capabilities with remarkable breakthroughs on various vision-language tasks. However, the application of LLMs in generating reliable medical…

人工智能 · 计算机科学 2025-02-18 Xueshen Li , Xinlong Hou , Ziyi Huang , Yu Gan

Computer-assisted diagnostic and prognostic systems of the future should be capable of simultaneously processing multimodal data. Multimodal deep learning (MDL), which involves the integration of multiple sources of data, such as images and…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Zhaoyi Sun , Mingquan Lin , Qingqing Zhu , Qianqian Xie , Fei Wang , Zhiyong Lu , Yifan Peng

With the widespread application of artificial intelligence (AI), particularly deep learning (DL) and vision large language models (VLLMs), in skin disease diagnosis, the need for interpretability becomes crucial. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Yuhao Shen , Liyuan Sun , Yan Xu , Wenbin Liu , Shuping Zhang , Shawn Afvari , Zhongyi Han , Jiaoyan Song , Yongzhi Ji , Tao Lu , Xiaonan He , Xin Gao , Juexiao Zhou

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Wei Chen , Lin Li , Yongqi Yang , Bin Wen , Fan Yang , Tingting Gao , Yu Wu , Long Chen

Multi-modal learning has significantly advanced generative AI, especially in vision-language modeling. Innovations like GPT-4V and open-source projects such as LLaVA have enabled robust conversational agents capable of zero-shot task…

计算与语言 · 计算机科学 2024-06-17 Zekai Chen , Arda Pekis , Kevin Brown

Objective: Disease knowledge graphs are a way to connect, organize, and access disparate information about diseases with numerous benefits for artificial intelligence (AI). To create knowledge graphs, it is necessary to extract knowledge…

机器学习 · 计算机科学 2022-09-01 Yucong Lin , Keming Lu , Sheng Yu , Tianxi Cai , Marinka Zitnik

Compared with the domain-specific model, the vision-language pre-training models (VLPMs) have shown superior performance on downstream tasks with fast fine-tuning process. For example, ERNIE-ViL, Oscar and UNIMO trained VLPMs with a uniform…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Sha Yuan , Shuai Zhao , Jiahong Leng , Zhao Xue , Hanyu Zhao , Peiyu Liu , Zheng Gong , Wayne Xin Zhao , Junyi Li , Jie Tang

Rapid developments of AI tools are expected to offer unprecedented assistance to the research of natural science including chemistry. However, neither existing unimodal task-specific specialist models nor emerging general large multimodal…

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from satisfactory. This paper proposes FoodLMM, a versatile food…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Yuehao Yin , Huiyan Qi , Bin Zhu , Jingjing Chen , Yu-Gang Jiang , Chong-Wah Ngo

Recent studies on plant disease diagnosis using machine learning (ML) have highlighted concerns about the overestimated diagnostic performance due to inappropriate data partitioning, where training and test datasets are derived from the…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Yuji Arima , Satoshi Kagiwada , Hitoshi Iyatomi

Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular,…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Jing Hao , Yuxuan Fan , Yanpeng Sun , Kaixin Guo , Lizhuo Lin , Jinrong Yang , Qi Yong H. Ai , Lun M. Wong , Hao Tang , Kuo Feng Hung

Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual question answering. These models are trained on large-scale…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Xiaomei Zhang , Hanyu Zheng , Xiangyu Zhu , Jinghuan Wei , Junhong Zou , Zhen Lei , Zhaoxiang Zhang