English
Related papers

Related papers: PathAsst: A Generative Foundation AI Assistant Tow…

200 papers

Multimodal Large Language Models (MLLMs) show promise for medical applications, yet progress in dermatology lags due to limited training data, narrow task coverage, and lack of clinically-grounded supervision that mirrors expert diagnostic…

Computation and Language · Computer Science 2026-01-06 Jinghan Ru , Siyuan Yan , Yuguo Yin , Yuexian Zou , Zongyuan Ge

As natural image understanding moves towards the pretrain-finetune era, research in pathology imaging is concurrently evolving. Despite the predominant focus on pretraining pathological foundation models, how to adapt foundation models to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jiaxuan Lu , Fang Yan , Xiaofan Zhang , Yue Gao , Shaoting Zhang

Artificial intelligence (AI)-enabled diagnostics in maxillofacial pathology require structured, high-quality multimodal datasets. However, existing resources provide limited ameloblastoma coverage and lack the format consistency needed for…

Artificial Intelligence · Computer Science 2026-02-06 Ajo Babu George , Anna Mariam John , Athul Anoop , Balu Bhasuran

Whole slide images (WSIs) enable weakly supervised prognostic modeling via multiple instance learning (MIL). Spatial transcriptomics (ST) preserves in situ gene expression, providing a spatial molecular context that complements morphology.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Lihe Liu , Xiaoxi Pan , Yinyin Yuan , Lulu Shang

Accurate image classification and retrieval are of importance for clinical diagnosis and treatment decision-making. The recent contrastive language-image pretraining (CLIP) model has shown remarkable proficiency in understanding natural…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Sunyi Zheng , Xiaonan Cui , Yuxuan Sun , Jingxiong Li , Honglin Li , Yunlong Zhang , Pingyi Chen , Xueping Jing , Zhaoxiang Ye , Lin Yang

AI tools in pathology have improved screening throughput, standardized quantification, and revealed prognostic patterns that inform treatment. However, adoption remains limited because most systems still lack the human-readable reasoning…

Artificial Intelligence · Computer Science 2025-11-18 Yunqi Hong , Johnson Kao , Liam Edwards , Nein-Tzu Liu , Chung-Yen Huang , Alex Oliveira-Kowaleski , Cho-Jui Hsieh , Neil Y. C. Lin

Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency: pervasive noise…

Image and Video Processing · Electrical Eng. & Systems 2025-07-25 Minxi Ouyang , Lianghui Zhu , Yaqing Bao , Qiang Huang , Jingli Ouyang , Tian Guan , Xitong Ling , Jiawen Li , Song Duan , Wenbin Dai , Li Zheng , Xuemei Zhang , Yonghong He

The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR),…

Image and Video Processing · Electrical Eng. & Systems 2025-02-11 Nicholas Evans , Stephen Baker , Miles Reed

In medical imaging, access to data is commonly limited due to patient privacy restrictions and the issue that it can be difficult to acquire enough data in the case of rare diseases.[1] The purpose of this investigation was to develop a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 John R. McNulty , Lee Kho , Alexandria L. Case , Charlie Fornaca , Drew Johnston , David Slater , Joshua M. Abzug , Sybil A. Russell

Pathologists diagnose cancer using gigapixel whole-slide images (WSIs), but the current digital workflow is fragmented. These multiscale datasets often exceed 100,000 x 100,000 pixels, yet standard 2D monitors restrict the field of view.…

Human-Computer Interaction · Computer Science 2026-02-27 Jai Prakash Veerla , Partha Sai Guttikonda , Helen H. Shang , Mohammad Sadegh Nasr , Cesar Torres , Jacob M. Luber

[Purpose] The pathology is decisive for disease diagnosis, but relies heavily on the experienced pathologists. Recently, pathological artificial intelligence (PAI) is thought to improve diagnostic accuracy and efficiency. However, the high…

Image and Video Processing · Electrical Eng. & Systems 2022-05-25 Yuanqing Yang , Kai Sun , Yanhua Gao , Kuangsong Wang , Gang Yu

Artificial Intelligence (AI) has great potential to improve health outcomes by training systems on vast digitized clinical datasets. Computational Pathology, with its massive amounts of microscopy image data and impact on diagnostics and…

Image and Video Processing · Electrical Eng. & Systems 2024-05-24 Gabriele Campanella , Eugene Fluder , Jennifer Zeng , Chad Vanderbilt , Thomas J. Fuchs

Driven by the recent advances in deep learning methods and, in particular, by the development of modern self-supervised learning algorithms, increased interest and efforts have been devoted to build foundation models (FMs) for medical…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 kaiko. ai , Nanne Aben , Edwin D. de Jong , Ioannis Gatopoulos , Nicolas Känzig , Mikhail Karasikov , Axel Lagré , Roman Moser , Joost van Doorn , Fei Tang

Vision-language models have emerged as a powerful tool for previously challenging multi-modal classification problem in the medical domain. This development has led to the exploration of automated image description generation for…

Computer Vision and Pattern Recognition · Computer Science 2024-06-03 Mansi Kakkar , Dattesh Shanbhag , Chandan Aladahalli , Gurunath Reddy M

Recent advancements in mixed-modal generative have opened new avenues for developing unified biomedical assistants capable of analyzing biomedical images, answering complex questions about them, and generating multimodal patient reports.…

Artificial Intelligence · Computer Science 2025-04-24 Hritik Bansal , Daniel Israel , Siyan Zhao , Shufan Li , Tung Nguyen , Aditya Grover

Retrieval-augmented generation (RAG) improves the response quality of large language models (LLMs) by retrieving knowledge from external databases. Typical RAG approaches split the text database into chunks, organizing them in a flat…

Computation and Language · Computer Science 2025-11-18 Boyu Chen , Zirui Guo , Zidan Yang , Yuluo Chen , Junze Chen , Zhenghao Liu , Chuan Shi , Cheng Yang

Automated pathology report generation from Whole Slide Images (WSIs) faces two key challenges: (1) lack of semantic content in visual features and (2) inherent information redundancy in WSIs. To address these issues, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ling Zhang , Boxiang Yun , Qingli Li , Yan Wang

Forensic pathology is critical in analyzing death manner and time from the microscopic aspect to assist in the establishment of reliable factual bases for criminal investigation. In practice, even the manual differentiation between…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Chen Shen , Jun Zhang , Xinggong Liang , Zeyi Hao , Kehan Li , Fan Wang , Zhenyuan Wang , Chunfeng Lian

Traditional biomedical artificial intelligence (AI) models, designed for specific tasks or modalities, often exhibit limited flexibility in real-world deployment and struggle to utilize holistic information. Generalist AI holds the…

Foundation models, first introduced in 2021, refer to large-scale pretrained models (e.g., large language models (LLMs) and vision-language models (VLMs)) that learn from extensive unlabeled datasets through unsupervised methods, enabling…