English
Related papers

Related papers: SlideChat: A Large Vision-Language Assistant for W…

200 papers

Pathological captioning of Whole Slide Images (WSIs), though is essential in computer-aided pathological diagnosis, has rarely been studied due to the limitations in datasets and model training efficacy. In this paper, we propose a new…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Wenkang Qin , Rui Xu , Peixiang Huang , Xiaomin Wu , Heyu Zhang , Lin Luo

Whole slide pathology image classification presents challenges due to gigapixel image sizes and limited annotation labels, hindering model generalization. This paper introduces a prompt learning method to adapt large vision-language models…

Most of the existing multi-modal models, hindered by their incapacity to adeptly manage interleaved image-and-text inputs in multi-image, multi-round dialogues, face substantial constraints in resource allocation for training and data…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Zhewei Yao , Xiaoxia Wu , Conglong Li , Minjia Zhang , Heyang Qin , Olatunji Ruwase , Ammar Ahmad Awan , Samyam Rajbhandari , Yuxiong He

Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiveness in joint video and language understanding has not been…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Ruipu Luo , Ziwang Zhao , Min Yang , Zheming Yang , Minghui Qiu , Tao Wang , Zhongyu Wei , Yanhao Wang , Cen Chen

Large language models exhibit enhanced zero-shot performance on various tasks when fine-tuned with instruction-following data. Multimodal instruction-following models extend these capabilities by integrating both text and images. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Yupan Huang , Zaiqiao Meng , Fangyu Liu , Yixuan Su , Nigel Collier , Yutong Lu

Automating crash video analysis is essential to leverage the growing availability of driving video data for traffic safety research and accountability attribution in autonomous driving. Crash video analysis is a challenging multitask…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Kaidi Liang , Ke Li , Xianbiao Hu , Ruwen Qin

Multimodal Large Language Models (MLLM) have made significant progress in the field of document analysis. Despite this, existing benchmarks typically focus only on extracting text and simple layout information, neglecting the complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Lei Chen , Feng Yan , Yujie Zhong , Shaoxiang Chen , Zequn Jie , Lin Ma

Diagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology multi-modal models.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Mehmet Saygin Seyfioglu , Wisdom O. Ikezogwo , Fatemeh Ghezloo , Ranjay Krishna , Linda Shapiro

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Munan Ning , Bin Zhu , Yujia Xie , Bin Lin , Jiaxi Cui , Lu Yuan , Dongdong Chen , Li Yuan

Whole slide imaging is fundamental to biomedical microscopy and computational pathology. Previously, learning representations for gigapixel-sized whole slide images (WSIs) has relied on multiple instance learning with weak labels, which do…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Xinhai Hou , Cheng Jiang , Akhil Kondepudi , Yiwei Lyu , Asadur Chowdury , Honglak Lee , Todd C. Hollon

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Fanqing Meng , Jin Wang , Chuanhao Li , Quanfeng Lu , Hao Tian , Jiaqi Liao , Xizhou Zhu , Jifeng Dai , Yu Qiao , Ping Luo , Kaipeng Zhang , Wenqi Shao

Multimodal Large Language Models (MLLMs) have shown success in various general image processing tasks, yet their application in medical imaging is nascent, lacking tailored models. This study investigates the potential of MLLMs in improving…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Ling Yang , Zhanyu Wang , Zhenghao Chen , Xinyu Liang , Luping Zhou

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial…

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remains underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Xikai Yang , Juzheng Miao , Yuchen Yuan , Jiaze Wang , Qi Dou , Jinpeng Li , Pheng-Ann Heng

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Chunyuan Li , Cliff Wong , Sheng Zhang , Naoto Usuyama , Haotian Liu , Jianwei Yang , Tristan Naumann , Hoifung Poon , Jianfeng Gao

Large Language Models (LLMs) have shown immense potential in education, automating tasks like quiz generation and content summarization. However, generating effective presentation slides introduces unique challenges due to the complexity of…

Artificial Intelligence · Computer Science 2025-11-14 Eric Xie , Danielle Waterfield , Michael Kennedy , Aidong Zhang

Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant involuntary movements in neurological disorders remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Lina Zhang , Tonmoy Monsoor , Mehmet Efe Lorasdagi , Prateik Sinha , Chong Han , Peizheng Li , Yuan Wang , Jessica Pasqua , Colin McCrimmon , Rajarshi Mazumder , Vwani Roychowdhury

Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Jiangbo Shi , Chen Li , Tieliang Gong , Yefeng Zheng , Huazhu Fu

The rapid digitization of histopathology slides has opened up new possibilities for computational tools in clinical and research workflows. Among these, content-based slide retrieval stands out, enabling pathologists to identify…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hongyi Wang , Zhengjie Zhu , Jiabo Ma , Fang Wang , Yue Shi , Bo Luo , Jili Wang , Qiuyu Cai , Xiuming Zhang , Yen-Wei Chen , Lanfen Lin , Hao Chen

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Tianyu Yu , Jinyi Hu , Yuan Yao , Haoye Zhang , Yue Zhao , Chongyi Wang , Shan Wang , Yinxv Pan , Jiao Xue , Dahai Li , Zhiyuan Liu , Hai-Tao Zheng , Maosong Sun