English
Related papers

Related papers: Compact Multimodal Language Models as Robust OCR A…

200 papers

Surgical procedures unfold in complex environments demanding coordination between surgical teams, tools, imaging and increasingly, intelligent robotic systems. Ensuring safety and efficiency in ORs of the future requires intelligent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Ege Özsoy , Chantal Pellegrini , David Bani-Harouni , Kun Yuan , Matthias Keicher , Nassir Navab

Radiology reports are invaluable for clinical decision-making and hold great potential for automated analysis when structured into machine-readable formats. These reports often contain uncertainty, which we categorize into two distinct…

Computation and Language · Computer Science 2026-03-02 Paloma Rabaey , Jong Hak Moon , Jung-Oh Lee , Min Gwan Kim , Hangyul Yoon , Thomas Demeester , Edward Choi

Contrastive language-image Pre-training (CLIP) [13] can leverage large datasets of unlabeled Image-Text pairs, which have demonstrated impressive performance in various downstream tasks. Given that annotating medical data is time-consuming…

Image and Video Processing · Electrical Eng. & Systems 2023-07-13 Yuhao Wang

Generating radiology reports automatically reduces the workload of radiologists and helps the diagnoses of specific diseases. Many existing methods take this task as modality transfer process. However, since the key information related to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yitian Tao , Liyan Ma , Jing Yu , Han Zhang

Computed tomography (CT) is extensively used for accurate visualization and segmentation of organs and lesions. While deep learning models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have significantly…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Yuheng Li , Yuxiang Lai , Maria Thor , Deborah Marshall , Zachary Buchwald , David S. Yu , Xiaofeng Yang

While recent advancements in Image Super-Resolution (SR) using diffusion models have shown promise in improving overall image quality, their application to scene text images has revealed limitations. These models often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Keren Ye , Ignacio Garcia Dorado , Michalis Raptis , Mauricio Delbracio , Irene Zhu , Peyman Milanfar , Hossein Talebi

Recent research advances achieve human-level accuracy for de-identifying free-text clinical notes on research datasets, but gaps remain in reproducing this in large real-world settings. This paper summarizes lessons learned from building a…

Computation and Language · Computer Science 2023-12-15 Veysel Kocaman , Hasham Ul Haq , David Talby

An important topic in medical research is the process of improving the images obtained from medical devices. As a consequence, there is also a need to improve medical image resolution and analysis. Another issue in this field is the large…

Image and Video Processing · Electrical Eng. & Systems 2023-05-26 Elena-Simona Apostol , Ciprian-Octavian Truică

Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

The analysis of physiological time series, such as electrocardiograms (ECG) and photoplethysmograms (PPG), is persistently hindered by modality and frequency gaps stemming from heterogeneous recording devices. Existing foundation models…

Signal Processing · Electrical Eng. & Systems 2026-05-14 Bo Cui , Xiaowen Song , Yaowen Zhang , Shunzhe Zhang , B. J. F. van Beijnum , Monique Tabak , Ying Wang

Vision-language models pre-trained on large scale of unlabeled biomedical images and associated reports learn generalizable semantic representations. These multi-modal representations can benefit various downstream tasks in the biomedical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Xinliu Zhong , Kayhan Batmanghelich , Li Sun

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Thomas Buckley , James A. Diao , Pranav Rajpurkar , Adam Rodman , Arjun K. Manrai

Radiology reports are critical for clinical decision-making but often lack a standardized format, limiting both human interpretability and machine learning (ML) applications. While large language models (LLMs) have shown strong capabilities…

Computation and Language · Computer Science 2025-07-15 Johannes Moll , Louisa Fay , Asfandyar Azhar , Sophie Ostmeier , Tim Lueth , Sergios Gatidis , Curtis Langlotz , Jean-Benoit Delbrouck

3D medical image analysis is of great importance in disease diagnosis and treatment. Recently, multimodal large language models (MLLMs) have exhibited robust perceptual capacity, strong cross-modal alignment, and promising generalizability.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yang Yu , Dunyuan Xu , Yaoqian Li , Xiaomeng Li , Jinpeng Li , Pheng-Ann Heng

Recent advancements in Large Multimodal Models (LMMs) have attracted interest in their generalization capability with only a few samples in the prompt. This progress is particularly relevant to the medical domain, where the quality and…

Computation and Language · Computer Science 2024-05-06 Seonhee Cho , Choonghan Kim , Jiho Lee , Chetan Chilkunda , Sujin Choi , Joo Heung Yoon

Healthcare data now span EHRs, medical imaging, genomics, and wearable sensors, but most diagnostic models still process these modalities in isolation. This limits their ability to capture early, cross-modal disease signatures. This paper…

Machine Learning · Computer Science 2025-12-18 Md Talha Mohsin , Ismail Abdulrashid

Vulnerability to lexical perturbation is a critical weakness of automatic evaluation metrics for image captioning. This paper proposes Perturbation Robust Multi-Lingual CLIPScore(PR-MCS), which exhibits robustness to such perturbations, as…

Computation and Language · Computer Science 2023-03-16 Yongil Kim , Yerin Hwang , Hyeongu Yun , Seunghyun Yoon , Trung Bui , Kyomin Jung

This paper presents a comparative analysis of Large Language Models (LLMs) and traditional Optical Character Recognition (OCR) systems on Urdu newspapers, addressing challenges posed by complex multi-column layouts, low-resolution scans,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Samee Arif , Sualeha Farid

We perform a comprehensive benchmarking of contrastive frameworks for learning multimodal representations in the medical domain. Through this study, we aim to answer the following research questions: (i) How transferable are general-domain…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Shuvendu Roy , Yasaman Parhizkar , Franklin Ogidi , Vahid Reza Khazaie , Michael Colacci , Ali Etemad , Elham Dolatabadi , Arash Afkanpour

The digitisation of historical print media archives is crucial for increasing accessibility to contemporary records. However, the process of Optical Character Recognition (OCR) used to convert physical records to digital text is prone to…

Computation and Language · Computer Science 2025-01-23 Jonathan Bourne