English
Related papers

Related papers: RetFiner: A Vision-Language Refinement Scheme for …

200 papers

The advent of foundation models (FMs) is transforming medical domain. In ophthalmology, RETFound, a retina-specific FM pre-trained sequentially on 1.4 million natural images and 1.6 million retinal images, has demonstrated high adaptability…

In medical science, the use of computer science in disease detection and diagnosis is gaining popularity. Previously, the detection of disease used to take a significant amount of time and was less reliable. Machine learning (ML) techniques…

Image and Video Processing · Electrical Eng. & Systems 2019-12-18 Nowshin Tasnim , Mahmudul Hasan , Ishrak Islam

In this work, we present a simple yet theoretically motivated improvement to Supervised Fine-Tuning (SFT) for the Large Language Model (LLM), addressing its limited generalization compared to reinforcement learning (RL). Through…

Machine Learning · Computer Science 2026-03-02 Yongliang Wu , Yizhou Zhou , Zhou Ziheng , Yingzhe Peng , Xinyu Ye , Xinting Hu , Wenbo Zhu , Lu Qi , Ming-Hsuan Yang , Xu Yang

Supervised fine-tuning (SFT) is a standard approach to adapting large language models (LLMs) to new domains. In this work, we improve the statistical efficiency of SFT by selecting an informative subset of training examples. Specifically,…

Machine Learning · Computer Science 2025-05-22 Rohan Deb , Kiran Thekumparampil , Kousha Kalantari , Gaurush Hiranandani , Shoham Sabach , Branislav Kveton

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language models (MLLMs) integrate LLMs with visual models to process…

Computation and Language · Computer Science 2025-03-11 Xiaoyi Liang , Mouxiao Bian , Moxin Chen , Lihao Liu , Junjun He , Jie Xu , Lin Li

Optical coherence tomography (OCT) is a prevalent imaging technique for retina. However, it is affected by multiplicative speckle noise that can degrade the visibility of essential anatomical structures, including blood vessels and tissue…

Image and Video Processing · Electrical Eng. & Systems 2021-07-12 Dewei Hu , Joseph D. Malone , Yigit Atay , Yuankai K. Tao , Ipek Oguz

The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. Existing methods for synthesising Colour…

Image and Video Processing · Electrical Eng. & Systems 2025-07-18 Junzhi Ning , Cheng Tang , Kaijing Zhou , Diping Song , Lihao Liu , Ming Hu , Wei Li , Huihui Xu , Yanzhou Su , Tianbin Li , Jiyao Liu , Jin Ye , Sheng Zhang , Yuanfeng Ji , Junjun He

Optical Coherence Tomography Angiography (OCTA) and its derived en-face projections provide high-resolution visualization of the retinal and choroidal vasculature, which is critical for the rapid and accurate diagnosis of retinal diseases.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Pooya Khosravi , Kun Han , Anthony T. Wu , Arghavan Rezvani , Zexin Feng , Xiaohui Xie

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is…

Computation and Language · Computer Science 2024-06-14 Amit Meghanani , Thomas Hain

Vision-Language Foundation Models (VLFM) have shown a tremendous increase in performance in terms of generating high-resolution, photorealistic natural images. While VLFMs show a rich understanding of semantic content across modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Parham Saremi , Amar Kumar , Mohamed Mohamed , Zahra TehraniNasab , Tal Arbel

Semi-supervised learning (SSL) has witnessed remarkable progress, resulting in the emergence of numerous method variations. However, practitioners often encounter challenges when attempting to deploy these methods due to their subpar…

Machine Learning · Computer Science 2024-05-21 Kai Gan , Tong Wei

According to the World Health Organization, 285 million people worldwide live with visual impairment. The most commonly used imaging technique for diagnosis in ophthalmology is optical coherence tomography (OCT). However, analysis of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Max-Heinrich Laves , Sontje Ihler , Lüder A. Kahrs , Tobias Ortmaier

Most prior deepfake detection methods lack explainable outputs. With the growing interest in multimodal large language models (MLLMs), researchers have started exploring their use in interpretable deepfake detection. However, a major…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Ning Jiang , Dingheng Zeng , Yanhong Liu , Haiyang Yi , Shijie Yu , Minghe Weng , Haifeng Shen , Ying Li

Foundation deep learning (DL) models are general models, designed to learn general, robust and adaptable representations of their target modality, enabling finetuning across a range of downstream tasks. These models are pretrained on large,…

Signal Processing · Electrical Eng. & Systems 2024-11-18 Ahmed Aboulfotouh , Ashkan Eshaghbeigi , Hatem Abou-Zeid

Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limited interpretability,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Samuel Ofosu Mensah , Camila Roa , Kerol Djoumessi , Philipp Berens

Optical Coherence Tomography (OCT) is essential for diagnosing conditions such as glaucoma, diabetic retinopathy, and age-related macular degeneration. Accurate retinal layer segmentation enables quantitative biomarkers critical for…

Image and Video Processing · Electrical Eng. & Systems 2025-09-10 S M Asiful Islam Saky , Ugyen Tshering

Self-supervised Learning (SSL) has been widely applied to learn image representations through exploiting unlabeled images. However, it has not been fully explored in the medical image analysis field. In this work, Saliency-guided…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Yijin Huang , Junyan Lyu , Pujin Cheng , Roger Tam , Xiaoying Tang

Optical coherence tomography (OCT) is a non-invasive 3D modality widely used in ophthalmology for imaging the retina. Achieving automated, anatomically coherent retinal layer segmentation on OCT is important for the detection and monitoring…

Image and Video Processing · Electrical Eng. & Systems 2022-10-26 Botond Fazekas , Guilherme Aresta , Dmitrii Lachinov , Sophie Riedl , Julia Mai , Ursula Schmidt-Erfurth , Hrvoje Bogunovic

Self-Supervised Learning (SSL) Automatic Speech Recognition (ASR) models have shown great promise over Supervised Learning (SL) ones in low-resource settings. However, the advantages of SSL are gradually weakened when the amount of labeled…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Li Fu , Siqi Li , Qingtao Li , Fangzhu Li , Liping Deng , Lu Fan , Meng Chen , Youzheng Wu , Xiaodong He

30 million Optical Coherence Tomography (OCT) imaging tests are issued annually to diagnose various retinal diseases, but accurate diagnosis of OCT scans requires trained eye care professionals who are still prone to making errors. With…

Image and Video Processing · Electrical Eng. & Systems 2022-10-04 Evan Wen , Rebecca Sorenson , Max Ehrlich