English
Related papers

Related papers: I-ODA, Real-World Multi-modal Longitudinal Data fo…

200 papers

Eye-gaze tracking research offers significant promise in enhancing various healthcare-related tasks, above all in medical image analysis and interpretation. Eye tracking, a technology that monitors and records the movement of the eyes,…

Image and Video Processing · Electrical Eng. & Systems 2024-03-13 Sahar Moradizeyveh , Mehnaz Tabassum , Sidong Liu , Robert Ahadizad Newport , Amin Beheshti , Antonio Di Ieva

Healthcare data are inherently multimodal, including electronic health records (EHR), medical images, and multi-omics data. Combining these multimodal data sources contributes to a better understanding of human health and provides optimal…

Machine Learning · Computer Science 2022-10-28 Farida Mohsen , Hazrat Ali , Nady El Hajj , Zubair Shah

Artificial intelligence (AI) that can effectively learn ultrasound representations by integrating multi-source data holds significant promise for advancing clinical care. However, the scarcity of large labeled datasets in real-world…

Developing a generalist radiology diagnosis system can greatly enhance clinical diagnostics. In this paper, we introduce RadDiag, a foundational model supporting 2D and 3D inputs across various modalities and anatomies, using a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Qiaoyu Zheng , Weike Zhao , Chaoyi Wu , Xiaoman Zhang , Lisong Dai , Hengyu Guan , Yuehua Li , Ya Zhang , Yanfeng Wang , Weidi Xie

Vision-language-action models (VLAs) have become an increasingly popular approach for addressing robot manipulation problems in recent years. However, such models need to output actions at a rate suitable for robot control, which limits the…

Robotics · Computer Science 2025-09-30 Eric Hannus , Miika Malin , Tran Nguyen Le , Ville Kyrki

Stroke remains a leading cause of global morbidity and mortality, imposing a heavy socioeconomic burden. Advances in endovascular reperfusion therapy and CT and MR imaging for treatment guidance have significantly improved patient outcomes.…

Despite the revolutionary impact of AI and the development of locally trained algorithms, achieving widespread generalized learning from multi-modal data in medical AI remains a significant challenge. This gap hinders the practical…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Fatema-E Jannat , Sina Gholami , Minhaj Nur Alam , Hamed Tabkhi

The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modalities, including 2D/3D images and videos, yet existing models typically employ separate…

Computation and Language · Computer Science 2025-04-22 Songtao Jiang , Yuan Wang , Sibo Song , Yan Zhang , Zijie Meng , Bohan Lei , Jian Wu , Jimeng Sun , Zuozhu Liu

Long-tailed learning is considered to be an extremely challenging problem in data imbalance learning. It aims to train well-generalized models from a large number of images that follow a long-tailed class distribution. In the medical field,…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Lie Ju , Siyuan Yan , Yukun Zhou , Yang Nan , Xiaodan Xing , Peibo Duan , Zongyuan Ge

In this paper, we propose a framework that incorporates experts diagnostics and insights into the analysis of Optical Coherence Tomography (OCT) using multi-modal learning. To demonstrate the effectiveness of this approach, we create a…

Image and Video Processing · Electrical Eng. & Systems 2022-03-22 Y. Logan , K. Kokilepersaud , G. Kwon , G. AlRegib , C. Wykoff , H. Yu

Medical Large Vision-Language Models (Med-LVLMs) demonstrate significant potential in healthcare, but their reliance on general medical data and coarse-grained global visual understanding limits them in intelligent ophthalmic diagnosis.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Sijing Li , Tianwei Lin , Lingshuai Lin , Wenqiao Zhang , Jiang Liu , Xiaoda Yang , Juncheng Li , Yucheng He , Xiaohui Song , Jun Xiao , Yueting Zhuang , Beng Chin Ooi

Animal re-identification (ReID) faces critical challenges due to viewpoint variations, particularly in Aerial-Ground (AG-ReID) settings where models must match individuals across drastic elevation changes. However, existing datasets lack…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 William Grolleau , Achraf Chaouch , Astrid Sabourin , Guillaume Lapouge , Catherine Achard

Dermatological care via telemedicine often lacks the rich context of in-person visits. Clinicians must make diagnoses based on a handful of images and brief descriptions, without the benefit of physical exams, second opinions, or reference…

Artificial Intelligence · Computer Science 2025-08-27 Karishma Thakrar , Shreyas Basavatia , Akshay Daftardar

The fast growing application of omnidirectional images calls for effective approaches for omnidirectional image quality assessment (OIQA). Existing OIQA methods have been developed and tested on homogeneously distorted omnidirectional…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Jiebin Yan , Ziwen Tan , Yuming Fang , Junjie Chen , Wenhui Jiang , Zhou Wang

Early ophthalmic screening in low-resource and remote settings is constrained by access to specialized equipment and trained practitioners. We present SKINOPATHY AI, a smartphone-first web application that delivers five complementary,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 S. Kalaycioglu , C. Hong , M. Zhu , H. Xie

Diagnosing and managing oral diseases necessitate advanced visual interpretation across diverse imaging modalities and integrated information synthesis. While current AI models excel at isolated tasks, they often fall short in addressing…

The scarcity and complexity of voxel-level annotations in 3D medical imaging present significant challenges, particularly due to the domain gap between labeled datasets from well-resourced centers and unlabeled datasets from less-resourced…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Haifan Gong , Yitao Wang , Yihan Wang , Jiashun Xiao , Xiang Wan , Haofeng Li

Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are significantly more…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Maya Varma , Jean-Benoit Delbrouck , Sarah Hooper , Akshay Chaudhari , Curtis Langlotz

Person re-identification (re-id) is a pivotal task within an intelligent surveillance pipeline and there exist numerous re-id frameworks that achieve satisfactory performance in challenging benchmarks. However, these systems struggle to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-17 Tharindu Fernando , Clinton Fookes , Sridha Sridharan , Dana Michalski

We endeavor on a rarely explored task named Insubstantial Object Detection (IOD), which aims to localize the object with following characteristics: (1) amorphous shape with indistinct boundary; (2) similarity to surroundings; (3) absence in…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Kailai Zhou , Yibo Wang , Tao Lv , Yunqian Li , Linsen Chen , Qiu Shen , Xun Cao