English
Related papers

Related papers: I-ODA, Real-World Multi-modal Longitudinal Data fo…

200 papers

Autonomous driving is a popular research area within the computer vision research community. Since autonomous vehicles are highly safety-critical, ensuring robustness is essential for real-world deployment. While several public multimodal…

Open Learning Analytics (OLA) is an emerging research area that aims at improving learning efficiency and effectiveness in lifelong learning environments. OLA employs multiple methods to draw value from a wide range of educational data…

Computers and Society · Computer Science 2023-03-23 Arham Muslim , Mohamed Amine Chatti , Mouadh Guesmi

We envision the "virtual eye" as a next-generation, AI-powered platform that uses interconnected foundation models to simulate the eye's intricate structure and biological function across all scales. Advances in AI, imaging, and multiomics…

Tissues and Organs · Quantitative Biology 2025-05-12 Yue Wu , Yibo Guo , Yulong Yan , Jiancheng Yang , Xin Zhou , Ching-Yu Cheng , Danli Shi , Mingguang He

In recent years, pre-trained large language models (LLMs) have achieved tremendous success in the field of Natural Language Processing (NLP). Prior studies have primarily focused on general and generic domains, with relatively less research…

Eye diseases are common in older Americans and can lead to decreased vision and blindness. Recent advancements in imaging technologies allow clinicians to capture high-quality images of the retinal blood vessels via Optical Coherence…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Jacob Thrasher , Annahita Amireskandari , Prashnna Gyawali

Learning bimanual manipulation is challenging due to its high dimensionality and tight coordination required between two arms. Eye-in-hand imitation learning, which uses wrist-mounted cameras, simplifies perception by focusing on…

Robotics · Computer Science 2025-08-19 I-Chun Arthur Liu , Jason Chen , Gaurav Sukhatme , Daniel Seita

We introduce a novel Image Quality Assessment (IQA) dataset comprising 6073 UHD-1 (4K) images, annotated at a fixed width of 3840 pixels. Contrary to existing No-Reference (NR) IQA datasets, ours focuses on highly aesthetic photos of high…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Vlad Hosu , Lorenzo Agnolucci , Oliver Wiedemann , Daisuke Iso , Dietmar Saupe

Optical coherence tomography (OCT) is a non-invasive imaging modality which is widely used in clinical ophthalmology. OCT images are capable of visualizing deep retinal layers which is crucial for early diagnosis of retinal diseases. In…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Peyman Gholami , Priyanka Roy , Mohana Kuppuswamy Parthasarathy , Vasudevan Lakshminarayanan

AI is the workhorse of modern data analytics and omnipresent across many sectors. Large Language Models and multi-modal foundation models are today capable of generating code, charts, visualizations, etc. How will these massive developments…

Human-Computer Interaction · Computer Science 2025-09-08 Katja Bühler , Thomas Höllt , Thomas Schulz , Pere-Pau Vázquez

Multimodal intent understanding is a significant research area that requires effective leveraging of multiple modalities to analyze human language. Existing methods face two main challenges in this domain. Firstly, they have limitations in…

Multimedia · Computer Science 2025-05-26 Hanlei Zhang , Qianrui Zhou , Hua Xu , Jianhua Su , Roberto Evans , Kai Gao

Ocular disease affects billions of individuals unevenly worldwide. It continues to increase in prevalence with trends of growing populations of diabetic people, increasing life expectancies, decreasing ophthalmologist availability, and…

Computers and Society · Computer Science 2025-07-29 Shiv Garg , Ginny Berkemeier

Person identification in the wild is very challenging due to great variation in poses, face quality, clothes, makeup and so on. Traditional research, such as face recognition, person re-identification, and speaker recognition, often focuses…

Computer Vision and Pattern Recognition · Computer Science 2019-04-23 Yuanliu Liu , Bo Peng , Peipei Shi , He Yan , Yong Zhou , Bing Han , Yi Zheng , Chao Lin , Jianbin Jiang , Yin Fan , Tingwei Gao , Ganwen Wang , Jian Liu , Xiangju Lu , Danming Xie

This study introduces an evaluation framework for multimodal models in medical imaging diagnostics. We developed a pipeline incorporating data preprocessing, model inference, and preference-based evaluation, expanding an initial set of 500…

Image and Video Processing · Electrical Eng. & Systems 2024-12-10 Cailian Ruan , Chengyue Huang , Yahe Yang

The astounding success made by artificial intelligence (AI) in healthcare and other fields proves that AI can achieve human-like performance. However, success always comes with challenges. Deep learning algorithms are data-dependent and…

Image and Video Processing · Electrical Eng. & Systems 2021-06-25 Johann Li , Guangming Zhu , Cong Hua , Mingtao Feng , BasheerBennamoun , Ping Li , Xiaoyuan Lu , Juan Song , Peiyi Shen , Xu Xu , Lin Mei , Liang Zhang , Syed Afaq Ali Shah , Mohammed Bennamoun

With advanced imaging, sequencing, and profiling technologies, multiple omics data become increasingly available and hold promises for many healthcare applications such as cancer diagnosis and treatment. Multimodal learning for integrative…

Genomics · Quantitative Biology 2022-12-20 Sina Tabakhi , Mohammod Naimul Islam Suvon , Pegah Ahadian , Haiping Lu

Vision impairment affects millions globally, and early detection is critical to preventing irreversible vision loss. Ophthalmology workflows require clinicians to integrate medical images, structured clinical data, and free-text notes to…

Object compositing, the task of placing and harmonizing objects in images of diverse visual scenes, has become an important task in computer vision with the rise of generative models. However, existing datasets lack the diversity and scale…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Jinwoo Kim , Sangmin Han , Jinho Jeong , Jiwoo Choi , Dongyoung Kim , Seon Joo Kim

Omnidirectional image (ODI) data is captured with a 360x180 field-of-view, which is much wider than the pinhole cameras and contains richer spatial information than the conventional planar images. Accordingly, omnidirectional vision has…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Hao Ai , Zidong Cao , Jinjing Zhu , Haotian Bai , Yucheng Chen , Lin Wang

Recent work has demonstrated that imaging systems can be evaluated through the information content of their measurements alone, enabling application-agnostic optical design that avoids computational decoding challenges. Information-Driven…

Image and Video Processing · Electrical Eng. & Systems 2025-07-14 Eric Markley , Henry Pinkard , Leyla Kabuli , Nalini Singh , Laura Waller