English
Related papers

Related papers: sleep2vec: Unified Cross-Modal Alignment for Heter…

200 papers

Human in-bed pose estimation has huge practical values in medical and healthcare applications yet still mainly relies on expensive pressure mapping (PM) solutions. In this paper, we introduce our novel physics inspired vision-based approach…

Computer Vision and Pattern Recognition · Computer Science 2019-09-23 Shuangjun Liu , Sarah Ostadabbas

Modeling effective representations using multiple views that positively influence each other is challenging, and the existing methods perform poorly on Electroencephalogram (EEG) signals for sleep-staging tasks. In this paper, we propose a…

Background and objective: Expert annotations limit large-scale supervised pretraining in medical imaging, while ubiquitous metadata (modality, anatomical region) remain underused. We introduce ModAn-MulSupCon, a modality- and anatomy-aware…

Image and Video Processing · Electrical Eng. & Systems 2025-08-27 Eichi Takaya , Ryusei Inamori

Large language models have recently shown promise for multimodal recommendation, particularly with text and image inputs. Yet real-world recommendation signals extend far beyond these modalities. To reflect this, we formalize recommendation…

Information Retrieval · Computer Science 2026-05-01 Zijie Lei , Tao Feng , Zhigang Hua , Yan Xie , Guanyu Lin , Shuang Yang , Ge Liu , Jiaxuan You

Simultaneous electrocardiography (ECG) and phonocardiogram (PCG) provide a comprehensive, multimodal perspective on cardiac function by capturing the heart's electrical and mechanical activities, respectively. However, the distinct and…

Machine Learning · Computer Science 2025-06-13 Sajjad Karimi , Amit J. Shah , Gari D. Clifford , Reza Sameni

Cross-modal alignment aims to map heterogeneous modalities into a shared latent space, as exemplified by models like CLIP, which benefit from large-scale image-text pretraining for strong recognition capabilities. However, when operating in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Jiaxiang Liu , Yuan Wang , Jiawei Du , Joey Tianyi Zhou , Mingkun Xu , Zuozhu Liu

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, text, audio, action,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiasen Lu , Christopher Clark , Sangho Lee , Zichen Zhang , Savya Khosla , Ryan Marten , Derek Hoiem , Aniruddha Kembhavi

Subtle visual signals, although difficult to perceive with the naked eye, contain important information that can reveal hidden patterns in visual data. These signals play a key role in many applications, including biometric security,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Dongliang Zhu , Zhiyi Niu , Bo Zhao , Jiajian Huang , Shuo Ye , Xun Lin , Hui Ma , Taorui Wang , Jiayu Zhang , Chunmei Zhu , Junzhe Cao , Yingjie Ma , Rencheng Song , Albert Clapés , Sergio Escalera , Dan Guo , Zitong Yu

A multi-modal machine learning system uses multiple unique data sources and types to improve its performance. This article proposes a system that combines results from several types of models, all of which are trained on different data…

Machine Learning · Computer Science 2024-02-05 Aaron Mullen , Samuel E. Armstrong , Jasmine Perdeh , Bjorn Bauer , Jeffrey Talbert , V. K. Cody Bumgardner

We present a novel multi-modal bio-sensing platform capable of integrating multiple data streams for use in real-time applications. The system is composed of a central compute module and a companion headset. The compute node collects,…

Human-Computer Interaction · Computer Science 2018-02-23 Siddharth , Aashish Patel , Tzyy-Ping Jung , Terrence J. Sejnowski

Representation learning provides new and powerful graph analytical approaches and tools for the highly valued data science challenge of mining knowledge graphs. Since previous graph analytical methods have mostly focused on homogeneous…

Information Retrieval · Computer Science 2019-05-29 Zheng Gao , Gang Fu , Chunping Ouyang , Satoshi Tsutsui , Xiaozhong Liu , Jeremy Yang , Christopher Gessner , Brian Foote , David Wild , Qi Yu , Ying Ding

Understanding sleep and activity patterns plays a crucial role in physical and mental health. This study introduces a novel approach for sleep detection using weakly supervised learning for scenarios where reliable ground truth labels are…

Machine Learning · Computer Science 2024-07-09 Matthias Boeker , Vajira Thambawita , Michael Riegler , Pål Halvorsen , Hugo L. Hammer

Objective: To develop and validate an automated method for bedside monitoring of sleep state fluctuations in neonatal intensive care units. Methods: A deep learning -based algorithm was designed and trained using 53 EEG recordings from a…

Signal Processing · Electrical Eng. & Systems 2022-08-26 Saeed Montazeri Moghadam , Päivi Nevalainen , Nathan J. Stevenson , Sampsa Vanhatalo

Multimodal alignment of histopathology encoders with transcriptomic and genomic data has been shown to significantly improve performance in downstream diagnostic tasks. Hematological cytology is unique in that visual single-cell evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Muhammed Furkan Dasdelen , Fatih Ozlugedik , Ilaria Looser , Rao Muhammad Umer , Christian Pohlkamp , Carsten Marr

In clinical practice, imaging modalities with functional characteristics, such as positron emission tomography (PET) and fractional anisotropy (FA), are often aligned with a structural reference (e.g., MRI, CT) for accurate interpretation…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Kyobin Choo , Hyunkyung Han , Jinyeong Kim , Chanyong Yoon , Seong Jae Hwang

Managing fluid balance in dialysis patients is crucial, as improper management can lead to severe complications. In this paper, we propose a multimodal approach that integrates visual features from lung ultrasound images with clinical data…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Tianqi Yang , Nantheera Anantrasirichai , Oktay Karakuş , Marco Allinovi , Alin Achim

Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Frank Li , Bardia Khosravi , Mohammadreza Chavoshi , Young Seok Jeon , Theo Dapamede , Hari Trivedi , Janice Newsome , Judy Gichoya

High annotation costs are a substantial bottleneck in applying modern deep learning architectures to clinically relevant medical use cases, substantiating the need for novel algorithms to learn from unlabeled data. In this work, we propose…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Aiham Taleb , Matthias Kirchler , Remo Monti , Christoph Lippert

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiao Xu , Junwei Liu , Jiangwei Lao , Qi Zhu , Yunpeng Zhao , Congyun Jin , Shinan Liu , Zhihong Lu , Lihe Zhang , Xin Chen , Jian Wang , Ping Wang