English
Related papers

Related papers: XxaCT-NN: Structure Agnostic Multimodal Learning f…

200 papers

Although diffusion models have achieved remarkable progress in multi-modal magnetic resonance imaging (MRI) translation tasks, existing methods still tend to suffer from anatomical inconsistencies or degraded texture details when handling…

Image and Video Processing · Electrical Eng. & Systems 2026-03-16 Jianqiang Lin , Zhiqiang Shen , Peng Cao , Jinzhu Yang , Osmar R. Zaiane , Xiaoli Liu

Real-world multimodal data usually exhibit complex structural relationships beyond traditional one-to-one mappings like image-caption pairs. Entities across modalities interact in intricate ways, with images and text forming diverse…

Machine Learning · Computer Science 2025-10-21 Xuying Ning , Dongqi Fu , Tianxin Wei , Wujiang Xu , Jingrui He

Image-text multimodal representation learning aligns data across modalities and enables important medical applications, e.g., image classification, visual grounding, and cross-modal retrieval. In this work, we establish a connection between…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Peiqi Wang , William M. Wells , Seth Berkowitz , Steven Horng , Polina Golland

We introduce MxDiffusion, a hybrid physics- and data-driven diffusion-based framework that enables efficient and highly accurate generation of photonic structures from target optical properties. The improved accuracy is achieved through a…

Optics · Physics 2026-02-20 Sujoy Mondal , Taehyuk Park , Sudipta Biswas , Alan X. Wang , Wenshan Cai

Many healthcare applications are inherently multimodal, involving several physiological signals. As sensors for these signals become more common, improving machine learning methods for multimodal healthcare data is crucial. Pretraining…

Machine Learning · Computer Science 2024-10-23 Ching Fang , Christopher Sandino , Behrooz Mahasseni , Juri Minxha , Hadi Pouransari , Erdrin Azemi , Ali Moin , Ellen Zippi

Recent advances in deep learning have enabled the generation of realistic data by training generative models on large datasets of text, images, and audio. While these models have demonstrated exceptional performance in generating novel and…

Materials Science · Physics 2024-06-17 Izumi Takahara , Kiyou Shibata , Teruyasu Mizoguchi

Recent advances in molecular representation integrates molecular topological and visual modalities, opening new avenues for precise Molecular Relational Learning (MRL). Existing MRL methods focus on intra-domain modeling, and their inherent…

Machine Learning · Computer Science 2026-05-25 Peiliang Zhang , Jingling Yuan , Shiqing Wu , Mengqing Hu , Chao Che , Yongjun Zhu , Lin Li

Radiology report generation (RRG) aims to describe automatically a radiology image with human-like language and could potentially support the work of radiologists, reducing the burden of manual reporting. Previous approaches often adopt an…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Jun Wang , Abhir Bhalerao , Yulan He

With the increasing attention to pre-trained vision-language models (VLMs), \eg, CLIP, substantial efforts have been devoted to many downstream tasks, especially in test-time adaptation (TTA). However, previous works focus on learning…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Xingyu Zhu , Shuo Wang , Beier Zhu , Miaoge Li , Yunfan Li , Junfeng Fang , Zhicai Wang , Dongsheng Wang , Hanwang Zhang

Modeling the atomic structure of amorphous materials has long been a critical challenge in materials science. Recent advances in monolayer amorphous materials enable direct observation of their atomic structures, paving the way for a better…

Materials Science · Physics 2026-05-05 Le-Ye Zhu , Xi Zhang , Yun-Peng Wang , Jieheng Shi , Junwei Zhang , Shixuan Du , Yu-Yang Zhang

In the field of chemical structure recognition, the task of converting molecular images into machine-readable data formats such as SMILES string stands as a significant challenge, primarily due to the varied drawing styles and conventions…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Yufan Chen , Ching Ting Leung , Yong Huang , Jianwei Sun , Hao Chen , Hanyu Gao

Traditional analysis of highly distorted micro-X-ray diffraction ({\mu}-XRD) patterns from hydrothermal fluid environments is a time-consuming process, often requiring substantial data preprocessing and labeled experimental data. This study…

Materials Science · Physics 2024-03-18 Yanfei Li , Juejing Liu , Xiaodong Zhao , Wenjun Liu , Tong Geng , Ang Li , Xin Zhang

Accurate crystal structure determination is critical across all scientific disciplines involving crystalline materials. However, solving and refining inorganic crystal structures from powder X-ray diffraction (PXRD) data is traditionally a…

Materials Science · Physics 2024-09-10 Qi Li , Rui Jiao , Liming Wu , Tiannian Zhu , Wenbing Huang , Shifeng Jin , Yang Liu , Hongming Weng , Xiaolong Chen

Machine learning (ML) is becoming increasingly popular for predicting material properties to accelerate materials discovery. Because material properties are strongly affected by its crystal structure, a key issue is converting the crystal…

Materials Science · Physics 2023-10-12 Hirofumi Tsuruta , Yukari Katsura , Masaya Kumagai

Building generalizable medical AI systems requires pretraining strategies that are data-efficient and domain-aware. Unlike internet-scale corpora, clinical datasets such as MIMIC-CXR offer limited image counts and scarce annotations, but…

The development of high-performance materials for microelectronics, energy storage, and extreme environments depends on our ability to describe and direct property-defining microstructural order. Our present understanding is typically…

We present a simplified, task-agnostic multi-modal pre-training approach that can accept either video or text input, or both for a variety of end tasks. Existing pre-training are task-specific by adopting either a single cross-modal encoder…

Computer Vision and Pattern Recognition · Computer Science 2021-10-04 Hu Xu , Gargi Ghosh , Po-Yao Huang , Prahal Arora , Masoumeh Aminzadeh , Christoph Feichtenhofer , Florian Metze , Luke Zettlemoyer

Detecting structures at the particle scale within plastically deformed crystalline materials allows a better understanding of the occurring phenomena. While previous approaches mostly relied on applying hand-chosen criteria on different…

Materials Science · Physics 2024-05-15 Armand Barbot , Riccardo Gatti

Multimodal pre-training with text, layout, and image has achieved SOTA performance for visually-rich document understanding tasks recently, which demonstrates the great potential for joint learning across different modalities. In this…

Computation and Language · Computer Science 2021-09-10 Yiheng Xu , Tengchao Lv , Lei Cui , Guoxin Wang , Yijuan Lu , Dinei Florencio , Cha Zhang , Furu Wei

Recent research has achieved significant advancements in visual reasoning tasks through learning image-to-language projections and leveraging the impressive reasoning abilities of Large Language Models (LLMs). This paper introduces an…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Artemis Panagopoulou , Le Xue , Ning Yu , Junnan Li , Dongxu Li , Shafiq Joty , Ran Xu , Silvio Savarese , Caiming Xiong , Juan Carlos Niebles