中文
相关论文

相关论文: Towards General Text-guided Image Synthesis for Cu…

200 篇论文

Recent advances in generative models for medical imaging have shown promise in representing multiple modalities. However, the variability in modality availability across datasets limits the general applicability of the synthetic data they…

图像与视频处理 · 电气工程与系统科学 2024-10-02 Sven Lüpke , Yousef Yeganeh , Ehsan Adeli , Nassir Navab , Azade Farshad

Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal processing. Inspired by…

计算与语言 · 计算机科学 2025-08-12 Chunyu Ye , Yunhao Zhang , Jingyuan Sun , Chong Li , Chengqing Zong , Shaonan Wang

We present Brain Harmony (BrainHarmonix), the first multimodal brain foundation model that unifies structural morphology and functional dynamics into compact 1D token representations. The model was pretrained on two of the largest…

Availability of large and diverse medical datasets is often challenged by privacy and data sharing restrictions. For successful application of machine learning techniques for disease diagnosis, prognosis, and precision medicine, large…

Image-to-image translation plays a vital role in tackling various medical imaging tasks such as attenuation correction, motion correction, undersampled reconstruction, and denoising. Generative adversarial networks have been shown to…

计算机视觉与模式识别 · 计算机科学 2021-07-05 Uddeshya Upadhyay , Yanbei Chen , Tobias Hepp , Sergios Gatidis , Zeynep Akata

Generative AI models hold great potential in creating synthetic brain MRIs that advance neuroimaging studies by, for example, enriching data diversity. However, the mainstay of AI research only focuses on optimizing the visual quality (such…

图像与视频处理 · 电气工程与系统科学 2023-10-10 Wei Peng , Tomas Bosschieter , Jiahong Ouyang , Robert Paul , Ehsan Adeli , Qingyu Zhao , Kilian M. Pohl

Brain imaging plays a crucial role in the diagnosis and treatment of various neurological disorders, providing valuable insights into the structure and function of the brain. Techniques such as magnetic resonance imaging (MRI) and computed…

图像与视频处理 · 电气工程与系统科学 2025-01-23 Fatima Haimour , Rizik Al-Sayyed , Waleed Mahafza , Omar S. Al-Kadi

The existing text-guided image synthesis methods can only produce limited quality results with at most \mbox{$\text{256}^2$} resolution and the textual instructions are constrained in a small Corpus. In this work, we propose a unified…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Weihao Xia , Yujiu Yang , Jing-Hao Xue , Baoyuan Wu

The human brain is a complex system requiring both macroscopic and microscopic components for comprehensive understanding. However, mapping nonlinear relationships between these scales remains challenging due to technical limitations and…

图像与视频处理 · 电气工程与系统科学 2025-10-28 Sooyoung Kim , Joonwoo Kwon , Junbeom Kwon , Jungyoun Janice Min , Sangyoon Bae , Yuewei Lin , Shinjae Yoo , Jiook Cha

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Hozaifa Kassab , Ahmed Mahmoud , Mohamed Bahaa , Ammar Mohamed , Ali Hamdi

Simulating prospective magnetic resonance imaging (MRI) scans from a given individual brain image is challenging, as it requires accounting for canonical changes in aging and/or disease progression while also considering the individual…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Jingru Fu , Yuqi Zheng , Neel Dey , Daniel Ferreira , Rodrigo Moreno

Benefiting from large-scale pre-trained text-to-image (T2I) generative models, impressive progress has been achieved in customized image generation, which aims to generate user-specified concepts. Existing approaches have extensively…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Ganggui Ding , Canyu Zhao , Wen Wang , Zhen Yang , Zide Liu , Hao Chen , Chunhua Shen

Insufficiency of training data is a persistent issue in medical image analysis, especially for task-based functional magnetic resonance images (fMRI) with spatio-temporal imaging data acquired using specific cognitive tasks. In this paper,…

图像与视频处理 · 电气工程与系统科学 2023-08-31 Jiyao Wang , Nicha C. Dvornek , Lawrence H. Staib , James S. Duncan

The work proposes a novel deep-learning framework for the synthesis of three-dimensional MRI volumes from corresponding 3D ultrasound images of the brain, leveraging a modified iteration of the Pix2Pix Generative Adversarial Network (GAN)…

图像与视频处理 · 电气工程与系统科学 2024-07-19 Shubham Singh , Mrunal Bewoor , Ammar Ranapurwala , Satyam Rai , Sheetal Patil

Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Petru-Daniel Tudosiu , Yongxin Yang , Shifeng Zhang , Fei Chen , Steven McDonagh , Gerasimos Lampouras , Ignacio Iacobacci , Sarah Parisot

LLMs have demonstrated remarkable capabilities in linguistic reasoning and are increasingly adept at vision-language tasks. The integration of image tokens into transformers has enabled direct visual input and output, advancing research…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Jonghun Kim , Sinyoung Ra , Hyunjin Park

In this paper, we focus on generating realistic images from text descriptions. Current methods first generate an initial image with rough shape and color, and then refine the initial image to a high-resolution one. Most existing…

计算机视觉与模式识别 · 计算机科学 2019-04-03 Minfeng Zhu , Pingbo Pan , Wei Chen , Yi Yang

Multi-modal medical images provide complementary soft-tissue characteristics that aid in the screening and diagnosis of diseases. However, limited scanning time, image corruption and various imaging protocols often result in incomplete…

计算机视觉与模式识别 · 计算机科学 2024-07-10 Yue Zhang , Chengtao Peng , Qiuli Wang , Dan Song , Kaiyan Li , S. Kevin Zhou

This study explores the use of text-prompted MRI image generation with the Stable Diffusion (SD) model to address challenges in acquiring real MRI datasets, such as high costs, limited rare case samples, and privacy concerns. The SD model,…

图像与视频处理 · 电气工程与系统科学 2025-05-30 Xinxian Fan , Mengye Lyu