中文
相关论文

相关论文: Towards General Text-guided Image Synthesis for Cu…

200 篇论文

After pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Tiancheng Gu , Kaicheng Yang , Chaoyi Zhang , Yin Xie , Xiang An , Ziyong Feng , Dongnan Liu , Weidong Cai , Jiankang Deng

Understanding how the brain encodes visual information is a central challenge in neuroscience and machine learning. A promising approach is to reconstruct visual stimuli, essentially images, from functional Magnetic Resonance Imaging (fMRI)…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Zheng Huang , Enpei Zhang , Weikang Qiu , Yinghao Cai , Carl Yang , Elynn Chen , Xiang Zhang , Rex Ying , Dawei Zhou , Yujun Yan

The reconstruction of images observed by subjects from fMRI data collected during visual stimuli has made strong progress in the past decade, thanks to the availability of extensive fMRI datasets and advancements in generative models for…

Training vision-language models (VLMs) typically requires large-scale, high-quality image-text pairs, but collecting or synthesizing such data is costly. In contrast, text data is abundant and inexpensive, prompting the question: can…

人工智能 · 计算机科学 2026-05-28 Xiaomin Yu , Wenjie Zhang , Ziyue Qiao , Chengwei Qin , Hui Xiong

We propose a multimodal latent diffusion model that jointly synthesizes volumetric magnetic resonance imaging (MRI) and tabular clinical data within a shared latent space via cross-attention. This approach enables coherent joint…

图像与视频处理 · 电气工程与系统科学 2026-05-11 Daniel Mensing , Jan Kapar , Jochen G. Hirsch , Matthias Günther , Horst Hahn , Marvin N. Wright

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven dance video…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Zhikang Dong , Weituo Hao , Ju-Chiang Wang , Peng Zhang , Pawel Polak

Multimodal large language models (MLLMs) are rapidly evolving, presenting increasingly complex safety challenges. However, current dataset construction methods, which are risk-oriented, fail to cover the growing complexity of real-world…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Jingen Qu , Lijun Li , Bo Zhang , Yichen Yan , Jing Shao

Magnetic Resonance (MR) images of different modalities can provide complementary information for clinical diagnosis, but whole modalities are often costly to access. Most existing methods only focus on synthesizing missing images between…

计算机视觉与模式识别 · 计算机科学 2020-05-05 Bingyu Xin , Yifan Hu , Yefeng Zheng , Hongen Liao

Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. As the number of reference identities…

机器学习 · 计算机科学 2026-04-10 Yucheng Zhou , Dubing Chen , Huan Zheng , Jianbing Shen

We present Meissonic, which elevates non-autoregressive masked image modeling (MIM) text-to-image to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations,…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Jinbin Bai , Tian Ye , Wei Chow , Enxin Song , Xiangtai Li , Zhen Dong , Lei Zhu , Shuicheng Yan

Functional MRI (fMRI) is a powerful technique that has allowed us to characterize visual cortex responses to stimuli, yet such experiments are by nature constructed based on a priori hypotheses, limited to the set of images presented to the…

Multimodal machine learning (MML) is rapidly reshaping the way mental-health disorders are detected, characterized, and longitudinally monitored. Whereas early studies relied on isolated data streams -- such as speech, text, or wearable…

机器学习 · 计算机科学 2025-06-25 Zahraa Al Sahili , Ioannis Patras , Matthew Purver

Deep generative models have emerged as a transformative tool in medical imaging, offering substantial potential for synthetic data generation. However, recent empirical studies highlight a critical vulnerability: these models can memorize…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Antonio Scardace , Lemuel Puglisi , Francesco Guarnera , Sebastiano Battiato , Daniele Ravì

3D structural Magnetic Resonance Imaging (MRI) brain scans are commonly acquired in clinical settings to monitor a wide range of neurological conditions, including neurodegenerative disorders and stroke. While deep learning models have…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Emily Kaczmarek , Justin Szeto , Brennan Nichyporuk , Tal Arbel

The recently emerging conditional diffusion models seem promising for mitigating the labor and expenses in building large 3D medical imaging datasets. However, previous studies on 3D CT generation primarily focus on specific organs…

图像与视频处理 · 电气工程与系统科学 2025-12-02 Linrui Dai , Rongzhao Zhang , Yongrui Yu , Xiaofan Zhang

Recent advancements in text-to-video (T2V) diffusion models have significantly enhanced the visual quality of the generated videos. However, even recent T2V models find it challenging to follow text descriptions accurately, especially when…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Jialu Li , Shoubin Yu , Han Lin , Jaemin Cho , Jaehong Yoon , Mohit Bansal

Composed Image Retrieval (CIR) retrieves target images using a multi-modal query that combines a reference image with text describing desired modifications. The primary challenge is effectively fusing this visual and textual information.…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Chaoyang Wang , Zeyu Zhang , Long Teng , Zijun Li , Shichao Kan

This review paper delves into the present state of medical imaging, with a specific focus on the use of deep learning techniques for brain image synthesis. The need for medical image synthesis to improve diagnostic accuracy and decrease…

图像与视频处理 · 电气工程与系统科学 2023-09-20 Shubham Singh , Ammar Ranapurwala , Mrunal Bewoor , Sheetal Patil , Satyam Rai

Multi-modality magnetic resonance imaging (MRI) is essential for the diagnosis and treatment of brain tumors. However, missing modalities are commonly observed due to limitations in scan time, scan corruption, artifacts, motion, and…

图像与视频处理 · 电气工程与系统科学 2025-01-08 Xiaojiao Xiao , Qinmin Vivian Hu , Guanghui Wang

Segmenting the fine structure of the mouse brain on magnetic resonance (MR) images is critical for delineating morphological regions, analyzing brain function, and understanding their relationships. Compared to a single MRI modality,…

图像与视频处理 · 电气工程与系统科学 2022-12-06 Ziqi Yu , Xiaoyang Han , Shengjie Zhang , Jianfeng Feng , Tingying Peng , Xiao-Yong Zhang