中文
相关论文

相关论文: Towards General Text-guided Image Synthesis for Cu…

200 篇论文

Ultrasound (US) imaging is widely used for anatomical structure inspection in clinical diagnosis. The training of new sonographers and deep learning based algorithms for US image analysis usually requires a large amount of data. However,…

图像与视频处理 · 电气工程与系统科学 2022-05-26 Jiamin Liang , Xin Yang , Yuhao Huang , Haoming Li , Shuangchi He , Xindi Hu , Zejian Chen , Wufeng Xue , Jun Cheng , Dong Ni

The integration of multimodal medical imaging can provide complementary and comprehensive information for the diagnosis of Alzheimer's disease (AD). However, in clinical practice, since positron emission tomography (PET) is often missing,…

计算工程、金融与科学 · 计算机科学 2024-12-03 Fuyou Mao , Lixin Lin , Ming Jiang , Dong Dai , Chao Yang , Hao Zhang , Yan Tang

Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on single-image…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Nupur Kumari , Xi Yin , Jun-Yan Zhu , Ishan Misra , Samaneh Azadi

High-resolution magnetic resonance images can provide fine-grained anatomical information, but acquiring such data requires a long scanning time. In this paper, a framework called the Fused Attentive Generative Adversarial Networks(FA-GAN)…

图像与视频处理 · 电气工程与系统科学 2021-08-29 Mingfeng Jiang , Minghao Zhi , Liying Wei , Xiaocheng Yang , Jucheng Zhang , Yongming Li , Pin Wang , Jiahao Huang , Guang Yang

Multi-modal information retrieval (MMIR) is a rapidly evolving field, where significant progress, particularly in image-text pairing, has been made through advanced representation learning and cross-modality alignment research. However,…

Generating synthetic text addresses the challenge of data availability in privacy-sensitive domains such as healthcare. This study explores the applicability of synthetic data in real-world medical settings. We introduce MedSyn, a novel…

Mammographic screening is an effective method for detecting breast cancer, facilitating early diagnosis. However, the current need to manually inspect images places a heavy burden on healthcare systems, spurring a desire for automated…

图像与视频处理 · 电气工程与系统科学 2025-01-30 Ciaran Bench , Emir Ahmed , Spencer A. Thomas

Tumor synthesis can generate examples that AI often misses or over-detects, improving AI performance by training on these challenging cases. However, existing synthesis methods, which are typically unconditional -- generating images from…

图像与视频处理 · 电气工程与系统科学 2024-12-25 Xinran Li , Yi Shuai , Chen Liu , Qi Chen , Qilong Wu , Pengfei Guo , Dong Yang , Can Zhao , Pedro R. A. S. Bassi , Daguang Xu , Kang Wang , Yang Yang , Alan Yuille , Zongwei Zhou

Computer-assisted diagnosis (CAD) based on deep learning has become a crucial diagnostic technology in the medical industry, effectively improving diagnosis accuracy. However, the scarcity of brain tumor Magnetic Resonance (MR) image…

计算机视觉与模式识别 · 计算机科学 2021-11-30 Panjian Huang , Xu Liu , Yongzhen Huang

Magnetic resonance imaging (MRI) is a widely used medical imaging modality. However, due to the limitations in hardware, scan time, and throughput, it is often clinically challenging to obtain high-quality MR images. The super-resolution…

图像与视频处理 · 电气工程与系统科学 2020-02-20 Qing Lyu , Hongming Shan , Ge Wang

Multi-modal brain MRI provides essential complementary information for clinical diagnosis. However, acquiring all modalities in practice is often constrained by time and cost. To address this, various methods have been proposed to generate…

图像与视频处理 · 电气工程与系统科学 2026-05-07 Hanyeol Yang , Sunggyu Kim , Mi Kyung Kim , Yongseon Yoo , Yu-Mi Kim , Min-Ho Shin , Insung Chung , Sang Baek Koh , Hyeon Chang Kim , Jong-Min Lee

Accurate brain tumor diagnosis requires models to not only detect lesions but also generate clinically interpretable reasoning grounded in imaging manifestations, yet existing public datasets remain limited in annotation richness and…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Feng Guo , Jiaxiang Liu , Yang Li , Qianqian Shi , Mingkun Xu

The reconstruction of human visual inputs from brain activity, particularly through functional Magnetic Resonance Imaging (fMRI), holds promising avenues for unraveling the mechanisms of the human visual system. Despite the significant…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Yujian Xiong , Wenhui Zhu , Zhong-Lin Lu , Yalin Wang

This paper presents instruct-imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce *multi-modal instruction* for image generation, a task representation articulating a range of…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Hexiang Hu , Kelvin C. K. Chan , Yu-Chuan Su , Wenhu Chen , Yandong Li , Kihyuk Sohn , Yang Zhao , Xue Ben , Boqing Gong , William Cohen , Ming-Wei Chang , Xuhui Jia

Neural Radiance Fields (NeRF) have shown impressive performances in the rendering of 3D scenes from arbitrary viewpoints. While RGB images are widely preferred for training volume rendering models, the interest in other radiance modalities…

图形学 · 计算机科学 2025-03-26 Federico Lincetto , Gianluca Agresti , Mattia Rossi , Pietro Zanuttigh

Image fusion typically employs non-invertible neural networks to merge multiple source images into a single fused image. However, for clinical experts, solely relying on fused images may be insufficient for making diagnostic decisions, as…

图像与视频处理 · 电气工程与系统科学 2024-06-11 Nishant Kumar , Ziyan Tao , Jaikirat Singh , Yang Li , Peiwen Sun , Binghui Zhao , Stefan Gumhold

Image-text retrieval is a central problem for understanding the semantic relationship between vision and language, and serves as the basis for various visual and language tasks. Most previous works either simply learn coarse-grained…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Chong Liu , Yuqi Zhang , Hongsong Wang , Weihua Chen , Fan Wang , Yan Huang , Yi-Dong Shen , Liang Wang

Text to Image Synthesis refers to the process of automatic generation of a photo-realistic image starting from a given text and is revolutionizing many real-world applications. In order to perform such process it is necessary to exploit…

Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Mingrui Wu , Hang Liu , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

For low-level computer vision and image processing ML tasks, training on large datasets is critical for generalization. However, the standard practice of relying on real-world images primarily from the Internet comes with image quality,…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Gyeongmin Choe , Beibei Du , Seonghyeon Nam , Xiaoyu Xiang , Bo Zhu , Rakesh Ranjan