English
Related papers

Related papers: MCAD: Multi-modal Conditioned Adversarial Diffusio…

200 papers

Computed tomography is a widely used imaging modality with applications ranging from medical imaging to material analysis. One major challenge arises from the lack of scanning information at certain angles, resulting in distortion or…

Reconstructing CT images from incomplete projection data remains challenging due to the ill-posed nature of the problem. Diffusion bridge models have recently shown promise in restoring clean images from their corresponding Filtered Back…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yuang Wang , Pengfei Jin , Siyeop Yoon , Matthew Tivnan , Shaoyang Zhang , Li Zhang , Quanzheng Li , Zhiqiang Chen , Dufan Wu

Objective: Cone-beam computed tomography (CBCT) provides a low-dose imaging alternative to conventional CT, but suffers from noise, scatter, and artifacts that degrade image quality. Synthetic CT (sCT) aims to translate CBCT to high-quality…

Medical Physics · Physics 2025-09-23 Alzahra Altalib , Chunhui Li , Alessandro Perelli

Medical image segmentation is crucial for computer-aided diagnosis, which necessitates understanding both coarse morphological and semantic structures, as well as carving fine boundaries. The morphological and semantic structures in medical…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Siyuan Song , Guyue Hu , Chenglong Li , Dengdi Sun , Zhe Jin , Jin Tang

The efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and patient metadata) and addressing the challenges of multi-label…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Yuan Zhang , Yutong Xie , Hu Wang , Jodie C Avery , M Louise Hull , Gustavo Carneiro

Large high-quality medical image datasets are difficult to acquire but necessary for many deep learning applications. For positron emission tomography (PET), reconstructed image quality is limited by inherent Poisson noise. We propose a…

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Shilong Zhang , He Zhang , Zhifei Zhang , Chongjian Ge , Shuchen Xue , Shaoteng Liu , Mengwei Ren , Soo Ye Kim , Yuqian Zhou , Qing Liu , Daniil Pakhomov , Kai Zhang , Zhe Lin , Ping Luo

Multimodal variational autoencoders have demonstrated their ability to learn the relationships between different modalities by mapping them into a latent representation. Their design and capacity to perform any-to-any conditional and…

Machine Learning · Computer Science 2025-02-04 Daniel Wesego , Pedram Rooshenas

Diffusion models have shown remarkable flexibility for solving inverse problems without task-specific retraining. However, existing approaches such as Manifold Preserving Guided Diffusion (MPGD) apply only a single gradient update per…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Aditya Chakravarty

Computational tomography (CT) provides high-resolution medical imaging, but it can expose patients to high radiation. X-ray scanners have low radiation exposure, but their resolutions are low. This paper proposes a new conditional diffusion…

Image and Video Processing · Electrical Eng. & Systems 2025-01-20 Yun Su Jeong , Hye Bin Yoo , Il Yong Chun

Limited-Angle Computed Tomography (LACT) is a challenging inverse problem where missing angular projections lead to incomplete sinograms and severe artifacts in the reconstructed images. While recent learning-based methods have demonstrated…

Image and Video Processing · Electrical Eng. & Systems 2025-07-09 Jiaqi Guo , Santiago López-Tapia

To obtain high-quality Positron emission tomography (PET) images while minimizing radiation exposure, numerous methods have been proposed to reconstruct standard-dose PET (SPET) images from the corresponding low-dose PET (LPET) images.…

Image and Video Processing · Electrical Eng. & Systems 2024-04-10 Jiaqi Cui , Yan Wang , Lu Wen , Pinxian Zeng , Xi Wu , Jiliu Zhou , Dinggang Shen

Therapeutic peptides represent a unique class of pharmaceutical agents crucial for the treatment of human diseases. Recently, deep generative models have exhibited remarkable potential for generating therapeutic peptides, but they only…

Quantitative Methods · Quantitative Biology 2024-01-05 Yongkang Wang , Xuan Liu , Feng Huang , Zhankun Xiong , Wen Zhang

Neurobiological and neurodegenerative diseases are inherently multifactorial, arising from coupled influences spanning genetic susceptibility, brain alterations, and environmental and behavioral factors. Multimodal modeling has therefore…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Shaowen Wan , Yanjun Lv , Lu Zhang , Dajiang Zhu , Bharat Biswal , Tianming Liu , Xiaobo Li , Lin Zhao

Positron Emission Tomography (PET) image reconstruction is inherently challenged by Poisson noise and physical degradation factors, which are further exacerbated in limited-angle acquisitions. While deep learning methods demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Rüveyda Yilmaz , Yuli Wu , Johannes Stegmaier , Volkmar Schulz

Computed Tomography (CT) is a medical imaging modality that can generate more informative 3D images than 2D X-rays. However, this advantage comes at the expense of more radiation exposure, higher costs, and longer acquisition time. Hence,…

Image and Video Processing · Electrical Eng. & Systems 2023-09-12 Shuangqin Cheng , Qingliang Chen , Qiyi Zhang , Ming Li , Yamuhanmode Alike , Kaile Su , Pengcheng Wen

Different from a unimodal model whose input is from a single modality, the input (called multi-modal input) of a multi-modal model is from multiple modalities such as image, 3D points, audio, text, etc. Similar to unimodal models, many…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Yanting Wang , Hongye Fu , Wei Zou , Jinyuan Jia

Cross-modality magnetic resonance (MR) image synthesis can be used to generate missing modalities from given ones. Existing (supervised learning) methods often require a large number of paired multi-modal data to train an effective…

Image and Video Processing · Electrical Eng. & Systems 2023-06-21 Yonghao Li , Tao Zhou , Kelei He , Yi Zhou , Dinggang Shen

Deep learning has significantly advanced PET image re-construction, achieving remarkable improvements in image quality through direct training on sinogram or image data. Traditional methods often utilize masks for inpainting tasks, but…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Bin Huang , Binzhong He , Yanhan Chen , Zhili Liu , Xinyue Wang , Binxuan Li , Qiegen Liu

Esophageal cancer is one of the most common types of cancer worldwide and ranks sixth in cancer-related mortality. Accurate computer-assisted diagnosis of cancer progression can help physicians effectively customize personalized treatment…

Image and Video Processing · Electrical Eng. & Systems 2024-05-17 Chengyu Wu , Chengkai Wang , Yaqi Wang , Huiyu Zhou , Yatao Zhang , Qifeng Wang , Shuai Wang