English
Related papers

Related papers: UniX: Unifying Autoregression and Diffusion for Ch…

200 papers

This paper introduces TBAC-UniImage, a novel unified model for multimodal understanding and generation. We achieve this by deeply integrating a pre-trained Diffusion Model, acting as a generative ladder, with a Multimodal Large Language…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Junzhe Xu , Yuyang Yin , Xi Chen

The recently developed discrete diffusion models perform extraordinarily well in the text-to-image task, showing significant promise for handling the multi-modality signals. In this work, we harness these traits and present a unified…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Minghui Hu , Chuanxia Zheng , Heliang Zheng , Tat-Jen Cham , Chaoyue Wang , Zuopeng Yang , Dacheng Tao , Ponnuthurai N. Suganthan

Developing robust and versatile deep-learning models is essential for enhancing diagnostic accuracy and guiding clinical interventions in medical imaging, but it requires a large amount of annotated data. The advancement of deep learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Nahid Ul Islam , DongAo Ma , Jiaxuan Pang , Shivasakthi Senthil Velan , Michael Gotway , Jianming Liang

Automated question-answering (QA) systems increasingly rely on retrieval-augmented generation (RAG) to ground large language models (LLMs) in authoritative medical knowledge, ensuring clinical accuracy and patient safety in Artificial…

Computation and Language · Computer Science 2026-03-05 Aswini Sivakumar , Vijayan Sugumaran , Yao Qiang

Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal understanding and reasoning especially in high-value verticals…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Qianchu Liu , Sheng Zhang , Guanghui Qin , Yu Gu , Ying Jin , Sam Preston , Yanbo Xu , Sid Kiblawi , Wen-wai Yim , Tim Ossowski , Tristan Naumann , Mu Wei , Hoifung Poon

Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM backbone handles both visual understanding and generation. Despite the architectural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Sucheng Ren , Chen Chen , Zhenbang Wang , Liangchen Song , Xiangxin Zhu , Alan Yuille , Liang-Chieh Chen , Jiasen Lu

We present UniTEX, a novel two-stage 3D texture generation framework to create high-quality, consistent textures for 3D assets. Existing approaches predominantly rely on UV-based inpainting to refine textures after reprojecting the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Yixun Liang , Kunming Luo , Xiao Chen , Rui Chen , Hongyu Yan , Weiyu Li , Jiarui Liu , Ping Tan

Counterfactual medical image generation have emerged as a critical tool for enhancing AI-driven systems in medical domain by answering "what-if" questions. However, existing approaches face two fundamental limitations: First, they fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hyungi Min , Taeseung You , Hangyeul Lee , Yeongjae Cho , Sungzoon Cho

We study the joint learning of image-to-text and text-to-image generations, which are naturally bi-directional tasks. Typical existing works design two separate task-specific models for each task, which impose expensive design efforts. In…

Computer Vision and Pattern Recognition · Computer Science 2021-10-20 Yupan Huang , Hongwei Xue , Bei Liu , Yutong Lu

Recent progress in controllable image generation and editing is largely driven by diffusion-based methods. Although diffusion models perform exceptionally well in specific tasks with tailored designs, establishing a unified model is still…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Jiteng Mu , Nuno Vasconcelos , Xiaolong Wang

We introduce CheXGenBench, a rigorous and multifaceted evaluation framework for synthetic chest radiograph generation that simultaneously assesses fidelity, privacy risks, and clinical utility across state-of-the-art text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Raman Dutt , Pedro Sanchez , Yongchen Yao , Steven McDonagh , Sotirios A. Tsaftaris , Timothy Hospedales

We introduce UGen, a unified autoregressive multimodal model that demonstrates strong performance across text processing, image understanding, and image generation tasks simultaneously. UGen converts both texts and images into discrete…

Computation and Language · Computer Science 2025-03-28 Hongxuan Tang , Hao Liu , Xinyan Xiao

Chest X-ray radiography (CXR) is an essential medical imaging technique for disease diagnosis. However, as 2D projectional images, CXRs are limited by structural superposition and hence fail to capture 3D anatomies. This limitation makes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Zefan Yang , Ge Wang , James Hendler , Mannudeep K. Kalra , Pingkun Yan

While Unified Multimodal Models (UMMs) have achieved remarkable success in cross-modal comprehension, a significant gap persists in their ability to leverage such internal knowledge for high-quality generation. We formalize this discrepancy…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Ruiyan Han , Zhen Fang , XinYu Sun , Yuchen Ma , Ziheng Wang , Yu Zeng , Zehui Chen , Lin Chen , Wenxuan Huang , Wei-Jie Xu , Yi Cao , Feng Zhao

Foundation models leveraging vision-language pretraining have shown promise in chest X-ray (CXR) interpretation, yet their real-world performance across diverse populations and diagnostic tasks remains insufficiently evaluated. This study…

Text-To-Image (TTI) generation is significant for controlled and diverse image generation with broad potential applications. Although current medical TTI methods have made some progress in report-to-Chest-Xray (CXR) generation, their…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Peng Huang , Bowen Guo , Shuyu Liang , Junhu Fu , Yuanyuan Wang , Yi Guo

Chest radiography is climacteric in identifying different pulmonary diseases, yet radiologist workload and inefficiency can lead to misdiagnoses. Automatic, accurate, and efficient segmentation of lung from X-ray images of chest is…

Image and Video Processing · Electrical Eng. & Systems 2024-12-17 Sharmin Akter

VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Yecheng Wu , Zhuoyang Zhang , Junyu Chen , Haotian Tang , Dacheng Li , Yunhao Fang , Ligeng Zhu , Enze Xie , Hongxu Yin , Li Yi , Song Han , Yao Lu

Purpose: To explore best-practice approaches for generating synthetic chest X-ray images and augmenting medical imaging datasets to optimize the performance of deep learning models in downstream tasks like classification and segmentation.…

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Xi Chen , Zhifei Zhang , He Zhang , Yuqian Zhou , Soo Ye Kim , Qing Liu , Yijun Li , Jianming Zhang , Nanxuan Zhao , Yilin Wang , Hui Ding , Zhe Lin , Hengshuang Zhao
‹ Prev 1 3 4 5 6 7 10 Next ›