English
Related papers

Related papers: VLM-Guided Group Preference Alignment for Diffusio…

200 papers

Estimating 3D mesh of the human body from a single 2D image is an important task with many applications such as augmented reality and Human-Robot interaction. However, prior works reconstructed 3D mesh from global image feature extracted by…

Computer Vision and Pattern Recognition · Computer Science 2021-06-08 Wang Zeng , Wanli Ouyang , Ping Luo , Wentao Liu , Xiaogang Wang

Diffusion models, while trained for image generation, have emerged as powerful foundational feature extractors for downstream tasks. We find that off-the-shelf diffusion models, trained exclusively to generate natural RGB images, can…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Nurislam Tursynbek , Hastings Greer , Basar Demir , Marc Niethammer

Without using explicit reward, direct preference optimization (DPO) employs paired human preference data to fine-tune generative models, a method that has garnered considerable attention in large language models (LLMs). However, exploration…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yunhong Lu , Qichao Wang , Hengyuan Cao , Xierui Wang , Xiaoyin Xu , Min Zhang

Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model. Although existing methods are derived from different perspectives, we show…

Machine Learning · Computer Science 2026-05-08 Jeongjae Lee , Jinho Chang , Jeongsol Kim , Jong Chul Ye

Modelling the diffusion-relaxation magnetic resonance (MR) signal obtained from multi-parametric sequences has recently gained immense interest in the community due to new techniques significantly reducing data acquisition time. A preferred…

Medical Physics · Physics 2025-01-28 Fabian Bogusz , Tomasz Pieciak

While recent works on blind face image restoration have successfully produced impressive high-quality (HQ) images with abundant details from low-quality (LQ) input images, the generated content may not accurately reflect the real appearance…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Chi-Wei Hsiao , Yu-Lun Liu , Cheng-Kun Yang , Sheng-Po Kuo , Kevin Jou , Chia-Ping Chen

As a dominant force in text-to-image generation tasks, Diffusion Probabilistic Models (DPMs) face a critical challenge in controllability, struggling to adhere strictly to complex, multi-faceted instructions. In this work, we aim to address…

Machine Learning · Computer Science 2024-02-27 Xuantong Liu , Tianyang Hu , Wenjia Wang , Kenji Kawaguchi , Yuan Yao

The challenge in fine-grained visual categorization lies in how to explore the subtle differences between different subclasses and achieve accurate discrimination. Previous research has relied on large-scale annotated data and pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Tianxu Wu , Shuo Ye , Shuhuang Chen , Qinmu Peng , Xinge You

In this paper we propose RecFusion, which comprise a set of diffusion models for recommendation. Unlike image data which contain spatial correlations, a user-item interaction matrix, commonly utilized in recommendation, lacks spatial…

Information Retrieval · Computer Science 2023-09-11 Gabriel Bénédict , Olivier Jeunen , Samuele Papa , Samarth Bhargav , Daan Odijk , Maarten de Rijke

Diffusion-based human animation aims to animate a human character based on a source human image as well as driving signals such as a sequence of poses. Leveraging the generative capacity of diffusion model, existing approaches are able to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Fa-Ting Hong , Zhan Xu , Haiyang Liu , Qinjie Lin , Luchuan Song , Zhixin Shu , Yang Zhou , Duygu Ceylan , Dan Xu

Diffusion models are state-of-the-art generative models, yet their samples often fail to satisfy application objectives such as safety constraints or domain-specific validity. Existing techniques for alignment require gradients, internal…

Single LDR to HDR reconstruction remains challenging for over-exposed regions where traditional methods often fail due to complete information loss. We present a training-free approach that enhances existing indirect and direct HDR…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yo-Tin Lin , Su-Kai Chen , Hou-Ning Hu , Yen-Yu Lin , Yu-Lun Liu

Multi-behavior sequential recommendation (MBSR) aims to learn the dynamic and heterogeneous interactions of users' multi-behavior sequences, so as to capture user preferences under target behavior for the next interacted item prediction.…

Information Retrieval · Computer Science 2026-02-27 Ruochen Yang , Xiaodong Li , Jiawei Sheng , Jiangxia Cao , Xinkui Lin , Shen Wang , Shuang Yang , Zhaojie Liu , Tingwen Liu

Diffusion models (DMs) have achieved promising performance in image restoration but haven't been explored for stereo images. The application of DM in stereo image restoration is confronted with a series of challenges. The need to…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Huiyun Cao , Yuan Shi , Bin Xia , Xiaoyu Jin , Wenming Yang

Image synthesis approaches, e.g., generative adversarial networks, have been popular as a form of data augmentation in medical image analysis tasks. It is primarily beneficial to overcome the shortage of publicly accessible data and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Shiyi Du , Xiaosong Wang , Yongyi Lu , Yuyin Zhou , Shaoting Zhang , Alan Yuille , Kang Li , Zongwei Zhou

Instruction-based image editing has made a great process in using natural human language to manipulate the visual content of images. However, existing models are limited by the quality of the dataset and cannot accurately localize editing…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Tiancheng Li , Jinxiu Liu , Huajun Chen , Qi Liu

With the rapid advancement of technologies such as virtual reality, augmented reality, and gesture control, users expect interactions with computer interfaces to be more natural and intuitive. Existing visual algorithms often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Haonan Li , Patrick P. K. Chen , Yitong Zhou

Diffusion models face significant challenges when employed for large-scale medical image reconstruction in real practice such as 3D Computed Tomography (CT). Due to the demanding memory, time, and data requirements, it is difficult to train…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Bowen Song , Jason Hu , Zhaoxu Luo , Jeffrey A. Fessler , Liyue Shen

We propose DiMeR, a novel geometry-texture disentangled feed-forward model with 3D supervision for sparse-view mesh reconstruction. Existing methods confront two persistent obstacles: (i) textures can conceal geometric errors, i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Lutao Jiang , Jiantao Lin , Kanghao Chen , Wenhang Ge , Xin Yang , Yifan Jiang , Yuanhuiyi Lyu , Xu Zheng , Yinchuan Li , Yingcong Chen

Diffusion Models represent a significant advancement in generative modeling, employing a dual-phase process that first degrades domain-specific information via Gaussian noise and restores it through a trainable model. This framework enables…

Neural and Evolutionary Computing · Computer Science 2024-11-21 Benedikt Hartl , Yanbo Zhang , Hananel Hazan , Michael Levin