English
Related papers

Related papers: HPSv3: Towards Wide-Spectrum Human Preference Scor…

200 papers

Large Visual Language Models (LVLMs) increasingly rely on preference alignment to ensure reliability, which steers the model behavior via preference fine-tuning on preference data structured as ``image - winner text - loser text'' triplets.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Kejia Chen , Jiawen Zhang , Jiacong Hu , Jiazhen Yang , Jian Lou , Zunlei Feng , Mingli Song

Capturing the diversity of people in images is challenging: recent literature tends to focus on diversifying one or two attributes, requiring expensive attribute labels or building classifiers. We introduce a diverse people image ranking…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Hansa Srinivasan , Candice Schumann , Aradhana Sinha , David Madras , Gbolahan Oluwafemi Olanubi , Alex Beutel , Susanna Ricco , Jilin Chen

To prevent misinformation and social issues arising from trustworthy-looking content generated by LLMs, it is crucial to develop efficient and reliable methods for identifying the source of texts. Previous approaches have demonstrated…

Computation and Language · Computer Science 2025-12-03 Fangqi Dai , Xingjian Jiang , Zizhuang Deng

Recent advances in zero-shot text-to-3D human generation, which employ the human model prior (eg, SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Jianhui Yu , Hao Zhu , Liming Jiang , Chen Change Loy , Weidong Cai , Wayne Wu

Direct Preference Optimization (DPO) has emerged as a powerful approach to align text-to-image (T2I) models with human feedback. Unfortunately, successful application of DPO to T2I models requires a huge amount of resources to collect and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Shyamgopal Karthik , Huseyin Coskun , Zeynep Akata , Sergey Tulyakov , Jian Ren , Anil Kag

Recent advances in Score Distillation Sampling (SDS) have improved 3D human generation from textual descriptions. However, existing methods still face challenges in accurately aligning 3D models with long and complex textual inputs. To…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Pengfei Zhou , Xukun Shen , Yong Hu

Can Visual Language Models (VLMs) effectively capture human visual preferences? This work addresses this question by training VLMs to think about preferences at test time, employing reinforcement learning methods inspired by DeepSeek R1 and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Alexander Gambashidze , Konstantin Sobolev , Andrey Kuznetsov , Ivan Oseledets

Generating high-quality, photorealistic textures for 3D human avatars remains a fundamental yet challenging task in computer vision and multimedia field. However, real paired front and back images of human subjects are rarely available with…

Graphics · Computer Science 2026-04-10 Mingxiao Tu , Shuchang Ye , Hoijoon Jung , Jinman Kim

Graphic layouts serve as an important and engaging medium for visual communication across different channels. While recent layout generation models have demonstrated impressive capabilities, they frequently fail to align with nuanced human…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Varun Gopal , Rishabh Jain , Aradhya Mathur , Nikitha SR , Sohan Patnaik , Sudhir Yarram , Mayur Hemani , Balaji Krishnamurthy , Mausoom Sarkar

Human body restoration, as a specific application of image restoration, is widely applied in practice and plays a vital role across diverse fields. However, thorough research remains difficult, particularly due to the lack of benchmark…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Jue Gong , Jingkai Wang , Zheng Chen , Xing Liu , Hong Gu , Yulun Zhang , Xiaokang Yang

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Mude Hui , Siwei Yang , Bingchen Zhao , Yichun Shi , Heng Wang , Peng Wang , Yuyin Zhou , Cihang Xie

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human preferences, limiting…

Computation and Language · Computer Science 2025-02-10 Zhenglin Zhou , Xiaobo Xia , Fan Ma , Hehe Fan , Yi Yang , Tat-Seng Chua

The increasing availability of image-text pairs has largely fueled the rapid advancement in vision-language foundation models. However, the vast scale of these datasets inevitably introduces significant variability in data quality, which…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Lei Zhang , Fangxun Shu , Tianyang Liu , Sucheng Ren , Hao Jiang , Cihang Xie

Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generated images with latent visual preferences a fundamental challenge in text-to-image diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Wenxi Wang , Hongbin Liu , Mingqian Li , Junyan Yuan , Junqi Zhang

Human pose estimation on medium and small scales has long been a significant challenge in this field. Most existing methods focus on restoring high-resolution feature maps by stacking multiple costly deconvolutional layers or by…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Zhoujie Xu

The goal of aligning language models to human preferences requires data that reveal these preferences. Ideally, time and money can be spent carefully collecting and tailoring bespoke preference data to each downstream application. However,…

Artificial Intelligence · Computer Science 2024-09-17 Judy Hanwen Shen , Archit Sharma , Jun Qin

Recent studies have demonstrated the exceptional potentials of leveraging human preference datasets to refine text-to-image generative models, enhancing the alignment between generated images and textual prompts. Despite these advances,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Xun Wu , Shaohan Huang , Furu Wei

Human pose and shape (HPS) estimation with lensless imaging is not only beneficial to privacy protection but also can be used in covert surveillance scenarios due to the small size and simple structure of this device. However, this task…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Haoyang Ge , Qiao Feng , Hailong Jia , Xiongzheng Li , Xiangjun Yin , You Zhou , Jingyu Yang , Kun Li

Large vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Chenglong Wang , Yang Gan , Yifu Huo , Yongyu Mu , Murun Yang , Qiaozhi He , Tong Xiao , Chunliang Zhang , Tongran Liu , Quan Du , Di Yang , Jingbo Zhu

RLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model generation with population-level preferences, neglecting the…

Machine Learning · Computer Science 2025-01-14 Meihua Dang , Anikait Singh , Linqi Zhou , Stefano Ermon , Jiaming Song