中文
相关论文

相关论文: HumanRefiner: Benchmarking Abnormal Human Generati…

200 篇论文

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Boyuan Wang , Xiaofeng Wang , Chaojun Ni , Guosheng Zhao , Zhiqin Yang , Zheng Zhu , Muyang Zhang , Yukun Zhou , Xinze Chen , Guan Huang , Lihong Liu , Xingang Wang

Recovering a 3D human mesh from a single RGB image is a challenging task due to depth ambiguity and self-occlusion, resulting in a high degree of uncertainty. Meanwhile, diffusion models have recently seen much success in generating…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Lin Geng Foo , Jia Gong , Hossein Rahmani , Jun Liu

Generative models are prone to hallucinations: plausible but incorrect structures absent in the ground truth. This issue is problematic in image restoration for safety-critical domains such as medical imaging, industrial inspection, and…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Seunghoi Kim , Henry F. J. Tregidgo , Chen Jin , Matteo Figini , Daniel C. Alexander

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential to revolutionize the…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Shiyue Zhang , Zheng Chong , Xi Lu , Wenqing Zhang , Haoxiang Li , Xujie Zhang , Jiehui Huang , Xiao Dong , Xiaodan Liang

Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Yizhi Song , Liu He , Zhifei Zhang , Soo Ye Kim , He Zhang , Wei Xiong , Zhe Lin , Brian Price , Scott Cohen , Jianming Zhang , Daniel Aliaga

Although diffusion models can generate high-quality human images, their applications are limited by the instability in generating hands with correct structures. In this paper, we introduce RHanDS, a conditional diffusion-based framework…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Chengrui Wang , Pengfei Liu , Min Zhou , Ming Zeng , Xubin Li , Tiezheng Ge , Bo zheng

3D human generation is an important problem with a wide range of applications in computer vision and graphics. Despite recent progress in generative AI such as diffusion models or rendering methods like Neural Radiance Fields or Gaussian…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction. However, prevailing systems impose an intact-limb prior and…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Jiaying Ying , Heming Du , Kaihao Zhang , Sean M. Tweedy , Xin Yu

In this research, we introduce RefineNet, a novel architecture designed to address resolution limitations in text-to-image conversion systems. We explore the challenges of generating high-resolution images from textual descriptions,…

计算机视觉与模式识别 · 计算机科学 2024-01-01 Fan Shi

Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such metrics frequently…

Accurately generating images of human bodies from text remains a challenging problem for state of the art text-to-image models. Commonly observed body-related artifacts include extra or missing limbs, unrealistic poses, blurred body parts,…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Nefeli Andreou , Varsha Vivek , Ying Wang , Alex Vorobiov , Tiffany Deng , Raja Bala , Larry Davis , Betty Mohler Tesch

Human mesh recovery (HMR) from a single image is inherently ill-posed due to depth ambiguity and occlusions. Probabilistic methods have tried to solve this by generating numerous plausible 3D human mesh predictions, but they often exhibit…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Wenhao Shen , Wanqi Yin , Xiaofeng Yang , Cheng Chen , Chaoyue Song , Zhongang Cai , Lei Yang , Hao Wang , Guosheng Lin

Image harmonization, which involves adjusting the foreground of a composite image to attain a unified visual consistency with the background, can be conceptualized as an image-to-image translation task. Diffusion models have recently…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Pengfei Zhou , Fangxiang Feng , Xiaojie Wang

Evaluating text-to-image generation models requires alignment with human perception, yet existing human-centric metrics are constrained by limited data coverage, suboptimal feature extraction, and inefficient loss functions. To address…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Yuhang Ma , Yunhao Shui , Xiaoshi Wu , Keqiang Sun , Hongsheng Li

We tackle the problem of Human Mesh Recovery (HMR) from a single RGB image, formulating it as an image-conditioned human pose and shape generation. While recovering 3D human pose from 2D observations is inherently ambiguous, most existing…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Donghwan Kim , Tae-Kyun Kim

This work presents a novel deep-learning-based pipeline for the inverse problem of image deblurring, leveraging augmentation and pre-training with synthetic data. Our results build on our winning submission to the recent Helsinki Deblur…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Theophil Trippe , Martin Genzel , Jan Macdonald , Maximilian März

Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper proposes Diffiner, a novel solution that utilizes the powerful…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Masato Hirano , Ryosuke Sawata , Naoki Murata , Shusuke Takahashi , Yuki Mitsufuji

3D human pose estimation from 2D images is a challenging problem due to depth ambiguity and occlusion. Because of these challenges the task is underdetermined, where there exists multiple -- possibly infinite -- poses that are plausible…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Francis Snelgar , Ming Xu , Stephen Gould , Liang Zheng , Akshay Asthana

In this paper, we aim to address the challenge of novel view rendering of human performers who wear clothes with complex texture patterns using a sparse set of camera views. Although some recent works have achieved remarkable rendering…

计算机视觉与模式识别 · 计算机科学 2023-10-24 Tiansong Zhou , Jing Huang , Tao Yu , Ruizhi Shao , Kun Li

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critical challenges: the…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Wanjiang Weng , Xiaofeng Tan , Hongsong Wang , Pan Zhou