English
Related papers

Related papers: Controlling Human Shape and Pose in Text-to-Image …

200 papers

Vanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs.Existing methods address this issue mostly by fine-tuning…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Junyan Wang , Zhenhong Sun , Zhiyu Tan , Xuanbai Chen , Weihua Chen , Hao Li , Cheng Zhang , Yang Song

Recently, a few self-supervised representation learning (SSL) methods have outperformed the ImageNet classification pre-training for vision tasks such as object detection. However, its effects on 3D human body pose and shape estimation…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Hongsuk Choi , Hyeongjin Nam , Taeryung Lee , Gyeongsik Moon , Kyoung Mu Lee

Recently, diffusion models have exhibited superior performance in the area of image inpainting. Inpainting methods based on diffusion models can usually generate realistic, high-quality image content for masked areas. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Ruichen Wang , Junliang Zhang , Qingsong Xie , Chen Chen , Haonan Lu

Our goal is to capture the pose of neuroscience model organisms, without using any manual supervision, to be able to study how neural circuits orchestrate behaviour. Human pose estimation attains remarkable accuracy when trained on real or…

Computer Vision and Pattern Recognition · Computer Science 2020-01-24 Siyuan Li , Semih Günel , Mirela Ostrek , Pavan Ramdya , Pascal Fua , Helge Rhodin

Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Jogendra Nath Kundu , Siddharth Seth , Anirudh Jamkhandi , Pradyumna YM , Varun Jampani , Anirban Chakraborty , R. Venkatesh Babu

We present PersonaCraft, a framework for controllable and occlusion-robust full-body personalized image synthesis of multiple individuals in complex scenes. Current methods struggle with occlusion-heavy scenarios and complete body…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Gwanghyun Kim , Suh Yoon Jeon , Seunggyu Lee , Se Young Chun

Previous probabilistic models for 3D Human Pose Estimation (3DHPE) aimed to enhance pose accuracy by generating multiple hypotheses. However, most of the hypotheses generated deviate substantially from the true pose. Compared to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Hongbo Kang , Yong Wang , Mengyuan Liu , Doudou Wu , Peng Liu , Xinlin Yuan , Wenming Yang

Text-conditioned video diffusion models have emerged as a powerful tool in the realm of video generation and editing. But their ability to capture the nuances of human movement remains under-explored. Indeed the ability of these models to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Paul Janson , Tiberiu Popa , Eugene Belilovsky

We propose an approach to domain adaptation for semantic segmentation that is both practical and highly accurate. In contrast to previous work, we abandon the use of computationally involved adversarial objectives, network ensembles and…

Computer Vision and Pattern Recognition · Computer Science 2021-05-04 Nikita Araslanov , Stefan Roth

Current state-of-the-art in 3D human pose and shape recovery relies on deep neural networks and statistical morphable body models, such as the Skinned Multi-Person Linear model (SMPL). However, regardless of the advantages of having both…

Computer Vision and Pattern Recognition · Computer Science 2019-08-09 Meysam Madadi , Hugo Bertiche , Sergio Escalera

Human pose transfer, as a misaligned image generation task, is very challenging. Existing methods cannot effectively utilize the input information, which often fail to preserve the style and shape of hair and clothes. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Jinsong Zhang , Xingzi Liu , Kun Li

Since annotating pixel-level labels for semantic segmentation is laborious, leveraging synthetic data is an attractive solution. However, due to the domain gap between synthetic domain and real domain, it is challenging for a model trained…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Myeongjin Kim , Hyeran Byun

Text-to-image models (T2I) such as StableDiffusion have been used to generate high quality images of people. However, due to the random nature of the generation process, the person has a different appearance e.g. pose, face, and clothing,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Soon Yau Cheong , Armin Mustafa , Andrew Gilbert

A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Senmao Li , Joost van de Weijer , Taihang Hu , Fahad Shahbaz Khan , Qibin Hou , Yaxing Wang , Jian Yang , Ming-Ming Cheng

Deep generative modelling for human body analysis is an emerging problem with many interesting applications. However, the latent space learned by such approaches is typically not interpretable, resulting in less flexibility. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-18 Rodrigo de Bem , Arnab Ghosh , Thalaiyasingam Ajanthan , Ondrej Miksik , Adnane Boukhayma , N. Siddharth , Philip Torr

With diffusion transformer (DiT) excelling in video generation, its use in specific tasks has drawn increasing attention. However, adapting DiT for pose-guided human image animation faces two core challenges: (a) existing U-Net-based pose…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Haoyu Zhao , Zhongang Qi , Cong Wang , Qingping Zheng , Guansong Lu , Fei Chen , Hang Xu , Zuxuan Wu

Image extrapolation aims at expanding the narrow field of view of a given image patch. Existing models mainly deal with natural scene images of homogeneous regions and have no control of the content generation process. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2019-12-30 Yijun Li , Lu Jiang , Ming-Hsuan Yang

The robustness of gaze and head pose estimation models is highly dependent on the amount of labeled data. Recently, generative modeling has shown excellent results in generating photo-realistic images, which can alleviate the need for…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Swati Jindal , Xin Eric Wang

While current talking head models are capable of generating photorealistic talking head videos, they provide limited pose controllability. Most methods require specific video sequences that should exactly contain the head pose desired,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Kwangho Lee , Patrick Kwon , Myung Ki Lee , Namhyuk Ahn , Junsoo Lee

Multi-modal foundation models are typically trained on millions of pairs of natural images and text captions, frequently obtained through web-crawling approaches. Although such models depict excellent generative capabilities, they do not…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Pierre Chambon , Christian Bluethgen , Curtis P. Langlotz , Akshay Chaudhari