English
Related papers

Related papers: Information-Regularized Constrained Inversion for …

200 papers

Face attribute editing aims to generate faces with one or multiple desired face attributes manipulated while other details are preserved. Unlike prior works such as GAN inversion, which has an expensive reverse mapping process, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Zhiliang Xu , Xiyu Yu , Zhibin Hong , Zhen Zhu , Junyu Han , Jingtuo Liu , Errui Ding , Xiang Bai

Recent advances in large pretrained text-to-image models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Qinghe Wang , Xu Jia , Xiaomin Li , Taiqing Li , Liqian Ma , Yunzhi Zhuge , Huchuan Lu

Identity preserving editing of faces is a generative task that enables modifying the illumination, adding/removing eyeglasses, face aging, editing hairstyles, modifying expression etc., while preserving the identity of the face. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Vishal Vinod

Existing solutions to image editing tasks suffer from several issues. Though achieving remarkably satisfying generated results, some supervised methods require huge amounts of paired training data, which greatly limits their usages. The…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Jinshu Chen , Bingchuan Li , Miao Hua , Panpan Xu , Qian He

Stylized 3D avatars have become increasingly prominent in our modern life. Creating these avatars manually usually involves laborious selection and adjustment of continuous and discrete parameters and is time-consuming for average users.…

Computer Vision and Pattern Recognition · Computer Science 2022-11-16 Shen Sang , Tiancheng Zhi , Guoxian Song , Minghao Liu , Chunpong Lai , Jing Liu , Xiang Wen , James Davis , Linjie Luo

Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These representations enable flexible and generalizable solutions…

Robotics · Computer Science 2026-02-11 Po-Chen Ko , Jiayuan Mao , Yu-Hsiang Fu , Hsien-Jeng Yeh , Chu-Rong Chen , Wei-Chiu Ma , Yilun Du , Shao-Hua Sun

We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Hyunsoo Lee , Minsoo Kang , Bohyung Han

Despite progress in human motion capture, existing multi-view methods often face challenges in estimating the 3D pose and shape of multiple closely interacting people. This difficulty arises from reliance on accurate 2D joint estimations,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Feichi Lu , Zijian Dong , Jie Song , Otmar Hilliges

Recent advances in video diffusion models have enabled realistic and controllable human image animation with temporal coherence. Although generating reasonable results, existing methods often overlook the need for regional supervision in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Zhongcong Xu , Chaoyue Song , Guoxian Song , Jianfeng Zhang , Jun Hao Liew , Hongyi Xu , You Xie , Linjie Luo , Guosheng Lin , Jiashi Feng , Mike Zheng Shou

Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Huanghao Yin , Shenkun Xu , Kanle Shi , Junhai Yong , Bin Wang

Current diffusion models for human image animation often struggle to maintain identity (ID) consistency, especially when the reference image and driving video differ significantly in body size or position. We introduce StableAnimator++, the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Shuyuan Tu , Zhen Xing , Xintong Han , Zhi-Qi Cheng , Qi Dai , Chong Luo , Zuxuan Wu , Yu-Gang Jiang

Neural Radiance Field (NeRF), as an implicit 3D scene representation, lacks inherent ability to accommodate changes made to the initial static scene. If objects are reconfigured, it is difficult to update the NeRF to reflect the new state…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Ziqi Lu , Jianbo Ye , Xiaohan Fei , Xiaolong Li , Jiawei Mo , Ashwin Swaminathan , Stefano Soatto

Exemplar-guided Image Editing (EIE) aims to modify a source image according to a visual reference. Existing approaches often require large-scale pre-training to learn relationships between the source and reference images, incurring high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yuke Li , Lianli Gao , Ji Zhang , Pengpeng Zeng , Lichuan Xiang , Hongkai Wen , Heng Tao Shen , Jingkuan Song

Recently, cross domain transfer has been applied for unsupervised image restoration tasks. However, directly applying existing frameworks would lead to domain-shift problems in translated images due to lack of effective supervision.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Wenchao Du , Hu Chen , Hongyu Yang

Diffusion models have recently enabled state-of-the-art reconstruction of positron emission tomography (PET) images while requiring only image training data. However, domain shift remains a key concern for clinical adoption: priors trained…

Medical Physics · Physics 2025-10-16 George Webber , Alexander Hammers , Andrew P. King , Andrew J. Reader

Diffusion models have shown great success in generating high-quality co-speech gestures for interactive humanoid robots or digital avatars from noisy input with the speech audio or text as conditions. However, they rarely focus on providing…

Human-Computer Interaction · Computer Science 2024-04-04 Zeyu Zhao , Nan Gao , Zhi Zeng , Guixuan Zhang , Jie Liu , Shuwu Zhang

Neural face avatars that are trained from multi-view data captured in camera domes can produce photo-realistic 3D reconstructions. However, at inference time, they must be driven by limited inputs such as partial views recorded by…

Computer Vision and Pattern Recognition · Computer Science 2022-03-16 Emre Aksan , Shugao Ma , Akin Caliskan , Stanislav Pidhorskyi , Alexander Richard , Shih-En Wei , Jason Saragih , Otmar Hilliges

A fundamental problem faced by object recognition systems is that objects and their features can appear in different locations, scales and orientations. Current deep learning methods attempt to achieve invariance to local translations via…

Computer Vision and Pattern Recognition · Computer Science 2017-12-12 Dimitrios C. Gklezakos , Rajesh P. N. Rao

The diversity, quantity, and quality of manipulation data are critical for training effective robot policies. However, due to hardware and physical setup constraints, collecting large-scale real-world manipulation data remains difficult to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Boyang Wang , Haoran Zhang , Shujie Zhang , Jinkun Hao , Mingda Jia , Qi Lv , Yucheng Mao , Zhaoyang Lyu , Jia Zeng , Xudong Xu , Jiangmiao Pang

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centered images, novel challenges arise with a nuanced task of "identity fine editing": precisely modifying specific features of a subject…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Haonan Lin , Mengmeng Wang , Yan Chen , Wenbin An , Yuzhe Yao , Guang Dai , Qianying Wang , Yong Liu , Jingdong Wang