English
Related papers

Related papers: Information-Regularized Constrained Inversion for …

200 papers

Due to lack of fully publicly available text-to-video models, current video editing methods tend to build on pre-trained text-to-image generation models, however, they still face grand challenges in dealing with the local editing of video…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Deyin Liu , Lin Yuanbo Wu , Xianghua Xie

Recent advances in the field of generative models and in particular generative adversarial networks (GANs) have lead to substantial progress for controlled image editing, especially compared with the pre-deep learning era. Despite their…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Gwilherm Lesné , Yann Gousseau , Saïd Ladjal , Alasdair Newson

We present a novel image inversion framework and a training pipeline to achieve high-fidelity image inversion with high-quality attribute editing. Inverting real images into StyleGAN's latent space is an extensively studied problem, yet the…

Computer Vision and Pattern Recognition · Computer Science 2023-01-02 Hamza Pehlivan , Yusuf Dalva , Aysegul Dundar

Recent advancements in large-scale text-to-image diffusion models have enabled many applications in image editing. However, none of these methods have been able to edit the layout of single existing images. To address this gap, we propose…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Zhiyuan Zhang , Zhitong Huang , Jing Liao

Current text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for…

Graphics · Computer Science 2025-05-01 Duotun Wang , Hengyu Meng , Zeyu Cai , Zhijing Shao , Qianxi Liu , Lin Wang , Mingming Fan , Xiaohang Zhan , Zeyu Wang

Many imaging technologies rely on tomographic reconstruction, which requires solving a multidimensional inverse problem given a finite number of projections. Backprojection is a popular class of algorithm for tomographic reconstruction,…

Image and Video Processing · Electrical Eng. & Systems 2020-06-03 Xueqing Liu , Paul Sajda

Recent works for face editing usually manipulate the latent space of StyleGAN via the linear semantic directions. However, they usually suffer from the entanglement of facial attributes, need to tune the optimal editing strength, and are…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Zhizhong Huang , Siteng Ma , Junping Zhang , Hongming Shan

Unsupervised text attribute transfer automatically transforms a text to alter a specific attribute (e.g. sentiment) without using any parallel data, while simultaneously preserving its attribute-independent content. The dominant approaches…

Computation and Language · Computer Science 2019-12-13 Ke Wang , Hang Hua , Xiaojun Wan

This paper proposes an efficient 3D avatar coding framework that leverages compact human priors and canonical-to-target transformation to enable high-quality 3D human avatar video compression at ultra-low bit rates. The framework begins by…

Image and Video Processing · Electrical Eng. & Systems 2025-10-14 Shanzhi Yin , Bolin Chen , Xinju Wu , Ru-Ling Liao , Jie Chen , Shiqi Wang , Yan Ye

A framework for unsupervised group activity analysis from a single video is here presented. Our working hypothesis is that human actions lie on a union of low-dimensional subspaces, and thus can be efficiently modeled as sparse linear…

Computer Vision and Pattern Recognition · Computer Science 2012-08-28 Zhongwei Tang , Alexey Castrodad , Mariano Tepper , Guillermo Sapiro

Sparse-view computed tomography (CT) enables fast and low-dose CT imaging, an essential feature for patient-save medical imaging and rapid non-destructive testing. In sparse-view CT, only a few projection views are acquired, causing…

Image and Video Processing · Electrical Eng. & Systems 2024-02-28 Nadja Gruber , Johannes Schwab , Elke Gizewski , Markus Haltmeier

Sparse auto-encoders (SAEs) have re-emerged as a prominent method for mechanistic interpretability, yet they face two significant challenges: the non-smoothness of the $L_1$ penalty, which hinders reconstruction and scalability, and a lack…

Artificial Intelligence · Computer Science 2026-05-19 Ouns El Harzli , Hugo Wallner , Yoonsoo Nam , Haixuan Xavier Tao

The problem of modeling an animatable 3D human head avatar under light-weight setups is of significant importance but has not been well solved. Existing 3D representations either perform well in the realism of portrait images synthesis or…

Computer Vision and Pattern Recognition · Computer Science 2023-10-02 Xiaochen Zhao , Lizhen Wang , Jingxiang Sun , Hongwen Zhang , Jinli Suo , Yebin Liu

While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entanglement that cause attribute leakage and temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Fei Shen , Weihao Xu , Rui Yan , Dong Zhang , Xiangbo Shu , Jinhui Tang , Maocheng Zhao

Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Xulu Zhang , Xiao-Yong Wei , Jinlin Wu , Tianyi Zhang , Zhaoxiang Zhang , Zhen Lei , Qing Li

Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information is implicitly entangled across their dense hidden states, providing no explicit handle for…

Machine Learning · Computer Science 2026-05-12 Mingxu Zhang , Yuhan Li , Lujundong Li , Dazhong Shen , Hui Xiong , Ying Sun

Modeling animatable human avatars from monocular or multi-view videos has been widely studied, with recent approaches leveraging neural radiance fields (NeRFs) or 3D Gaussian Splatting (3DGS) achieving impressive results in novel-view and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Yahui Li , Zhi Zeng , Liming Pang , Guixuan Zhang , Shuwu Zhang

Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limits scalability. We introduce FFAvatar, a generalizable feed-forward framework that…

Graphics · Computer Science 2026-05-18 Thuan Hoang Nguyen , Jiahao Luo , Yinyu Nie , Hao Li , Gordon Guocheng Qian , Jian Wang

Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Yue Han , Junwei Zhu , Keke He , Xu Chen , Yanhao Ge , Wei Li , Xiangtai Li , Jiangning Zhang , Chengjie Wang , Yong Liu

Face reenactment and portrait relighting are essential tasks in portrait editing, yet they are typically addressed independently, without much synergy. Most face reenactment methods prioritize motion control and multiview consistency, while…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Yizhou Zhao , Chunjiang Liu , Haoyu Chen , Bhiksha Raj , Min Xu , Tadas Baltrusaitis , Mitch Rundle , HsiangTao Wu , Kamran Ghasedi