中文
相关论文

相关论文: Optimizing ID Consistency in Multimodal Large Mode…

200 篇论文

Recent face reenactment works are limited by the coarse reference landmarks, leading to unsatisfactory identity preserving performance due to the distribution gap between the manipulated landmarks and those sampled from a real person. To…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Haichao Zhang , Youcheng Ben , Weixi Zhang , Tao Chen , Gang Yu , Bin Fu

Traditional methods for image-based 3D face reconstruction and facial motion retargeting fit a 3D morphable model (3DMM) to the face, which has limited modeling capacity and fail to generalize well to in-the-wild data. Use of deformation…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Bindita Chaudhuri , Noranart Vesdapunt , Linda Shapiro , Baoyuan Wang

Recent successes of deep learning-based recognition rely on maintaining the content related to the main-task label. However, how to explicitly dispel the noisy signals for better generalization in a controllable manner remains an open…

计算机视觉与模式识别 · 计算机科学 2020-06-16 Xiaofeng Liu

Despite the fact that DeepFake forgery detection algorithms have achieved impressive performance on known manipulations, they often face disastrous performance degradation when generalized to an unseen manipulation. Some recent works show…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Chuer Yu , Xuhong Zhang , Yuxuan Duan , Senbo Yan , Zonghui Wang , Yang Xiang , Shouling Ji , Wenzhi Chen

Unsupervised person re-identification (re-ID) has become an important topic due to its potential to resolve the scalability problem of supervised re-ID models. However, existing methods simply utilize pseudo labels from clustering for…

计算机视觉与模式识别 · 计算机科学 2021-04-02 Junhui Yin , Jiayan Qiu , Siqing Zhang , Jiyang Xie , Zhanyu Ma , Jun Guo

This paper discusses how ophthalmologists often rely on multimodal data to improve diagnostic accuracy. However, complete multimodal data is rare in real-world applications due to a lack of medical equipment and concerns about data privacy.…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Xinkun Wang , Yifang Wang , Senwei Liang , Feilong Tang , Chengzhi Liu , Ming Hu , Chao Hu , Junjun He , Zongyuan Ge , Imran Razzak

Prior research on out-of-distribution detection (OoDD) has primarily focused on single-modality models. Recently, with the advent of large-scale pretrained vision-language models such as CLIP, OoDD methods utilizing such multi-modal…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Jeonghyeon Kim , Sangheum Hwang

Large Vision-Language Models (LVLMs) have exhibited impressive capabilities across various visual tasks, yet they remain hindered by the persistent challenge of hallucinations. To address this critical issue, we propose Mixture of Decoding…

计算与语言 · 计算机科学 2025-06-11 Xinlong Chen , Yuanxing Zhang , Qiang Liu , Junfei Wu , Fuzheng Zhang , Tieniu Tan

Face anonymization aims to conceal identity information while preserving non-identity attributes. Mainstream diffusion models rely on inference-time interventions such as negative guidance or energy-based optimization, which are applied…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Haoxin Yang , Yihong Lin , Jingdan Kang , Xuemiao Xu , Yue Li , Cheng Xu , Shengfeng He

Generating consistent human images with controllable pose and appearance is essential for applications in virtual try on, image editing, and digital human creation. Current methods often suffer from occlusions, garment style drift, and pose…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Ziyu Shang , Haoran Liu , Rongchao Zhang , Zhiqian Wei , Tongtong Feng

Recent advances in generative modeling have enabled the generation of high-quality synthetic data that is applicable in a variety of domains, including face recognition. Here, state-of-the-art generative models typically rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Darian Tomašević , Fadi Boutros , Chenhao Lin , Naser Damer , Vitomir Štruc , Peter Peer

Although face swapping has attracted much attention in recent years, it remains a challenging problem. Existing methods leverage a large number of data samples to explore the intrinsic properties of face swapping without considering the…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Qi Li , Weining Wang , Chengzhong Xu , Zhenan Sun , Ming-Hsuan Yang

Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Guiyu Zhang , Chen Shi , Zijian Jiang , Xunzhi Xiang , Jingjing Qian , Shaoshuai Shi , Li Jiang

While text-to-image models have achieved impressive capabilities in image generation and editing, their application across various modalities often necessitates training separate models. Inspired by existing method of single image editing…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Gihyun Kwon , Jangho Park , Jong Chul Ye

In this paper, we address the problem of face aging: generating past or future facial images by incorporating age-related changes to the given face. Previous aging methods rely solely on human facial image datasets and are thus constrained…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Xiangyi Chen , Stéphane Lathuilière

Person re-identification (ReID) plays a critical role in applications such as security surveillance and criminal investigations. Most traditional image-based ReID methods face challenges including occlusions and lighting changes, while text…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Jincheng Yan , Yun Wang , Xiaoyan Luo , Yu-Wing Tai

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Jiaqi Liu , Jichao Zhang , Paolo Rota , Nicu Sebe

Artistic image stylization aims to render the content provided by text or image with the target style, where content and style decoupling is the key to achieve satisfactory results. However, current methods for content and style…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Ma Zhuoqi , Zhang Yixuan , You Zejun , Tian Long , Liu Xiyang

Face reenactment and portrait relighting are essential tasks in portrait editing, yet they are typically addressed independently, without much synergy. Most face reenactment methods prioritize motion control and multiview consistency, while…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Yizhou Zhao , Chunjiang Liu , Haoyu Chen , Bhiksha Raj , Min Xu , Tadas Baltrusaitis , Mitch Rundle , HsiangTao Wu , Kamran Ghasedi

3D editing - the task of locally modifying the geometry or appearance of a 3D asset - has wide applications in immersive content creation, digital entertainment, and AR/VR. However, unlike 2D editing, it remains challenging due to the need…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Ruihao Xia , Yang Tang , Pan Zhou