中文
相关论文

相关论文: PoseGuard: Pose-Guided Generation with Safety Guar…

200 篇论文

Producing prompt-faithful videos that preserve a user-specified identity remains challenging: models need to extrapolate facial dynamics from sparse reference while balancing the tension between identity preservation and motion naturalness.…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yixuan Lai , He Wang , Kun Zhou , Tianjia Shao

The widespread use of face recognition technology has given rise to privacy concerns, as many individuals are worried about the collection and utilization of their facial data. To address these concerns, researchers are actively exploring…

密码学与安全 · 计算机科学 2023-10-26 Zhiling Zhang , Jie Zhang , Kui Zhang , Wenbo Zhou , Weiming Zhang , Nenghai Yu

The misuse of deep learning-based facial manipulation poses a significant threat to civil rights. To prevent this fraud at its source, proactive defense has been proposed to disrupt the manipulation process by adding invisible adversarial…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Zuomin Qu , Wei Lu , Xiangyang Luo , Qian Wang , Xiaochun Cao

Human pose transfer, which aims at transferring the appearance of a given person to a target pose, is very challenging and important in many applications. Previous work ignores the guidance of pose features or only uses local attention…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Kun Li , Jinsong Zhang , Yebin Liu , Yu-Kun Lai , Qionghai Dai

Existing 3D human pose estimators suffer poor generalization performance to new datasets, largely due to the limited diversity of 2D-3D pose pairs in the training data. To address this problem, we present PoseAug, a new auto-augmentation…

计算机视觉与模式识别 · 计算机科学 2021-05-07 Kehong Gong , Jianfeng Zhang , Jiashi Feng

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Hongxiang Li , Yaowei Li , Yuhang Yang , Junjie Cao , Zhihong Zhu , Xuxin Cheng , Long Chen

Grasping user-specified objects is crucial for robotic assistants; however, most current 6-DoF grasp detection methods are object-agnostic, making it challenging to grasp specific targets from a scene. To achieve that, we present GoalGrasp,…

机器人学 · 计算机科学 2025-04-23 Shun Gui , Kai Gui , Yan Luximon

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and biased behavior. This, in turn, hampers their applicability…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Samuele Poppi , Tobia Poppi , Federico Cocchi , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

In this paper, we address the problem of generating person images conditioned on both pose and appearance information. Specifically, given an image xa of a person and a target pose P(xb), extracted from a different image xb, we synthesize a…

计算机视觉与模式识别 · 计算机科学 2019-10-15 Aliaksandr Siarohin , Stéphane Lathuilière , Enver Sangineto , Nicu Sebe

Estimating the 6D pose of arbitrary unseen objects from a single reference image is critical for robotics operating in the long-tail of real-world instances. However, this setting is notoriously challenging: 3D models are rarely available,…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Zheng Geng , Nan Wang , Shaocong Xu , Chongjie Ye , Bohan Li , Zhaoxi Chen , Sida Peng , Hao Zhao

In this paper we address the problem of generating person images conditioned on a given pose. Specifically, given an image of a person and a target pose, we synthesize a new image of that person in the novel pose. In order to deal with…

计算机视觉与模式识别 · 计算机科学 2018-04-09 Aliaksandr Siarohin , Enver Sangineto , Stephane Lathuiliere , Nicu Sebe

Video diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Jiahao Wang , Hualian Sheng , Sijia Cai , Weizhan Zhang , Caixia Yan , Yachuang Feng , Bing Deng , Jieping Ye

Image-to-video (I2V) generation aims to create a video sequence from a single image, which requires high temporal coherence and visual fidelity. However, existing approaches suffer from inconsistency of character appearances and poor…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Bingwen Zhu , Fanyi Wang , Tianyi Lu , Peng Liu , Jingwen Su , Jinxiu Liu , Yanhao Zhang , Zuxuan Wu , Guo-Jun Qi , Yu-Gang Jiang

Grasp detection of novel objects in unstructured environments is a key capability in robotic manipulation. For 2D grasp detection problems where grasps are assumed to lie in the plane, it is common to design a fully convolutional neural…

机器人学 · 计算机科学 2022-04-05 Andreas ten Pas , Colin Keil , Robert Platt

Text-to-image models are increasingly popular and impactful, yet concerns regarding their safety and fairness remain. This study investigates the ability of ten popular Stable Diffusion models to generate harmful images, including NSFW,…

计算机与社会 · 计算机科学 2025-08-29 Matthias Schneider , Thilo Hagendorff

Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Cheeun Hong , German Barquero , Fadime Sener , Markos Georgopoulos , Edgar Schönfeld , Stefan Popov , Yuming Du , Oscar Mañas , Albert Pumarola

Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a single modality of control signals and operate in isolation,…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yiheng Li , Ruibing Hou , Hong Chang , Shiguang Shan , Xilin Chen

Fine-tuning text-to-image diffusion models is widely used for personalization and adaptation for new domains. In this paper, we identify a critical vulnerability of fine-tuning: safety alignment methods designed to filter harmful content…

人工智能 · 计算机科学 2024-12-03 Sanghyun Kim , Moonseok Choi , Jinwoo Shin , Juho Lee

The success of deep learning based face recognition systems has given rise to serious privacy concerns due to their ability to enable unauthorized tracking of users in the digital world. Existing methods for enhancing privacy fail to…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Fahad Shamshad , Muzammal Naseer , Karthik Nandakumar