中文
相关论文

相关论文: DicFace: Dirichlet-Constrained Variational Codeboo…

200 篇论文

In recent years, the task of video prediction-forecasting future video given past video frames-has attracted attention in the research community. In this paper we propose a novel approach to this problem with Vector Quantized Variational…

计算机视觉与模式识别 · 计算机科学 2021-03-03 Jacob Walker , Ali Razavi , Aäron van den Oord

This work introduces \textbf{VideoMark}, a distortion-free robust watermarking framework for video diffusion models. As diffusion models excel in generating realistic videos, reliable content attribution is increasingly critical. However,…

密码学与安全 · 计算机科学 2025-11-18 Xuming Hu , Hanqian Li , Jungang Li , Yu Huang , Shuliang Liu , Qi Zheng , Junhao Chen , Aiwei Liu

Continual learning for video--language understanding is increasingly important as models face non-stationary data, domains, and query styles, yet prevailing solutions blur what should stay stable versus what should adapt, rely on static…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Mengzhu Xu , Hanzhi Liu , Ningkang Peng , Qianyu Chen , Canran Xiao

Constructing a compressed latent space through a variational autoencoder (VAE) is the key for efficient 3D diffusion models. This paper introduces COD-VAE that encodes 3D shapes into a COmpact set of 1D latent vectors without sacrificing…

计算机视觉与模式识别 · 计算机科学 2025-07-29 In Cho , Youngbeom Yoo , Subin Jeon , Seon Joo Kim

Face editing methods, essential for tasks like virtual avatars, digital human synthesis and identity preservation, have traditionally been built upon GAN-based techniques, while recent focus has shifted to diffusion-based models due to…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Mengting Wei , Tuomas Varanka , Yante Li , Xingxun Jiang , Huai-Qian Khor , Guoying Zhao

Facial appearance editing is crucial for digital avatars, AR/VR, and personalized content creation, driving realistic user experiences. However, preserving identity with generative models is challenging, especially in scenarios with limited…

计算机视觉与模式识别 · 计算机科学 2025-03-10 MD Wahiduzzaman Khan , Mingshan Jia , Xiaolin Zhang , En Yu , Caifeng Shan , Kaska Musial-Gabrys

Vector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis, which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Baoquan Zhang , Huaibin Wang , Luo Chuyao , Xutao Li , Liang Guotao , Yunming Ye , Xiaochen Qi , Yao He

Vision Transformers (ViTs) have revolutionized large-scale visual modeling, yet remain underexplored in face recognition (FR) where CNNs still dominate. We identify a critical bottleneck: CNN-inspired training paradigms fail to unlock ViT's…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Jinghan You , Shanglin Li , Yuanrui Sun , Jiangchuan Wei , Mingyu Guo , Chao Feng , Jiao Ran

Leveraging the generative ability of image diffusion models offers great potential for zero-shot video-to-video translation. The key lies in how to maintain temporal consistency across generated video frames by image diffusion models.…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Yuxiang Bao , Di Qiu , Guoliang Kang , Baochang Zhang , Bo Jin , Kaiye Wang , Pengfei Yan

tmospheric turbulence presents a significant challenge in long-range imaging. Current restoration algorithms often struggle with temporal inconsistency, as well as limited generalization ability across varying turbulence levels and scene…

图像与视频处理 · 电气工程与系统科学 2023-12-11 Haoming Cai , Jingxi Chen , Brandon Y. Feng , Weiyun Jiang , Mingyang Xie , Kevin Zhang , Ashok Veeraraghavan , Christopher Metzler

Video restoration aims at restoring multiple high-quality frames from multiple low-quality frames. Existing video restoration methods generally fall into two extreme cases, i.e., they either restore all frames in parallel or restore the…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Jingyun Liang , Yuchen Fan , Xiaoyu Xiang , Rakesh Ranjan , Eddy Ilg , Simon Green , Jiezhang Cao , Kai Zhang , Radu Timofte , Luc Van Gool

Large Language Models (LLMs) have demonstrated remarkable success across diverse fields, establishing a powerful paradigm for complex information processing. This has inspired the integration of speech into LLM frameworks, often by…

音频与语音处理 · 电气工程与系统科学 2025-12-30 Xiangyu Zhang , Fuming Fang , Peng Gao , Bin Qin , Beena Ahmed , Julien Epps

Free-viewpoint video conferencing allows a participant to observe the remote 3D scene from any freely chosen viewpoint. An intermediate virtual viewpoint image is commonly synthesized using two pairs of transmitted texture and depth maps…

多媒体 · 计算机科学 2025-05-06 Bruno Macchiavello , Camilo Dorea , Edson M. Hung , Gene Cheung , Wai-tian Tan

Ultra-high-definition (UHD) image restoration often faces computational bottlenecks and information loss due to its extremely high resolution. Existing studies based on Variational Autoencoders (VAE) improve efficiency by transferring the…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Yidi Liu , Dong Li , Yuxin Ma , Jie Huang , Wenlong Zhang , Xueyang Fu , Zheng-jun Zha

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Yuying Ge , Yizhuo Li , Yixiao Ge , Ying Shan

Human motion data is inherently rich and complex, containing both semantic content and subtle stylistic features that are challenging to model. We propose a novel method for effective disentanglement of the style and content in human motion…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Fatemeh Zargarbashi , Dhruv Agrawal , Jakob Buhmann , Martin Guay , Stelian Coros , Robert W. Sumner

Latent Video Diffusion Models (LVDMs) rely on Variational Autoencoders (VAEs) to compress videos into compact latent representations. For continuous Variational Autoencoders (VAEs), achieving higher compression rates is desirable; yet, the…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yubo Dong , Linchao Zhu

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Yao-Chih Lee , Ji-Ze Genevieve Jang , Yi-Ting Chen , Elizabeth Qiu , Jia-Bin Huang

Deep generative models can synthesize photorealistic images of human faces with novel identities. However, a key challenge to the wide applicability of such techniques is to provide independent control over semantically meaningful…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Marcel C. Bühler , Abhimitra Meka , Gengyan Li , Thabo Beeler , Otmar Hilliges

Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Mallikarjun B. R. , Fei Yin , Vikram Voleti , Nikita Drobyshev , Maksim Lapin , Aaryaman Vasishta , Varun Jampani