中文
相关论文

相关论文: ActCam: Zero-Shot Joint Camera and 3D Motion Contr…

200 篇论文

The risk of unauthorized remote access of streaming video from networked cameras underlines the need for stronger privacy safeguards. We propose a lens-free coded aperture camera system for human action recognition that is…

计算机视觉与模式识别 · 计算机科学 2019-04-18 Zihao W. Wang , Vibhav Vineet , Francesco Pittaluga , Sudipta Sinha , Oliver Cossairt , Sing Bing Kang

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

Motion-controllable image animation is a fundamental task with a wide range of potential applications. Recent works have made progress in controlling camera or object motion via various motion representations, while they still struggle to…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yingjie Chen , Yifang Men , Yuan Yao , Miaomiao Cui , Liefeng Bo

Diffusion models have made significant advances in generating high-quality images, but their application to video generation has remained challenging due to the complexity of temporal motion. Zero-shot video editing offers a solution by…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Xirui Li , Chao Ma , Xiaokang Yang , Ming-Hsuan Yang

Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Youyuan Zhang , Xuan Ju , James J. Clark

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Zeqi Xiao , Wenqi Ouyang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

Current diffusion-based text-to-video methods are limited to producing short video clips of a single shot and lack the capability to generate multi-shot videos with discrete transitions where the same character performs distinct activities…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Ozgur Kara , Krishna Kumar Singh , Feng Liu , Duygu Ceylan , James M. Rehg , Tobias Hinz

In video understanding tasks, particularly those involving human motion, synthetic data generation often suffers from uncanny features, diminishing its effectiveness for training. Tasks such as sign language translation, gesture…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Vaclav Knapp , Matyas Bohacek

World models based on video generation demonstrate remarkable potential for simulating interactive environments but face persistent difficulties in two key areas: maintaining long-term content consistency when scenes are revisited and…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Tianxing Xu , Zixuan Wang , Guangyuan Wang , Li Hu , Zhongyi Zhang , Peng Zhang , Bang Zhang , Song-Hai Zhang

DuctTake is a system designed to enable practical compositing of multiple takes of a scene into a single video. Current industry solutions are based around object segmentation, a hard problem that requires extensive manual input and…

计算机视觉与模式识别 · 计算机科学 2021-01-14 Jan Rueegg , Oliver Wang , Aljoscha Smolic , Markus Gross

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

We present a target-aware video diffusion model that generates videos from an input image, in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask, and the action…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Taeksoo Kim , Hanbyul Joo

Purpose: In order to produce a surgical gesture recognition system that can support a wide variety of procedures, either a very large annotated dataset must be acquired, or fitted models must generalize to new labels (so called "zero-shot"…

计算机视觉与模式识别 · 计算机科学 2024-08-23 Mingxing Rao , Yinhong Qin , Soheil Kolouri , Jie Ying Wu , Daniel Moyer

Recent advancements in text-to-image generation have enabled significant progress in zero-shot 3D shape generation. This is achieved by score distillation, a methodology that uses pre-trained text-to-image diffusion models to optimize the…

计算机视觉与模式识别 · 计算机科学 2023-05-29 Zhenzhen Weng , Zeyu Wang , Serena Yeung

We propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model. We condition its UNet denoiser on the camera trajectory, using four…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Stefan Popov , Amit Raj , Michael Krainin , Yuanzhen Li , William T. Freeman , Michael Rubinstein

Attributes such as style, fine-grained text, and trajectory are specific conditions for describing motion. However, existing methods often lack precise user control over motion attributes and suffer from limited generalizability to unseen…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Mingjie Wei , Xuemei Xie , Guangming Shi

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Zhenhui Ye , Tianyun Zhong , Yi Ren , Jiaqi Yang , Weichuang Li , Jiawei Huang , Ziyue Jiang , Jinzheng He , Rongjie Huang , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

In this work, we present CineMaster, a novel framework for 3D-aware and controllable text-to-video generation. Our goal is to empower users with comparable controllability as professional film directors: precise placement of objects within…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Qinghe Wang , Yawen Luo , Xiaoyu Shi , Xu Jia , Huchuan Lu , Tianfan Xue , Xintao Wang , Pengfei Wan , Di Zhang , Kun Gai

Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, existing diffusion-based video editing approaches lack the ability to offer precise control over generated content that…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Paul Couairon , Clément Rambour , Jean-Emmanuel Haugeard , Nicolas Thome
‹ 上一页 1 8 9 10 下一页 ›