中文
相关论文

相关论文: ActAnywhere: Subject-Aware Video Background Genera…

200 篇论文

Current datasets for action recognition tasks face limitations stemming from traditional collection and generation methods, including the constrained range of action classes, absence of multi-viewpoint recordings, limited diversity, poor…

计算机视觉与模式识别 · 计算机科学 2024-01-25 Xingyu Song , Zhan Li , Shi Chen , Kazuyuki Demachi

Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Ziyan Guo , Haoxuan Qu , Hossein Rahmani , Dewen Soh , Ping Hu , Qiuhong Ke , Jun Liu

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying 3D scene that gives…

计算机视觉与模式识别 · 计算机科学 2021-03-26 Paul Henderson , Christoph H. Lampert

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Xiaolin Hong , Hongwei Yi , Fazhi He , Qiong Cao

Video prediction aims to generate realistic future frames by learning dynamic visual patterns. One fundamental challenge is to deal with future uncertainty: How should a model behave when there are multiple correct, equally probable future?…

计算机视觉与模式识别 · 计算机科学 2018-07-10 Yunseok Jang , Gunhee Kim , Yale Song

Computer vision is increasingly effective at segmenting objects in images and videos; however, scene effects related to the objects -- shadows, reflections, generated smoke, etc -- are typically overlooked. Identifying such scene effects…

计算机视觉与模式识别 · 计算机科学 2021-10-04 Erika Lu , Forrester Cole , Tali Dekel , Andrew Zisserman , William T. Freeman , Michael Rubinstein

Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plausible locations and scales without interpenetration, aligning…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Qingyang Liu , Bingjie Gao , Weiheng Huang , Jun Zhang , Zhongqian Sun , Yang Wei , Fengrui Liu , Zelin Peng , Qianli Ma , Shuai Yang , Zhaohe Liao , Haonan Zhao , Li Niu

Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical frames and current…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Haoyu Wu , Jiwen Yu , Yingtian Zou , Xihui Liu

This study explores the potential of open-source video conditional generation models as encoders for downstream tasks, focusing on instance segmentation using the BAIR Robot Pushing Dataset. The researchers propose using video prediction…

计算机视觉与模式识别 · 计算机科学 2024-05-28 James Maier , Nishanth Mohankumar

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Mingyu Dong , Chong Xia , Mingyuan Jia , Weichen Lyu , Long Xu , Zheng Zhu , Yueqi Duan

Current video generation models cannot simulate physical consequences of 3D actions like forces and robotic manipulations, as they lack structural understanding of how actions affect 3D scenes. We present RealWonder, the first real-time…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Wei Liu , Ziyu Chen , Zizhang Li , Yue Wang , Hong-Xing Yu , Jiajun Wu

Auxiliary-free human video matting methods, which rely solely on input frames, often struggle with complex or ambiguous backgrounds. To address this, we propose MatAnyone, a robust framework tailored for target-assigned video matting.…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Peiqing Yang , Shangchen Zhou , Jixin Zhao , Qingyi Tao , Chen Change Loy

With the rapid advancement of text-conditioned Video Generation Models (VGMs), the quality of generated videos has significantly improved, bringing these models closer to functioning as ``*world simulators*'' and making real-world-level…

人工智能 · 计算机科学 2025-04-22 Haotong Yang , Qingyuan Zheng , Yunjian Gao , Yongkun Yang , Yangbo He , Zhouchen Lin , Muhan Zhang

Existing subject-driven text-to-image generation models suffer from tedious fine-tuning steps and struggle to maintain both text-image alignment and subject fidelity. For generating compositional subjects, it often encounters problems such…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Shengyuan Liu , Bo Wang , Ye Ma , Te Yang , Xipeng Cao , Quan Chen , Han Li , Di Dong , Peng Jiang

Video generation aims to produce temporally coherent sequences of visual frames, representing a pivotal advancement in Artificial Intelligence Generated Content (AIGC). Compared to static image generation, video generation poses unique…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Zhiyu Yin , Kehai Chen , Xuefeng Bai , Ruili Jiang , Juntao Li , Hongdong Li , Jin Liu , Yang Xiang , Jun Yu , Min Zhang

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Xindi Wu , Despoina Paschalidou , Jun Gao , Antonio Torralba , Laura Leal-Taixé , Olga Russakovsky , Sanja Fidler , Jonathan Lorraine

Recent video action recognition methods have shown excellent performance by adapting large-scale pre-trained language-image models to the video domain. However, language models contain rich common sense priors - the scene contexts that…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Xiaodan Hu , Chuhang Zou , Suchen Wang , Jaechul Kim , Narendra Ahuja

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Ziyang Chen , Prem Seetharaman , Bryan Russell , Oriol Nieto , David Bourgin , Andrew Owens , Justin Salamon

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and…

计算机视觉与模式识别 · 计算机科学 2026-02-26 JiaKui Hu , Jialun Liu , Liying Yang , Xinliang Zhang , Kaiwen Li , Shuang Zeng , Yuanwei Li , Haibin Huang , Chi Zhang , Yanye Lu