中文
相关论文

相关论文: 3D Video Loops from Asynchronous Input

200 篇论文

Generic motion understanding from video involves not only tracking objects, but also perceiving how their surfaces deform and move. This information is useful to make inferences about 3D shape, physical properties and object interactions.…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Carl Doersch , Ankush Gupta , Larisa Markeeva , Adrià Recasens , Lucas Smaira , Yusuf Aytar , João Carreira , Andrew Zisserman , Yi Yang

Mask-free video object insertion has emerged as a challenging task, requiring harmonious integration of reference objects into source videos. However, existing methods struggle when references exhibit severe stylistic domain gaps with the…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Xiao Cao , Yansong Qu , Xiangzhen , Chang , Wen Xiao , Jiakui Hu , Heyuan Li , Jialun Liu , Zhiyong Huang , Xuelong Li

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions…

机器人学 · 计算机科学 2026-01-01 Karthik Dharmarajan , Wenlong Huang , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens of minutes has become both feasible and increasingly…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Yuxiao Chen , Jue Wang , Zhikang Zhang , Jingru Yi , Xu Zhang , Yang Zou , Zhaowei Cai , Jianbo Yuan , Xinyu Li , Hao Yang , Davide Modolo

We present \emph{Video-in-the-Loop} (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first \emph{localizing} question-relevant interval(s) with a low-fps skim and then \emph{answering} via span-aware…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Chendong Wang , Donglin Bai , Yifan Yang , Xiao Jin , Anlan Zhang , Rui Wang , Shiqi Jiang , Yuqing Yang , Hao Wu , Qi Dai , Chong Luo , Ting Cao , Lili Qiu , Suman Banerjee

We present an algorithm for generating novel views at arbitrary viewpoints and any input time step given a monocular video of a dynamic scene. Our work builds upon recent advances in neural implicit representation and uses continuous and…

计算机视觉与模式识别 · 计算机科学 2021-05-14 Chen Gao , Ayush Saraf , Johannes Kopf , Jia-Bin Huang

Generating synthetic multi-view images from a text prompt is an essential bridge to generating synthetic 3D assets. In this work, we introduce RapidMV, a novel text-to-multi-view generative model that can produce 32 multi-view synthetic…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Seungwook Kim , Yichun Shi , Kejie Li , Minsu Cho , Peng Wang

In this paper we present a framework for the rendering of dynamic 3D virtual environments which can be integrated in the development of videogames. It includes methods to manage sounds and particle effects, paged static geometries, the…

图形学 · 计算机科学 2013-06-06 Salvatore Catanese , Emilio Ferrara , Giacomo Fiumara , Francesco Pagano

Novel view synthesis from a single image has recently attracted a lot of attention, and it has been primarily advanced by 3D deep learning and rendering techniques. However, most work is still limited by synthesizing new views within…

计算机视觉与模式识别 · 计算机科学 2022-03-18 Xuanchi Ren , Xiaolong Wang

Virtual 3D meetings offer the potential to enhance copresence, increase engagement and thus improve effectiveness of remote meetings compared to standard 2D video calls. However, representing people in 3D meetings remains a challenge;…

We present an algorithm for producing a seamless animated loop from a single image. The algorithm detects periodic structures, such as the windows of a building or the steps of a staircase, and generates a non-trivial displacement vector…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Tavi Halperin , Hanit Hakim , Orestis Vantzos , Gershon Hochman , Netai Benaim , Lior Sassy , Michael Kupchik , Ofir Bibi , Ohad Fried

Video representation is a long-standing problem that is crucial for various down-stream tasks, such as tracking,depth prediction,segmentation,view synthesis,and editing. However, current methods either struggle to model complex motions due…

计算机视觉与模式识别 · 计算机科学 2024-06-27 Yang-Tian Sun , Yi-Hua Huang , Lin Ma , Xiaoyang Lyu , Yan-Pei Cao , Xiaojuan Qi

Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV, E-NeRV). While achieving promising results, existing INR-based methods…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Bo He , Xitong Yang , Hanyu Wang , Zuxuan Wu , Hao Chen , Shuaiyi Huang , Yixuan Ren , Ser-Nam Lim , Abhinav Shrivastava

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Zizhong Li , Haopeng Zhang , Jiawei Zhang

We present a fully automatic approach to real-time 3D face reconstruction from monocular in-the-wild videos. With the use of a cascaded-regressor based face tracking and a 3D Morphable Face Model shape fitting, we obtain a semi-dense 3D…

计算机视觉与模式识别 · 计算机科学 2017-08-28 Patrik Huber , Philipp Kopp , Matthias Rätsch , William Christmas , Josef Kittler

Dynamic novel view synthesis aims to capture the temporal evolution of visual content within videos. Existing methods struggle to distinguishing between motion and structure, particularly in scenarios where camera poses are either unknown…

计算机视觉与模式识别 · 计算机科学 2024-01-12 Chaoyang Wang , Peiye Zhuang , Aliaksandr Siarohin , Junli Cao , Guocheng Qian , Hsin-Ying Lee , Sergey Tulyakov

Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or external visual-language models (VLMs), which face issues like…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Yuan Xie , Tianshui Chen , Zheng Ge , Lionel Ni

Today, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross-camera streams remains…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Shaowei Liu , David Yifan Yao , Saurabh Gupta , Shenlong Wang

To facilitate depth-based 3D action recognition, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation. With 3D space voxelization, the key idea of 3DV is to encode 3D motion information within depth video into a regular…

计算机视觉与模式识别 · 计算机科学 2020-05-13 Yancheng Wang , Yang Xiao , Fu Xiong , Wenxiang Jiang , Zhiguo Cao , Joey Tianyi Zhou , Junsong Yuan

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained…