中文
相关论文

相关论文: Spatia: Video Generation with Updatable Spatial Me…

200 篇论文

Depth estimation is an important step in many computer vision problems such as 3D reconstruction, novel view synthesis, and computational photography. Most existing work focuses on depth estimation from single frames. When applied to…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Numair Khan , Eric Penner , Douglas Lanman , Lei Xiao

We capitalize on large amounts of unlabeled video in order to learn a model of scene dynamics for both video recognition tasks (e.g. action classification) and video generation tasks (e.g. future prediction). We propose a generative…

计算机视觉与模式识别 · 计算机科学 2016-10-27 Carl Vondrick , Hamed Pirsiavash , Antonio Torralba

We address the problem of generating video features for action recognition. The spatial pyramid and its variants have been very popular feature models due to their success in balancing spatial location encoding and spatial invariance.…

计算机视觉与模式识别 · 计算机科学 2015-10-16 Zhenzhong Lan , Alexander G. Hauptmann

Video-to-Video synthesis (Vid2Vid) has achieved remarkable results in generating a photo-realistic video from a sequence of semantic maps. However, this pipeline suffers from high computational cost and long inference latency, which largely…

计算机视觉与模式识别 · 计算机科学 2022-07-12 Long Zhuo , Guangcong Wang , Shikai Li , Wayne Wu , Ziwei Liu

Integrating motion into static images not only enhances visual expressiveness but also creates a sense of immersion and temporal depth, establishing it as a longstanding and impactful theme in artistic expression. Fluid elements such as…

图形学 · 计算机科学 2025-10-21 Hao Jin , Haoran Xie

Stochastic video generation is particularly challenging when the camera is mounted on a moving platform, as camera motion interacts with observed image pixels, creating complex spatio-temporal dynamics and making the problem partially…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Meenakshi Sarkar , Devansh Bhardwaj , Debasish Ghose

Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jan Ackermann , Shengqu Cai , Boyang Deng , Zhengfei Kuang , Songyou Peng , Gordon Wetzstein

We present Split-then-Merge (StM), a novel framework designed to enhance control in generative video composition and address its data scarcity problem. Unlike conventional methods relying on annotated datasets or handcrafted rules, StM…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Ozgur Kara , Yujia Chen , Ming-Hsuan Yang , James M. Rehg , Wen-Sheng Chu , Du Tran

The framework of dominant learned video compression methods is usually composed of motion prediction modules as well as motion vector and residual image compression modules, suffering from its complex structure and error propagation…

图像与视频处理 · 电气工程与系统科学 2021-04-14 Zhenhong Sun , Zhiyu Tan , Xiuyu Sun , Fangyi Zhang , Dongyang Li , Yichen Qian , Hao Li

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Chenhan Jiang , Yu Chen , Qingwen Zhang , Jifei Song , Songcen Xu , Dit-Yan Yeung , Jiankang Deng

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced…

计算机视觉与模式识别 · 计算机科学 2018-03-26 Jiawei He , Andreas Lehrmann , Joseph Marino , Greg Mori , Leonid Sigal

Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by-detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in…

计算机视觉与模式识别 · 计算机科学 2023-09-04 Ali Athar , Sabarinath Mahadevan , Aljoša Ošep , Laura Leal-Taixé , Bastian Leibe

Existing person video generation methods either lack the flexibility in controlling both the appearance and motion, or fail to preserve detailed appearance and temporal consistency. In this paper, we tackle the problem of motion transfer…

计算机视觉与模式识别 · 计算机科学 2019-08-13 Kun Cheng , Hao-Zhi Huang , Chun Yuan , Lingyiqing Zhou , Wei Liu

Text to video generation has emerged as a critical frontier in generative artificial intelligence, yet existing approaches struggle with maintaining temporal consistency, compositional understanding, and fine grained control over visual…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Piyushkumar Patel

Recent advances in imitation learning have shown significant promise for robotic control and embodied intelligence. However, achieving robust generalization across diverse mounted camera observations remains a critical challenge. In this…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Travis Davies , Jiahuan Yan , Xiang Chen , Yu Tian , Yueting Zhuang , Yiqi Huang , Luhui Hu

Recently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby improving the processing efficiency of video data while…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Jiahuan Zhou , Kai Zhu , Zhenyu Cui , Zichen Liu , Xu Zou , Gang Hua

As Artificial Intelligence Generated Content (AIGC) advances, a variety of methods have been developed to generate text, images, videos, and 3D objects from single or multimodal inputs, contributing efforts to emulate human-like cognitive…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Yiying Yang , Fukun Yin , Jiayuan Fan , Xin Chen , Wanzhang Li , Gang Yu

Although powerful for image generation, consistent and controllable video is a longstanding problem for diffusion models. Video models require extensive training and computational resources, leading to high costs and large environmental…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Muhammad Haaris Khan , Hadrien Reynaud , Bernhard Kainz

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

Generative models that can model and predict sequences of future events can, in principle, learn to capture complex real-world phenomena, such as physical interactions. However, a central challenge in video prediction is that the future is…

计算机视觉与模式识别 · 计算机科学 2020-02-13 Manoj Kumar , Mohammad Babaeizadeh , Dumitru Erhan , Chelsea Finn , Sergey Levine , Laurent Dinh , Durk Kingma