English
Related papers

Related papers: MoCoGAN: Decomposing Motion and Content for Video …

200 papers

A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yu Yuan , Xijun Wang , Tharindu Wickremasinghe , Zeeshan Nadir , Bole Ma , Stanley H. Chan

A class of recent approaches for generating images, called Generative Adversarial Networks (GAN), have been used to generate impressively realistic images of objects, bedrooms, handwritten digits and a variety of other image modalities.…

Computer Vision and Pattern Recognition · Computer Science 2017-06-08 Swaminathan Gurumurthy , Ravi Kiran Sarvadevabhatla , Venkatesh Babu Radhakrishnan

We are interested in learning visual representations which allow for 3D manipulations of visual objects based on a single 2D image. We cast this into an image-to-image transformation task, and propose Iterative Generative Adversarial…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Ysbrand Galama , Thomas Mensink

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

In this paper we formulate structure from motion as a learning problem. We train a convolutional network end-to-end to compute depth and camera motion from successive, unconstrained image pairs. The architecture is composed of multiple…

Computer Vision and Pattern Recognition · Computer Science 2018-01-18 Benjamin Ummenhofer , Huizhong Zhou , Jonas Uhrig , Nikolaus Mayer , Eddy Ilg , Alexey Dosovitskiy , Thomas Brox

Gatys et al. (2015) showed that optimizing pixels to match features in a convolutional network with respect reference image features is a way to render images of high visual quality. We show that unrolling this gradient-based optimization…

Machine Learning · Computer Science 2016-12-14 Daniel Jiwoong Im , Chris Dongjoo Kim , Hui Jiang , Roland Memisevic

Video segmentation consists of a frame-by-frame selection process of meaningful areas related to foreground moving objects. Some applications include traffic monitoring, human tracking, action recognition, efficient video surveillance, and…

Computer Vision and Pattern Recognition · Computer Science 2022-12-22 Daniel F. S. Santos , Rafael G. Pires , Danilo Colombo , João P. Papa

Current vision models typically maintain a fixed correspondence between their representation structure and image space. Each layer comprises a set of tokens arranged "on-the-grid," which biases patches or tokens to encode information at a…

Recent advances in video diffusion models shows promise for generating robotic decision-making data, with trajectory conditions further enabling fine-grained control. However, existing methods primarily focus on individual object motion and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Xiao Fu , Xintao Wang , Xian Liu , Jianhong Bai , Runsen Xu , Pengfei Wan , Di Zhang , Dahua Lin

Motion is a salient cue to recognize actions in video. Modern action recognition models leverage motion information either explicitly by using optical flow as input or implicitly by means of 3D convolutional filters that simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Heng Wang , Du Tran , Lorenzo Torresani , Matt Feiszli

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond…

Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands two capabilities: (1) disentangled motion control, allowing users to separately control the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Shaowei Liu , Xuanchi Ren , Tianchang Shen , Huan Ling , Saurabh Gupta , Shenlong Wang , Sanja Fidler , Jun Gao

Well-trained generative neural networks (GNN) are very efficient at compressing visual information for static images in their learned parameters but not as efficient as inter- and intra-prediction for most video content. However, for…

Image and Video Processing · Electrical Eng. & Systems 2020-10-07 Jonah Probell

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ruiyan Wang , Teng Hu , Kaihui Huang , Zihan Su , Ran Yi , Lizhuang Ma

Existing deep learning methods of video recognition usually require a large number of labeled videos for training. But for a new task, videos are often unlabeled and it is also time-consuming and labor-intensive to annotate them. Instead of…

Computer Vision and Pattern Recognition · Computer Science 2018-05-14 Feiwu Yu , Xinxiao Wu , Yuchao Sun , Lixin Duan

Prior works about text-to-image synthesis typically concatenated the sentence embedding with the noise vector, while the sentence embedding and the noise vector are two different factors, which control the different aspects of the…

Multimedia · Computer Science 2023-03-27 Jiguo Li , Xiaobin Liu , Lirong Zheng

In this paper, we propose a novel 3D-RecGAN approach, which reconstructs the complete 3D structure of a given object from a single arbitrary depth view using generative adversarial networks. Unlike the existing work which typically requires…

Computer Vision and Pattern Recognition · Computer Science 2017-08-29 Bo Yang , Hongkai Wen , Sen Wang , Ronald Clark , Andrew Markham , Niki Trigoni

Generative Adversarial Networks (GANs) have recently achieved impressive results for many real-world applications, and many GAN variants have emerged with improvements in sample quality and training stability. However, they have not been…

Computer Vision and Pattern Recognition · Computer Science 2018-12-11 David Bau , Jun-Yan Zhu , Hendrik Strobelt , Bolei Zhou , Joshua B. Tenenbaum , William T. Freeman , Antonio Torralba

We present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Yanhui Wang , Jianmin Bao , Wenming Weng , Ruoyu Feng , Dacheng Yin , Tao Yang , Jingxu Zhang , Qi Dai Zhiyuan Zhao , Chunyu Wang , Kai Qiu , Yuhui Yuan , Chuanxin Tang , Xiaoyan Sun , Chong Luo , Baining Guo

We present a novel generative model for human motion modeling using Generative Adversarial Networks (GANs). We formulate the GAN discriminator using dense validation at each time-scale and perturb the discriminator input to make it…

Computer Vision and Pattern Recognition · Computer Science 2018-05-21 Xiao Lin , Mohamed R. Amer
‹ Prev 1 8 9 10 Next ›