中文
相关论文

相关论文: DDLP: Unsupervised Object-Centric Video Prediction…

200 篇论文

We present differentiable particle filters (DPFs): a differentiable implementation of the particle filter algorithm with learnable motion and measurement models. Since DPFs are end-to-end differentiable, we can efficiently train their…

机器学习 · 计算机科学 2018-05-31 Rico Jonschkowski , Divyam Rastogi , Oliver Brock

Multi-object tracking aims to maintain object identities over time by associating detections across video frames. Two dominant paradigms exist in literature: tracking-by-detection methods, which are computationally efficient but rely on…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Momir Adžemović

Generalization has been one of the major challenges for learning dynamics models in model-based reinforcement learning. However, previous work on action-conditioned dynamics prediction focuses on learning the pixel-level motion and thus…

计算机视觉与模式识别 · 计算机科学 2018-10-31 Guangxiang Zhu , Zhiao Huang , Chongjie Zhang

We present a new model DrNET that learns disentangled image representations from video. Our approach leverages the temporal coherence of video and a novel adversarial loss to learn a representation that factorizes each frame into a…

机器学习 · 计算机科学 2024-03-15 Remi Denton , Vighnesh Birodkar

Unsupervised video Object-Centric Learning (OCL) is promising as it enables object-level scene representation and understanding as we humans do. Mainstream video OCL methods adopt a recurrent architecture: An aggregator aggregates current…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Rongzhen Zhao , Jian Li , Juho Kannala , Joni Pajarinen

We propose a framework to continuously learn object-centric representations for visual learning and understanding. Existing object-centric representations either rely on supervisions that individualize objects in the scene, or perform…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Chuanyu Pan , Yanchao Yang , Kaichun Mo , Yueqi Duan , Leonidas Guibas

Human perception involves decomposing complex multi-object scenes into time-static object appearance (i.e., size, shape, color) and time-varying object motion (i.e., position, velocity, acceleration). For machines to achieve human-like…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Yeon-Ji Song , Jaein Kim , Suhyung Choi , Jin-Hwa Kim , Byoung-Tak Zhang

Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a…

计算机视觉与模式识别 · 计算机科学 2021-12-14 Jinyang Yuan , Bin Li , Xiangyang Xue

Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built. Recent work on simple 2D and 3D datasets has shown that models…

Our goal is to predict future video frames given a sequence of input frames. Despite large amounts of video data, this remains a challenging task because of the high-dimensionality of video frames. We address this challenge by proposing the…

机器学习 · 计算机科学 2018-10-19 Jun-Ting Hsieh , Bingbin Liu , De-An Huang , Li Fei-Fei , Juan Carlos Niebles

We present High-Density Visual Particle Dynamics (HD-VPD), a learned world model that can emulate the physical dynamics of real scenes by processing massive latent point clouds containing 100K+ particles. To enable efficiency at this scale,…

机器学习 · 计算机科学 2024-07-01 William F. Whitney , Jacob Varley , Deepali Jain , Krzysztof Choromanski , Sumeet Singh , Vikas Sindhwani

Humans easily recognize object parts and their hierarchical structure by watching how they move; they can then predict how each part moves in the future. In this paper, we propose a novel formulation that simultaneously learns a…

计算机视觉与模式识别 · 计算机科学 2019-03-14 Zhenjia Xu , Zhijian Liu , Chen Sun , Kevin Murphy , William T. Freeman , Joshua B. Tenenbaum , Jiajun Wu

Object-centric learning aims to decompose an input image into a set of meaningful object files (slots). These latent object representations enable a variety of downstream tasks. Yet, object-centric learning struggles on real-world datasets,…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Krishnakant Singh , Simone Schaub-Meyer , Stefan Roth

In this paper, we propose a novel object detection algorithm named "Deep Regionlets" by integrating deep neural networks and a conventional detection schema for accurate generic object detection. Motivated by the effectiveness of regionlets…

计算机视觉与模式识别 · 计算机科学 2019-12-04 Hongyu Xu , Xutao Lv , Xiaoyu Wang , Zhou Ren , Navaneeth Bodla , Rama Chellappa

In order to interact with the world, agents must be able to predict the results of the world's dynamics. A natural approach to learn about these dynamics is through video prediction, as cameras are ubiquitous and powerful sensors. Direct…

计算机视觉与模式识别 · 计算机科学 2021-05-07 Karl Schmeckpeper , Georgios Georgakis , Kostas Daniilidis

We investigate a new task in human motion prediction, which is predicting motions under unexpected physical perturbation potentially involving multiple people. Compared with existing research, this task involves predicting less controlled,…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Jiangbei Yue , Baiyi Li , Julien Pettré , Armin Seyfried , He Wang

Particle filters flexibly represent multiple posterior modes nonparametrically, via a collection of weighted samples, but have classically been applied to tracking problems with known dynamics and observation likelihoods. Such generative…

机器学习 · 计算机科学 2024-04-16 Ali Younis , Erik Sudderth

A diffusion probabilistic model (DPM), which constructs a forward diffusion process by gradually adding noise to data points and learns the reverse denoising process to generate new samples, has been shown to handle complex data…

计算机视觉与模式识别 · 计算机科学 2023-10-16 Zhengxiong Luo , Dayou Chen , Yingya Zhang , Yan Huang , Liang Wang , Yujun Shen , Deli Zhao , Jingren Zhou , Tieniu Tan

We study the problem of dynamic visual reasoning on raw videos. This is a challenging problem; currently, state-of-the-art models often require dense supervision on physical object properties and events from simulation, which are…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Zhenfang Chen , Jiayuan Mao , Jiajun Wu , Kwan-Yee Kenneth Wong , Joshua B. Tenenbaum , Chuang Gan

We present a method to estimate depth of a dynamic scene, containing arbitrary moving objects, from an ordinary video captured with a moving camera. We seek a geometrically and temporally consistent solution to this underconstrained…

计算机视觉与模式识别 · 计算机科学 2021-08-04 Zhoutong Zhang , Forrester Cole , Richard Tucker , William T. Freeman , Tali Dekel