English
Related papers

Related papers: SIMONe: View-Invariant, Temporally-Abstracted Obje…

200 papers

We present a slot-wise, object-based transition model that decomposes a scene into objects, aligns them (with respect to a slot-wise object memory) to maintain a consistent order across time, and predicts how those objects evolve over…

Computer Vision and Pattern Recognition · Computer Science 2021-03-09 Antonia Creswell , Rishabh Kabra , Chris Burgess , Murray Shanahan

Event camera, a novel neuromorphic vision sensor, records data with high temporal resolution and wide dynamic range, offering new possibilities for accurate visual representation in challenging scenarios. However, event data is inherently…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Lin Zhu , Ruonan Liu , Xiao Wang , Lizhi Wang , Hua Huang

We propose novel motion representations for animating articulated objects consisting of distinct parts. In a completely unsupervised manner, our method identifies object parts, tracks them in a driving video, and infers their motions by…

Computer Vision and Pattern Recognition · Computer Science 2021-04-26 Aliaksandr Siarohin , Oliver J. Woodford , Jian Ren , Menglei Chai , Sergey Tulyakov

Representing a scene and its constituent objects from raw sensory data is a core ability for enabling robots to interact with their environment. In this paper, we propose a novel approach for scene understanding, leveraging a hierarchical…

Robotics · Computer Science 2023-02-08 Toon Van de Maele , Tim Verbelen , Pietro Mazzaglia , Stefano Ferraro , Bart Dhoedt

Retinal image of surrounding objects varies tremendously due to the changes in position, size, pose, illumination condition, background context, occlusion, noise, and nonrigid deformations. But despite these huge variations, our visual…

Computer Vision and Pattern Recognition · Computer Science 2017-02-14 Saeed Reza Kheradpisheh , Mohammad Ganjtabesh , Timothée Masquelier

We propose an unsupervised, mid-level representation for a generative model of scenes. The representation is mid-level in that it is neither per-pixel nor per-image; rather, scenes are modeled as a collection of spatial, depth-ordered…

Computer Vision and Pattern Recognition · Computer Science 2022-08-02 Dave Epstein , Taesung Park , Richard Zhang , Eli Shechtman , Alexei A. Efros

We introduce a novel framework to build a model that can learn how to segment objects from a collection of images without any human annotation. Our method builds on the observation that the location of object segments can be perturbed…

Computer Vision and Pattern Recognition · Computer Science 2019-11-05 Adam Bielski , Paolo Favaro

Computer vision is increasingly effective at segmenting objects in images and videos; however, scene effects related to the objects -- shadows, reflections, generated smoke, etc -- are typically overlooked. Identifying such scene effects…

Computer Vision and Pattern Recognition · Computer Science 2021-10-04 Erika Lu , Forrester Cole , Tali Dekel , Andrew Zisserman , William T. Freeman , Michael Rubinstein

In this paper, we consider the task of unsupervised object discovery in videos. Previous works have shown promising results via processing optical flows to segment objects. However, taking flow as input brings about two drawbacks. First,…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Shuangrui Ding , Weidi Xie , Yabo Chen , Rui Qian , Xiaopeng Zhang , Hongkai Xiong , Qi Tian

We construct an unsupervised learning model that achieves nonlinear disentanglement of underlying factors of variation in naturalistic videos. Previous work suggests that representations can be disentangled if all but a few factors in the…

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into background, characters, and their animations. Our method capitalizes on moving characters being equivariant with respect…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Cinjon Resnick , Or Litany , Cosmas Heiß , Hugo Larochelle , Joan Bruna , Kyunghyun Cho

We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simple objective that can easily generalize all these methods to…

Computer Vision and Pattern Recognition · Computer Science 2021-04-30 Christoph Feichtenhofer , Haoqi Fan , Bo Xiong , Ross Girshick , Kaiming He

Representational learning forms the backbone of most deep learning applications, and the value of a learned representation is intimately tied to its information content regarding different factors of variation. Finding good representations…

Machine Learning · Computer Science 2022-03-31 Kieran A. Murphy , Varun Jampani , Srikumar Ramalingam , Ameesh Makadia

Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 James Oldfield , Christos Tzelepis , Yannis Panagakis , Mihalis A. Nicolaou , Ioannis Patras

We present a generative model of images that explicitly reasons over the set of objects they show. Our model learns a structured latent representation that separates objects from each other and from the background; unlike prior works, it…

Machine Learning · Computer Science 2020-04-03 Titas Anciukevicius , Christoph H. Lampert , Paul Henderson

Agents that are aware of the separation between themselves and their environments can leverage this understanding to form effective representations of visual input. We propose an approach for learning such structured representations for RL…

Machine Learning · Computer Science 2023-09-06 Kevin Gmelin , Shikhar Bahl , Russell Mendonca , Deepak Pathak

Disentangled representations seek to recover latent factors of variation underlying observed data, yet their identifiability is still not fully understood. We introduce a unified framework in which disentanglement is achieved through…

Machine Learning · Computer Science 2026-05-12 Stefan Matthes , Zhiwei Han , Hao Shen

Learning latent representations from complex data is central to modern machine learning, spanning temporal, multimodal, and partially observed systems. In such settings, representations are better understood as latent states capturing…

Machine Learning · Computer Science 2026-05-18 Gwenolé Quellec

We present an approach to learn the dynamics of multiple objects from image sequences in an unsupervised way. We introduce a probabilistic model that first generate noisy positions for each object through a separate linear state-space…

Computer Vision and Pattern Recognition · Computer Science 2019-07-31 Silvia Chiappa , Ulrich Paquet

Recent advances in visual representation learning allowed to build an abundance of powerful off-the-shelf features that are ready-to-use for numerous downstream tasks. This work aims to assess how well these features preserve information…

Computer Vision and Pattern Recognition · Computer Science 2022-12-21 Monika Wysoczańska , Tom Monnier , Tomasz Trzciński , David Picard