English
Related papers

Related papers: Self-Supervised Keypoint Discovery in Behavioral V…

200 papers

Skeleton-based motion representations are robust for action localization and understanding for their invariance to perspective, lighting, and occlusion, compared with images. Yet, they are often ambiguous and incomplete when taken out of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Qihang Fang , Chengcheng Tang , Shugao Ma , Yanchao Yang

Unsupervised multi-object discovery (MOD) aims to detect and localize distinct object instances in visual scenes without any form of human supervision. Recent approaches leverage object-centric learning (OCL) and motion cues from video to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Xinrui Gong , Oliver Hahn , Christoph Reich , Krishnakant Singh , Simone Schaub-Meyer , Daniel Cremers , Stefan Roth

Skeleton-based temporal action segmentation is a fundamental yet challenging task, playing a crucial role in enabling intelligent systems to perceive and respond to human activities. While fully-supervised methods achieve satisfactory…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hongsong Wang , Yiqin Shen , Pengbo Yan , Jie Gui

Learning from visual data opens the potential to accrue a large range of manipulation behaviors by leveraging human demonstrations without specifying each of them mathematically, but rather through natural task specification. In this paper,…

Robotics · Computer Science 2021-11-16 Haoyu Xiong , Quanzhou Li , Yun-Chun Chen , Homanga Bharadhwaj , Samarth Sinha , Animesh Garg

A robot's ability to act is fundamentally constrained by what it can perceive. Many existing approaches to visual representation learning utilize general-purpose training criteria, e.g. image reconstruction, smoothness in latent space, or…

Self-supervised tasks have been utilized to build useful representations that can be used in downstream tasks when the annotation is unavailable. In this paper, we introduce a self-supervised video representation learning method based on…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Duc Quang Vu , Ngan T. H. Le , Jia-Ching Wang

Structured representations such as keypoints are widely used in pose transfer, conditional image generation, animation, and 3D reconstruction. However, their supervised learning requires expensive annotation for each target domain. We…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Xingzhe He , Bastian Wandt , Helge Rhodin

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. We present an unsupervised approach to learn atomic actions…

Computer Vision and Pattern Recognition · Computer Science 2021-06-08 AJ Piergiovanni , Anelia Angelova , Michael S. Ryoo , Irfan Essa

Behavioral annotation using signal processing and machine learning is highly dependent on training data and manual annotations of behavioral labels. Previous studies have shown that speech information encodes significant behavioral…

Machine Learning · Computer Science 2017-01-13 Haoqi Li , Brian Baucom , Panayiotis Georgiou

We study the problem of unsupervised physical object discovery. While existing frameworks aim to decompose scenes into 2D segments based off each object's appearance, we explore how physics, especially object interactions, facilitates…

Computer Vision and Pattern Recognition · Computer Science 2021-03-24 Yilun Du , Kevin Smith , Tomer Ulman , Joshua Tenenbaum , Jiajun Wu

Gesture is an important mean of non-verbal communication, with visual modality allows human to convey information during interaction, facilitating peoples and human-machine interactions. However, it is considered difficult to automatically…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Fabien Allemand , Alessio Mazzela , Jun Villette , Decky Aspandi , Titus Zaharia

We present Neural Marionette, an unsupervised approach that discovers the skeletal structure from a dynamic sequence and learns to generate diverse motions that are consistent with the observed motion dynamics. Given a video stream of point…

Computer Vision and Pattern Recognition · Computer Science 2022-02-18 Jinseok Bae , Hojun Jang , Cheol-Hui Min , Hyungun Choi , Young Min Kim

Researchers have attempted utilizing deep neural network (DNN) to learn novel local features from images inspired by its recent successes on a variety of vision tasks. However, existing DNN-based algorithms have not achieved such remarkable…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Yafei Song , Ling Cai , Jia Li , Yonghong Tian , Mingyang Li

We introduce a novel self-supervised learning approach to learn representations of videos that are responsive to changes in the motion dynamics. Our representations can be learned from data without human annotation and provide a substantial…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Simon Jenni , Givi Meishvili , Paolo Favaro

Weakly-supervised temporal action localization aims to identify and localize the action instances in the untrimmed videos with only video-level action labels. When humans watch videos, we can adapt our abstract-level knowledge about actions…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Xijun Wang , Aggelos K. Katsaggelos

Children learn powerful internal models of the world around them from a few years of egocentric visual experience. Can such internal models be learned from a child's visual experience with highly generic learning algorithms or do they…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 A. Emin Orhan , Wentao Wang , Alex N. Wang , Mengye Ren , Brenden M. Lake

Dog owners are typically capable of recognizing behavioral cues that reveal subjective states of their dogs, such as pain. But automatic recognition of the pain state is very challenging. This paper proposes a novel video-based, two-stream…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Hongyi Zhu , Yasemin Salgırlı , Pınar Can , Durmuş Atılgan , Albert Ali Salah

This paper proposes a human activity recognition method which is based on features learned from 3D video data without incorporating domain knowledge. The experiments on data collected by RGBD cameras produce results outperforming other…

Computer Vision and Pattern Recognition · Computer Science 2015-08-11 Ngu Nguyen

Advances in deep learning have enabled the development of models that have exhibited a remarkable tendency to recognize and even localize actions in videos. However, they tend to experience errors when faced with scenes or examples beyond…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Sathyanarayanan N. Aakur , Sanjoy Kundu , Nikhil Gunti

Despite their irresistible success, deep learning algorithms still heavily rely on annotated data. On the other hand, unsupervised settings pose many challenges, especially about determining the right inductive bias in diverse scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2021-03-11 Beril Besbinar , Pascal Frossard