English
Related papers

Related papers: Learning semantical dynamics and spatiotemporal co…

200 papers

Multi-person motion prediction remains a challenging problem, especially in the joint representation learning of individual motion and social interactions. Most prior methods only involve learning local pose dynamics for individual motion…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Xiaogang Peng , Yaodi Shen , Haoran Wang , Binling Nie , Yigang Wang , Zizhao Wu

Transformers have significantly advanced the field of 3D human pose estimation (HPE). However, existing transformer-based methods primarily use self-attention mechanisms for spatio-temporal modeling, leading to a quadratic complexity,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yunlong Huang , Junshuo Liu , Ke Xian , Robert Caiming Qiu

Data-driven modeling of human motions is ubiquitous in computer graphics and computer vision applications, such as synthesizing realistic motions or recognizing actions. Recent research has shown that such problems can be approached by…

Graphics · Computer Science 2019-08-21 He Wang , Edmond S. L. Ho , Hubert P. H. Shum , Zhanxing Zhu

Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Zhongwei Qiu , Yang Qiansheng , Jian Wang , Haocheng Feng , Junyu Han , Errui Ding , Chang Xu , Dongmei Fu , Jingdong Wang

We propose a semantics-driven unsupervised learning approach for monocular depth and ego-motion estimation from videos in this paper. Recent unsupervised learning methods employ photometric errors between synthetic view and actual image as…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Xiaobin Wei , Jianjiang Feng , Jie Zhou

Human pose estimation in two-dimensional images videos has been a hot topic in the computer vision problem recently due to its vast benefits and potential applications for improving human life, such as behaviors recognition, motion capture…

Computer Vision and Pattern Recognition · Computer Science 2022-02-08 Thong Duy Nguyen , Milan Kresovic

We explore the importance of spatial contextual information in human pose estimation. Most state-of-the-art pose networks are trained in a multi-stage manner and produce several auxiliary predictions for deep supervision. With this…

Computer Vision and Pattern Recognition · Computer Science 2019-01-08 Hong Zhang , Hao Ouyang , Shu Liu , Xiaojuan Qi , Xiaoyong Shen , Ruigang Yang , Jiaya Jia

Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge. This is because of the under-determined nature of frame prediction; multiple potential futures can arise…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Huiwon Jang , Dongyoung Kim , Junsu Kim , Jinwoo Shin , Pieter Abbeel , Younggyo Seo

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-channel) audio…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision due to the depth ambiguity of 2D-to-3D lifting. To improve accuracy and address occlusion issues, inertial sensor has been…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yiming Bao , Xu Zhao , Dahong Qian

Modern approaches for multi-person pose estimation in video require large amounts of dense annotations. However, labeling every frame in a video is costly and labor intensive. To reduce the need for dense annotations, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2019-12-12 Gedas Bertasius , Christoph Feichtenhofer , Du Tran , Jianbo Shi , Lorenzo Torresani

In this work, we focus on the challenge of temporally consistent human-centric dense prediction across video sequences. Existing models achieve strong per-frame accuracy but often flicker under motion, occlusion, and lighting changes, and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Xingyu Miao , Junting Dong , Qin Zhao , Yuhang Yang , Junhao Chen , Yang Long

Bottom-up approaches for image-based multi-person pose estimation consist of two stages: (1) keypoint detection and (2) grouping of the detected keypoints to form person instances. Current grouping approaches rely on learned embedding from…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Jiahao Lin , Gim Hee Lee

The key of Human-Object Interaction(HOI) recognition is to infer the relationship between human and objects. Recently, the image's Human-Object Interaction(HOI) detection has made significant progress. However, there is still room for…

Computer Vision and Pattern Recognition · Computer Science 2022-06-08 Hongsheng Li , Guangming Zhu , Wu Zhen , Lan Ni , Peiyi Shen , Liang Zhang , Ning Wang , Cong Hua

In monocular videos that capture dynamic scenes, estimating the 3D geometry of video contents has been a fundamental challenge in computer vision. Specifically, the task is significantly challenged by the object motion, where existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Seong Hyeon Park , Jinwoo Shin

Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Wen-Li Wei , Jen-Chun Lin , Tyng-Luh Liu , Hong-Yuan Mark Liao

Self-supervised multi-frame monocular depth estimation relies on the geometric consistency between successive frames under the assumption of a static scene. However, the presence of moving objects in dynamic scenes introduces inevitable…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Sungmin Woo , Wonjoon Lee , Woo Jin Kim , Dogyoon Lee , Sangyoun Lee

This paper proposes to learn reliable dense correspondence from videos in a self-supervised manner. Our learning process integrates two highly related tasks: tracking large image regions \emph{and} establishing fine-grained pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2019-09-27 Xueting Li , Sifei Liu , Shalini De Mello , Xiaolong Wang , Jan Kautz , Ming-Hsuan Yang

We present a framework designed to learn the underlying dynamics between two images observed at consecutive time steps. The complex nature of image data and the lack of temporal information pose significant challenges in capturing the…

Machine Learning · Computer Science 2023-10-17 Jihun Han , Yoonsang Lee , Anne Gelb

The importance and demands of visual scene understanding have been steadily increasing along with the active development of autonomous systems. Consequently, there has been a large amount of research dedicated to semantic segmentation and…

Computer Vision and Pattern Recognition · Computer Science 2016-07-27 Junhwa Hur , Stefan Roth