English
Related papers

Related papers: HomE: Homography-Equivariant Video Representation …

200 papers

This paper proposes a novel pretext task to address the self-supervised video representation learning problem. Specifically, given an unlabeled video clip, we compute a series of spatio-temporal statistical summaries, such as the spatial…

Computer Vision and Pattern Recognition · Computer Science 2021-02-01 Jiangliu Wang , Jianbo Jiao , Linchao Bao , Shengfeng He , Wei Liu , Yun-hui Liu

In this paper, we propose a novel learning scheme for self-supervised video representation learning. Motivated by how humans understand videos, we propose to first learn general visual concepts then attend to discriminative local areas for…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Rui Qian , Shuangrui Ding , Xian Liu , Dahua Lin

In this work we employ multitask learning to capitalize on the structure that exists in related supervised tasks to train complex neural networks. It allows training a network for multiple objectives in parallel, in order to improve…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas Noldus , Remco Veltkamp

Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is the lack of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Antoine Miech , Ivan Laptev , Josef Sivic

In light of the success of contrastive learning in the image domain, current self-supervised video representation learning methods usually employ contrastive loss to facilitate video representation learning. When naively pulling two…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Shuangrui Ding , Maomao Li , Tianyu Yang , Rui Qian , Haohang Xu , Qingyi Chen , Jue Wang , Hongkai Xiong

Human pose analysis is presently dominated by deep convolutional networks trained with extensive manual annotations of joint locations and beyond. To avoid the need for expensive labeling, we exploit spatiotemporal relations in training…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Ömer Sümer , Tobias Dencker , Björn Ommer

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Gang Li , Heliang Zheng , Daqing Liu , Chaoyue Wang , Bing Su , Changwen Zheng

Human pose estimation (HPE) usually requires large-scale training data to reach high performance. However, it is rather time-consuming to collect high-quality and fine-grained annotations for human body. To alleviate this issue, we revisit…

Computer Vision and Pattern Recognition · Computer Science 2022-05-26 Xixia Xu , Yingguo Gao , Ke Yan , Xue Lin , Qi Zou

Self-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Yiping Wei , Kunyu Peng , Alina Roitberg , Jiaming Zhang , Junwei Zheng , Ruiping Liu , Yufan Chen , Kailun Yang , Rainer Stiefelhagen

Pairwise pose estimation from images with little or no overlap is an open challenge in computer vision. Existing methods, even those trained on large-scale datasets, struggle in these scenarios due to the lack of identifiable…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Ruojin Cai , Jason Y. Zhang , Philipp Henzler , Zhengqi Li , Noah Snavely , Ricardo Martin-Brualla

Despite the success of deep learning in video understanding tasks, processing every frame in a video is computationally expensive and often unnecessary in real-time applications. Frame selection aims to extract the most informative and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Mingjun Zhao , Yakun Yu , Xiaoli Wang , Lei Yang , Di Niu

How can agents learn internal models that veridically represent interactions with the real world is a largely open question. As machine learning is moving towards representations containing not just observational but also interventional…

Machine Learning · Computer Science 2024-07-03 Hamza Keurti , Hsiao-Ru Pan , Michel Besserve , Benjamin F. Grewe , Bernhard Schölkopf

We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constraints on the embedding space using paired text and video…

Computer Vision and Pattern Recognition · Computer Science 2020-07-08 AJ Piergiovanni , Michael S. Ryoo

Current Transformer-based methods for small object detection continue emerging, yet they have still exhibited significant shortcomings. This paper introduces HeatMap Position Embedding (HMPE), a novel Transformer Optimization technique that…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 YangChen Zeng

Recent advancements in 3D hand pose estimation have shown promising results, but its effectiveness has primarily relied on the availability of large-scale annotated datasets, the creation of which is a laborious and costly process. To…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Xiaozheng Zheng , Chao Wen , Zhou Xue , Pengfei Ren , Jingyu Wang

We propose a novel unsupervised cross-modal homography estimation framework based on intra-modal Self-supervised learning, Correlation, and consistent feature map Projection, namely SCPNet. The concept of intra-modal self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Runmin Zhang , Jun Ma , Si-Yuan Cao , Lun Luo , Beinan Yu , Shu-Jie Chen , Junwei Li , Hui-Liang Shen

Audio-visual video parsing is the task of categorizing a video at the segment level with weak labels, and predicting them as audible or visible events. Recent methods for this task leverage the attention mechanism to capture the semantic…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Yaru Chen , Ruohao Guo , Xubo Liu , Peipei Wu , Guangyao Li , Zhenbo Li , Wenwu Wang

Supervised learning in large discriminative models is a mainstay for modern computer vision. Such an approach necessitates investing in large-scale human-annotated datasets for achieving state-of-the-art results. In turn, the efficacy of…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Liang-Chieh Chen , Raphael Gontijo Lopes , Bowen Cheng , Maxwell D. Collins , Ekin D. Cubuk , Barret Zoph , Hartwig Adam , Jonathon Shlens

Existing research on action recognition treats activities as monolithic events occurring in videos. Recently, the benefits of formulating actions as a combination of atomic-actions have shown promise in improving action understanding with…

Computer Vision and Pattern Recognition · Computer Science 2021-05-12 Nishant Rai , Haofeng Chen , Jingwei Ji , Rishi Desai , Kazuki Kozuka , Shun Ishizaka , Ehsan Adeli , Juan Carlos Niebles

Most recent view-invariant action recognition and performance assessment approaches rely on a large amount of annotated 3D skeleton data to extract view-invariant features. However, acquiring 3D skeleton data can be cumbersome, if not…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Faegheh Sardari , Björn Ommer , Majid Mirmehdi