English
Related papers

Related papers: Learning Spatial Features from Audio-Visual Corres…

200 papers

Self-supervised learning holds the promise of learning good representations from real-world continuous uncurated data streams. However, most existing works in visual self-supervised learning focus on static images or artificial data…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Yanlai Yang , Mengye Ren

Face performance capture and reenactment techniques use multiple cameras and sensors, positioned at a distance from the face or mounted on heavy wearable devices. This limits their applications in mobile and outdoor environments. We present…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Mohamed Elgharib , Mallikarjun BR , Ayush Tewari , Hyeongwoo Kim , Wentao Liu , Hans-Peter Seidel , Christian Theobalt

Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically generate diverse audio for image captioning datasets. This…

Computer Vision and Pattern Recognition · Computer Science 2019-09-20 Gabriel Ilharco , Yuan Zhang , Jason Baldridge

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the…

Computer Vision and Pattern Recognition · Computer Science 2017-11-10 Marc Bolaños , Álvaro Peris , Francisco Casacuberta , Sergi Soler , Petia Radeva

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

Sound · Computer Science 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

Contrastive learning of auditory and visual perception has been extremely successful when investigated individually. However, there are still major questions on how we could integrate principles learned from both domains to attain effective…

Computer Vision and Pattern Recognition · Computer Science 2021-10-15 Haider Al-Tahan , Yalda Mohsenzadeh

We propose a self-supervised method to learn feature representations from videos. A standard approach in traditional self-supervised methods uses positive-negative data pairs to train with contrastive learning strategy. In such a case,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Li Tao , Xueting Wang , Toshihiko Yamasaki

Video analysis tasks rely heavily on identifying the pixels from different frames that correspond to the same visual target. To tackle this problem, recent studies have advocated feature learning methods that aim to learn distinctive…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Rui Li , Shenglong Zhou , Dong Liu

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three…

Computer Vision and Pattern Recognition · Computer Science 2019-08-23 Evangelos Kazakos , Arsha Nagrani , Andrew Zisserman , Dima Damen

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-13 Otavio Braga , Olivier Siohan

Recently, many efforts have been made to explore how the brain processes speech using electroencephalographic (EEG) signals, where deep learning-based approaches were shown to be applicable in this field. In order to decode speech signals…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-24 Qiushi Zhu , Xiaoying Zhao , Jie Zhang , Yu Gu , Chao Weng , Yuchen Hu

This paper studies a conceptually simple extension of Masked Autoencoders (MAE) to spatiotemporal representation learning from videos. We randomly mask out spacetime patches in videos and learn an autoencoder to reconstruct them in pixels.…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Christoph Feichtenhofer , Haoqi Fan , Yanghao Li , Kaiming He

Learning representations from videos requires understanding continuous motion and visual correspondences between frames. In this paper, we introduce the Concatenated Masked Autoencoders (CatMAE) as a spatial-temporal learner for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Zhouqiang Jiang , Bowen Wang , Tong Xiang , Zhaofeng Niu , Hong Tang , Guangshun Li , Liangzhi Li

For reliable autonomous robot navigation in urban settings, the robot must have the ability to identify semantically traversable terrains in the image based on the semantic understanding of the scene. This reasoning ability is based on…

Robotics · Computer Science 2024-12-30 Yunho Kim , Jeong Hyun Lee , Choongin Lee , Juhyeok Mun , Donghoon Youm , Jeongsoo Park , Jemin Hwangbo

While self-supervised learning has enabled effective representation learning in the absence of labels, for vision, video remains a relatively untapped source of supervision. To address this, we propose Pixel-level Correspondence (PiCo), a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-11 Yash Sharma , Yi Zhu , Chris Russell , Thomas Brox

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Shervin Ardeshir , Ali Borji

We propose a semantics-driven unsupervised learning approach for monocular depth and ego-motion estimation from videos in this paper. Recent unsupervised learning methods employ photometric errors between synthetic view and actual image as…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Xiaobin Wei , Jianjiang Feng , Jie Zhou

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Qiushi Zhu , Jie Zhang , Yu Gu , Yuchen Hu , Lirong Dai

The way people look in terms of facial attributes (ethnicity, hair color, facial hair, etc.) and the clothes or accessories they wear (sunglasses, hat, hoodies, etc.) is highly dependent on geo-location and weather condition, respectively.…

Computer Vision and Pattern Recognition · Computer Science 2016-06-24 Jing Wang , Yu Cheng , Rogerio Schmidt Feris