English
Related papers

Related papers: Egocentric Auditory Attention Localization in Conv…

200 papers

Estimating camera wearer's body pose from an egocentric view (egopose) is a vital task in augmented and virtual reality. Existing approaches either use a narrow field of view front facing camera that barely captures the wearer, or an…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Hao Jiang , Vamsi Krishna Ithapu

In this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in…

Sound · Computer Science 2021-07-21 Siqi Zheng , Weilong Huang , Xianliang Wang , Hongbin Suo , Jinwei Feng , Zhijie Yan

In this work we use deep reinforcement learning to create an autonomous agent that can navigate in a two-dimensional space using only raw auditory sensory information from the environment, a problem that has received very little attention…

Sound · Computer Science 2021-05-17 Petros Giannakopoulos , Aggelos Pikrakis , Yannis Cotronis

In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extremely simple approach to generating (weak) speech…

Multimedia · Computer Science 2017-06-02 Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , Ian Sturdy

Wearable cameras stand out as one of the most promising devices for the upcoming years, and as a consequence, the demand of computer algorithms to automatically understand the videos recorded with them is increasing quickly. An automatic…

Computer Vision and Pattern Recognition · Computer Science 2017-03-29 Alejandro Betancourt , Natalia Díaz-Rodríguez , Emilia Barakova , Lucio Marcenaro , Matthias Rauterberg , Carlo Regazzoni

Predicting the future location of vehicles is essential for safety-critical applications such as advanced driver assistance systems (ADAS) and autonomous driving. This paper introduces a novel approach to simultaneously predict both the…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Yu Yao , Mingze Xu , Chiho Choi , David J. Crandall , Ella M. Atkins , Behzad Dariush

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

Sound · Computer Science 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand. Good synchronization between…

Computer Vision and Pattern Recognition · Computer Science 2018-12-17 Naji Khosravan , Shervin Ardeshir , Rohit Puri

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual…

Sound · Computer Science 2022-10-06 Yinfeng Yu , Lele Cao , Fuchun Sun , Xiaohong Liu , Liejun Wang

Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we introduce…

This paper presents an experimental study on deep speaker embedding with an attention mechanism that has been found to be a powerful representation learning technique in speaker recognition. In this framework, an attention model works as a…

Sound · Computer Science 2018-09-26 Qiongqiong Wang , Koji Okabe , Kong Aik Lee , Hitoshi Yamamoto , Takafumi Koshinaka

We propose an algorithm to separate simultaneously speaking persons from each other, the "cocktail party problem", using a single microphone. Our approach involves a deep recurrent neural networks regression to a vector space that is…

Sound · Computer Science 2017-05-22 Cory Stephenson , Patrick Callier , Abhinav Ganesh , Karl Ni

Egocentric vision aims to capture and analyse the world from the first-person perspective. We explore the possibilities for egocentric wearable devices to improve and enhance industrial use cases w.r.t. data collection, annotation,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Vivek Chavan , Oliver Heimann , Jörg Krüger

Self-supervised learning has been used to leverage unlabelled data, improving accuracy and generalisation of speech systems through the training of representation models. While many recent works have sought to produce effective…

Computation and Language · Computer Science 2023-10-18 Antoni Dimitriadis , Siqi Pan , Vidhyasaharan Sethu , Beena Ahmed

Sound Event Localization and Detection refers to the problem of identifying the presence of independent or temporally-overlapped sound sources, correctly identifying to which sound class it belongs, estimating their spatial directions while…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Daniel Arteaga , Andrés Pérez-López

Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-27 Michael Neri , Archontis Politis , Daniel Krause , Marco Carli , Tuomas Virtanen

Acoustic scene classification systems using deep neural networks classify given recordings into pre-defined classes. In this study, we propose a novel scheme for acoustic scene classification which adopts an audio tagging system inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-21 Jee-weon Jung , Hye-jin Shim , Ju-ho Kim , Seung-bin Kim , Ha-Jin Yu

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Sanjoy Chowdhury , Subrata Biswas , Sayan Nag , Tushar Nagarajan , Calvin Murdock , Ishwarya Ananthabhotla , Yijun Qian , Vamsi Krishna Ithapu , Dinesh Manocha , Ruohan Gao

A speaker naming task, which finds and identifies the active speaker in a certain movie or drama scene, is crucial for dealing with high-level video analysis applications such as automatic subtitle labeling and video summarization. Modern…

Multimedia · Computer Science 2019-12-03 Jungwoo Pyo , Joohyun Lee , Youngjune Park , Tien-Cuong Bui , Sang Kyun Cha

Wearable collaborative robots stand to assist human wearers who need fall prevention assistance or wear exoskeletons. Such a robot needs to be able to constantly adapt to the surrounding scene based on egocentric vision, and predict the ego…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Weizhuo Wang , C. Karen Liu , Monroe Kennedy