English
Related papers

Related papers: Egocentric Auditory Attention Localization in Conv…

200 papers

Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial…

Computation and Language · Computer Science 2018-06-19 Matthias Sperber , Jan Niehues , Graham Neubig , Sebastian Stüker , Alex Waibel

Egocentric visual context detection can support intelligence augmentation applications. We created a wearable system, called PAL, for wearable, personalized, and privacy-preserving egocentric visual context detection. PAL has a wearable…

Computer Vision and Pattern Recognition · Computer Science 2021-05-25 Mina Khan , Pattie Maes

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are…

Signal Processing · Electrical Eng. & Systems 2024-09-16 Haolin Zhu , Yujie Yan , Xiran Xu , Zhongshu Ge , Pei Tian , Xihong Wu , Jing Chen

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the quality of life of the…

Computer Vision and Pattern Recognition · Computer Science 2017-11-10 Marc Bolaños , Álvaro Peris , Francisco Casacuberta , Sergi Soler , Petia Radeva

This paper presents a self-supervised method for visual detection of the active speaker in a multi-person spoken interaction scenario. Active speaker detection is a fundamental prerequisite for any artificial cognitive system attempting to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-19 Kalin Stefanov , Jonas Beskow , Giampiero Salvi

Active speaker detection (ASD) in egocentric videos presents unique challenges due to unstable viewpoints, motion blur, and off-screen speech sources - conditions under which traditional visual-centric methods degrade significantly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yu Wang , Juhyung Ha , David J. Crandall

In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order to improve recognition performance. To incorporate the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Evangelos Kazakos , Jaesung Huh , Arsha Nagrani , Andrew Zisserman , Dima Damen

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Shervin Ardeshir , Ali Borji

Speaker counting is the task of estimating the number of people that are simultaneously speaking in an audio recording. For several audio processing tasks such as speaker diarization, separation, localization and tracking, knowing the…

Sound · Computer Science 2021-01-07 Pierre-Amaury Grumiaux , Srdan Kitic , Laurent Girin , Alexandre Guérin

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to modeling multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-05-05 Yanbei Chen , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bingwen Zhu , Yuqian Fu , Qiaole Dong , Guolei Sun , Tianwen Qian , Yuzheng Wu , Danda Pani Paudel , Xiangyang Xue , Yanwei Fu

A key function of auditory cognition is the association of characteristic sounds with their corresponding semantics over time. Humans attempting to discriminate between fine-grained audio categories, often replay the same discriminative…

Sound · Computer Science 2023-03-14 Alexandros Stergiou , Dima Damen

Acoustic sensing has proved effective as a foundation for numerous applications in health and human behavior analysis. In this work, we focus on the problem of detecting in-person social interactions in naturalistic settings from audio…

Sound · Computer Science 2022-03-23 Dawei Liang , Zifan Xu , Yinuo Chen , Rebecca Adaimi , David Harwath , Edison Thomaz

Individuals regularly experience Hearing Difficulty Moments in everyday conversation. Identifying these moments of hearing difficulty has particular significance in the field of hearing assistive technology where timely interventions are…

This paper presents the External Attention Vision Transformer (EAViT) model, a novel approach designed to enhance audio classification accuracy. As digital audio resources proliferate, the demand for precise and efficient audio…

Effective communication requires adapting to the idiosyncrasies of each communicative context--such as the common ground shared with each partner. Humans demonstrate this ability to specialize to their audience in many contexts, such as the…

Machine Learning · Computer Science 2023-05-03 Aaditya K. Singh , David Ding , Andrew Saxe , Felix Hill , Andrew K. Lampinen

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury

Speaker diarization relies on the assumption that speech segments corresponding to a particular speaker are concentrated in a specific region of the speaker space; a region which represents that speaker's identity. These identities are not…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Nikolaos Flemotomos , Panayiotis Georgiou , Shrikanth Narayanan

Speaker diarization is a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify "who spoke when". In the early years, speaker diarization algorithms were developed for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-29 Tae Jin Park , Naoyuki Kanda , Dimitrios Dimitriadis , Kyu J. Han , Shinji Watanabe , Shrikanth Narayanan