English
Related papers

Related papers: Low-latency auditory spatial attention detection b…

200 papers

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

Segmentation for continuous Automatic Speech Recognition (ASR) has traditionally used silence timeouts or voice activity detectors (VADs), which are both limited to acoustic features. This segmentation is often overly aggressive, given that…

Computation and Language · Computer Science 2022-10-28 Piyush Behre , Naveen Parihar , Sharman Tan , Amy Shah , Eva Sharma , Geoffrey Liu , Shuangyu Chang , Hosam Khalil , Chris Basoglu , Sayan Pathak

Resting-state brain networks represent the intrinsic state of the brain during the majority of cognitive and sensorimotor tasks. However, no study has yet presented concise predictors of task-induced vigilance variability from…

Signal Processing · Electrical Eng. & Systems 2019-10-23 Mastaneh Torkamani-Azar , Sumeyra Demir Kanik , Serap Aydin , Mujdat Cetin

The electroencephalography (EEG) signal is a non-stationary, stochastic, and highly non-linear bioelectric signal for which achieving high classification accuracy is challenging, especially when the number of subjects is limited. As…

Signal Processing · Electrical Eng. & Systems 2021-08-03 Xiangyun Li , Peng Chen , Zhanpeng Bao

Despite the advancement in the domain of audio and audio-visual speech recognition, visual speech recognition systems are still quite under-explored due to the visual ambiguity of some phonemes. In this work, we propose a new lip-reading…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Shahd Elashmawy , Marian Ramsis , Hesham M. Eraqi , Farah Eldeshnawy , Hadeel Mabrouk , Omar Abugabal , Nourhan Sakr

In recent years, multiple noninvasive imaging modalities have been used to develop a better understanding of the human brain functionality, including positron emission tomography, single-photon emission computed tomography, and functional…

Signal Processing · Electrical Eng. & Systems 2019-10-18 Shiva Asadzadeh , Tohid Yousefi Rezaii , Soosan Beheshti , Azra Delpak , Saeed Meshgini

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous…

Sound · Computer Science 2025-03-18 Chen Liu , Peike Li , Liying Yang , Dadong Wang , Lincheng Li , Xin Yu

In this paper, we compare the performance of using binaural audio features in place of single-channel features for sound event detection. Three different binaural features are studied and evaluated on the publicly available TUT Sound Events…

Sound · Computer Science 2017-10-10 Sharath Adavanne , Tuomas Virtanen

Sound Event Localization and Detection refers to the problem of identifying the presence of independent or temporally-overlapped sound sources, correctly identifying to which sound class it belongs, estimating their spatial directions while…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Daniel Arteaga , Andrés Pérez-López

Existing EEG-driven image reconstruction methods often overlook spatial attention mechanisms, limiting fidelity and semantic coherence. To address this, we propose a dual-conditioning framework that combines EEG embeddings with spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Igor Abramov , Ilya Makarov

Recognizing the feelings of human beings plays a critical role in our daily communication. Neuroscience has demonstrated that different emotion states present different degrees of activation in different brain regions, EEG frequency bands…

Signal Processing · Electrical Eng. & Systems 2021-11-09 Jiyao Liu , Yanxi Zhao , Hao Wu , Dongmei Jiang

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine to have the same…

Data augmentation methods have shown great importance in diverse supervised learning problems where labeled data is scarce or costly to obtain. For sound event localization and detection (SELD) tasks several augmentation methods have been…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-20 Ricardo Falcon-Perez , Kazuki Shimada , Yuichiro Koyama , Shusuke Takahashi , Yuki Mitsufuji

Standardized tests play a crucial role in the detection of cognitive impairment. Previous work demonstrated that automatic detection of cognitive impairment is possible using audio data from a standardized picture description task. The…

In the status quo, dementia is yet to be cured. Precise diagnosis prior to the onset of the symptoms can prevent the rapid progression of the emerging cognitive impairment. Recent progress has shown that Electroencephalography (EEG) is the…

Electroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDANet) to classify the correct auditory stimulus evoking the…

Sound · Computer Science 2023-03-21 Fan Cui , Liyong Guo , Lang He , Jiyao Liu , ErCheng Pei , Yujun Wang , Dongmei Jiang

Electroencephalography (EEG) is a tool that allows us to analyze brain activity with high temporal resolution. These measures, combined with deep learning and digital signal processing, are widely used in neurological disorder detection and…

Signal Processing · Electrical Eng. & Systems 2024-11-20 Isaac Ariza , Lorenzo J. Tardon , Ana M. Barbancho , Irene De-Torres , Isabel Barbancho

Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural…

In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism often allocates a portion of attention weights to irrelevant…

Sound · Computer Science 2025-07-04 Junyu Wang , Tianrui Wang , Meng Ge , Longbiao Wang , Jianwu Dang

In conventional multichannel audio signal enhancement, spatial and spectral filtering are often performed sequentially. In contrast, it has been shown that for neural spatial filtering a joint approach of spectro-spatial filtering is more…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Annika Briegleb , Mhd Modar Halimeh , Walter Kellermann