English
Related papers

Related papers: A contrastive-learning approach for auditory atten…

200 papers

Transformers are groundbreaking architectures that have changed a flow of deep learning, and many high-performance models are developing based on transformer architectures. Transformers implemented only with attention with encoder-decoder…

Human-Computer Interaction · Computer Science 2021-12-20 Young-Eun Lee , Seo-Hyun Lee

Speech enhancement is widely used as a front-end to improve the speech quality in many audio systems, while it is hard to extract the target speech in multi-talker conditions without prior information on the speaker identity. It was shown…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-26 Jie Zhang , Qing-Tian Xu , Zhen-Hua Ling , Haizhou Li

Art has long played a profound role in shaping human emotion, cognition, and behavior. While visual arts such as painting and architecture have been studied through eye tracking, revealing distinct gaze patterns between experts and novices,…

Neurons and Cognition · Quantitative Biology 2025-12-08 Taketo Akama , Zhuohao Zhang , Tsukasa Nagashima , Takagi Yutaka , Shun Minamikawa , Natalia Polouliakh

We propose a novel neural network-based end-to-end acoustic echo cancellation (E2E-AEC) method capable of streaming inference, which operates effectively without reliance on traditional linear AEC (LAEC) techniques and time delay…

Sound · Computer Science 2026-01-26 Yiheng Jiang , Biao Tian , Haoxu Wang , Shengkui Zhao , Bin Ma , Daren Chen , Xiangang Li

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-13 Otavio Braga , Olivier Siohan

This paper proposes a novel framework for automatically capturing the time-frequency nature of electroencephalogram (EEG) signals of human sleep based on the authoritative sleep medicine guidance. The framework consists of two parts: the…

Machine Learning · Computer Science 2023-01-13 Zheng Chen , Ziwei Yang , Lingwei Zhu , Wei Chen , Toshiyo Tamura , Naoaki Ono , MD Altaf-Ul-Amin , Shigehiko Kanaya , Ming Huang

The supervised learning paradigm is limited by the cost - and sometimes the impracticality - of data collection and labeling in multiple domains. Self-supervised learning, a paradigm which exploits the structure of unlabeled data to create…

Decoding linguistic information from non-invasive brain signals using EEG has gained increasing research attention due to its vast applicational potential. Recently, a number of works have adopted a generative-based framework to decode…

Computation and Language · Computer Science 2024-08-12 Jinzhao Zhou , Yiqun Duan , Ziyi Zhao , Yu-Cheng Chang , Yu-Kai Wang , Thomas Do , Chin-Teng Lin

With the advancement of science and technology, the importance of emotion research has become increasingly evident. Electroencephalography (EEG)-based emotion recognition has emerged as an active research area in recent years, owing to its…

Human-Computer Interaction · Computer Science 2026-05-22 Ying Xie , Yi Zheng , Zehui Xiao , Wenkai Lu , Mengting Liu

A great challenge in speaker representation learning using deep models is to design learning objectives that can enhance the discrimination of unseen speakers under unseen domains. This work proposes a supervised contrastive learning…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Zhe Li , Man-Wai Mak

We present a model for separating a set of voices out of a sound mixture containing an unknown number of sources. Our Attentional Gating Network (AGN) uses a variable attentional context to specify which speakers in the mixture are of…

Sound · Computer Science 2019-05-28 Shariq Mobin , Bruno Olshausen

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-23 Xiao-Min Zeng , Yan Song , Zhu Zhuo , Yu Zhou , Yu-Hong Li , Hui Xue , Li-Rong Dai , Ian McLoughlin

In this paper, we propose a novel information theoretic model to interpret the entire "transmission chain" comprising stimulus generation, brain processing by the human subject, and the electroencephalograph (EEG) response measurements as a…

Information Theory · Computer Science 2015-09-15 Ketan Mehta , Jörg Kliewer

Electroencephalography (EEG)-based auditory attention detection (AAD) offers a non-invasive way to enhance hearing aids, but conventional methods rely on too many electrodes, limiting wearability and comfort. This paper presents SHAP-AAD, a…

Signal Processing · Electrical Eng. & Systems 2025-07-08 Rayan Salmi , Guorui Lu , Qinyu Chen

Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial…

Computation and Language · Computer Science 2018-06-19 Matthias Sperber , Jan Niehues , Graham Neubig , Sebastian Stüker , Alex Waibel

Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated paired data. However, obtaining high-quality paired data in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-24 Runwu Shi , Chang Li , Jiang Wang , Rui Zhang , Nabeela Khan , Benjamin Yen , Takeshi Ashizawa , Kazuhiro Nakadai

Cued Speech (CS) is a communication system for deaf people or hearing impaired people, in which a speaker uses it to aid a lipreader in phonetic level by clarifying potentially ambiguous mouth movements with hand shape and positions.…

Multimedia · Computer Science 2021-06-29 Jianrong Wang , Nan Gu , Mei Yu , Xuewei Li , Qiang Fang , Li Liu

Acoustic word embeddings (AWEs) are fixed-dimensional representations of variable-length speech segments. For zero-resource languages where labelled data is not available, one AWE approach is to use unsupervised autoencoder-based recurrent…

Computation and Language · Computer Science 2021-03-22 Christiaan Jacobs , Yevgen Matusevych , Herman Kamper

Decoding speech from non-invasive brain signals, such as electroencephalography (EEG), has the potential to advance brain-computer interfaces (BCIs), with applications in silent communication and assistive technologies for individuals with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-30 Terrance Yu-Hao Chen , Yulin Chen , Pontus Soederhaell , Sadrishya Agrawal , Kateryna Shapovalenko

Machine unlearning is a complex process that necessitates the model to diminish the influence of the training data while keeping the loss of accuracy to a minimum. Despite the numerous studies on machine unlearning in recent years, the…

Machine Learning · Computer Science 2024-05-14 Zixin Wang , Kongyang Chen
‹ Prev 1 8 9 10 Next ›