中文
相关论文

相关论文: Improved Feature Extraction Network for Neuro-Orie…

200 篇论文

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on…

声音 · 计算机科学 2018-09-12 Mandar Gogate , Ahsan Adeel , Ricard Marxer , Jon Barker , Amir Hussain

Feature engineering for generalized seizure detection models remains a significant challenge. Recently proposed models show variable performance depending on the training data and remain ineffective at accurately distinguishing artifacts…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Kazi Mahmudul Hassan , Xuyang Zhao , Hidenori Sugano , Toshihisa Tanaka

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and…

声音 · 计算机科学 2025-05-29 Zhenghai You , Zhenyu Zhou , Lantian Li , Dong Wang

Deep speaker embeddings have been shown effective for assessing cognitive impairments aside from their original purpose of speaker verification. However, the research found that speaker embeddings encode speaker identity and an array of…

音频与语音处理 · 电气工程与系统科学 2022-03-22 Dongseok Heo , Cheul Young Park , Jaemin Cheun , Myung Jin Ko

Depression, a common mental disorder, significantly influences individuals and imposes considerable societal impacts. The complexity and heterogeneity of the disorder necessitate prompt and effective detection, which nonetheless, poses a…

声音 · 计算机科学 2023-08-25 Xiao Xu , Yang Wang , Xinru Wei , Fei Wang , Xizhe Zhang

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded…

声音 · 计算机科学 2026-04-28 Ziyang Jiang , Jiahe Lei , Xueyan Chen , Yifan Zhang , Zexu Pan , Wei Xue , Xinyuan Qian

Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech…

声音 · 计算机科学 2021-03-01 Jiatong Shi , Chunlei Zhang , Chao Weng , Shinji Watanabe , Meng Yu , Dong Yu

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which depends on the video…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Yidi Jiang , Ruijie Tao , Zexu Pan , Haizhou Li

Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or…

音频与语音处理 · 电气工程与系统科学 2025-10-20 Tongtao Ling , Shulin He , Pengjie Shen , Zhong-Qiu Wang

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement.…

声音 · 计算机科学 2022-10-31 Shulin He , Wei Rao , Jinjiang Liu , Jun Chen , Yukai Ju , Xueliang Zhang , Yannan Wang , Shidong Shang

While deep learning based speech enhancement systems have made rapid progress in improving the quality of speech signals, they can still produce outputs that contain artifacts and can sound unnatural. We propose a novel approach to speech…

声音 · 计算机科学 2022-07-12 Muqiao Yang , Joseph Konan , David Bick , Anurag Kumar , Shinji Watanabe , Bhiksha Raj

Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Haoshuai Zhou , Changgeng Mo , Boxuan Cao , Linkai Li , Shan Xiang Wang

Speech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of…

声音 · 计算机科学 2021-03-05 Panagiotis Tzirakis , Anh Nguyen , Stefanos Zafeiriou , Björn W. Schuller

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

We introduce EffiFusion-GAN (Efficient Fusion Generative Adversarial Network), a lightweight yet powerful model for speech enhancement. The model integrates depthwise separable convolutions within a multi-scale block to capture diverse…

声音 · 计算机科学 2025-08-21 Bin Wen , Tien-Ping Tan

Speaker embeddings achieve promising results on many speaker verification tasks. Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings. In this paper, we introduce phonetic…

声音 · 计算机科学 2018-06-15 Yi Liu , Liang He , Jia Liu , Michael T. Johnson

In the past decade, Convolutional Neural Networks (CNNs) and Transformers have achieved wide applicaiton in semantic segmentation tasks. Although CNNs with Transformer models greatly improve performance, the global context modeling remains…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Feixiang Du , Shengkun Wu

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

多媒体 · 计算机科学 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neural diarization (EEND)…

音频与语音处理 · 电气工程与系统科学 2022-12-19 Soumi Maiti , Yushi Ueda , Shinji Watanabe , Chunlei Zhang , Meng Yu , Shi-Xiong Zhang , Yong Xu

Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when the loudspeaker plays back. In this paper, we present an…

音频与语音处理 · 电气工程与系统科学 2022-05-24 Meng Yu , Yong Xu , Chunlei Zhang , Shi-Xiong Zhang , Dong Yu