中文
相关论文

相关论文: Enhancing End-to-End Multi-channel Speech Separati…

200 篇论文

We investigate the potential of stochastic neural networks for learning effective waveform-based acoustic models. The waveform-based setting, inherent to fully end-to-end speech recognition systems, is motivated by several comparative…

机器学习 · 统计学 2021-08-17 Dino Oglic , Zoran Cvetkovic , Peter Sollich

Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Taous Iatariene , Alexandre Guérin , Romain Serizel

Robust local feature representations are essential for spatial intelligence tasks such as robot navigation and augmented reality. Establishing reliable correspondences requires descriptors that provide both high discriminative power and…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Haodi Yao , Fenghua He , Ning Hao , Yao Su

Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task.…

音频与语音处理 · 电气工程与系统科学 2025-07-11 Hiroshi Sato , Tsubasa Ochiai , Marc Delcroix , Takafumi Moriya , Takanori Ashihara , Ryo Masumura

Contextual information is crucial for semantic segmentation. However, finding the optimal trade-off between keeping desired fine details and at the same time providing sufficiently large receptive fields is non trivial. This is even more…

计算机视觉与模式识别 · 计算机科学 2017-09-08 Yang He , Margret Keuper , Bernt Schiele , Mario Fritz

In recent years, deep learning has presented a great advance in hyperspectral image (HSI) classification. Particularly, long short-term memory (LSTM), as a special deep learning structure, has shown great ability in modeling long-term…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Wen-Shuai Hu , Heng-Chao Li , Lei Pan , Wei Li , Ran Tao , Qian Du

We address the problem of acoustic source separation in a deep learning framework we call "deep clustering." Rather than directly estimating signals or masking functions, we train a deep network to produce spectrogram embeddings that are…

神经与进化计算 · 计算机科学 2015-08-19 John R. Hershey , Zhuo Chen , Jonathan Le Roux , Shinji Watanabe

Deep superpixel algorithms have made remarkable strides by substituting hand-crafted features with learnable ones. Nevertheless, we observe that existing deep superpixel methods, serving as mid-level representation operations, remain…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Sen Xu , Shikui Wei , Tao Ruan , Lixin Liao

In this work, we propose a new approach for language identification using multi-head self-attention combined with raw waveform based 1D convolutional neural networks for Indian languages. Our approach uses an encoder, multi-head…

音频与语音处理 · 电气工程与系统科学 2021-02-02 Krishna D N , Ankita Patil

The current dominant approach for neural speech enhancement is based on supervised learning by using simulated training data. The trained models, however, often exhibit limited generalizability to real-recorded data. To address this, this…

音频与语音处理 · 电气工程与系统科学 2025-03-25 Zhong-Qiu Wang

Self-supervised learning has demonstrated impressive performance in speech tasks, yet there remains ample opportunity for advancement in the realm of speech enhancement research. In addressing speech tasks, confining the attention mechanism…

音频与语音处理 · 电气工程与系统科学 2024-08-14 Tao Zheng , Liejun Wang , Yinfeng Yu

Discriminative segmental models offer a way to incorporate flexible feature functions into speech recognition. However, their appeal has been limited by their computational requirements, due to the large number of possible segments to…

计算与语言 · 计算机科学 2016-08-03 Hao Tang , Weiran Wang , Kevin Gimpel , Karen Livescu

Automatic recognition of disordered speech remains a highly challenging task to date. Sources of variability commonly found in normal speech including accent, age or gender, when further compounded with the underlying causes of speech…

声音 · 计算机科学 2022-01-20 Mengzhe Geng , Shansong Liu , Jianwei Yu , Xurong Xie , Shoukang Hu , Zi Ye , Zengrui Jin , Xunying Liu , Helen Meng

Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned,…

声音 · 计算机科学 2023-02-28 Xue Jiang , Xiulian Peng , Yuan Zhang , Yan Lu

Speech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enhancement, e.g. the…

音频与语音处理 · 电气工程与系统科学 2024-02-23 Changjiang Zhao , Shulin He , Xueliang Zhang

In this paper, we propose a new wireless video communication scheme to achieve high-efficiency video transmission over noisy channels. It exploits the idea of model division multiple access (MDMA) and extracts common semantic features…

多媒体 · 计算机科学 2023-05-26 Zhicheng Bao , Haotai Liang , Chen Dong , Xiaodong Xu , Geng Liu

A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal statistics for time-invariant minimum variance…

声音 · 计算机科学 2021-10-04 Zhong-Qiu Wang , Gordon Wichern , Jonathan Le Roux

Neural networks have recently become the dominant approach to sound separation. Their good performance relies on large datasets of isolated recordings. For speech and music, isolated single channel data are readily available; however the…

声音 · 计算机科学 2024-10-02 Jacob Kealey , John Hershey , François Grondin

Automatic Speech Recognition (ASR) has shown remarkable progress, yet it still faces challenges in real-world distant scenarios across various array topologies each with multiple recording devices. The focal point of the CHiME-7 Distant ASR…

声音 · 计算机科学 2023-12-18 Bingshen Mu , Pengcheng Guo , Dake Guo , Pan Zhou , Wei Chen , Lei Xie

This work presents a noise reduction method with perceptually relevant preservation of the interaural time difference (ITD) of the residual noise in binaural hearing aids. The interaural coherence (IC) concept, previously applied to the…

音频与语音处理 · 电气工程与系统科学 2018-06-26 Fábio P. Itturriet , Márcio H. Costa