中文
相关论文

相关论文: Radically Old Way of Computing Spectra: Applicatio…

200 篇论文

Conventional Frequency Domain Linear Prediction (FDLP) technique models the squared Hilbert envelope of speech with varied degrees of approximation which can be sampled at the required frame rate and used as features for Automatic Speech…

声音 · 计算机科学 2022-04-04 Samik Sadhu , Hynek Hermansky

The task of speech recognition in far-field environments is adversely affected by the reverberant artifacts that elicit as the temporal smearing of the sub-band envelopes. In this paper, we develop a neural model for speech dereverberation…

音频与语音处理 · 电气工程与系统科学 2021-08-21 Anurenjan Purushothaman , Anirudh Sreeram , Rohit Kumar , Sriram Ganapathy

Prediction of late reverberation component using multi-channel linear prediction (MCLP) in short-time Fourier transform (STFT) domain is an effective means to enhance reverberant speech. Traditionally, a speech power spectral density (PSD)…

音频与语音处理 · 电气工程与系统科学 2018-12-05 Srikanth Raj Chetupalli , Thippur V. Sreenivas

The end-to-end (E2E) automatic speech recognition (ASR) systems are often required to operate in reverberant conditions, where the long-term sub-band envelopes of the speech are temporally smeared. In this paper, we develop a feature…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Rohit Kumar , Anurenjan Purushothaman , Anirudh Sreeram , Sriram Ganapathy

We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody. We present a generic multi-scale spectrogram prediction mechanism where the system first predicts…

音频与语音处理 · 电气工程与系统科学 2021-07-01 Ammar Abbas , Bajibabu Bollepalli , Alexis Moinet , Arnaud Joly , Penny Karanasou , Peter Makarov , Simon Slangens , Sri Karlapati , Thomas Drugman

Historically, most speech models in machine-learning have used the mel-spectrogram as a speech representation. Recently, discrete audio tokens produced by neural audio codecs have become a popular alternate speech representation for speech…

音频与语音处理 · 电气工程与系统科学 2025-06-05 Ryan Langman , Ante Jukić , Kunal Dhawan , Nithin Rao Koluguri , Jason Li

Automatic speech recognition in reverberant conditions is a challenging task as the long-term envelopes of the reverberant speech are temporally smeared. In this paper, we propose a neural model for enhancement of sub-band temporal…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Anurenjan Purushothaman , Anirudh Sreeram , Rohit Kumar , Sriram Ganapathy

We present a method to remove unknown convolutive noise introduced to speech by reverberations of recording environments, utilizing some amount of training speech data from the reverberant environment, and any available non-reverberant…

音频与语音处理 · 电气工程与系统科学 2022-10-04 Samik Sadhu , Hynek Hermansky

We propose multi-layer perceptron (MLP)-based architectures suitable for variable length input. MLP-based architectures, recently proposed for image classification, can only be used for inputs of a fixed, pre-defined size. However, many…

音频与语音处理 · 电气工程与系统科学 2022-02-18 Jin Sakuma , Tatsuya Komatsu , Robin Scheibler

Several methods have recently been proposed to analyze speech and automatically infer the personality of the speaker. These methods often rely on prosodic and other hand crafted speech processing features extracted with off-the-shelf…

计算机视觉与模式识别 · 计算机科学 2022-05-10 Marc-André Carbonneau , Eric Granger , Yazid Attabi , Ghyslain Gagnon

We propose an end-to-end Automatic Speech Recognition (ASR) system that can be trained on transcribed speech data, text-only data, or a mixture of both. The proposed model uses an integrated auxiliary block for text-based training. This…

音频与语音处理 · 电气工程与系统科学 2024-07-08 Vladimir Bataev , Roman Korostik , Evgeny Shabalin , Vitaly Lavrukhin , Boris Ginsburg

Frequency modulation (FM) is a form of radio broadcasting which is widely used nowadays and has been for almost a century. We suggest a software-defined-radio (SDR) receiver for FM demodulation that adopts an end-to-end learning based…

机器学习 · 计算机科学 2017-10-10 Dan Elbaz , Michael Zibulevsky

The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however,…

We compare a wide band sub-band speech coder using ADPCM schemes with linear prediction against the same scheme with nonlinear prediction based on multi-layer perceptrons. Exhaustive results are presented in each band, and the full signal.…

声音 · 计算机科学 2022-03-25 Marcos Faundez-Zanuy

Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…

声音 · 计算机科学 2019-04-17 Gene-Ping Yang , Chao-I Tuan , Hung-Yi Lee , Lin-shan Lee

Training a robust Automatic Speech Recognition (ASR) system for children's speech recognition is a challenging task due to inherent differences in acoustic attributes of adult and child speech and scarcity of publicly available children's…

音频与语音处理 · 电气工程与系统科学 2022-03-15 Vishwanath Pratap Singh , Hardik Sailor , Supratik Bhattacharya , Abhishek Pandey

Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care and to delay further progression. Speech based automatic AD screening systems provide a non-intrusive and more scalable alternative to other clinical…

计算与语言 · 计算机科学 2023-04-03 Yi Wang , Jiajun Deng , Tianzi Wang , Bo Zheng , Shoukang Hu , Xunying Liu , Helen Meng

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

音频与语音处理 · 电气工程与系统科学 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan

This paper addresses the problem of correctly formatting numeric expressions in automatic speech recognition (ASR) transcripts. This is challenging since the expected transcript format depends on the context, e.g., 1945 (year) vs. 19:45…

音频与语音处理 · 电气工程与系统科学 2025-06-24 Christian Huber , Alexander Waibel

Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (AuM) adopt square…

声音 · 计算机科学 2025-09-01 Aditya Makineni , Baocheng Geng , Qing Tian
‹ 上一页 1 2 3 10 下一页 ›