中文
相关论文

相关论文: Harmonic-aligned Frame Mask Based on Non-stationar…

200 篇论文

We consider sparseness properties of adaptive time-frequency representations obtained using nonstationary Gabor frames (NSGFs). NSGFs generalize classical Gabor frames by allowing for adaptivity in either time or frequency. It is known that…

泛函分析 · 数学 2018-01-03 Emil Solsbæk Ottosen , Morten Nielsen

Frame alignments can be computed by different methods in GMM-based speaker verification. By incorporating a phonetic Gaussian mixture model (PGMM), we are able to compare the performance using alignments extracted from the deep neural…

声音 · 计算机科学 2018-09-05 Yi Liu , Liang He , Weiqiang Zhang , Jia Liu , Michael T. Johnson

We propose a new algorithm for time stretching music signals based on the theory of nonstationary Gabor frames (NSGFs). The algorithm extends the techniques of the classical phase vocoder (PV) by incorporating adaptive time-frequency (TF)…

声音 · 计算机科学 2017-09-14 Emil Solsbæk Ottosen , Monika Dörfler

Audio signal processing frequently requires time-frequency representations and in many applications, a non-linear spacing of frequency-bands is preferable. This paper introduces a framework for efficient implementation of invertible signal…

泛函分析 · 数学 2013-05-17 Nicki Holighaus , Monika Dörfler , Gino Angelo Velasco , Thomas Grill

This letter proposes a novel blind acoustic mask (BAM) designed to adaptively detect noise components and preserve target speech segments in time-domain. A robust standard deviation estimator is applied to the non-stationary noisy speech to…

音频与语音处理 · 电气工程与系统科学 2021-07-07 F. Farias , R. Coelho

Certain signal classes such as audio signals call for signal representations with the ability to adapt to the signal's properties. In this article we introduce the new concept of quilted frames, which aim at adaptivity in time-frequency…

泛函分析 · 数学 2010-01-25 Monika Doerfler

In this paper, we apply star-Digital Gabor Transform in analysis Compressed Sensing and speech denoising. Based on assumptions on the ambient dimension, we produce a window vector that generates a spark deficient Gabor frame with many…

数值分析 · 数学 2021-10-19 Vasiliki Kouni , Holger Rauhut

Nonstationary Gabor frames, recently introduced in adaptive signal analysis, represent a natural generalization of classical Gabor frames by allowing for adaptivity of windows and lattice in either time or frequency. Due to the lack of a…

泛函分析 · 数学 2013-01-10 Monika Dörfler , Ewa Matusiak

This paper introduces Gabor scattering, a feature extractor based on Gabor frames and Mallat's scattering transform. By using a simple signal model for audio signals specific properties of Gabor scattering are studied. It is shown that for…

声音 · 计算机科学 2019-10-02 Roswitha Bammer , Monika Dörfler , Pavol Harar

In this study, we present an innovative technique for speaker adaptation in order to improve the accuracy of segmentation with application to unit-selection Text-To-Speech (TTS) systems. Unlike conventional techniques for speaker…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Claudio Zito , Fabio Tesser , Mauro Nicolao , Piero Cosi

In this paper, we address the speech denoising problem, where Gaussian, pink and blue additive noises are to be removed from a given speech signal. Our approach is based on a redundant, analysis-sparse representation of the original speech…

声音 · 计算机科学 2023-04-27 Vasiliki Kouni , Holger Rauhut , Theoharis Theoharis

The mechanism proposed here is for real-time speaker change detection in conversations, which firstly trains a neural network text-independent speaker classifier using in-domain speaker data. Through the network, features of conversational…

声音 · 计算机科学 2017-03-20 Zhenhao Ge , Ananth N. Iyer , Srinath Cheluvaraja , Aravind Ganapathiraju

Recent works of utilizing phonetic posteriograms (PPGs) for non-parallel voice conversion have significantly increased the usability of voice conversion since the source and target DBs are no longer required for matching contents. In this…

音频与语音处理 · 电气工程与系统科学 2020-06-15 Sunghee Jung , Youngjoo Suh , Yeunju Choi , Hoirin Kim

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

Aligning acoustic and linguistic representations is a central challenge to bridge the pre-trained models in knowledge transfer for automatic speech recognition (ASR). This alignment is inherently structured and asymmetric: while multiple…

计算与语言 · 计算机科学 2026-03-06 Xugang Lu , Peng Shen , Hisashi Kawai

We present Multiscale Audio Spectrogram Transformer (MAST) for audio classification, which brings the concept of multiscale feature hierarchies to the Audio Spectrogram Transformer (AST). Given an input audio spectrogram, we first patchify…

音频与语音处理 · 电气工程与系统科学 2023-05-19 Sreyan Ghosh , Ashish Seth , S. Umesh , Dinesh Manocha

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Hongyu Wang , Hui Li , Bo Li

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

声音 · 计算机科学 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

In recent years generative adversarial network (GAN) based models have been successfully applied for unsupervised speech-to-speech conversion.The rich compact harmonic view of the magnitude spectrogram is considered a suitable choice for…

音频与语音处理 · 电气工程与系统科学 2020-05-20 Mohammad Asif Khan , Fabien Cardinaux , Stefan Uhlich , Marc Ferras , Asja Fischer

Dereverberation of recorded speech signals is one of the most pertinent problems in speech processing. In the present work, the objective is to understand and implement dereverberation techniques that aim at enhancing the magnitude…

音频与语音处理 · 电气工程与系统科学 2025-12-30 Dhruv Nigam
‹ 上一页 1 2 3 10 下一页 ›