中文
相关论文

相关论文: Non-locally averaged pruned reassigned spectrogram…

200 篇论文

A phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due to their ability to disentangle pronunciation features from…

音频与语音处理 · 电气工程与系统科学 2024-02-28 Cameron Churchwell , Max Morrison , Bryan Pardo

Spurred by the demand for interpretable models, research on eXplainable AI for language technologies has experienced significant growth, with feature attribution methods emerging as a cornerstone of this progress. While prior work in NLP…

计算与语言 · 计算机科学 2025-03-18 Dennis Fucci , Marco Gaido , Beatrice Savoldi , Matteo Negri , Mauro Cettolo , Luisa Bentivogli

Source-tract decomposition (or glottal flow estimation) is one of the basic problems of speech processing. For this, several techniques have been proposed in the literature. However studies comparing different approaches are almost…

声音 · 计算机科学 2020-01-06 Thomas Drugman , Baris Bozkurt , Thierry Dutoit

This paper proposes APSS, a novel neural speech separation model with parallel amplitude and phase spectrum estimation. Unlike most existing speech separation methods, the APSS distinguishes itself by explicitly estimating the phase…

声音 · 计算机科学 2025-09-18 Fei Liu , Yang Ai , Zhen-Hua Ling

Fine-grained editing of speech attributes$\unicode{x2014}$such as prosody (i.e., the pitch, loudness, and phoneme durations), pronunciation, speaker identity, and formants$\unicode{x2014}$is useful for fine-tuning and fixing imperfections…

音频与语音处理 · 电气工程与系统科学 2024-07-09 Max Morrison , Cameron Churchwell , Nathan Pruyne , Bryan Pardo

The reconstruction of clipped speech signals is an important task in audio signal processing to achieve an enhanced audio quality for further processing. In this paper, Frequency Selective Extrapolation (FSE), which is commonly used for…

音频与语音处理 · 电气工程与系统科学 2022-04-11 Markus Jonscher , Jürgen Seiler , André Kaup

This paper proposes an efficient reconfigurable hardware design for speech enhancement based on multi band spectral subtraction algorithm and involving both magnitude and phase components. Our proposed design is novel as it estimates…

声音 · 计算机科学 2015-08-26 Tanmay Biswas , Sudhindu Bikash Mandal , Debasree Saha , Amlan Chakrabarti

The estimation of glottal flow from a speech waveform is a key method for speech analysis and parameterization. Significant research effort has been made to dissociate the first vocal tract resonance from the glottal formant (the…

音频与语音处理 · 电气工程与系统科学 2021-06-09 Olivier Perrotin , Ian Vince McLoughlin

Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features. In this study, we propose SpecGrad that adapts the diffusion noise so that…

音频与语音处理 · 电气工程与系统科学 2022-08-08 Yuma Koizumi , Heiga Zen , Kohei Yatabe , Nanxin Chen , Michiel Bacchiani

Ultrasound imaging is a cornerstone of non-invasive clinical diagnostics, yet its limited field of view poses challenges for novel view synthesis. We present UltraGS, a real-time framework that adapts Gaussian Splatting to sensorless…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Yuezhe Yang , Qingqing Ruan , Wenjie Cai , Yudang Dong , Dexin Yang , Xingbo Dong , Zhe Jin , Yong Dai

Tuning the frequency of a resonant element is of vital importance in both the macroscopic world, such as when tuning a musical instrument, as well as at the nanoscale. In particular, precisely controlling the resonance frequency of isolated…

介观与纳米尺度物理 · 物理学 2020-04-22 David J. Miller , Andrew Blaikie , Benjamin J. Aleman

The throat microphone is a body-attached transducer that is worn against the neck. It captures the signals that are transmitted through the vocal folds, along with the buzz tone of the larynx. Due to its skin contact, it is more robust to…

音频与语音处理 · 电气工程与系统科学 2018-04-18 Mehmet Ali Tugtekin Turan

Some glottal analysis approaches based upon linear prediction or complex cepstrum approaches have been proved to be effective to estimate glottal source from real speech utterances. We propose a new approach employing both an all-pole…

声音 · 计算机科学 2016-12-16 Yiqiao Chen , John N. Gowdy

Room Impulse Response (RIR) prediction at arbitrary receiver positions is essential for practical applications such as spatial audio rendering. We propose Neural Acoustic Multipole Splatting (NAMS), which synthesizes RIRs at unseen receiver…

音频与语音处理 · 电气工程与系统科学 2026-02-03 Geonwoo Baek , Jung-Woo Choi

Random pulse sequences are a powerful method for qubit noise spectroscopy, enabling efficient reconstruction of sparse noise spectra. Here, we advance this method in two complementary directions. First, we extend the method using a…

量子物理 · 物理学 2026-01-07 Kaixin Huang , Demitry Farfurnik , Dror Baron , Yi-Kai Liu

Detailed analysis of scanning probe microscopy (SPM) data acquired for faceted and non-flat surfaces is usually complicated due to the presence of a large number of surface areas tilted by large/variable angles relative to the scanning…

Pre-trained language models (PLM), for example BERT or RoBERTa, mark the state-of-the-art for natural language understanding task when fine-tuned on labeled data. However, their large size poses challenges in deploying them for inference in…

机器学习 · 计算机科学 2024-08-27 Aaron Klein , Jacek Golebiowski , Xingchen Ma , Valerio Perrone , Cedric Archambeau

We extend the APES ({\sl Amplitude and Phase Estimation}) method of spectral analysis to the case of non-quasi periodic signals like appears in OFDM.

数值分析 · 数学 2012-02-21 Jean-Philippe Préaux

A speech enhancement method based on probabilistic geometric approach to spectral subtraction (PGA) performed on short time magnitude spectrum is presented in this paper. A confidence parameter of noise estimation is introduced in the gain…

音频与语音处理 · 电气工程与系统科学 2018-02-15 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

Spatial audio is fundamental to immersive virtual experiences, yet synthesizing high-fidelity binaural audio from sparse observations remains a significant challenge. Existing methods typically rely on implicit neural representations…

声音 · 计算机科学 2026-04-13 Chunhao Bi , Houqiang Zhong , Zhixin Xu , Li Song , Zhengxue Cheng
‹ 上一页 1 2 3 10 下一页 ›