中文
相关论文

相关论文: Glottal Source Estimation using an Automatic Chirp…

200 篇论文

Open-vocabulary audio-language models, like CLAP, offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper,…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Sreyan Ghosh , Sonal Kumar , Chandra Kiran Reddy Evuru , Oriol Nieto , Ramani Duraiswami , Dinesh Manocha

While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation that unifies Foreign…

声音 · 计算机科学 2026-02-06 Jinzuomu Zhong , Korin Richmond , Zhiba Su , Siqi Sun

The great majority of current voice technology applications relies on acoustic features characterizing the vocal tract response, such as the widely used MFCC of LPC parameters. Nonetheless, the airflow passing through the vocal folds, and…

声音 · 计算机科学 2020-01-01 Thomas Drugman , Paavo Alku , Abeer Alwan , Bayya Yegnanarayana

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

Spike deconvolution is the problem of recovering point sources from their convolution with a known point spread function, playing a fundamental role in many sensing and imaging applications. This paper proposes a novel approach combining…

信号处理 · 电气工程与系统科学 2025-02-13 Joseph Gabet , Meghna Kalra , Maxime Ferreira Da Costa , Kiryung Lee

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based…

声音 · 计算机科学 2024-10-08 Kai Lia , Masato Akagia , Yongwei Lib , Masashi Unokia

In this work, we incorporated acoustically derived source features, aperiodicity, periodicity and pitch as additional targets to an acoustic-to-articulatory speech inversion (SI) system. We also propose a Temporal Convolution based SI…

音频与语音处理 · 电气工程与系统科学 2022-11-01 Yashish M. Siriwardena , Carol Espy-Wilson

The estimation of glottal flow from a speech waveform is a key method for speech analysis and parameterization. Significant research effort has been made to dissociate the first vocal tract resonance from the glottal formant (the…

音频与语音处理 · 电气工程与系统科学 2021-06-09 Olivier Perrotin , Ian Vince McLoughlin

An inverse scattering problem is analyzed for vowel articulation in the human vocal tract. When a unit amplitude, monochromatic, sinusoidal volume velocity is sent from the glottis towards the lips, various types of scattering data are used…

数学物理 · 物理学 2007-05-23 Tuncay Aktosun

A sound source was proposed for acoustic measurements of physical models of the human vocal tract. The physical models are produced by Fast Prototyping, based on Magnetic Resonance Imaging during prolonged vowel production. The sound…

仪器与探测器 · 物理学 2017-11-22 Antti Hannukainen , Juha Kuortti , Jarmo Malinen , Antti Ojalammi

Diffusion/score-based models have recently emerged as powerful generative priors for solving inverse problems, including accelerated MRI reconstruction. While their flexibility allows decoupling the measurement model from the learned prior,…

图像与视频处理 · 电气工程与系统科学 2025-09-15 Yaşar Utku Alçalar , Junno Yun , Mehmet Akçakaya

In this paper we present a blind deconvolution scheme based on statistical wavelet estimation. We assume no prior knowledge of the wavelet, and do not select a reflector from the signal. Instead, the wavelet (ultrasound pulse) is…

其他计算机科学 · 计算机科学 2015-06-01 Roberto H. Herrera , Zhaorui Liu , Natasha Raffa , Paul Christensen , Adrianus Elvers

In this paper, we propose a new method to estimate instantaneous frequency using a combined approach based on the discrete linear chirp transform (DLCT) and the Wigner distribution (WD). The DLCT locally represents a signal as a…

信号处理 · 电气工程与系统科学 2018-10-15 Osama A. Alkishriwo , Luis F. Chaparro

In this paper, a discrete LCT (DLCT) irrelevant to the sampling periods and without oversampling operation is developed. This DLCT is based on the well-known CM-CC-CM decomposition, that is, implemented by two discrete chirp multiplications…

信息论 · 计算机科学 2017-09-20 Soo-Chang Pei , Shih-Gu Huang

Recent advances in the chirplet transform and wavelet-chirplet transform (WCT) have enabled the estimation of instantaneous frequencies (IFs) and chirprates, as well as mode retrieval from multicomponent signals with crossover IF curves.…

信号处理 · 电气工程与系统科学 2025-08-26 Qingtang Jiang , Shuixin Li , Jiecheng Chen , Lin Li

Quantitative computed tomography (QCT) is a widely used tool for osteoporosis diagnosis and monitoring. The assessment of cortical markers like cortical bone mineral density (BMD) and thickness is a demanding task, mainly because of the…

图像与视频处理 · 电气工程与系统科学 2019-02-18 Stefan Reinhold. Timo Damm , Lukas Huber , Reimer Andresen , Reinhard Barkmann , Claus-C. Glüer , Reinhard Koch

Zero-Shot Cross-lingual Transfer (ZS-XLT) utilizes a model trained in a source language to make predictions in another language, often with a performance loss. To alleviate this, additional improvements can be achieved through subsequent…

计算与语言 · 计算机科学 2024-04-04 Emilio Villa-Cueva , A. Pastor López-Monroy , Fernando Sánchez-Vega , Thamar Solorio

Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Han Zhu , Wei Kang , Zengwei Yao , Liyong Guo , Fangjun Kuang , Zhaoqing Li , Weiji Zhuang , Long Lin , Daniel Povey

Denoising of broadband non--stationary signals is a challenging problem in communication systems. In this paper, we introduce a time-varying filter algorithm based on the discrete linear chirp transform (DLCT), which provides local signal…

信号处理 · 电气工程与系统科学 2018-10-15 Osama A. S. Alkishriwo , Ali A. Elghariani , Aydin Akan

The graph linear canonical transform (GLCT) is presented as an extension of the graph Fourier transform (GFT) and the graph fractional Fourier transform (GFrFT), offering more flexibility as an effective tool for graph signal processing. In…

信息论 · 计算机科学 2024-07-26 Na Li , Zhichao Zhang , Jie Han , Yunjie Chen , Chunzheng Cao