中文
相关论文

相关论文: Glottal Source Estimation using an Automatic Chirp…

200 篇论文

This paper addresses the problem of estimating entropy-regularized optimal transport (EOT) maps with squared-Euclidean cost between source and target measures that are subGaussian. In the case that the target measure is compactly supported…

机器学习 · 统计学 2023-11-21 Matthew Werenski , James M. Murphy , Shuchin Aeron

Acoustic articulatory inversion is a major processing challenge, with a wide range of applications from speech synthesis to feedback systems for language learning and rehabilitation. In recent years, deep learning methods have been applied…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

An efficient new approach to signal compression is presented based of a novel variation on the Gabor basis set. Following earlier work by Shimshovitz and Tannor, we convolve the conventional Gabor functions with Dirichlet functions to…

音频与语音处理 · 电气工程与系统科学 2025-04-01 Roger Alimi , David J. Tannor

Compositional Zero-shot Learning (CZSL) aims to recognize novel concepts composed of known knowledge without training samples. Standard CZSL either identifies visual primitives or enhances unseen composed entities, and as a result,…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Xiaocheng Lu , Ziming Liu , Song Guo , Jingcai Guo , Fushuo Huo , Sikai Bai , Tao Han

Zero-shot learning (ZSL) aims to classify objects that are not observed or seen during training. It relies on class semantic description to transfer knowledge from the seen classes to the unseen classes. Existing methods of obtaining class…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Fahimul Hoque Shubho , Townim Faisal Chowdhury , Ali Cheraghian , Morteza Saberi , Nabeel Mohammed , Shafin Rahman

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on…

声音 · 计算机科学 2025-05-27 Zhichao Wu , Yueteng Kang , Songjun Cao , Long Ma , Qiulin Li , Qun Yang

In recent years, high-speed videoendoscopy (HSV) has significantly aided the diagnosis of voice pathologies and furthered the understanding the voice production in recent years. As the first step of these studies, automatic segmentation of…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Xin Chen , Emma Marriott , Yuling Yan

Humans are good at compositional zero-shot reasoning; someone who has never seen a zebra before could nevertheless recognize one when we tell them it looks like a horse with black and white stripes. Machine learning systems, on the other…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Frank Ruis , Gertjan Burghouts , Doina Bucur

A method for analyzing sampling jitter in audio equipment is proposed. The method is based on the time-domain analysis where the time fluctuations of zero-crossing points in recorded sinusoidal waves are employed to characterize jitter.…

声音 · 计算机科学 2023-05-09 Makoto Takeuchi , Haruo Saito

The problem to establish not only the asymptotic distribution results for statistical estimators but also the moment convergence of the estimators has been recognized as an important issue in advanced theories of statistics. One of the main…

统计理论 · 数学 2012-07-02 Ilia Negri , Yoichi Nishiyama

This paper introduces GlOttal-flow LPC Filter (GOLF), a novel method for singing voice synthesis (SVS) that exploits the physical characteristics of the human voice using differentiable digital signal processing. GOLF employs a glottal…

音频与语音处理 · 电气工程与系统科学 2024-10-21 Chin-Yun Yu , György Fazekas

We study the performance of three different methods to automatically detect a chirp in background noise. (1) The standard deviation detector uses the computation of the signal to noise ratio. (2) The spectral covariance detector is based on…

数据分析、统计与概率 · 物理学 2007-05-23 M. Ignaccolo , T. Farges , E. Blanc , M. Fullekrug

We study stochastic zeroth-order (ZO) optimization of smooth nonconvex objectives under heavy-tailed sample-gradient noise. This regime is motivated by empirical evidence that gradient noise in modern machine learning can violate the…

最优化与控制 · 数学 2026-05-19 Taha El Bakkali , El Mahdi Chayti , Qiuyi Zhang , Imane Rahali , Omar Saadi

Linear time-varying (LTV) systems model radar scenes where each reflector/target applies a delay, Doppler shift and complex amplitude scaling to a transmitted waveform. The receiver processes the received signal using the transmitted signal…

信号处理 · 电气工程与系统科学 2025-03-25 Danish Nisar , Saif Khan Mohammed , Ronny Hadani , Ananthanarayanan Chockalingam , Robert Calderbank

Conventional automatic speech recognition (ASR) systems trained from frame-level alignments can easily leverage posterior fusion to improve ASR accuracy and build a better single model with knowledge distillation. End-to-end ASR systems…

计算与语言 · 计算机科学 2019-07-03 Gakuto Kurata , Kartik Audhkhasi

Pearson's chi-squared test is widely used to assess the uniformity of discrete histograms, typically relying on a continuous chi-squared distribution to approximate the test statistic, since computing the exact distribution is…

统计方法学 · 统计学 2025-07-01 Nikola Banić , Neven Elezović

Accurately characterizing the true redshift (true-$z$) distribution of a photometric redshift (photo-$z$) sample is critical for cosmological analyses in imaging surveys. Clustering-based techniques, which include clustering-redshift (CZ)…

宇宙学与河外天体物理 · 物理学 2024-12-18 Weilun Zheng , Kwan Chuen Chan , Haojie Xu , Le Zhang , Ruiyu Song

We present a new technique that we have defined as the z-scan confocal method to determine the absolute location and size of the focal spot in a tight focused ultrashort laser pulse. The method permits to accurately position a target in the…

仪器与探测器 · 物理学 2017-05-01 P. Castro-Marín , G. Castro-Olvera , C. Ruíz , J. Garduño-Mejía , M. Rosete-Aguilar , N. C. Bruce

Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result, synthesizing speech with a desired style often requires…

音频与语音处理 · 电气工程与系统科学 2026-04-21 Haitao Li , Chunxiang Jin , Chenglin Li , Wenhao Guan , Zhengxing Huang , Xie Chen