中文
相关论文

相关论文: Causal-Anticausal Decomposition of Speech using Co…

200 篇论文

This paper presents an Expert Decision Support System for the identification of time-invariant, aeroacoustic source types. The system comprises two steps: first, acoustic properties are calculated based on spectral and spatial information.…

声音 · 计算机科学 2022-03-09 Armin Goudarzi , Carsten Spehr , Steffen Herbold

The increasing size and complexity of medical imaging datasets, particularly in 3D formats, present significant barriers to collaborative research and transferability. This study investigates whether the ZFP compression technique can…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Shimaa Elbana , Ahmad Kamal , Shahd Ahmed Ali , Ahmad Al-Kabbany

Glottal Closure Instants (GCIs) correspond to the temporal locations of significant excitation to the vocal tract occurring during the production of voiced speech. GCI detection from speech signals is a well-studied problem given its…

声音 · 计算机科学 2019-07-11 Mohit Goyal , Varun Srivastava , Prathosh A. P

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based…

声音 · 计算机科学 2024-10-08 Kai Lia , Masato Akagia , Yongwei Lib , Masashi Unokia

Sonar systems are frequently used to classify objects at a distance by using the structure of the echoes of acoustic waves as a proxy for the object's shape and composition. Traditional synthetic aperture processing is highly effective in…

计算工程、金融与科学 · 计算机科学 2021-12-14 Michael Robinson

Glottal Closure Instants (GCI) detection consists in automatically detecting temporal locations of most significant excitation of the vocal tract from the speech signal. It is used in many speech analysis and processing applications, and…

音频与语音处理 · 电气工程与系统科学 2020-02-21 Luc Ardaillon , Axel Roebel

We introduce DeCaFlow, a deconfounding causal generative model. Training once per dataset using just observational data and the underlying causal graph, DeCaFlow enables accurate causal inference on continuous variables under the presence…

机器学习 · 计算机科学 2025-10-27 Alejandro Almodóvar , Adrián Javaloy , Juan Parras , Santiago Zazo , Isabel Valera

Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and/or prosodic channels independently is…

计算与语言 · 计算机科学 2024-08-21 Waris Quamer , Ricardo Gutierrez-Osuna

Formant tracking is one of the most fundamental problems in speech processing. Traditionally, formants are estimated using signal processing methods. Recent studies showed that generic convolutional architectures can outperform recurrent…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Wang Dai , Jinsong Zhang , Yingming Gao , Wei Wei , Dengfeng Ke , Binghuai Lin , Yanlu Xie

This paper introduces a computationally efficient technique for estimating high-resolution Doppler blood flow from an ultrafast ultrasound image sequence. More precisely, it consists in a new fast alternating minimization algorithm that…

图像与视频处理 · 电气工程与系统科学 2020-11-04 Duong-Hung Pham , Adrian Basarab , Jean-Pierre Remenieras , Paul Rodríguez , Denis Kouamé

We propose to combine cepstrum and nonlinear time-frequency (TF) analysis to study mutiple component oscillatory signals with time-varying frequency and amplitude and with time-varying non-sinusoidal oscillatory pattern. The concept of…

数据分析、统计与概率 · 物理学 2016-11-23 Chen-Yun Lin , Li Su , Hau-tieng Wu

The effectiveness of zero-shot classification in large vision-language models (VLMs), such as Contrastive Language-Image Pre-training (CLIP), depends on access to extensive, well-aligned text-image datasets. In this work, we introduce two…

计算机视觉与模式识别 · 计算机科学 2025-02-27 Anju Rani , Daniel O. Arroyo , Petar Durdevic

Zero-resource speech technology is a growing research area that aims to develop methods for speech processing in the absence of transcriptions, lexicons, or language modelling text. Early term discovery systems focused on identifying…

计算与语言 · 计算机科学 2017-09-19 Herman Kamper , Aren Jansen , Sharon Goldwater

While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation that unifies Foreign…

声音 · 计算机科学 2026-02-06 Jinzuomu Zhong , Korin Richmond , Zhiba Su , Siqi Sun

Traditional speech enhancement techniques modify the magnitude of a speech in time-frequency domain, and use the phase of a noisy speech to resynthesize a time domain speech. This work proposes a complex-valued Gaussian process latent…

声音 · 计算机科学 2017-01-02 Sih-Huei Chen , Yuan-Shan Lee , Jia-Ching Wang

Music source separation is important for applications such as karaoke and remixing. Much of previous research focuses on estimating short-time Fourier transform (STFT) magnitude and discarding phase information. We observe that, for singing…

音频与语音处理 · 电气工程与系统科学 2020-11-05 Yixuan Zhang , Yuzhou Liu , DeLiang Wang

Flow cytometry mainly used for detecting the characteristics of a number of biochemical substances based on the expression of specific markers in cells. It is particularly useful for detecting membrane surface receptors, antigens, ions, or…

机器学习 · 计算机科学 2023-03-17 Yanhua Xu

Purpose: This work explores the use of external phrase break prediction models to enhance listener comprehension in End-to-End Text-to-Speech (TTS) systems. Methods: The effectiveness of these models is evaluated based on listener…

音频与语音处理 · 电气工程与系统科学 2025-01-31 Anandaswarup Vadapalli

This paper introduces GlOttal-flow LPC Filter (GOLF), a novel method for singing voice synthesis (SVS) that exploits the physical characteristics of the human voice using differentiable digital signal processing. GOLF employs a glottal…

音频与语音处理 · 电气工程与系统科学 2024-10-21 Chin-Yun Yu , György Fazekas

Voiced segments of speech are assumed to be composed of non-stationary acoustic objects which can be described as stationary response of a non-stationary fundamental drive (FD) process and which are furthermore suited to reconstruct the…

声音 · 计算机科学 2007-05-23 Friedhelm R. Drepper