中文
相关论文

相关论文: Complex Cepstrum-based Decomposition of Speech for…

200 篇论文

This paper proposes a simple and effective approach for automatic recognition of Cued Speech (CS), a visual communication tool that helps people with hearing impairment to understand spoken language with the help of hand gestures that can…

计算与语言 · 计算机科学 2022-04-12 Sanjana Sankar , Denis Beautemps , Thomas Hueber

A novel text-independent speaker identification (SI) method is proposed. This method uses the Mel-frequency Cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature sets to capture speaker's…

声音 · 计算机科学 2020-02-04 Zhanyu Ma , Hong Yu

Synthesized speech is common today due to the prevalence of virtual assistants, easy-to-use tools for generating and modifying speech signals, and remote work practices. Synthesized speech can also be used for nefarious purposes, including…

声音 · 计算机科学 2022-05-05 Emily R. Bartusiak , Edward J. Delp

Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encouraging models to…

计算与语言 · 计算机科学 2026-01-05 Yuxin Li , Xiangyu Zhang , Yifei Li , Zhiwei Guo , Haoyang Zhang , Eng Siong Chng , Cuntai Guan

Systolic murmurs are extra heart sounds that occur during the contraction phase of the cardiac cycle, often indicating heart abnormalities caused by turbulent blood flow. Their intensity, pitch, and quality vary, requiring precise…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Mahmoud Fakhry , Abeer FathAllah Brery

Conventional NMF methods for source separation factorize the matrix of spectral magnitudes. Spectral Phase is not included in the decomposition process of these methods. However, phase of the speech mixture is generally used in…

声音 · 计算机科学 2014-11-26 Chaitanya Ahuja , Karan Nathwani , Rajesh M. Hegde

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide…

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio…

声音 · 计算机科学 2024-01-19 Yimin Deng , Huaizhen Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Pitch and Formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are essentially resonance…

音频与语音处理 · 电气工程与系统科学 2022-09-09 Seyedamiryousef Hosseini Goki , Mahdieh Ghazvini , Sajad Hamzenejadi

Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to…

音频与语音处理 · 电气工程与系统科学 2025-01-08 Wei Zhang , Tian-Hao Zhang , Chao Luo , Hui Zhou , Chao Yang , Xinyuan Qian , Xu-Cheng Yin

A novel feature, based on the chirp z-transform, that offers an improved representation of the underlying true spectrum is proposed. This feature, the chirp MFCC, is derived by computing the Mel frequency cepstral coefficients from the…

信号处理 · 电气工程与系统科学 2024-08-27 S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

Singing techniques are used for expressive vocal performances by employing temporal fluctuations of the timbre, the pitch, and other components of the voice. Their classification is a challenging task, because of mainly two factors: 1) the…

声音 · 计算机科学 2022-06-27 Yuya Yamamoto , Juhan Nam , Hiroko Terasawa

The Glottal Source is an important component of voice as it can be considered as the excitation signal to the voice apparatus. Nowadays, new techniques of speech processing such as speech recognition and speech synthesis use the glottal…

其他计算机科学 · 计算机科学 2010-04-20 Nikhil Raj , R. K. Sharma

This paper addresses the problem of high-resolution Doppler blood flow estimation from an ultrafast sequence of ultrasound images. Formulating the separation of clutter and blood components as an inverse problem has been shown in the…

图像与视频处理 · 电气工程与系统科学 2020-07-13 Duong-Hung Pham , Adrian Basarab , Ilyess Zemmoura , Jean-Pierre Remenieras , Denis Kouame

Mask-based blind speech separation (BSS) estimates source-wise time-frequency (TF) masks by clustering multichannel observations using spatial information. The directional statistical approach clusters normalized multichannel observations…

音频与语音处理 · 电气工程与系统科学 2026-05-26 Nobutaka Ito

In multi-speaker speech synthesis, data from a number of speakers usually tend to have great diversity due to the fact that the speakers may differ largely in ages, speaking styles, emotions, and so on. It is important but challenging to…

声音 · 计算机科学 2022-02-14 Qinghua Wu , Quanbo Shen , Jian Luan , YuJun Wang

Flow cytometry is a cornerstone technique in medical and biological research, providing crucial information about cell size and granularity through forward scatter (FSC) and side scatter (SSC) signals. Despite its widespread use, the…

光学 · 物理学 2024-08-29 Jaepil Jo , Herve Hugonnet , Mahn Jae Lee , YongKeun Park

Voiced segments of speech are assumed to be composed of non-stationary acoustic objects which can be described as stationary response of a non-stationary fundamental drive (FD) process and which are furthermore suited to reconstruct the…

声音 · 计算机科学 2007-05-23 Friedhelm R. Drepper

When a motor-powered vessel travels past a fixed hydrophone in a multipath environment, a Lloyd's mirror constructive/destructive interference pattern is observed in the output spectrogram. The power cepstrum detects the periodic structure…

声音 · 计算机科学 2018-10-30 Eric L. Ferguson , Stefan B. Williams , Craig T. Jin

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

声音 · 计算机科学 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud