中文
相关论文

相关论文: Complex Cepstrum-based Decomposition of Speech for…

200 篇论文

Many turbulent flows exhibit time-periodic statistics. These include turbomachinery flows, flows with external harmonic forcing, and the wakes of bluff bodies. Many existing techniques for identifying turbulent coherent structures, however,…

流体动力学 · 物理学 2024-05-01 Liam Heidt , Tim Colonius

Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain.…

声音 · 计算机科学 2019-03-19 Ziqiang Shi , Huibin Lin , Liu Liu , Rujie Liu , Shoji Hayakawa , Shouji Harada , Jiqing Han

The speech feature extraction has been a key focus in robust speech recognition research; it significantly affects the recognition performance. In this paper, we first study a set of different features extraction methods such as linear…

计算与语言 · 计算机科学 2014-07-01 Imen Trabelsi , Dorra Ben Ayed

A weakly-supervised semantic segmentation framework with a tied deconvolutional neural network is presented. Each deconvolution layer in the framework consists of unpooling and deconvolution operations. 'Unpooling' upsamples the input…

计算机视觉与模式识别 · 计算机科学 2016-03-15 Hyo-Eun Kim , Sangheum Hwang

Open-vocabulary segmentation is a challenging task requiring segmenting and recognizing objects from an open set of categories. One way to address this challenge is to leverage multi-modal models, such as CLIP, to provide image and text…

计算机视觉与模式识别 · 计算机科学 2023-11-16 Qihang Yu , Ju He , Xueqing Deng , Xiaohui Shen , Liang-Chieh Chen

Blind Source Separation (BSS) has proven to be a powerful tool for the analysis of composite patterns in engineering and science. We introduce Convex Analysis of Mixtures (CAM) for separating non-negative well-grounded sources, which learns…

机器学习 · 统计学 2015-12-14 Yitan Zhu , Niya Wang , David J. Miller , Yue Wang

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

声音 · 计算机科学 2022-06-20 Yikang Wang , Hiromitsu Nishizaki

The application of machine learning to medical ultrasound videos of the heart, i.e., echocardiography, has recently gained traction with the availability of large public datasets. Traditional supervised tasks, such as ejection fraction…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Grégoire Petit , Nathan Palluau , Axel Bauer , Clemens Dlaska

Objective: This study introduces ChatSchema, an effective method for extracting and structuring information from unstructured data in medical paper reports using a combination of Large Multimodal Models (LMMs) and Optical Character…

计算与语言 · 计算机科学 2024-07-29 Fei Wang , Yuewen Zheng , Qin Li , Jingyi Wu , Pengfei Li , Luxia Zhang

This article investigates into recently emerging approaches that use deep neural networks for the estimation of glottal closure instants (GCI). We build upon our previous approach that used synthetic speech exclusively to create perfectly…

音频与语音处理 · 电气工程与系统科学 2020-03-04 Frederik Bous , Luc Ardaillon , Axel Roebel

Over 70 million people worldwide experience stuttering, yet most automatic speech systems misinterpret disfluent utterances or fail to transcribe them accurately. Existing methods for stutter correction rely on handcrafted feature…

音频与语音处理 · 电气工程与系统科学 2025-11-06 Qianheng Xu

Graph-based Transform (GT) has been recently leveraged successfully in the signal processing domain, specifically for compression purposes. In this paper, we employ the GBT, as well as the Singular Value Decomposition (SVD) with the goal to…

音频与语音处理 · 电气工程与系统科学 2020-03-19 Majid Farzaneh , Rahil Mahdian Toroghi

Early identification of respiratory irregularities is critical for improving lung health and reducing global mortality rates. The analysis of respiratory sounds plays a significant role in characterizing the respiratory system's condition…

音频与语音处理 · 电气工程与系统科学 2024-11-12 Loredana Daria Mang , Francisco David Gonzalez Martinez , Damian Martinez Munoz , Sebastian Garcia Galan , Raquel Cortina

There are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs)trained to discriminate speakers, pass-phrases and triphone states for improving the performance of text-dependent speaker…

声音 · 计算机科学 2019-05-14 Achintya kr. Sarkar , Zheng-Hua Tan , Hao Tang , Suwon Shon , James Glass

Commonly used features in spoken language identification (LID), such as mel-spectrogram or MFCC, lose high-frequency information due to windowing. The loss further increases for longer temporal contexts. To improve generalization of the…

音频与语音处理 · 电气工程与系统科学 2023-10-04 Spandan Dey , Premjeet Singh , Goutam Saha

The capability of the human to pay attention to both coarse and fine-grained regions has been applied to computer vision tasks. Motivated by that, we propose a collaborative learning framework in the complex domain for monaural noise…

声音 · 计算机科学 2021-06-23 Andong Li , Chengshi Zheng , Lu Zhang , Xiaodong Li

Audio-Language Models (ALMs) are making strides in understanding speech and non-speech audio. However, domain-specialist Foundation Models (FMs) remain the best for closed-ended speech processing tasks such as Speech Emotion Recognition…

音频与语音处理 · 电气工程与系统科学 2026-03-25 Saurabh Kataria , Xiao Hu

A comprehensive understanding of molecular clumps is essential for investigating star formation. We present an algorithm for molecular clump detection, called FacetClumps. This algorithm uses a morphological approach to extract signal…

天体物理仪器与方法 · 物理学 2023-08-09 Yu Jiang , Zhiwei Chen , Sheng Zheng , Zhibo Jiang , Yao Huang , Shuguang Zeng , Xiangyun Zeng , Xiaoyu Luo

In this study, we propose a modulation decoupling based single channel speech enhancement subspace framework, in which the spectrogram of noisy speech is decoupled as the product of a spectral envelop subspace and a spectral details…

声音 · 计算机科学 2017-02-24 Pengfei Sun , Jun Qin

Precise detection of speech endpoints is an important factor which affects the performance of the systems where speech utterances need to be extracted from the speech signal such as Automatic Speech Recognition (ASR) system. Existing…

音频与语音处理 · 电气工程与系统科学 2018-09-26 Tanmoy Roy , Tshilidzi Marwala , Snehashish Chakraverty
‹ 上一页 1 8 9 10 下一页 ›