中文
相关论文

相关论文: Deep Factorization for Speech Signal

200 篇论文

Recent studies have shown that frame-level deep speaker features can be derived from a deep neural network with the training target set to discriminate speakers by a short speech segment. By pooling the frame-level features, utterance-level…

音频与语音处理 · 电气工程与系统科学 2018-11-09 Lantian Li , Zhiyuan Tang , Ying Shi , Dong Wang

In response to the rapid growth of Internet of Things (IoT) devices and rising security risks, Radio Frequency Fingerprint (RFF) has become key for device identification and authentication. However, various changing factors - beyond the RFF…

信号处理 · 电气工程与系统科学 2025-08-19 Yezhuo Zhang , Zinan Zhou , Guangyu Li , Xuanpeng Li

Speech emotion recognition (SER) has gained significant attention due to its several application fields, such as mental health, education, and human-computer interaction. However, the accuracy of SER systems is hindered by high-dimensional…

音频与语音处理 · 电气工程与系统科学 2024-06-07 Alaa Nfissi , Wassim Bouachir , Nizar Bouguila , Brian Mishara

Speaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising…

音频与语音处理 · 电气工程与系统科学 2020-02-18 Saurabh Kataria , Phani Sankar Nidadavolu , Jesús Villalba , Nanxin Chen , Paola García , Najim Dehak

One-shot voice conversion has received significant attention since only one utterance from source speaker and target speaker respectively is required. Moreover, source speaker and target speaker do not need to be seen during training.…

声音 · 计算机科学 2021-06-22 Hongqiang Du , Lei Xie

A deep learning approach has been proposed recently to derive speaker identifies (d-vector) by a deep neural network (DNN). This approach has been applied to text-dependent speaker recognition tasks and shows reasonable performance gains…

计算与语言 · 计算机科学 2015-06-30 Lantian Li , Yiye Lin , Zhiyong Zhang , Dong Wang

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network…

Neural speech models build deeply entangled internal representations, which capture a variety of features (e.g., fundamental frequency, loudness, syntactic category, or semantic content of a word) in a distributed encoding. This complexity…

计算与语言 · 计算机科学 2024-10-07 Hosein Mohebbi , Grzegorz Chrupała , Willem Zuidema , Afra Alishahi , Ivan Titov

Speech deepfake detection (DFD) has benefited from diverse acoustic and semantic speech representations, many of which encode valuable speech information and are costly to train. Existing approaches typically enhance DFD by tuning the…

声音 · 计算机科学 2026-02-26 Yupei Li , Chenyang Lyu , Longyue Wang , Weihua Luo , Kaifu Zhang , Björn W. Schuller

We introduce deep switching auto-regressive factorization (DSARF), a deep generative model for spatio-temporal data with the capability to unravel recurring patterns in the data and perform robust short- and long-term predictions. Similar…

机器学习 · 计算机科学 2020-09-14 Amirreza Farnoosh , Bahar Azari , Sarah Ostadabbas

Speaker diarization is a task to label an audio or video recording with the identity of the speaker at each given time stamp. In this work, we propose a novel machine learning framework to conduct real-time multi-speaker diarization and…

声音 · 计算机科学 2023-02-23 Baihan Lin , Xinxin Zhang

Speaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech's content. However, this paper challenges this by highlighting the…

声音 · 计算机科学 2024-09-10 Massa Baali , Abdulhamid Aldoobi , Hira Dhamyal , Rita Singh , Bhiksha Raj

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different…

信号处理 · 电气工程与系统科学 2021-02-09 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision networks operate. Meanwhile, deep neural networks have also…

Large language models have shown a remarkable ability to extract meaning from unstructured data, offering new ways to interpret biomedical signals beyond traditional numerical methods. In this study, we present a matrix factorization…

音频与语音处理 · 电气工程与系统科学 2025-07-15 Yasaman Torabi , Shahram Shirani , James P. Reilly

In conversational speech, the acoustic signal provides cues that help listeners disambiguate difficult parses. For automatically parsing spoken utterances, we introduce a model that integrates transcribed text and acoustic-prosodic features…

计算与语言 · 计算机科学 2018-04-17 Trang Tran , Shubham Toshniwal , Mohit Bansal , Kevin Gimpel , Karen Livescu , Mari Ostendorf

Most state-of-the-art speech systems are using Deep Neural Networks (DNNs). Those systems require a large amount of data to be learned. Hence, learning state-of-the-art frameworks on under-resourced speech languages/problems is a difficult…

音频与语音处理 · 电气工程与系统科学 2020-03-10 Vincent Roger , Jérôme Farinas , Julien Pinquier

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural…

音频与语音处理 · 电气工程与系统科学 2020-02-17 Yuma Koizumi , Kohei Yatabe , Marc Delcroix , Yoshiki Masuyama , Daiki Takeuchi

Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In this paper, we solve…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Yusuke Fujita , Shinji Watanabe , Shota Horiguchi , Yawen Xue , Jing Shi , Kenji Nagamatsu

Recently, advanced technologies have unlimited potential in solving various problems with a large amount of data. However, these technologies have yet to show competitive performance in brain-computer interfaces (BCIs) which deal with brain…

人工智能 · 计算机科学 2022-06-20 Byeong-Hoo Lee , Jeong-Hyun Cho , Byoung-Hee Kwon , Seong-Whan Lee