中文
相关论文

相关论文: Composing General Audio Representation by Fusing M…

200 篇论文

Enhancing noisy speech is an important task to restore its quality and to improve its intelligibility. In traditional non-machine-learning (ML) based approaches the parameters required for noise reduction are estimated blindly from the…

声音 · 计算机科学 2018-01-16 Robert Rehr , Timo Gerkmann

Recent studies have shown that deep neural networks (DNNs) perform significantly better than shallow networks and Gaussian mixture models (GMMs) on large vocabulary speech recognition tasks. In this paper, we argue that the improved…

机器学习 · 计算机科学 2018-12-06 Dong Yu , Michael L. Seltzer , Jinyu Li , Jui-Ting Huang , Frank Seide

Multi-speaker speech synthesis is a technique for modeling multiple speakers' voices with a single model. Although many approaches using deep neural networks (DNNs) have been proposed, DNNs are prone to overfitting when the amount of…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Kentaro Mitsui , Tomoki Koriyama , Hiroshi Saruwatari

This paper investigates different trade-offs between the number of model parameters and enhanced speech qualities by employing several deep tensor-to-vector regression models for speech enhancement. We find that a hybrid architecture,…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Jun Qi , Hu Hu , Yannan Wang , Chao-Han Huck Yang , Sabato Marco Siniscalchi , Chin-Hui Lee

With the advancement of generative modeling techniques, synthetic human speech becomes increasingly indistinguishable from real, and tricky challenges are elicited for the audio deepfake detection (ADD) system. In this paper, we exploit…

声音 · 计算机科学 2024-03-05 Yujie Yang , Haochen Qin , Hang Zhou , Chengcheng Wang , Tianyu Guo , Kai Han , Yunhe Wang

Modern automatic speaker verification relies largely on deep neural networks (DNNs) trained on mel-frequency cepstral coefficient (MFCC) features. While there are alternative feature extraction methods based on phase, prosody and long-term…

音频与语音处理 · 电气工程与系统科学 2020-07-31 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Self-supervised learning general-purpose audio representations have demonstrated high performance in a variety of tasks. Although they can be optimized for application by fine-tuning, even higher performance can be expected if they can be…

音频与语音处理 · 电气工程与系统科学 2023-08-04 Daisuke Niizumi , Daiki Takeuchi , Yasunori Ohishi , Noboru Harada , Kunio Kashino

Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating high-quality samples. Efficient sampling for this class of models has however…

Pre-trained wav2vec2.0 model has been proved its effectiveness for speaker recognition. However, current feature processing methods are focusing on classical pooling on the output features of the pre-trained wav2vec2.0 model, such as mean…

音频与语音处理 · 电气工程与系统科学 2023-03-21 Zirui Ge , Haiyan Guo , Zhen Yang

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

To simplify the generation process, several text-to-speech (TTS) systems implicitly learn intermediate latent representations instead of relying on predefined features (e.g., mel-spectrogram). However, their generation quality is…

声音 · 计算机科学 2023-08-29 Hyungchan Yoon , Seyun Um , Changwhan Kim , Hong-Goo Kang

This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus emphasize semantic abstraction while losing fine grained…

声音 · 计算机科学 2025-09-30 Zeyu Xie , Chenxing Li , Xuenan Xu , Mengyue Wu , Wenfu Wang , Ruibo Fu , Meng Yu , Dong Yu , Yuexian Zou

In this paper, we present a novel deep fusion architecture for audio classification tasks. The multi-channel model presented is formed using deep convolution layers where different acoustic features are passed through each channel. To…

声音 · 计算机科学 2018-11-05 Gaurav Bhatt , Akshita Gupta , Aditya Arora , Balasubramanian Raman

One of the most important parts of an end-to-end speaker verification system is the speaker embedding generation. In our previous paper, we reported that shortcut connections-based multi-layer aggregation improves the representational power…

音频与语音处理 · 电气工程与系统科学 2020-07-29 Soonshin Seo , Ji-Hwan Kim

Generalization is a main issue for current audio deepfake detectors, which struggle to provide reliable results on out-of-distribution data. Given the speed at which more and more accurate synthesis methods are developed, it is very…

声音 · 计算机科学 2024-07-02 Alessandro Pianese , Davide Cozzolino , Giovanni Poggi , Luisa Verdoliva

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Pin-Jui Ku , Alexander H. Liu , Roman Korostik , Sung-Feng Huang , Szu-Wei Fu , Ante Jukić

We propose an algorithm to extract noise-robust acoustic features from noisy speech. We use Total Variability Modeling in combination with Non-negative Matrix Factorization (NMF) to learn a total variability subspace and adapt NMF…

音频与语音处理 · 电气工程与系统科学 2019-07-17 Kunal Dhawan , Colin Vaz , Ruchir Travadi , Shrikanth Narayanan

Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features. For achieving high performance, DNNs often need a large amount of annotated data which can be difficult and…

机器学习 · 计算机科学 2020-07-09 Xavier Favory , Konstantinos Drossos , Tuomas Virtanen , Xavier Serra

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

声音 · 计算机科学 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

In computational bioacoustics, deep learning models are composed of feature extractors and classifiers. The feature extractors generate vector representations of the input sound segments, called embeddings, which can be input to a…

机器学习 · 计算机科学 2025-04-10 Vincent S. Kather , Burooj Ghani , Dan Stowell