中文
相关论文

相关论文: Deep Learning-Based Single-Ended Objective Quality…

200 篇论文

Large Audio-Language Models (ALMs) have recently demonstrated remarkable capabilities in holistic audio understanding, yet they remain unreliable for temporal grounding, i.e., the task of pinpointing exactly when an event occurs within…

声音 · 计算机科学 2026-04-15 Luoyi Sun , Xiao Zhou , Zeqian Li , Ya Zhang , Yanfeng Wang , Weidi Xie

This paper investigates the optimization of Truncated Backpropagation Through Time (TBPTT) for training neural networks in digital audio effect modeling, with a focus on dynamic range compression. The study evaluates key TBPTT…

机器学习 · 计算机科学 2025-12-09 Yann Bourdin , Pierrick Legrand , Fanny Roche

Time-Scale Modification (TSM) of speech aims to alter the playback rate of audio without changing its pitch. While classical methods like Waveform Similarity-based Overlap-Add (WSOLA) provide strong baselines, they often introduce artifacts…

音频与语音处理 · 电气工程与系统科学 2025-10-06 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Fo-Rui Li , Yan-Tsung Peng , Hsin-Min Wang , Yu Tsao

Recent years have seen considerable advances in audio synthesis with deep generative models. However, the state-of-the-art is very difficult to quantify; different studies often use different evaluation methodologies and different metrics…

声音 · 计算机科学 2022-09-02 Ashvala Vinay , Alexander Lerch

The evaluation of machine learning algorithms in biomedical fields for applications involving sequential data lacks standardization. Common quantitative scalar evaluation metrics such as sensitivity and specificity can often be misleading…

机器学习 · 计算机科学 2019-12-03 Saeedeh Ziyabari , Vinit Shah , Meysam Golmohammadi , Iyad Obeid , Joseph Picone

We consider the problem of decentralized estimation using wireless sensor networks. Specifically, we propose a novel framework based on level-triggered sampling, a non-uniform sampling strategy, and sequential estimation. The proposed…

应用统计 · 统计学 2013-09-24 Yasin Yilmaz , Xiaodong Wang

Recently, zero-shot text-to-speech (TTS) systems, capable of synthesizing any speaker's voice from a short audio prompt, have made rapid advancements. However, the quality of the generated speech significantly deteriorates when the audio…

音频与语音处理 · 电气工程与系统科学 2024-06-11 Xiaofei Wang , Sefik Emre Eskimez , Manthan Thakker , Hemin Yang , Zirun Zhu , Min Tang , Yufei Xia , Jinzhu Li , Sheng Zhao , Jinyu Li , Naoyuki Kanda

When the input signal is correlated input signals, and the input and output signal is contaminated by Gaussian noise, the total least squares normalized subband adaptive filter (TLS-NSAF) algorithm shows good performance. However, when it…

信号处理 · 电气工程与系统科学 2023-07-21 Haiquan Zhao , Zian Cao , Yida Chen

Purpose: To develop a deep learning approach to de-noise optical coherence tomography (OCT) B-scans of the optic nerve head (ONH). Methods: Volume scans consisting of 97 horizontal B-scans were acquired through the center of the ONH using a…

Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-training method that tailors the foundational Audio Large…

音频与语音处理 · 电气工程与系统科学 2026-03-12 Elizaveta Kostenok , Mathieu Salzmann , Milos Cernak

Recent research has explored using neural networks to reconstruct undersampled magnetic resonance imaging (MRI) data. Because of the complexity of the artifacts in the reconstructed images, there is a need to develop task-based approaches…

图像与视频处理 · 电气工程与系统科学 2022-10-25 Joshua D. Herman , Rachel E. Roca , Alexandra G. O'Neill , Marcus L. Wong , Sajan G. Lingala , Angel R. Pineda

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech…

声音 · 计算机科学 2024-04-02 Xiang Li , Fan Bu , Ambuj Mehrish , Yingting Li , Jiale Han , Bo Cheng , Soujanya Poria

Objective: Ultrasound elastography is gaining traction as an accessible and useful diagnostic tool for such things as cancer detection and differentiation and thyroid disease diagnostics. Unfortunately, state of the art shear wave imaging…

机器学习 · 计算机科学 2019-07-31 Micha Feigin , Daniel Freedman , Brian W. Anthony

There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~reference-free, metrics to assess the performance of SE systems,…

声音 · 计算机科学 2025-08-05 George Close , Kris Hong , Thomas Hain , Stefan Goetze

Most of the deep learning based speech enhancement (SE) methods rely on estimating the magnitude spectrum of the clean speech signal from the observed noisy speech signal, either by magnitude spectral masking or regression. These methods…

音频与语音处理 · 电气工程与系统科学 2020-10-28 Raktim Gautam Goswami , Sivaganesh Andhavarapu , K Sri Rama Murty

In this article, we adapted five recent SSL methods to the task of audio classification. The first two methods, namely Deep Co-Training (DCT) and Mean Teacher (MT), involve two collaborative neural networks. The three other algorithms,…

声音 · 计算机科学 2023-03-09 Léo Cances , Etienne Labbé , Thomas Pellegrini

Spatial audio quality is a highly multifaceted concept, with many interactions between environmental, geometrical, anatomical, psychological, and contextual considerations. Methods for characterization or evaluation of the geometrical…

音频与语音处理 · 电气工程与系统科学 2024-08-27 Karn N. Watcharasupat , Alexander Lerch

Deepfake audio detection has progressed rapidly with strong pre-trained encoders (e.g., WavLM, Wav2Vec2, MMS). However, performance in realistic capture conditions - background noise (domestic/office/transport), room reverberation, and…

声音 · 计算机科学 2025-12-17 Udayon Sen , Alka Luqman , Anupam Chattopadhyay

The aim of speech enhancement is to improve speech signal quality and intelligibility from a noisy microphone signal. In many applications, it is crucial to enable processing with small computational complexity and minimal requirements…

音频与语音处理 · 电气工程与系统科学 2023-09-08 Julitta Bartolewska , Stanisław Kacprzak , Konrad Kowalczyk

Deep representation learning has gained significant momentum in advancing text-dependent speaker verification (TD-SV) systems. When designing deep neural networks (DNN) for extracting bottleneck features, key considerations include training…

声音 · 计算机科学 2022-01-19 Achintya kr. Sarkar , Zheng-Hua Tan