中文
相关论文

相关论文: Aliasing Reduction in Neural Amp Modeling by Smoot…

200 篇论文

Recent transformer-based ASR models have achieved word-error rates (WER) below 4%, surpassing human annotator accuracy, yet they demand extensive server resources, contributing to significant carbon footprints. The traditional server-based…

声音 · 计算机科学 2024-05-03 Aditya Chakravarty

Facial Expression Recognition (FER) is a classification task that points to face variants. Hence, there are certain affinity features between facial expressions, receiving little attention in the FER literature. Convolution padding, despite…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Jiawei Shi , Songhao Zhu , Zhiwei Liang

In multi-channel speech enhancement and robust automatic speech recognition (ASR), beamforming can typically improve the signal-to-noise ratio (SNR) of the target speaker and produce reliable enhancement with little distortion to target…

音频与语音处理 · 电气工程与系统科学 2025-07-22 Zhong-Qiu Wang , Ruizhe Pang

An efficient algorithm for adaptive kernel smoothing (AKS) of two-dimensional imaging data has been developed and implemented using the Interactive Data Language (IDL). The functional form of the kernel can be varied (top-hat, Gaussian…

天体物理学 · 物理学 2009-11-11 H. Ebeling , D. A. White , F. V. N. Rangarajan

Undersampling can accelerate the signal acquisition but at the cost of bringing in artifacts. Removing these artifacts is a fundamental problem in signal processing and this task is also called signal reconstruction. Through modeling…

信号处理 · 电气工程与系统科学 2024-08-15 Yihui Huang , Zi Wang , Xinlin Zhang , Jian Cao , Zhangren Tu , Meijin Lin , Di Guo , Xiaobo Qu

Automatic Speech Recognition (ASR) systems suffer considerably when source speech is corrupted with noise or room impulse responses (RIR). Typically, speech enhancement is applied in both mismatched and matched scenario training and…

音频与语音处理 · 电气工程与系统科学 2022-04-26 Shashi Kumar , Shakti P. Rath , Abhishek Pandey

Digital audio watermarking consists in inserting a message into audio signals in a transparent way and can be used to allow automatic recognition of audio material and management of the copyrights. We propose a perceptual loss function to…

音频与语音处理 · 电气工程与系统科学 2024-11-05 Martin Moritz , Toni Olán , Tuomas Virtanen

Regression problems are pervasive in real-world applications. Generally a substantial amount of labeled samples are needed to build a regression model with good generalization ability. However, many times it is relatively easy to collect a…

机器学习 · 计算机科学 2018-08-14 Dongrui Wu , Chin-Teng Lin , Jian Huang

Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most current ASR systems still follow a single-pass paradigm, which…

人工智能 · 计算机科学 2026-05-29 Zixuan Jiang , Yanqiao Zhu , Peng Wang , Qinyuan Chen , Xinjian Zhao , Xipeng Qiu , Wupeng Wang , Zhifu Gao , Xiangang Li , Kai Yu , Xie Chen

Environmental noises and reverberation have a detrimental effect on the performance of automatic speech recognition (ASR) systems. Multi-condition training of neural network-based acoustic models is used to deal with this problem, but it…

音频与语音处理 · 电气工程与系统科学 2021-02-03 Desh Raj , Jesus Villalba , Daniel Povey , Sanjeev Khudanpur

Speech-based virtual assistants, such as Amazon Alexa, Google assistant, and Apple Siri, typically convert users' audio signals to text data through automatic speech recognition (ASR) and feed the text to downstream dialog models for…

计算与语言 · 计算机科学 2020-06-11 Longshaokan Wang , Maryam Fazel-Zarandi , Aditya Tiwari , Spyros Matsoukas , Lazaros Polymenakos

Stochastic resonance (SR) is a coherence enhancement effect due to noise that occurs in periodically-driven nonlinear dynamical systems. A very broad range of physical and biological systems present this effect such as climate change,…

斑图形成与孤子 · 物理学 2021-06-08 Adriano A. Batista , A. A. Lisboa de Souza , Raoni S. N. Moreira

In spectroscopic analysis, the peak-based signal-to-noise ratio (pSNR) is commonly used but suffers from limitations such as sensitivity to noise spikes and reduced effectiveness for broader peaks. We introduce the area-based…

信号处理 · 电气工程与系统科学 2025-12-25 Alex Yu , Huaqing Zhao , Lin Z. Li

Arterial spin labeling perfusion MRI is a noninvasive technique for measuring quantitative cerebral blood flow (CBF), but the measurement is subject to a low signal-to-noise-ratio(SNR). Various post-processing methods have been proposed to…

计算机视觉与模式识别 · 计算机科学 2018-01-31 Danfeng Xie , Li Bai , Ze Wang

This paper introduces an active learning (AL) framework for anomalous sound detection (ASD) in machine condition monitoring system. Typically, ASD models are trained solely on normal samples due to the scarcity of anomalous data, leading to…

声音 · 计算机科学 2024-08-13 Tuan Vu Ho , Kota Dohi , Yohei Kawaguchi

Overfitting is one of the critical problems in deep neural networks. Many regularization schemes try to prevent overfitting blindly. However, they decrease the convergence speed of training algorithms. Adaptive regularization schemes can…

机器学习 · 计算机科学 2021-06-18 Mohammad Mahdi Bejani , Mehdi Ghatee

Growing interest in automatic speaker verification (ASV)systems has lead to significant quality improvement of spoofing attackson them. Many research works confirm that despite the low equal er-ror rate (EER) ASV systems are still…

声音 · 计算机科学 2017-05-25 Galina Lavrentyeva , Sergey Novoselov , Konstantin Simonchik

While Automatic Speech Recognition has been shown to be vulnerable to adversarial attacks, defenses against these attacks are still lagging. Existing, naive defenses can be partially broken with an adaptive attack. In classification tasks,…

计算与语言 · 计算机科学 2022-01-12 Raphael Olivier , Bhiksha Raj

A new method for designing non-uniform filter-banks for acoustic echo cancellation is proposed. In the method, the analysis prototype filter design is framed as a convex optimization problem that maximizes the signal-to-alias ratio (SAR) in…

声音 · 计算机科学 2014-02-19 R. C. Nongpiur , D. J. Shpak

Nowadays, research in speech technologies has gotten a lot out thanks to recently created public domain corpora that contain thousands of recording hours. These large amounts of data are very helpful for training the new complex models…

音频与语音处理 · 电气工程与系统科学 2021-05-12 Guillermo Cámbara , Alex Peiró-Lilja , Mireia Farrús , Jordi Luque