中文
相关论文

相关论文: Aliasing Reduction in Neural Amp Modeling by Smoot…

200 篇论文

Neural Radiance Fields (NeRF) have shown remarkable success in representing 3D scenes and generating novel views. However, they often struggle with aliasing artifacts, especially when rendering images from different camera distances from…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Youngin Park , Seungtae Nam , Cheul-hee Hahm , Eunbyung Park

We present our experiments in training robust to noise an end-to-end automatic speech recognition (ASR) model using intensive data augmentation. We explore the efficacy of fine-tuning a pre-trained model to improve noise robustness, and we…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Jagadeesh Balam , Jocelyn Huang , Vitaly Lavrukhin , Slyne Deng , Somshubra Majumdar , Boris Ginsburg

Stochastic resonance (SR) could amplify weak electric-field signals in nonlinear systems by means of the externally injected noises. Here we propose and experimentally demonstrate a modified SR method, termed squeezing-induced SR,…

High-resolution simulations often rely on the Adaptive Mesh Resolution (AMR) technique to optimize memory consumption versus attainable precision. While this technique allows for dramatic improvements in terms of computing performance, the…

数据结构与算法 · 计算机科学 2013-01-03 Marc Labadens , Daniel Pomarède , Damien Chapon , Romain Teyssier , Frédéric Bournaud , Florent Renaud , Nicolas Grandjouan

Next-generation communication and localization systems increasingly rely on extremely large-scale arrays (XL-arrays), which promise unprecedented spatial resolution and new functionalities. These gains arise from their inherent operation in…

信号处理 · 电气工程与系统科学 2026-05-22 Gilles Monnoyer , Jérôme Louveaux , Laurence Defraigne , Baptiste Sambon , Luc Vandendorpe

Most existing text-to-image person retrieval methods usually assume that the training image-text pairs are perfectly aligned; however, the noisy correspondence(NC) issue (i.e., incorrect or unreliable alignment) exists due to poor image…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Runqing Zhang , Xue Zhou

Past studies on end-to-end meeting transcription have focused on model architecture and have mostly been evaluated on simulated meeting data. We present a novel study aiming to optimize the use of a Speaker-Attributed ASR (SA-ASR) system in…

计算与语言 · 计算机科学 2024-09-06 Can Cui , Imran Ahamad Sheikh , Mostafa Sadeghi , Emmanuel Vincent

Reducing noise interference is crucial for automatic speech recognition (ASR) in a real-world scenario. However, most single-channel speech enhancement (SE) generates "processing artifacts" that negatively affect ASR performance. Hence, in…

声音 · 计算机科学 2023-08-25 Kuan-Hsun Ho , En-Lun Yu , Jeih-weih Hung , Berlin Chen

Arterial spin labeling (ASL) allows to quantify the cerebral blood flow (CBF) by magnetic labeling of the arterial blood water. ASL is increasingly used in clinical studies due to its noninvasiveness, repeatability and benefits in…

计算机视觉与模式识别 · 计算机科学 2018-06-13 Cagdas Ulas , Giles Tetteh , Stephan Kaczmarz , Christine Preibisch , Bjoern H. Menze

End-to-end (E2E) multi-channel ASR systems show state-of-the-art performance in far-field ASR tasks by joint training of a multi-channel front-end along with the ASR model. The main limitation of such systems is that they are usually…

音频与语音处理 · 电气工程与系统科学 2021-09-24 Marco Gaudesi , Felix Weninger , Dushyant Sharma , Puming Zhan

Deep learning models achieve strong performance in chest radiograph (CXR) interpretation, yet fairness and reliability concerns persist. Models often show uneven accuracy across patient subgroups, leading to hidden failures not reflected in…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Han-Jay Shu , Wei-Ning Chiu , Shun-Ting Chang , Meng-Ping Huang , Takeshi Tohyama , Ahram Han , Po-Chih Kuo

Speech models are often trained on sensitive data in order to improve model performance, leading to potential privacy leakage. Our work considers noise masking attacks, introduced by Amid et al. 2022, which attack automatic speech…

机器学习 · 计算机科学 2024-04-03 Matthew Jagielski , Om Thakkar , Lun Wang

Automatic speech recognition (ASR) models trained on large amounts of audio data are now widely used to convert speech to written text in a variety of applications from video captioning to automated assistants used in healthcare and other…

计算与语言 · 计算机科学 2024-07-22 Changye Li , Trevor Cohen , Serguei Pakhomov

The performance of automatic speech recognition systems under noisy environments still leaves room for improvement. Speech enhancement or feature enhancement techniques for increasing noise robustness of these systems usually add components…

计算与语言 · 计算机科学 2016-09-19 Stefan Braun , Daniel Neil , Shih-Chii Liu

The transcriptions used to train an Automatic Speech Recognition (ASR) system may contain errors. Usually, either a quality control stage discards transcriptions with too many errors, or the noisy transcriptions are used as is. We introduce…

计算与语言 · 计算机科学 2019-10-17 Adrien Dufraux , Emmanuel Vincent , Awni Hannun , Armelle Brun , Matthijs Douze

The signal to noise ratio (SNR) is one of the important measures for reducing the noise.A technique that uses a linear prediction error filter (LPEF) and an adaptive digital filter (ADF) to achieve noise reduction in a speech and image…

网络与互联网体系结构 · 计算机科学 2011-10-12 R. Seshadri , N. Penchalaiah

Automatic speech recognition (ASR) outcomes serve as input for downstream tasks, substantially impacting the satisfaction level of end-users. Hence, the diagnosis and enhancement of the vulnerabilities present in the ASR model bear…

计算与语言 · 计算机科学 2024-01-29 Seonmin Koo , Chanjun Park , Jinsung Kim , Jaehyung Seo , Sugyeong Eo , Hyeonseok Moon , Heuiseok Lim

In recent years, significant progress has been made in deep model-based automatic speech recognition (ASR), leading to its widespread deployment in the real world. At the same time, adversarial attacks against deep ASR systems are highly…

音频与语音处理 · 电气工程与系统科学 2022-11-04 Christian Heider Nielsen , Zheng-Hua Tan

Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While…

声音 · 计算机科学 2025-08-25 Kuan-Tang Huang , Li-Wei Chen , Hung-Shin Lee , Berlin Chen , Hsin-Min Wang

It is well-known that neural networks can unintentionally memorize their training examples, causing privacy concerns. However, auditing memorization in large non-auto-regressive automatic speech recognition (ASR) models has been challenging…

机器学习 · 计算机科学 2023-10-19 Lun Wang , Om Thakkar , Rajiv Mathews