English
Related papers

Related papers: Modelling Noise-Resilient Single-Switch Scanning S…

200 papers

Single Image Super-Resolution (SISR) is one of the low-level computer vision problems that has received increased attention in the last few years. Current approaches are primarily based on harnessing the power of deep learning models and…

Computer Vision and Pattern Recognition · Computer Science 2021-05-19 Santiago López-Tapia , Nicolás Pérez de la Blanca

The past decade has witnessed substantial growth of data-driven speech enhancement (SE) techniques thanks to deep learning. While existing approaches have shown impressive performance in some common datasets, most of them are designed only…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-19 Wangyou Zhang , Kohei Saijo , Zhong-Qiu Wang , Shinji Watanabe , Yanmin Qian

Along with recent diffusion models, randomized smoothing has become one of a few tangible approaches that offers adversarial robustness to models at scale, e.g., those of large pre-trained models. Specifically, one can perform randomized…

Machine Learning · Computer Science 2023-10-30 Jongheon Jeong , Jinwoo Shin

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle…

Blind identification is popular for modeling a system without the input information, such as in the research areas of structural health monitoring and audio signal processing. Existing blind identification methods have both advantages and…

Systems and Control · Electrical Eng. & Systems 2021-08-20 Runzhe Han , Christian Bohn , Georg Bauer

The goal of this work is to develop a meeting transcription system that can recognize speech even when utterances of different speakers are overlapped. While speech overlaps have been regarded as a major obstacle in accurately transcribing…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-10 Takuya Yoshioka , Hakan Erdogan , Zhuo Chen , Xiong Xiao , Fil Alleva

One of the fundamental challenges affecting the performance of communication systems is the undesired impact of noise on a signal. Noise distorts the signal and originates due to several sources including, system non-linearity and noise…

Signal Processing · Electrical Eng. & Systems 2018-01-04 Adnan Quadri

Robust speaker verification under noisy conditions remains an open challenge. Conventional deep learning methods learn a robust unified speaker representation space against diverse background noise and achieve significant improvement. In…

Sound · Computer Science 2026-03-11 Bin Gu , Haitao Zhao , Jibo Wei

With the continuing advances in scientific instrumentation, scanning microscopes are now able to image physical systems with up to sub-atomic-level spatial resolutions and sub-picosecond time resolutions. Commensurately, they are generating…

Face-swapping models have been drawing attention for their compelling generation quality, but their complex architectures and loss functions often require careful tuning for successful training. We propose a new face-swapping model called…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Jiseob Kim , Jihoon Lee , Byoung-Tak Zhang

We tackle a challenging blind image denoising problem, in which only single distinct noisy images are available for training a denoiser, and no information about noise is known, except for it being zero-mean, additive, and independent of…

Image and Video Processing · Electrical Eng. & Systems 2021-07-06 Sungmin Cha , Taeeon Park , Byeongjoon Kim , Jongduk Baek , Taesup Moon

Speech recognition system performance degrades in noisy environments. If the acoustic models are built using features of clean utterances, the features of a noisy test utterance would be acoustically mismatched with the trained model. This…

Computation and Language · Computer Science 2015-07-16 D. S. Pavan Kumar

In this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired…

Sound · Computer Science 2022-09-21 Jisi Zhang , Catalin Zorila , Rama Doddipatla , Jon Barker

Scoring systems are linear classification models that only require users to add or subtract a few small numbers in order to make a prediction. They are used for example by clinicians to assess the risk of medical conditions. This work…

Human-Computer Interaction · Computer Science 2019-12-02 Hans-Jürgen Profitlich , Daniel Sonntag

Improving the spatial resolution of CT images is a meaningful yet challenging task, often accompanied by the issue of noise amplification. This article introduces an innovative framework for noise-controlled CT super-resolution utilizing…

Computer Vision and Pattern Recognition · Computer Science 2025-02-17 Yuang Wang , Siyeop Yoon , Rui Hu , Baihui Yu , Duhgoon Lee , Rajiv Gupta , Li Zhang , Zhiqiang Chen , Dufan Wu

Evaluating state-of-the-art speech systems necessitates scalable and sensitive evaluation methods to detect subtle but unacceptable artifacts. Standard MUSHRA is sensitive but lacks scalability, while ACR scales well but loses sensitivity…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Laura Lechler , Ivana Balic

For people with noise sensitivity, everyday soundscapes can be overwhelming. Existing tools such as active noise cancellation reduce discomfort by suppressing the entire acoustic environment, often at the cost of awareness of surrounding…

Sound · Computer Science 2026-04-02 Jeremy Zhengqi Huang , Emani Hicks , Sidharth , Gillian R. Hayes , Dhruv Jain

Joint training of speech enhancement model (SE) and speech recognition model (ASR) is a common solution for robust ASR in noisy environments. SE focuses on improving the auditory quality of speech, but the enhanced feature distribution is…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-04 Tianrui Wang , Weibin Zhu , Yingying Gao , Junlan Feng , Shilei Zhang

Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior work, this paper examines the conditions and model…

Sound · Computer Science 2025-07-28 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

The practical implementation of maximum likelihood detection is limited by its high complexity as well as requiring perfect channel state information. Although conventional blind detection techniques reduce complexity, they degrade…

Signal Processing · Electrical Eng. & Systems 2019-09-12 M. A. Amirabadi