English
Related papers

Related papers: Single-sided Real-time PESQ Score Estimation

200 papers

Due to the diversity of assessment requirements in various application scenarios for the IQA task, existing IQA methods struggle to directly adapt to these varied requirements after training. Thus, when facing new requirements, a typical…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Zewen Chen , Haina Qin , Juan Wang , Chunfeng Yuan , Bing Li , Weiming Hu , Liang Wang

Speech enhancement concerns the processes required to remove unwanted background sounds from the target speech to improve its quality and intelligibility. In this paper, a novel approach for single-channel speech enhancement is presented,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-27 Sania Gul , Muhammad Salman Khan , Muhammad Fazeel

Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. However, the quality of these models is often measured by the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-18 Alexander H. Liu , Sung-Lin Yeh , James Glass

Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging metric scores across languages, yet this is suspicious since metrics may suffer from…

Computation and Language · Computer Science 2026-04-21 Jingxuan Liu , Zhi Qu , Jin Tei , Hidetaka Kamigaito , Lemao Liu , Taro Watanabe

In a subjective experiment to evaluate the perceptual audiovisual quality of multimedia and television services, raw opinion scores collected from test subjects are often noisy and unreliable. To produce the final mean opinion scores (MOS),…

Multimedia · Computer Science 2021-05-10 Zhi Li , Christos G. Bampis , Lukáš Krasula , Lucjan Janowski , Ioannis Katsavounidis

Objective evaluation of audio processed with Time-Scale Modification (TSM) is seeing a resurgence of interest. Recently, a labelled time-scaled audio dataset was used to train an objective measure for TSM evaluation. This DE measure was an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-09 Timothy Roberts , Aaron Nicolson , Kuldip K. Paliwal

Text-to-speech (TTS) systems offer the opportunity to compensate for a hearing loss at the source rather than correcting for it at the receiving end. This removes limitations such as time constraints for algorithms that amplify a sound in a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-23 Josef Schlittenlacher , Thomas Baer

While recent face recognition (FR) systems achieve excellent results in many deployment scenarios, their performance in challenging real-world settings is still under question. For this reason, face image quality assessment (FIQA)…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Žiga Babnik , Vitomir Štruc

Augmented Reality (AR) devices are commonly head-worn to overlay context-dependent information into the field of view of the device operators. One particular scenario is the overlay of still images, either in a traditional fashion, or as…

Multimedia · Computer Science 2017-05-04 Brian Bauman , Patrick Seeling

Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the…

Multimedia · Computer Science 2024-02-07 Xiongkuo Min , Huiyu Duan , Wei Sun , Yucheng Zhu , Guangtao Zhai

Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as…

Sound · Computer Science 2025-08-15 Iksoon Jeong , Kyung-Joong Kim , Kang-Hun Ahn

Self-supervised speech representations (SSSRs) have been successfully applied to a number of speech-processing tasks, e.g. as feature extractor for speech quality (SQ) prediction, which is, in turn, relevant for assessment and training…

Sound · Computer Science 2023-12-08 George Close , Thomas Hain , Stefan Goetze

In this paper, for real time enhancement of noisy speech, a method of threshold determination based on modeling of Teager energy (TE) operated perceptual wavelet packet (PWP) coefficients of the noisy speech and noise by an Erlang-2 PDF is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-02-13 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustrate pho- netic concepts that are otherwise difficult to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-10 Frederik Rautenberg , Fritz Seebauer , Jana Wiechmann , Michael Kuhlmann , Petra Wagner , Reinhold Haeb-Umbach

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded…

Sound · Computer Science 2026-04-28 Ziyang Jiang , Jiahe Lei , Xueyan Chen , Yifan Zhang , Zexu Pan , Wei Xue , Xinyuan Qian

Voice assistants (VAs) are typically evaluated through task performance metrics and self-report questionnaires, but people's voices themselves carry rich paralinguistic cues that reveal affect, effort, and interaction breakdowns. We present…

Human-Computer Interaction · Computer Science 2026-03-23 Yong Ma , Xuesong Zhang , Xuedong Zhang , Natalia Bartłomiejczyk , Seungwoo Je , Adrian Holzer , Morten Fjeld , Andreas Butz

Ever-growing scale of large language models (LLMs) is pushing for improved efficiency, favoring fully quantized training (FQT) over BF16. While FQT accelerates training, it faces consistency challenges and requires searching over an…

Machine Learning · Computer Science 2025-05-19 Myeonghwan Ahn , Sungjoo Yoo

We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervised pseudo speech…

Computation and Language · Computer Science 2022-05-03 Felix Wu , Kwangyoun Kim , Shinji Watanabe , Kyu Han , Ryan McDonald , Kilian Q. Weinberger , Yoav Artzi

Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transform source language…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-29 Jarod Duret , Yannick Estève , Titouan Parcollet

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to…

Sound · Computer Science 2024-06-17 Tanel Pärnamaa , Ando Saabas