English
Related papers

Related papers: Navigating PESQ: Up-to-Date Versions and Open Impl…

200 papers

Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in…

Computation and Language · Computer Science 2022-08-01 Suwon Shon , Ankita Pasad , Felix Wu , Pablo Brusco , Yoav Artzi , Karen Livescu , Kyu J. Han

One challenging problem of robust automatic speech recognition (ASR) is how to measure the goodness of a speech enhancement algorithm (SEA) without calculating the word error rate (WER) due to the high costs of manual transcriptions,…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-29 Li Chai , Jun Du , Chin-Hui Lee

Audio Question Answering (AQA) is a key task for evaluating Audio-Language Models (ALMs), yet assessing open-ended responses remains challenging. Existing metrics used for AQA such as BLEU, METEOR and BERTScore, mostly adapted from NLP and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Satvik Dixit , Soham Deshmukh , Bhiksha Raj

In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely…

Computation and Language · Computer Science 2025-06-23 Kexin Huang , Qian Tu , Liwei Fan , Chenchen Yang , Dong Zhang , Shimin Li , Zhaoye Fei , Qinyuan Cheng , Xipeng Qiu

As more speech technologies rely on a supervised deep learning approach with clean speech as the ground truth, a methodology to onboard said speech at scale is needed. However, this approach needs to minimize the dependency on human…

Sound · Computer Science 2024-02-21 Adam Sabra , Cyprian Wronka , Michelle Mao , Samer Hijazi

Digital images contain a lot of redundancies, therefore, compression techniques are applied to reduce the image size without loss of reasonable image quality. Same become more prominent in the case of videos which contains image sequences…

Computer Vision and Pattern Recognition · Computer Science 2023-05-17 Nisar Ahmed , Hafiz Muhammad Shahzad Asif , Hassan Khalid

Modern text simplification (TS) heavily relies on the availability of gold standard data to build machine learning models. However, existing studies show that parallel TS corpora contain inaccurate simplifications and incorrect alignments.…

Computation and Language · Computer Science 2021-07-30 Laura Vásquez-Rodríguez , Matthew Shardlow , Piotr Przybyła , Sophia Ananiadou

The development of Large Language Models (LLM) and Diffusion Models brings the boom of Artificial Intelligence Generated Content (AIGC). It is essential to build an effective quality assessment framework to provide a quantifiable evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Xi Fang , Weigang Wang , Xiaoxin Lv , Jun Yan

This paper introduces an updated and combined version of the bidirectional English-German EPIC-UdS (spoken) and EuroParl-UdS (written) corpora containing original European Parliament speeches as well as their translations and…

Computation and Language · Computer Science 2026-03-17 Maria Kunilovskaya , Christina Pollkläsener

The evaluation of Question Answering (QA) systems over Knowledge Graphs has historically suffered from fragmentation, inconsistency, and limited reproducibility. While significant progress has been made in semantic parsing and SPARQL query…

Information Retrieval · Computer Science 2026-05-01 Yousouf Taghzouti , Tao Jiang , Camille Juigné , Benjamin Navet , Fabien Gandon , Franck Michel , Louis-Felix Nothias

Recent trends in neural network based text-to-speech/speech synthesis pipelines have employed recurrent Seq2seq architectures that can synthesize realistic sounding speech directly from text characters. These systems however have complex…

Computation and Language · Computer Science 2019-03-19 Gary Wang

MSstatsQC [3] is an open-source software that provides longitudinal system suitability monitoring tools in the form of control charts for proteomic experiments. It includes simultaneous tools for the mean and dispersion of suitability…

Human-Computer Interaction · Computer Science 2020-03-27 Sara Mohammad Taheri , Omkar Terse , Eralp Dogu , Magy Seif El-Nasr , Olga Vitek

The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-19 Pranay Manocha , Buye Xu , Anurag Kumar

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis…

Single channel speech enhancement is a challenging task in speech community. Recently, various neural networks based methods have been applied to speech enhancement. Among these models, PHASEN and T-GSA achieve state-of-the-art performances…

Sound · Computer Science 2021-05-07 Dengfeng Ke , Jinsong Zhang , Yanlu Xie , Yanyan Xu , Binghuai Lin

Image Quality Assessment (IQA) metrics are widely used to quantitatively estimate the extent of image degradation following some forming, restoring, transforming, or enhancing algorithms. We present PyTorch Image Quality (PIQ), a…

Image and Video Processing · Electrical Eng. & Systems 2022-09-01 Sergey Kastryulin , Jamil Zakirov , Denis Prokopenko , Dmitry V. Dylov

Speech is a multiplexed signal displaying levels of complexity, organizational principles and perceptual units of analysis at distinct timescales. This critical acoustic signal for human communication is thus characterized at distinct…

Neurons and Cognition · Quantitative Biology 2024-07-10 Jérémy Giroud , Benjamin Morillon

The quality of daily spontaneous conversations is of importance towards both our well-being as well as the development of interactive social agents. Prior research directly studying the quality of social conversations has operationalized it…

Human-Computer Interaction · Computer Science 2022-07-14 Chirag Raman , Navin Raj Prabhu , Hayley Hung

Results reported in large-scale multilingual evaluations are often fragmented and confounded by factors such as target languages, differences in experimental setups, and model choices. We propose a framework that disentangles these…

Computation and Language · Computer Science 2025-08-26 Songbo Hu , Ivan Vulić , Anna Korhonen

Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transform source language…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-29 Jarod Duret , Yannick Estève , Titouan Parcollet