中文
相关论文

相关论文: Multi-dimensional Speech Quality Assessment in Cro…

200 篇论文

The ICASSP 2023 Speech Signal Improvement Challenge is intended to stimulate research in the area of improving the speech signal quality in communication systems. The speech signal quality can be measured with SIG in ITU-T P.835 and is…

音频与语音处理 · 电气工程与系统科学 2023-10-16 Ross Cutler , Ando Saabas , Babak Naderi , Nicolae-Cătălin Ristea , Sebastian Braun , Solomiya Branets

Subjective video quality assessment (VQA) is the gold standard for measuring end-user experience across communication, streaming, and UGC pipelines. Beyond high-validity lab studies, crowdsourcing offers accurate, reliable, faster, and…

图像与视频处理 · 电气工程与系统科学 2025-09-25 Babak Naderi , Ross Cutler

Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets…

The MUSHRA framework is widely used for detecting subtle audio quality differences but traditionally relies on expert listeners in controlled environments, making it costly and impractical for model development. As a result, objective…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Laura Lechler , Chamran Moradi , Ivana Balic

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion…

音频与语音处理 · 电气工程与系统科学 2021-04-06 Meng Yu , Chunlei Zhang , Yong Xu , Shixiong Zhang , Dong Yu

Machine Translation (MT) has achieved remarkable performance, with growing interest in speech translation and multimodal approaches. However, despite these advancements, MT quality assessment remains largely text centric, typically relying…

计算与语言 · 计算机科学 2025-09-18 Sami Ul Haq , Sheila Castilho , Yvette Graham

Estimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this…

音频与语音处理 · 电气工程与系统科学 2019-03-19 Anderson R. Avila , Hannes Gamper , Chandan Reddy , Ross Cutler , Ivan Tashev , Johannes Gehrke

The INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical…

Crowdsourcing enables one to leverage on the intelligence and wisdom of potentially large groups of individuals toward solving problems. Common problems approached with crowdsourcing are labeling images, translating or transcribing text,…

人机交互 · 计算机科学 2018-01-09 Florian Daniel , Pavel Kucherbaev , Cinzia Cappiello , Boualem Benatallah , Mohammad Allahbakhsh

In the last decade, crowdsourcing has become a popular method for conducting quantitative empirical studies in human-machine interaction. The remote work on a given task in crowdworking settings suits the character of typical…

人机交互 · 计算机科学 2024-11-19 Annalena Aicher , Stefan Hillmann , Isabel Feustel , Thilo Michael , Sebastian Möller , Wolfgang Minker

Audio captioning is a novel field of multi-modal translation and it is the task of creating a textual description of the content of an audio signal (e.g. "people talking in a big room"). The creation of a dataset for this task requires a…

声音 · 计算机科学 2019-07-23 Samuel Lipping , Konstantinos Drossos , Tuomas Virtanen

The creation of relevance assessments by human assessors (often nowadays crowdworkers) is a vital step when building IR test collections. Prior works have investigated assessor quality & behaviour, though into the impact of a document's…

信息检索 · 计算机科学 2023-04-24 Nirmal Roy , Agathe Balayn , David Maxwell , Claudia Hauff

Crowdsourcing with the intelligent agents carrying smart devices is becoming increasingly popular in recent years. It has opened up meeting an extensive list of real life applications such as measuring air pollution level, road traffic…

计算机科学与博弈论 · 计算机科学 2022-03-18 Vikash Kumar Singh , Anjani Samhitha Jasti , Sunil Kumar Singh , Sanket Mishra

In this paper, we present an update to the NISQA speech quality prediction model that is focused on distortions that occur in communication networks. In contrast to the previous version, the model is trained end-to-end and the…

音频与语音处理 · 电气工程与系统科学 2021-12-15 Gabriel Mittag , Babak Naderi , Assmaa Chehadi , Sebastian Möller

Many subjective experiments have been performed to develop objective speech intelligibility measures, but the novel coronavirus outbreak has made it very difficult to conduct experiments in a laboratory. One solution is to perform remote…

音频与语音处理 · 电气工程与系统科学 2022-03-30 Ayako Yamamoto , Toshio Irino , Kenichi Arai , Shoko Araki , Atsunori Ogawa , Keisuke Kinoshita , Tomohiro Nakatani

It is essential to perform speech intelligibility (SI) experiments with human listeners in order to evaluate objective intelligibility measures for developing effective speech enhancement and noise reduction algorithms. Recently,…

High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper presents an approach to overcome this limitation by implementing…

音频与语音处理 · 电气工程与系统科学 2024-10-18 José Giraldo , Martí Llopart-Font , Alex Peiró-Lilja , Carme Armentano-Oller , Gerard Sant , Baybars Külebi

For enhancement of noisy speech, a method of threshold determination based on modeling of Teager energy (TE) operated perceptual wavelet packet (PWP) coefficients of the noisy speech by exponential distribution is presented. A custom…

音频与语音处理 · 电气工程与系统科学 2018-02-19 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

This paper explores grading text-based audio retrieval relevances with crowdsourcing assessments. Given a free-form text (e.g., a caption) as a query, crowdworkers are asked to grade audio clips using numeric scores (between 0 and 100) to…

音频与语音处理 · 电气工程与系统科学 2023-08-16 Huang Xie , Khazar Khorrami , Okko Räsänen , Tuomas Virtanen

This letter introduces a novel speech enhancement method in the Hilbert-Huang Transform domain to mitigate the effects of acoustic impulsive noises. The estimation and selection of noise components is based on the impulsiveness index of…

音频与语音处理 · 电气工程与系统科学 2019-10-08 C. Medina , R. Coelho