中文
相关论文

相关论文: Study on the Correlation between Objective Evaluat…

200 篇论文

Although recent neural text-to-speech (TTS) systems have achieved high-quality speech synthesis, there are cases where a TTS system generates low-quality speech, mainly caused by limited training data or information loss during knowledge…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Yeunju Choi , Youngmoon Jung , Youngjoo Suh , Hoirin Kim

The increasing ubiquity of language technology necessitates a shift towards considering cultural diversity in the machine learning realm, particularly for subjective tasks that rely heavily on cultural nuances, such as Offensive Language…

计算与语言 · 计算机科学 2024-09-04 Li Zhou , Antonia Karamolegkou , Wenyu Chen , Daniel Hershcovich

This paper investigates the performance of Deep Learning for speech emotion classification when the speech is compounded with noise. It reports on the classification accuracy and concludes with the future directions for achieving greater…

人机交互 · 计算机科学 2016-04-13 Rajib Rana

It is widely accepted that information derived from analyzing speech (the acoustic signal) and language production (words and sentences) serves as a useful window into the health of an individual's cognitive ability. In fact, most…

计算与语言 · 计算机科学 2019-11-06 Rohit Voleti , Julie M. Liss , Visar Berisha

In black-box optimization, noise in the objective function is inevitable. Noise disrupts the ranking of candidate solutions in comparison-based optimization, possibly deteriorating the search performance compared with a noiseless scenario.…

神经与进化计算 · 计算机科学 2024-01-26 Daiki Morinaga , Youhei Akimoto

Recent studies in speech perception have been closely linked to fields of cognitive psychology, phonology, and phonetics in linguistics. During perceptual attunement, a critical and sensitive developmental trajectory has been examined in…

计算与语言 · 计算机科学 2021-10-14 Anuj Saraswat , Mehar Bhatia , Yaman Kumar Singla , Changyou Chen , Rajiv Ratn Shah

Algorithmic interpretability is necessary to build trust, ensure fairness, and track accountability. However, there is no existing formal measurement method for algorithmic interpretability. In this work, we build upon programming language…

人工智能 · 计算机科学 2022-05-23 John P. Lalor , Hong Guo

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Qiang Huang , Thomas Hain

Concept probing has recently garnered increasing interest as a way to help interpret artificial neural networks, dealing both with their typically large size and their subsymbolic nature, which ultimately renders them unfeasible for direct…

人工智能 · 计算机科学 2025-07-25 Manuel de Sousa Ribeiro , Afonso Leote , João Leite

Assessing the performance of interpreting services is a complex task, given the nuanced nature of spoken language translation, the strategies that interpreters apply, and the diverse expectations of users. The complexity of this task become…

计算与语言 · 计算机科学 2024-06-17 Xiaoman Wang , Claudio Fantinuoli

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

音频与语音处理 · 电气工程与系统科学 2025-11-05 Cedric Chan , Jianjing Kuang

Subjectivity is the expression of internal opinions or beliefs which cannot be objectively observed or verified, and has been shown to be important for sentiment analysis and word-sense disambiguation. Furthermore, subjectivity is an…

计算与语言 · 计算机科学 2020-10-07 Johannes Bjerva , Nikita Bhutani , Behzad Golshan , Wang-Chiew Tan , Isabelle Augenstein

Uncertainty quantification is a set of techniques that measure confidence in language models. They can be used, for example, to detect hallucinations or alert users to review uncertain predictions. To be useful, these confidence scores must…

计算与语言 · 计算机科学 2026-04-13 Lorenzo Jaime Yu Flores , Cesare Spinoso di-Piano , Jackie Chi Kit Cheung

Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This…

音频与语音处理 · 电气工程与系统科学 2025-01-22 Julius Richter , Danilo de Oliveira , Timo Gerkmann

The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Objective evaluation metrics which consider human perception…

声音 · 计算机科学 2021-06-07 Szu-Wei Fu , Cheng Yu , Tsun-An Hsieh , Peter Plantinga , Mirco Ravanelli , Xugang Lu , Yu Tsao

Realizing general-purpose language intelligence has been a longstanding goal for natural language processing, where standard evaluation benchmarks play a fundamental and guiding role. We argue that for general-purpose language intelligence…

Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization criterion. On the other hand, automated metrics are efficient…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Pranay Manocha , Adam Finkelstein , Richard Zhang , Nicholas J. Bryan , Gautham J. Mysore , Zeyu Jin

ODAQ (Open Dataset of Audio Quality) provides a comprehensive framework for exploring both monaural and binaural audio quality degradations across a range of distortion classes and signals, accompanied by subjective quality ratings. A…

音频与语音处理 · 电气工程与系统科学 2025-12-12 Pablo M. Delgado , Sascha Dick , Christoph Thompson , Chih-Wei Wu , Phillip A. Williams

Virtual human animations have a wide range of applications in virtual and augmented reality. While automatic generation methods of animated virtual humans have been developed, assessing their quality remains challenging. Recently,…

图形学 · 计算机科学 2025-11-17 Rim Rekik , Stefanie Wuhrer , Ludovic Hoyet , Katja Zibrek , Anne-Hélène Olivier

Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so…

音频与语音处理 · 电气工程与系统科学 2026-03-16 Jaden Pieper , Stephen D. Voran