English
Related papers

Related papers: MMMOS: Multi-domain Multi-axis Audio Quality Asses…

200 papers

Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently underperform on fine-grained acoustic perception. We attribute this gap to a fundamental…

Computation and Language · Computer Science 2026-04-15 Linhao Zhang , Yuhan Song , Aiwei Liu , Chuhan Wu , Sijun Zhang , Wei Jia , Yuan Liu , Houfeng Wang , Xiao Zhou

This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-30 Ryandhimas E. Zezario , Yu-Wen Chen , Szu-Wei Fu , Yu Tsao , Hsin-Min Wang , Chiou-Shann Fuh

The Voice Conversion Challenge 2020 is the third edition under its flagship that promotes intra-lingual semiparallel and cross-lingual voice conversion (VC). While the primary evaluation of the challenge submissions was done through…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-09 Rohan Kumar Das , Tomi Kinnunen , Wen-Chin Huang , Zhenhua Ling , Junichi Yamagishi , Yi Zhao , Xiaohai Tian , Tomoki Toda

Mean Opinion Score (MOS) prediction has made significant progress in specific domains. However, the unstable performance of MOS prediction models across diverse samples presents ongoing challenges in the practical application of these…

Machine Learning · Computer Science 2024-08-26 Hui Wang , Shiwan Zhao , Jiaming Zhou , Xiguang Zheng , Haoqin Sun , Xuechen Wang , Yong Qin

Classroom discourse is an essential vehicle through which teaching and learning take place. Assessing different characteristics of discursive practices and linking them to student learning achievement enhances the understanding of teaching…

Computers and Society · Computer Science 2025-05-14 Ruikun Hou , Babette Bühler , Tim Fütterer , Efe Bozkir , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci

Speech Enhancement (SE) systems typically operate on monaural input and are used for applications including voice communications and capture cleanup for user generated content. Recent advancements and changes in the devices used for these…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-29 Aaron Master , Lie Lu , Nathan Swedlow

We propose a novel application based on acoustic-to-articulatory inversion towards quality assessment of voice converted speech. The ability of humans to speak effortlessly requires coordinated movements of various articulators, muscles,…

Sound · Computer Science 2015-11-24 Avni Rajpal , Nirmesh J. Shah , Mohammadi Zaki , Hemant A. Patil

While speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is evident in their struggle to disentangle complex features like…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-13 Guojian Li , Zhixian Zhao , Zhennan Lin , Jingbin Hu , Qirui Zhan , Yuang Cao , Pengyuan Xie , Chuan Xie , Jie Liu , Qiang Zhang , Zhonghua Fu , Lei Xie

Aspect-based sentiment analysis (ABSA) garnered growing research interest in multilingual contexts in the past. However, the majority of the studies lack more robust feature alignment and finer aspect-level alignment. In this paper, we…

Computation and Language · Computer Science 2026-04-13 Chengyan Wu , Bolei Ma , Ningyuan Deng , Yanqing He , Yun Xue , Xiaoyong Liu

Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets…

Multi-modal semantic segmentation (MMSS) faces significant challenges in real-world applications due to incomplete, degraded, or missing sensor data. While current MMSS methods typically use self-distillation with modality dropout to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Jiaqi Tan , Xu Zheng , Yang Liu

The Consensus Auditory-Perceptual Evaluation of Voice is a widely employed tool in clinical voice quality assessment that is significant for streaming communication among clinical professionals and benchmarking for the determination of…

Sound · Computer Science 2023-11-28 Yi-Heng Lin , Wen-Hsuan Tseng , Li-Chin Chen , Ching-Ting Tan , Yu Tsao

Usually, hearing impaired people use hearing aids which are implemented with speech enhancement algorithms. Estimation of speech and estimation of nose are the components in single channel speech enhancement system. The main objective of…

Sound · Computer Science 2014-11-10 M. Ravichandra Kumar , B. Ravi Teja

This study investigates the evaluation of multimedia quality models, focusing on the inherent uncertainties in subjective Mean Opinion Score (MOS) ratings due to factors like rater inconsistency and bias. Traditional statistical measures…

Multimedia · Computer Science 2024-11-12 Alessandro Ragano , Helard Becerra Martinez , Andrew Hines

Automatic methods to predict Mean Opinion Score (MOS) of listeners have been researched to assure the quality of Text-to-Speech systems. Many previous studies focus on architectural advances (e.g. MBNet, LDNet, etc.) to capture relations…

Sound · Computer Science 2022-06-29 Aki Kunikoshi , Jaebok Kim , Wonsuk Jun , Kåre Sjölander

Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are…

While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality assessment address the problem by predicting human…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-12 Soumi Maiti , Yifan Peng , Takaaki Saeki , Shinji Watanabe

The majority of online reviews consist of plain-text feedback together with a single numeric score. However, there are multiple dimensions to products and opinions, and understanding the `aspects' that contribute to users' ratings may help…

Computation and Language · Computer Science 2012-11-01 Julian McAuley , Jure Leskovec , Dan Jurafsky

Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches often rely on pretrained encoders, semantic distillation,…

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Siyin Wang , Wenyi Yu , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Lu Lu , Yu Tsao , Junichi Yamagishi , Yuxuan Wang , Chao Zhang