English
Related papers

Related papers: Modeling Beyond MOS: Quality Assessment Models Mus…

200 papers

Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities are inconsistent.…

Artificial Intelligence · Computer Science 2025-10-14 Zhiyuan Han , Beier Zhu , Yanlong Xu , Peipei Song , Xun Yang

The application of visual instruction tuning and other post-training techniques has significantly enhanced the capabilities of Large Language Models (LLMs) in visual understanding, enriching Vision-Language Models (VLMs) with more…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Mingjie Xu , Andrew Estornell , Hongzheng Yang , Yuzhi Zhao , Zhaowei Zhu , Qi Xuan , Jiaheng Wei

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep…

Computation and Language · Computer Science 2016-11-29 Brian Patton , Yannis Agiomyrgiannakis , Michael Terry , Kevin Wilson , Rif A. Saurous , D. Sculley

Automated Essay Scoring (AES) is crucial for modern education, particularly with the increasing prevalence of multimodal assessments. However, traditional AES methods struggle with evaluation generalizability and multimodal perception,…

Computation and Language · Computer Science 2025-05-21 Jiamin Su , Yibo Yan , Zhuoran Gao , Han Zhang , Xiang Liu , Xuming Hu

Speech quality assessment (SQA) aims to evaluate the quality of speech samples without relying on time-consuming listener questionnaires. Recent efforts have focused on training neural-based SQA models to predict the mean opinion score…

Sound · Computer Science 2025-06-24 Yuto Kondo , Hirokazu Kameoka , Kou Tanaka , Takuhiro Kaneko

The acoustic environment can degrade speech quality during communication (e.g., video call, remote presentation, outside voice recording), and its impact is often unknown. Objective metrics for speech quality have proven challenging to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Karl El Hajal , Milos Cernak , Pablo Mainar

Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods generate features for modality missing from available ones, but…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Chenglizhao Chen , Yuchen Cao , Xinyu Liu , Mengke Song , Guisheng Zhang , Xiaomin Yu

Mean Opinion Score (MOS) is a popular measure for evaluating synthesized speech. However, the scores obtained in MOS tests are heavily dependent upon many contextual factors. One such factor is the overall range of quality of the samples…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Erica Cooper , Junichi Yamagishi

The ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-15 Gabriel Mittag , Saman Zadtootaghaj , Thilo Michael , Babak Naderi , Sebastian Möller

Multimodal Sentiment Analysis (MSA) endeavors to understand human sentiment by leveraging language, visual, and acoustic modalities. Despite the remarkable performance exhibited by previous MSA approaches, the presence of inherent…

Multimedia · Computer Science 2025-05-09 Weize Quan , Yunfei Feng , Ming Zhou , Yunzhen Zhao , Tong Wang , Dong-Ming Yan

With state-of-the-art models achieving high performance on standard benchmarks, contemporary research paradigms continue to emphasize general intelligence as an enduring objective. However, this pursuit overlooks the fundamental disparities…

Artificial Intelligence · Computer Science 2023-10-03 Nick DiSanto

Speech quality assessment has been a critical issue in speech processing for decades. Existing automatic evaluations usually require clean references or parallel ground truth data, which is infeasible when the amount of data soars.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-21 Wei-Cheng Tseng , Chien-yu Huang , Wei-Tsung Kao , Yist Y. Lin , Hung-yi Lee

Improving model robustness against potential modality noise, as an essential step for adapting multimodal models to real-world applications, has received increasing attention among researchers. For Multimodal Sentiment Analysis (MSA), there…

Multimedia · Computer Science 2022-11-28 Huisheng Mao , Baozheng Zhang , Hua Xu , Ziqi Yuan , Yihe Liu

Commonsense question-answering (QA) tasks, in the form of benchmarks, are constantly being introduced for challenging and comparing commonsense QA systems. The benchmarks provide question sets that systems' developers can use to train and…

Artificial Intelligence · Computer Science 2020-12-23 Henrique Santos , Minor Gordon , Zhicheng Liang , Gretchen Forbush , Deborah L. McGuinness

The rank correlation coefficients and the ranked-based statistical tests (as a subset of non-parametric techniques) might be misleading when they are applied to subjectively collected opinion scores. Those techniques assume that the data is…

Multimedia · Computer Science 2020-10-01 Babak Naderi , Sebastian Möller

This work discusses how to build more rational language and multimodal agents and what criteria define rationality in intelligent systems. Rationality is the quality of being guided by reason, characterized by decision-making that aligns…

Artificial Intelligence · Computer Science 2025-02-18 Bowen Jiang , Yangxinyu Xie , Xiaomeng Wang , Yuan Yuan , Zhuoqun Hao , Xinyi Bai , Weijie J. Su , Camillo J. Taylor , Tanwi Mallick

Document Visual Question Answering (VQA) models have evolved at an impressive rate over the past few years, coming close to or matching human performance on some benchmarks. We argue that common evaluation metrics used by popular benchmarks…

Computation and Language · Computer Science 2025-03-26 Armineh Nourbakhsh , Siddharth Parekh , Pranav Shetty , Zhao Jin , Sameena Shah , Carolyn Rose

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-05 Cedric Chan , Jianjing Kuang

Multimodal large language models (MLLMs) have shown great potential in perception and interpretation tasks, but their capabilities in predictive reasoning remain under-explored. To address this gap, we introduce a novel benchmark that…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Mingwei Zhu , Leigang Sha , Yu Shu , Kangjia Zhao , Tiancheng Zhao , Jianwei Yin

Foundation models (FMs) deployed in real-world tasks such as computer-use agents must integrate diverse modalities. How good are FMs at performing joint reasoning, simultaneously reasoning over multiple modalities, especially when the…

Artificial Intelligence · Computer Science 2025-10-07 Chen Henry Wu , Neil Kale , Aditi Raghunathan