English
Related papers

Related papers: AnimeScore: A Preference-Based Dataset and Framewo…

200 papers

In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the…

The area of automatic image caption evaluation is still undergoing intensive research to address the needs of generating captions which can meet adequacy and fluency requirements. Based on our past attempts at developing highly…

Computer Vision and Pattern Recognition · Computer Science 2020-12-25 Naeha Sharif , Lyndon White , Mohammed Bennamoun , Wei Liu , Syed Afaq Ali Shah

Human subjective evaluation is the gold standard to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. We have recently developed a non-intrusive speech quality…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-07 Chandan K A Reddy , Vishak Gopal , Ross Cutler

The task of automated code review has recently gained a lot of attention from the machine learning community. However, current review comment evaluation metrics rely on comparisons with a human-written reference for a given code change…

Software Engineering · Computer Science 2025-03-18 Atharva Naik , Marcus Alenius , Daniel Fried , Carolyn Rose

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficients often overlooks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Kai Zheng , Zejian Kang , Rui Mao , Hongyuan Zou , Yuanchen Fei , Xuanyang Xu , Xiangru Huang

Large language models (LLMs) have advanced conversational AI assistants. However, systematically evaluating how well these assistants apply personalization--adapting to individual user preferences while completing tasks--remains…

Computation and Language · Computer Science 2025-06-12 Zheng Zhao , Clara Vania , Subhradeep Kayal , Naila Khan , Shay B. Cohen , Emine Yilmaz

This paper addresses the automatic classification of X-rated videos by analyzing its obscene sounds. In this paper, obscene sounds refer to audio signals generated from sexual moans and screams during sexual scenes. By analyzing various…

Multimedia · Computer Science 2011-12-12 JaeDeok Lim , ByeongCheol Choi , SeungWan Han , ChoelHoon Lee

Personalized recommendation on new track releases has always been a challenging problem in the music industry. To combat this problem, we first explore user listening history and demographics to construct a user embedding representing the…

Sound · Computer Science 2021-03-31 Ke Chen , Beici Liang , Xiaoshuan Ma , Minwei Gu

Peer assessment is an efficient and effective learning assessment method that has been used widely in diverse fields in higher education. Despite its many benefits, a fundamental problem in peer assessment is that participants lack the…

Computers and Society · Computer Science 2015-06-19 Yanqing Wang , Yaowen Liang , Luning Liu , Ying Liu

This paper presents an experiment of automatically scoring handwritten descriptive answers in the trial tests for the new Japanese university entrance examination, which were made for about 120,000 examinees in 2017 and 2018. There are…

Machine Learning · Computer Science 2023-12-04 Hung Tuan Nguyen , Cuong Tuan Nguyen , Haruki Oka , Tsunenori Ishioka , Masaki Nakagawa

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs…

Cross-modal associations between voice and face from a person can be learnt algorithmically, which can benefit a lot of applications. The problem can be defined as voice-face matching and retrieval tasks. Much research attention has been…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Chuyuan Xiong , Deyuan Zhang , Tao Liu , Xiaoyong Du

Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the ASR transcript of a learner's speech. Recently, self-supervised learning (SSL) has shown stellar…

Sound · Computer Science 2025-03-04 Tien-Hong Lo , Fu-An Chao , Tzu-I Wu , Yao-Ting Sung , Berlin Chen

With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density…

Sound · Computer Science 2024-02-02 Huan Liao , Haonan Han , Kai Yang , Tianjiao Du , Rui Yang , Zunnan Xu , Qinmei Xu , Jingquan Liu , Jiasheng Lu , Xiu Li

Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts. However, the emotional fidelity of TTM systems remains largely underexplored compared to human…

Sound · Computer Science 2025-09-05 Gyehun Go , Satbyul Han , Ahyeon Choi , Eunjin Choi , Juhan Nam , Jeong Mi Park

We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of…

The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing models often fail to achieve a reliable…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-25 Jaemin Jung , Jaehun Kim , Inkyu Shin , Joon Son Chung

One challenging problem of robust automatic speech recognition (ASR) is how to measure the goodness of a speech enhancement algorithm (SEA) without calculating the word error rate (WER) due to the high costs of manual transcriptions,…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-29 Li Chai , Jun Du , Chin-Hui Lee

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep…

Computation and Language · Computer Science 2016-11-29 Brian Patton , Yannis Agiomyrgiannakis , Michael Terry , Kevin Wilson , Rif A. Saurous , D. Sculley