English
Related papers

Related papers: AnimeScore: A Preference-Based Dataset and Framewo…

200 papers

Automatic Speech Scoring (ASS) is the computer-assisted evaluation of a candidate's speaking proficiency in a language. ASS systems face many challenges like open grammar, variable pronunciations, and unstructured or semi-structured…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Yaman Kumar Singla , Avykat Gupta , Shaurya Bagga , Changyou Chen , Balaji Krishnamurthy , Rajiv Ratn Shah

Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which such studies arrive at this conclusion is inconsistent. To pave the wayfor more reliable…

Computation and Language · Computer Science 2026-05-12 Felix Herron , Ange Richard , François Portet , Alexandre Allauzen , Solange Rossato

The industrial recommender systems always pursue more than one business goals. The inherent intensions between objectives pose significant challenges for ranking stage. A popular solution is to build a multi-objective ensemble (ME) model to…

Information Retrieval · Computer Science 2026-02-10 Boyang Xia , Zhou Yu , Zhiliang Zhu , Hanxiao Sun , Biyun Han , Jun Wang , Runnan Liu , Wenwu Ou

The quality of the speech communication systems, which include noise suppression algorithms, are typically evaluated in laboratory experiments according to the ITU-T Rec. P.835, in which participants rate background noise, speech signal,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-19 Babak Naderi , Ross Cutler

The ITU-T Recommendation P.808 provides a crowdsourcing approach for conducting a subjective assessment of speech quality using the Absolute Category Rating (ACR) method. We provide an open-source implementation of the ITU-T Rec. P.808 that…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Babak Naderi , Ross Cutler

Speech language models have significantly advanced in generating realistic speech, with neural codec language models standing out. However, the integration of human feedback to align speech outputs to human preferences is often neglected.…

Computation and Language · Computer Science 2024-04-09 Dong Zhang , Zhaowei Li , Shimin Li , Xin Zhang , Pengyu Wang , Yaqian Zhou , Xipeng Qiu

We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly cartoon caption…

Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When faced with higher-rate audio at test time, these models can produce biased scores. We introduce HighRateMOS, the first…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-30 Wenze Ren , Yi-Cheng Lin , Wen-Chin Huang , Ryandhimas E. Zezario , Szu-Wei Fu , Sung-Feng Huang , Erica Cooper , Haibin Wu , Hung-Yu Wei , Hsin-Min Wang , Hung-yi Lee , Yu Tsao

The ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-15 Gabriel Mittag , Saman Zadtootaghaj , Thilo Michael , Babak Naderi , Sebastian Möller

Evaluating concept customization is challenging, as it requires a comprehensive assessment of fidelity to generative prompts and concept images. Moreover, evaluating multiple concepts is considerably more difficult than evaluating a single…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Reina Ishikawa , Ryo Fujii , Hideo Saito , Ryo Hachiuma

Suicide remains a public health challenge, necessitating improved detection methods to facilitate timely intervention and treatment. This systematic review evaluates the role of Artificial Intelligence (AI) and Machine Learning (ML) in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-29 Ambre Marie , Marine Garnier , Thomas Bertin , Laura Machart , Guillaume Dardenne , Gwenolé Quellec , Sofian Berrouiguet

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-27 Khandokar Md. Nayem , Donald S. Williamson

User preferences are increasingly used to personalize Large Language Model (LLM) responses, yet how to reliably leverage preference signals for answer generation remains under-explored. In practice, preferences can be noisy, incomplete, or…

Computation and Language · Computer Science 2026-04-09 Tianyu Zhao , Siqi Li , Yasser Shoukry , Salma Elmalaki

For d/Deaf and hard of hearing (DHH) people, captioning is an essential accessibility tool. Significant developments in artificial intelligence (AI) mean that Automatic Speech Recognition (ASR) is now a part of many popular applications.…

Computation and Language · Computer Science 2024-08-30 Korbinian Kuhn , Verena Kersken , Benedikt Reuter , Niklas Egger , Gottfried Zimmermann

The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while…

Computation and Language · Computer Science 2026-05-04 Woody Haosheng Gan , William Held , Diyi Yang

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can…

Sound · Computer Science 2025-12-01 Chen Li , Peiji Yang , Yicheng Zhong , Jianxing Yu , Zhisheng Wang , Zihao Gou , Wenqing Chen , Jian Yin

With the recent rapid increase in digitization across all major industries, acquiring programming skills has increased the demand for introductory programming courses. This has further resulted in universities integrating programming…

Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent…

Sound · Computer Science 2025-07-11 Haokun Tian , Stefan Lattner , Charalampos Saitis

Conversational search systems, such as Google Assistant and Microsoft Cortana, provide a new search paradigm where users are allowed, via natural language dialogues, to communicate with search systems. Evaluating such systems is very…

Information Retrieval · Computer Science 2021-09-08 Zeyang Liu , Ke Zhou , Jiaxin Mao , Max L. Wilson

This paper studies the quality of multimedia content focusing on 360 video and ambisonic spatial audio reproduced using a head-mounted display and a multichannel loudspeaker setup. Encoding parameters following basic video quality test…

Multimedia · Computer Science 2020-05-20 Randy Frans Fela , Nick Zacharov , Søren Forchhammer