English
Related papers

Related papers: SingMOS: An extensive Open-Source Singing Voice Da…

200 papers

Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existing in audio signals, beyond spoken content, such as emotion,…

Sound · Computer Science 2025-01-16 Gallil Maimon , Amit Roth , Yossi Adi

Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference…

Sound · Computer Science 2025-03-10 Yulin Song , Guorui Sang , Jing Yu , Chuangbai Xiao

Singing voice synthesis (SVS) is a task that aims to generate audio signals according to musical scores and lyrics. With its multifaceted nature concerning music and language, producing singing voices indistinguishable from that of human…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-07 Yin-Ping Cho , Fu-Rong Yang , Yung-Chuan Chang , Ching-Ting Cheng , Xiao-Han Wang , Yi-Wen Liu

In singing voice synthesis (SVS), generating singing voices from musical scores faces challenges due to limited data availability. This study proposes a unique strategy to address the data scarcity in SVS. We employ an existing singing…

Sound · Computer Science 2024-06-14 Jiatong Shi , Yueqian Lin , Xinyi Bai , Keyi Zhang , Yuning Wu , Yuxun Tang , Yifeng Yu , Qin Jin , Shinji Watanabe

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-19 Paul A. Bereuter , Benjamin Stahl , Mark D. Plumbley , Alois Sontacchi

This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-30 Ryandhimas E. Zezario , Yu-Wen Chen , Szu-Wei Fu , Yu Tsao , Hsin-Min Wang , Chiou-Shann Fuh

Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Jaden Pieper , Stephen D. Voran

Automatic methods to predict Mean Opinion Score (MOS) of listeners have been researched to assure the quality of Text-to-Speech systems. Many previous studies focus on architectural advances (e.g. MBNet, LDNet, etc.) to capture relations…

Sound · Computer Science 2022-06-29 Aki Kunikoshi , Jaebok Kim , Wonsuk Jun , Kåre Sjölander

Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge is that anime-likeness, unlike naturalness, lacks a shared absolute scale, making conventional Mean…

Sound · Computer Science 2026-03-13 Joonyong Park , Jerry Li

Previous methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-14 Jozef Coldenhoff , Andrew Harper , Paul Kendrick , Tijana Stojkovic , Milos Cernak

The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-27 Max Morrison , Pranav Pawar , Nathan Pruyne , Jennifer Cole , Bryan Pardo

Singing Voice Synthesis (SVS) remains constrained in practical deployment due to its strong dependence on accurate phoneme-level alignment and manually annotated melody contours, requirements that are resource-intensive and hinder…

Sound · Computer Science 2025-12-05 Junjie Zheng , Chunbo Hao , Guobin Ma , Xiaoyu Zhang , Gongyu Chen , Chaofan Ding , Zihao Chen , Lei Xie

Musical dynamics form a core part of expressive singing voice performances. However, automatic analysis of musical dynamics for singing voice has received limited attention partly due to the scarcity of suitable datasets and a lack of clear…

Sound · Computer Science 2024-10-29 Jyoti Narang , Nazif Can Tamer , Viviana De La Vega , Xavier Serra

Methods for automatically assessing speech quality in real world environments are critical for developing robust human language technologies and assistive devices. Behavioral ratings provided by human raters (e.g., mean opinion scores; MOS)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-09 Mattson Ogg , Caitlyn Bishop , Han Yi , Sarah Robinson

The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS…

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not meet requirements for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-21 Rongjie Huang , Feiyang Chen , Yi Ren , Jinglin Liu , Chenye Cui , Zhou Zhao

Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly…

Sound · Computer Science 2025-12-03 Xueyan Li , Yuxin Wang , Mengjie Jiang , Qingzi Zhu , Jiang Zhang , Zoey Kim , Yazhe Niu

Evaluating speech generation still relies heavily on human judgments, such as Mean Opinion Score (MOS), which are expensive, subjective, and difficult to reproduce at scale. While a few recent studies have begun to explore AudioLLM-based…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-25 Yuanyuan Wang , Dongchao Yang , Yayue Deng , Zhiyong Wu , Yiwen Guo , Helen Meng , Xixin Wu

Automatic Mean Opinion Score (MOS) prediction is crucial to evaluate the perceptual quality of the synthetic speech. While recent approaches using pre-trained self-supervised learning (SSL) models have shown promising results, they only…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-01 Hui Wang , Shiwan Zhao , Xiguang Zheng , Yong Qin

The quality of the speech communication systems, which include noise suppression algorithms, are typically evaluated in laboratory experiments according to the ITU-T Rec. P.835, in which participants rate background noise, speech signal,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-19 Babak Naderi , Ross Cutler
‹ Prev 1 3 4 5 6 7 10 Next ›