中文
相关论文

相关论文: Speech MOS multi-task learning and rater bias corr…

200 篇论文

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for cross modality…

Methods for automatically assessing speech quality in real world environments are critical for developing robust human language technologies and assistive devices. Behavioral ratings provided by human raters (e.g., mean opinion scores; MOS)…

音频与语音处理 · 电气工程与系统科学 2025-10-09 Mattson Ogg , Caitlyn Bishop , Han Yi , Sarah Robinson

Emotion plays a fundamental role in human interaction, and therefore systems capable of identifying emotions in speech are crucial in the context of human-computer interaction. Speech emotion recognition (SER) is a challenging problem,…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Lucas Ueda , João Lima , Leonardo Marques , Paula Costa

Although numerous recent studies have suggested new frameworks for zero-shot TTS using large-scale, real-world data, studies that focus on the intelligibility of zero-shot TTS are relatively scarce. Zero-shot TTS demands additional efforts…

音频与语音处理 · 电气工程与系统科学 2024-01-31 Sunghee Jung , Won Jang , Jaesam Yoon , Bongwan Kim

Speech intelligibility and quality assessment models are essential tools for researchers to evaluate and improve speech processing models. However, only a few studies have investigated multi-task models for intelligibility and quality…

声音 · 计算机科学 2022-07-04 Yu-Wen Chen , Yu Tsao

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal…

The analysis of speech in individuals with amyotrophic lateral sclerosis is a powerful tool to support clinicians in the assessment of bulbar dysfunction. However, current methods used in clinical practice consist of subjective evaluations…

音频与语音处理 · 电气工程与系统科学 2025-05-28 Francesco Pierotti , Andrea Bandini

When the available data of a target speaker is insufficient to train a high quality speaker-dependent neural text-to-speech (TTS) system, we can combine data from multiple speakers and train a multi-speaker TTS model instead. Many studies…

音频与语音处理 · 电气工程与系统科学 2019-04-09 Hieu-Thi Luong , Xin Wang , Junichi Yamagishi , Nobuyuki Nishizawa

We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all participating teams, excluding the official baseline. In…

声音 · 计算机科学 2024-12-24 Yu-Fei Shi , Yang Ai , Ye-Xin Lu , Hui-Peng Du , Zhen-Hua Ling

Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due…

声音 · 计算机科学 2025-10-07 Takashi Maekaku , Keita Goto , Jinchuan Tian , Yusuke Shinohara , Shinji Watanabe

A pooling mechanism is essential for mean opinion score (MOS) prediction, facilitating the transformation of variable-length audio features into a concise fixed-size representation that effectively encodes speech quality. Existing pooling…

声音 · 计算机科学 2025-09-01 Cheng-Yeh Yang , Kuan-Tang Huang , Chien-Chun Wang , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen

Mood recognition is an important problem in music informatics and has key applications in music discovery and recommendation. These applications have become even more relevant with the rise of music streaming. Our work investigates the…

声音 · 计算机科学 2021-10-12 Rajnish Kumar , Manjeet Dahiya

This paper aims to enhance low-resource TTS by reducing training data requirements using compact speech representations. A Multi-Stage Multi-Codebook (MSMC) VQ-GAN is trained to learn the representation, MSMCR, and decode it to waveforms.…

声音 · 计算机科学 2022-10-28 Haohan Guo , Fenglong Xie , Xixin Wu , Hui Lu , Helen Meng

Smart contract vulnerabilities pose significant security risks to blockchain systems, potentially leading to severe financial losses. Existing methods face several limitations: (1) Program analysis-based approaches rely on predefined…

软件工程 · 计算机科学 2025-04-17 Hang Yuan , Lei Yu , Zhirong Huang , Jingyuan Zhang , Junyi Lu , Shiqi Cheng , Li Yang , Fengjun Zhang , Jiajia Ma , Chun Zuo

The rapid proliferation of AI-Generated Content (AIGC) has necessitated robust metrics for perceptual quality assessment. However, automatic Mean Opinion Score (MOS) prediction models are often compromised by data scarcity, predisposing…

音频与语音处理 · 电气工程与系统科学 2026-03-18 Kuan-Tang Huang , Chien-Chun Wang , Cheng-Yeh Yang , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen

Recent works in end-to-end speech-to-text translation (ST) have proposed multi-tasking methods with soft parameter sharing which leverage machine translation (MT) data via secondary encoders that map text inputs to an eventual cross-modal…

计算与语言 · 计算机科学 2023-09-28 Brian Yan , Xuankai Chang , Antonios Anastasopoulos , Yuya Fujita , Shinji Watanabe

A major bottleneck in training end-to-end task-oriented dialog system is the lack of data. To utilize limited training data more efficiently, we propose Modular Supervision Network (MOSS), an encoder-decoder training framework that could…

人工智能 · 计算机科学 2019-09-13 Weixin Liang , Youzhi Tian , Chengcai Chen , Zhou Yu

Deep neural network based speech enhancement technique focuses on learning a noisy-to-clean transformation supervised by paired training data. However, the task-specific evaluation metric (e.g., PESQ) is usually non-differentiable and can…

声音 · 计算机科学 2023-02-24 Chen Chen , Yuchen Hu , Weiwei Weng , Eng Siong Chng

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs…

Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only…

声音 · 计算机科学 2025-05-28 Jiatong Shi , Hye-Jin Shim , Shinji Watanabe