English
Related papers

Related papers: TG-Critic: A Timbre-Guided Model for Reference-Ind…

200 papers

In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric.…

Sound · Computer Science 2024-06-21 Yuxun Tang , Jiatong Shi , Yuning Wu , Qin Jin

We propose a novel method to model hierarchical metrical structures for both symbolic music and audio signals in a self-supervised manner with minimal domain knowledge. The model trains and inferences on beat-aligned music signals and…

Sound · Computer Science 2023-01-26 Junyan Jiang , Gus Xia

Music contains hierarchical structures beyond beats and measures. While hierarchical structure annotations are helpful for music information retrieval and computer musicology, such annotations are scarce in current digital music databases.…

Sound · Computer Science 2022-09-22 Junyan Jiang , Daniel Chin , Yixiao Zhang , Gus Xia

Timbre, the sound's unique "color", is fundamental to how we perceive and appreciate music. This review explores the multifaceted world of timbre perception and representation. It begins by tracing the word's origin, offering an intuitive…

Sound · Computer Science 2024-05-24 Hong Zhang , Jie Lin , Shengxuan Chen

The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-09 Sumit Kumar , Suraj Jaiswal , Parampreet Singh , Vipul Arora

Since the vocal component plays a crucial role in popular music, singing voice detection has been an active research topic in music information retrieval. Although several proposed algorithms have shown high performances, we argue that…

Sound · Computer Science 2018-06-05 Kyungyun Lee , Keunwoo Choi , Juhan Nam

Neural text classification models typically treat output labels as categorical variables which lack description and semantics. This forces their parametrization to be dependent on the label set size, and, hence, they are unable to scale to…

Computation and Language · Computer Science 2019-01-31 Nikolaos Pappas , James Henderson

Recently deep learning based recommendation systems have been actively explored to solve the cold-start problem using a hybrid approach. However, the majority of previous studies proposed a hybrid model where collaborative filtering and…

Information Retrieval · Computer Science 2018-07-19 Jongpil Lee , Kyungyun Lee , Jiyoung Park , Jangyeon Park , Juhan Nam

Expressive zero-shot voice conversion (VC) is a critical and challenging task that aims to transform the source timbre into an arbitrary unseen speaker while preserving the original content and expressive qualities. Despite recent progress…

Sound · Computer Science 2025-01-13 Yuguang Yang , Yu Pan , Jixun Yao , Xiang Zhang , Jianhao Ye , Hongbin Zhou , Lei Xie , Lei Ma , Jianjun Zhao

Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity. However, singing contains more expressive…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-07 Xu Li , Shansong Liu , Ying Shan

We establish THumB, a rubric-based human evaluation protocol for image captioning models. Our scoring rubrics and their definitions are carefully developed based on machine- and human-generated captions on the MSCOCO dataset. Each caption…

Computation and Language · Computer Science 2022-05-20 Jungo Kasai , Keisuke Sakaguchi , Lavinia Dunagan , Jacob Morrison , Ronan Le Bras , Yejin Choi , Noah A. Smith

Current computational-emotion research has focused on applying acoustic properties to analyze how emotions are perceived mathematically or used in natural language processing machine learning models. While recent interest has focused on…

Sound · Computer Science 2021-07-06 Daniel Szelogowski

In recent years, text-to-audio systems have achieved remarkable success, enabling the generation of complete audio segments directly from text descriptions. While these systems also facilitate music creation, the element of human creativity…

Sound · Computer Science 2025-04-15 Weixuan Yuan , Qadeer Khan , Vladimir Golkov

In a typical voice conversion system, prior works utilize various acoustic features (e.g., the pitch, voiced/unvoiced flag, aperiodicity) of the source speech to control the prosody of generated waveform. However, the prosody is related…

Sound · Computer Science 2020-06-01 Zheng Lian , Zhengqi Wen

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles are easy to model:…

Traditionally, Machine Translation (MT) Evaluation has been treated as a regression problem -- producing an absolute translation-quality score. This approach has two limitations: i) the scores lack interpretability, and human annotators…

Computation and Language · Computer Science 2024-01-31 Ibraheem Muhammad Moosa , Rui Zhang , Wenpeng Yin

This paper proposes an audio-conditioned phonemic and prosodic annotation model for building text-to-speech (TTS) datasets from unlabeled speech samples. For creating a TTS dataset that consists of label-speech paired data, the proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yuma Shirahata , Byeongseon Park , Ryuichi Yamamoto , Kentaro Tachibana

In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Shahan Nercessian , Johannes Imort , Ninon Devis , Frederik Blang

This paper proposes a method for selecting training data for text-to-speech (TTS) synthesis from dark data. TTS models are typically trained on high-quality speech corpora that cost much time and money for data collection, which makes it…

Sound · Computer Science 2022-10-27 Kentaro Seki , Shinnosuke Takamichi , Takaaki Saeki , Hiroshi Saruwatari

Time Scale Modification (TSM) is a well-researched field; however, no effective objective measure of quality exists. This paper details the creation, subjective evaluation, and analysis of a dataset for use in the development of an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-17 Timothy Roberts , Kuldip K. Paliwal