English
Related papers

Related papers: Perceived Femininity in Singing Voice: Analysis an…

200 papers

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker identity. Because SC is constrained to a single audio stream, it offers a more direct setting for…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Shree Harsha Bokkahalli Satish , Harm Lameris , Olivier Perrotin , Gustav Eje Henter , Éva Székely

We introduce a neural auto-encoder that transforms the musical dynamic in recordings of singing voice via changes in voice level. Since most recordings of singing voice are not annotated with voice level we propose a means to estimate the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-06 Frederik Bous , Axel Roebel

Gender recognition is an essential component of automatic speech recognition and interactive voice response systems. Determining gender of the speaker reduces the computational burden of such systems for any further processing. Typical…

Sound · Computer Science 2016-01-08 Jamil Ahmad , Mustansar Fiaz , Soon-il Kwon , Maleerat Sodanil , Bay Vo , Sung Wook Baik

Recently, and under the umbrella of Responsible AI, efforts have been made to develop gender-ambiguous synthetic speech to represent with a single voice all individuals in the gender spectrum. However, research efforts have completely…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-18 Maria Koutsogiannaki , Shafel Mc Dowall , Ioannis Agiomyrgiannakis

This paper proposes singing voice synthesis (SVS) based on frame-level sequence-to-sequence models considering vocal timing deviation. In SVS, it is essential to synchronize the timing of singing with temporal structures represented by…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Miku Nishihara , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled…

The quality of human voice plays an important role across various fields like music, speech therapy, and communication, yet it lacks a universally accepted, objective definition. Instead, voice quality is referred to using subjective…

Sound · Computer Science 2024-10-15 Hira Dhamyal , Rita Singh

The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice conveys extensive personal characteristics. We aim to create a…

Sound · Computer Science 2023-05-10 Wei Xue , Yiwen Wang , Qifeng Liu , Yike Guo

Commercial voice assistants are largely feminized and associated with stereotypically feminine traits such as warmth and submissiveness. As these assistants continue to be adopted for everyday uses, it is imperative to understand how the…

Human-Computer Interaction · Computer Science 2024-09-25 Amama Mahmood , Chien-Ming Huang

Music information retrieval is currently an active research area that addresses the extraction of musically important information from audio signals, and the applications of such information. The extracted information can be used for search…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-08 Preeti Rao

The rapid growth of Speech Emotion Recognition (SER) has diverse global applications, from improving human-computer interactions to aiding mental health diagnostics. However, SER models might contain social bias toward gender, leading to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-06 Yi-Cheng Lin , Haibin Wu , Huang-Cheng Chou , Chi-Chun Lee , Hung-yi Lee

The Voice Conversion Challenge 2020 is the third edition under its flagship that promotes intra-lingual semiparallel and cross-lingual voice conversion (VC). While the primary evaluation of the challenge submissions was done through…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-09 Rohan Kumar Das , Tomi Kinnunen , Wen-Chin Huang , Zhenhua Ling , Junichi Yamagishi , Yi Zhao , Xiaohai Tian , Tomoki Toda

While there is no a priori definition of good singing voices, we tend to make consistent evaluations of the quality of singing almost instantaneously. Such an instantaneous evaluation might be based on the sound spectrum that can be…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-31 Akito Yoshida , Shigeru Shinomoto

Singing voice synthesis (SVS) has seen remarkable advancements in recent years. However, compared to speech and general audio data, publicly available singing datasets remain limited. In practice, this data scarcity often leads to…

Sound · Computer Science 2025-12-17 Yiwen Zhao , Jiatong Shi , Yuxun Tang , William Chen , Shinji Watanabe

The application of text mining methods is becoming increasingly prevalent, particularly within Humanities and Computational Social Sciences, as well as in a broader range of disciplines. This paper presents an analysis of gender bias in…

Computation and Language · Computer Science 2025-03-13 Danqing Chen , Adithi Satish , Rasul Khanbayov , Carolin M. Schuster , Georg Groh

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal quality of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Ruiqi Li , Rongjie Huang , Yongqi Wang , Zhiqing Hong , Zhou Zhao

Singing voice conversion is to convert a singer's voice to another one's voice without changing singing content. Recent work shows that unsupervised singing voice conversion can be achieved with an autoencoder-based approach [1]. However,…

Sound · Computer Science 2020-02-19 Chengqi Deng , Chengzhu Yu , Heng Lu , Chao Weng , Dong Yu

Voice assistants (VAs) are typically evaluated through task performance metrics and self-report questionnaires, but people's voices themselves carry rich paralinguistic cues that reveal affect, effort, and interaction breakdowns. We present…

Human-Computer Interaction · Computer Science 2026-03-23 Yong Ma , Xuesong Zhang , Xuedong Zhang , Natalia Bartłomiejczyk , Seungwoo Je , Adrian Holzer , Morten Fjeld , Andreas Butz

In this study, the notion of perceptual features is introduced for describing general music properties based on human perception. This is an attempt at rethinking the concept of features, in order to understand the underlying human…

Information Retrieval · Computer Science 2014-04-01 Anders Friberg , Erwin Schoonderwaldt , Anton Hedblad , Marco Fabiani , Anders Elowsson