English
Related papers

Related papers: Prosodic ABX: A Language-Agnostic Method for Measu…

200 papers

In expressive speech synthesis it is widely adopted to use latent prosody representations to deal with variability of the data during training. Same text may correspond to various acoustic realizations, which is known as a one-to-many…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-13 Mikolaj Babianski , Kamil Pokora , Raahil Shah , Rafal Sienkiewicz , Daniel Korzekwa , Viacheslav Klimkov

We investigate the possibility of forcing a self-supervised model trained using a contrastive predictive loss to extract slowly varying latent representations. Rather than producing individual predictions for each of the future…

In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained…

Computation and Language · Computer Science 2025-06-10 Rui Hu , Xiaolong Lin , Jiawang Liu , Shixi Huang , Zhenpeng Zhan

Existing fake audio detection systems perform well in in-domain testing, but still face many challenges in out-of-domain testing. This is due to the mismatch between the training and test data, as well as the poor generalizability of…

Sound · Computer Science 2023-05-24 Chenglong Wang , Jiangyan Yi , Jianhua Tao , Chuyuan Zhang , Shuai Zhang , Xun Chen

Vision-language models can encode societal biases and stereotypes, but there are challenges to measuring and mitigating these multimodal harms due to lacking measurement robustness and feature degradation. To address these challenges, we…

Machine Learning · Computer Science 2022-10-27 Hugo Berg , Siobhan Mackenzie Hall , Yash Bhalgat , Wonsuk Yang , Hannah Rose Kirk , Aleksandar Shtedritski , Max Bain

This paper presents methods for building speech recognizers tailored for Japanese speaking assessment tasks. Specifically, we build a speech recognizer that outputs phonemic labels with accent markers. Although Japanese is resource-rich,…

Computation and Language · Computer Science 2025-09-26 Yotaro Kubo , Richard Sproat , Chihiro Taguchi , Llion Jones

Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical…

Computation and Language · Computer Science 2025-03-25 Kwanghee Choi , Eunjung Yeo , Kalvin Chang , Shinji Watanabe , David Mortensen

In this paper, we explore automatic prediction of dialect density of the African American English (AAE) dialect, where dialect density is defined as the percentage of words in an utterance that contain characteristics of the non-standard…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-05 Alexander Johnson , Kevin Everson , Vijay Ravi , Anissa Gladney , Mari Ostendorf , Abeer Alwan

The prompt has become an effective linguistic tool for utilizing pre-trained language models. However, in few-shot scenarios, subtle changes in the prompt design always make the result widely different, and the prompt learning methods also…

Computation and Language · Computer Science 2024-03-13 Jinta Weng , Yifan Deng , d Donghao Li , Hao You , Yue Hu , Heyan Huang

Existing works on Aspect Sentiment Triplet Extraction (ASTE) explicitly focus on developing more efficient fine-tuning techniques for the task. Instead, our motivation is to come up with a generic approach that can improve the downstream…

Computation and Language · Computer Science 2023-10-25 Rajdeep Mukherjee , Nithish Kannen , Saurabh Kumar Pandey , Pawan Goyal

Self-supervised representation learning can mitigate the limitations in recognition tasks with few manually labeled data but abundant unlabeled data---a common scenario in sound event research. In this work, we explore unsupervised…

Sound · Computer Science 2020-11-17 Eduardo Fonseca , Diego Ortego , Kevin McGuinness , Noel E. O'Connor , Xavier Serra

Making decent multi-lingual sentence representations is critical to achieve high performances in cross-lingual downstream tasks. In this work, we propose a novel method to align multi-lingual embeddings based on the similarity of sentences…

Computation and Language · Computer Science 2024-05-29 Minsu Park , Seyeon Choi , Chanyeol Choi , Jun-Seong Kim , Jy-yong Sohn

Growing digital archives and improving algorithms for automatic analysis of text and speech create new research opportunities for fundamental research in phonetics. Such empirical approaches allow statistical evaluation of a much larger set…

Computation and Language · Computer Science 2017-06-05 Elodie Gauthier , Laurent Besacier , Sylvie Voisin

We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both…

Computation and Language · Computer Science 2024-10-15 Maureen de Seyssel , Antony D'Avirro , Adina Williams , Emmanuel Dupoux

Many self-supervised speech models (S3Ms) have been introduced over the last few years, improving performance and data efficiency on various speech tasks. However, these empirical successes alone do not give a complete picture of what is…

Computation and Language · Computer Science 2024-02-01 Ankita Pasad , Chung-Ming Chien , Shane Settle , Karen Livescu

Numerous examples in the literature proved that deep learning models have the ability to work well with multimodal data. Recently, CLIP has enabled deep learning systems to learn shared latent spaces between images and text descriptions,…

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-16 Aurosweta Mahapatra , Ismail Rasim Ulgen , Kong Aik Lee , Nicholas Andrews , Berrak Sisman

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

Tone is a prosodic feature used to distinguish words in many languages, some of which are endangered and scarcely documented. In this work, we use unsupervised representation learning to identify probable clusters of syllables that share…

Sound · Computer Science 2020-05-18 Bai Li , Jing Yi Xie , Frank Rudzicz

Our native language influences the way we perceive speech sounds, affecting our ability to discriminate non-native sounds. We compare two ideas about the influence of the native language on speech perception: the Perceptual Assimilation…

Computation and Language · Computer Science 2022-06-01 Juliette Millet , Ioana Chitoran , Ewan Dunbar