English
Related papers

Related papers: Zero resource speech synthesis using transcripts d…

200 papers

Most speech and language technologies are trained with massive amounts of speech and text information. However, most of the world languages do not have such resources or stable orthography. Systems constructed under these almost zero…

This study addresses the problem of unsupervised subword unit discovery from untranscribed speech. It forms the basis of the ultimate goal of ZeroSpeech 2019, building text-to-speech systems without text labels. In this work, unit discovery…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-29 Siyuan Feng , Tan Lee , Zhiyuan Peng

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-18 Francois G. Germain , Qifeng Chen , Vladlen Koltun

Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, classifying or indexing. Audio Word2Vec previously proposed was…

Computation and Language · Computer Science 2018-11-08 Sung-Feng Huang , Yi-Chen Chen , Hung-yi Lee , Lin-shan Lee

In this study, we present an innovative technique for speaker adaptation in order to improve the accuracy of segmentation with application to unit-selection Text-To-Speech (TTS) systems. Unlike conventional techniques for speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Claudio Zito , Fabio Tesser , Mauro Nicolao , Piero Cosi

We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-05 Jaeyeon Kim , Injune Hwang , Kyogu Lee

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional TTS for seen…

Sound · Computer Science 2023-05-24 Minki Kang , Wooseok Han , Sung Ju Hwang , Eunho Yang

Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance. However, existing LM-based VC models usually apply offline conversion from source semantics to acoustic features, demanding the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-22 Zhichao Wang , Yuanzhe Chen , Xinsheng Wang , Lei Xie , Yuping Wang

Recent efforts have aimed to utilize multilingual pretrained language models (mPLMs) to extend semantic parsing (SP) across multiple languages without requiring extensive annotations. However, achieving zero-shot cross-lingual transfer for…

Computation and Language · Computer Science 2024-10-02 Deokhyung Kang , Seonjeong Hwang , Yunsu Kim , Gary Geunbae Lee

Modern speech synthesis techniques can produce natural-sounding speech given sufficient high-quality data and compute resources. However, such data is not readily available for many languages. This paper focuses on speech synthesis for…

Computation and Language · Computer Science 2022-07-05 Perez Ogayo , Graham Neubig , Alan W Black

In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, close-talk mixture signals, in addition to using far-field…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-20 Zhong-Qiu Wang , Samuele Cornell

In this study, we tackle the challenge of inadequate and costly training data that has hindered the development of conversational question answering (ConvQA) systems. Enterprises have a large corpus of diverse internal documents. Instead of…

Computation and Language · Computer Science 2024-06-07 Fanyou Wu , Weijie Xu , Chandan K. Reddy , Srinivasan H. Sengamedu

Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis systems only consider the information within each sentence…

Sound · Computer Science 2023-05-19 Ya-Jie Zhang , Wei Song , Yanghao Yue , Zhengchen Zhang , Youzheng Wu , Xiaodong He

The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text-to-speech systems with voice cloning capabilities, that…

Sound · Computer Science 2024-09-20 Sebastião Quintas , Isabelle Ferrané , Thomas Pellegrini

The remarkable zero-shot reasoning capabilities of large-scale Visual Language Models (VLMs) on static images have yet to be fully translated to the video domain. Conventional video understanding models often rely on extensive,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Shihao Ji , Zihui Song

End-to-end (E2E) spoken language understanding (SLU) is constrained by the cost of collecting speech-semantics pairs, especially when label domains change. Hence, we explore \textit{zero-shot} E2E SLU, which learns E2E SLU without…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-06 Jianfeng He , Julian Salazar , Kaisheng Yao , Haoqi Li , Jinglun Cai

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only approaches at low…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Ahsan Adeel , Mandar Gogate , Amir Hussain

This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of naturalness, leaving…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-29 Junseok Ahn , Youkyum Kim , Yeunju Choi , Doyeop Kwak , Ji-Hoon Kim , Seongkyu Mun , Joon Son Chung

The human perception system is often assumed to recruit motor knowledge when processing auditory speech inputs. Using articulatory modeling and deep learning, this study examines how this articulatory information can be used for discovering…

Computation and Language · Computer Science 2022-06-20 Marc-Antoine Georges , Jean-Luc Schwartz , Thomas Hueber

Recent developments in neural speech synthesis and vocoding have sparked a renewed interest in voice conversion (VC). Beyond timbre transfer, achieving controllability on para-linguistic parameters such as pitch and Speed is critical in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-16 Meiying Chen , Zhiyao Duan