English
Related papers

Related papers: Teochew-Wild: The First In-the-wild Teochew Datase…

200 papers

In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos…

Multimedia · Computer Science 2024-03-05 Sindhu Hegde , Rudrabha Mukhopadhyay , C. V. Jawahar , Vinay Namboodiri

This paper introduces FT Speech, a new speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT). The corpus contains over 1,800 hours of transcribed speech by a total of 434 speakers.…

Computation and Language · Computer Science 2020-10-29 Andreas Kirkedal , Marija Stepanović , Barbara Plank

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech…

This paper reports on the development of a text-to-speech (TTS) system for Mizo, a low-resource, tonal, and Tibeto-Burman language spoken primarily in the Indian state of Mizoram. The TTS was built with only 5.18 hours of data; however, in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-06 Abhijit Mohanta , Remruatpuii , Priyankoo Sarmah , Rohit Sinha , Wendy Lalhminghlui

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spanning four language…

Computation and Language · Computer Science 2026-03-09 Mohammad Mamun Or Rashid

We present the first annotated corpus of nonverbal behaviors in receptionist interactions, and the first nonverbal corpus (excluding the original video and audio data) of service encounters freely available online. Native speakers of…

Computation and Language · Computer Science 2012-03-13 Maxim Makatchev , Reid Simmons , Majd Sakr

This paper introduces SpeeChain, an open-source Pytorch-based toolkit designed to develop the machine speech chain for large-scale use. This first release focuses on the TTS-to-ASR chain, a core component of the machine speech chain, that…

Computation and Language · Computer Science 2023-01-10 Heli Qi , Sashi Novitasari , Andros Tjandra , Sakriani Sakti , Satoshi Nakamura

In this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes advantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Michael Hassid , Michelle Tadmor Ramanovich , Brendan Shillingford , Miaosen Wang , Ye Jia , Tal Remez

Text Simplification is a task that has been minimally explored for low-resource languages. Consequently, there are only a few manually curated datasets. In this paper, we present a human curated sentence-level text simplification dataset…

Computation and Language · Computer Science 2024-12-03 Surangika Ranathunga , Rumesh Sirithunga , Himashi Rathnayake , Lahiru De Silva , Thamindu Aluthwala , Saman Peramuna , Ravi Shekhar

We introduce a text-to-speech (TTS) model called BASE TTS, which stands for $\textbf{B}$ig $\textbf{A}$daptive $\textbf{S}$treamable TTS with $\textbf{E}$mergent abilities. BASE TTS is the largest TTS model to-date, trained on 100K hours of…

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is urgently needed to…

Multimedia · Computer Science 2024-08-28 Zeyu Jin , Jia Jia , Qixin Wang , Kehan Li , Shuoyi Zhou , Songtao Zhou , Xiaoyu Qin , Zhiyong Wu

Building state-of-the-art text-to-speech (TTS) systems typically demands millions of hours of proprietary data and complex multi-stage architectures, creating substantial barriers for resource-constrained research teams. In this report, we…

Deaf or hard-of-hearing (DHH) speakers typically have atypical speech caused by deafness. With the growing support of speech-based devices and software applications, more work needs to be done to make these devices inclusive to everyone. To…

Sound · Computer Science 2023-06-27 Lester Phillip Violeta , Tomoki Toda

A Spoken dialogue system for an unseen language is referred to as Zero resource speech. It is especially beneficial for developing applications for languages that have low digital resources. Zero resource speech synthesis is the task of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-11 Karthik Pandia D S , Anusha Prakash , Mano Ranjith Kumar , Hema A Murthy

Abstract End-to-end text-to-speech (TTS) systems has proved its great success in the presence of a large amount of high-quality training data recorded in anechoic room with high-quality microphone. Another approach is to use available…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-22 Viet Lam Phung , Phan Huy Kinh , Anh Tuan Dinh , Quoc Bao Nguyen

Information surrounds people in modern life. Text is a very efficient type of information that people use for communication for centuries. However, automated text-in-the-wild recognition remains a challenging problem. The major limitation…

Computer Vision and Pattern Recognition · Computer Science 2023-03-30 Igor Markov , Sergey Nesteruk , Andrey Kuznetsov , Denis Dimitrov

We describe an effort to annotate a corpus of natural language instructions consisting of 622 wet lab protocols to facilitate automatic or semi-automatic conversion of protocols into a machine-readable format and benefit biological…

Computation and Language · Computer Science 2018-05-02 Chaitanya Kulkarni , Wei Xu , Alan Ritter , Raghu Machiraju

Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference…

Computation and Language · Computer Science 2025-10-29 Frederik Broy , Maike Züfle , Jan Niehues

Recent advances in text-to-speech (TTS) models show impressive speech naturalness and quality, yet the role of large-scale open data in driving this progress remains underexplored. In this work, we introduce Raon-OpenTTS, an open TTS model…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Semin Kim , Seungjun Chung , Taehong Moon , Sangheon Lee , Minyoung Ahn , Keon Lee , Nam Soo Kim , Jaewoong Cho , Ludwig Schmidt , Kangwook Lee , Dongmin Park

While deep learning-based text-to-speech (TTS) models such as VITS have shown excellent results, they typically require a sizable set of high-quality <text, audio> pairs to train, which is expensive to collect. So far, most languages in the…

Sound · Computer Science 2023-01-05 Xin Yuan , Robin Feng , Mingming Ye