English
Related papers

Related papers: Developing a Multi-Platform Speech Recording Syste…

200 papers

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and…

Building a natural language dataset requires caution since word semantics is vulnerable to subtle text change or the definition of the annotated concept. Such a tendency can be seen in generative tasks like question-answering and dialogue…

Computation and Language · Computer Science 2023-04-04 Won Ik Cho , Yoon Kyung Lee , Seoyeon Bae , Jihwan Kim , Sangah Park , Moosung Kim , Sowon Hahn , Nam Soo Kim

The construction of high-quality datasets is a cornerstone of modern text-to-speech (TTS) systems. However, the increasing scale of available data poses significant challenges, including storage constraints. To address these issues, we…

Sound · Computer Science 2025-07-14 Kentaro Seki , Shinnosuke Takamichi , Takaaki Saeki , Hiroshi Saruwatari

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice…

A balanced speech corpus is the basic need for any speech processing task. In this report we describe our effort on development of Assamese speech corpus. We mainly focused on some issues and challenges faced during development of the…

Computation and Language · Computer Science 2013-09-30 Himangshu Sarma , Navanath Saharia , Utpal Sharma , Smriti Kumar Sinha , Mancha Jyoti Malakar

To build Sounding Board, we develop a system architecture that is capable of accommodating dialog strategies that we designed for socialbot conversations. The architecture consists of a multi-dimensional language understanding module for…

Computation and Language · Computer Science 2020-05-07 Hao Fang

Most research on task oriented dialog modeling is based on written text input. However, users interact with practical dialog systems often using speech as input. Typically, systems convert speech into text using an Automatic Speech…

Artificial Intelligence · Computer Science 2022-12-20 Hagen Soltau , Izhak Shafran , Mingqiu Wang , Abhinav Rastogi , Jeffrey Zhao , Ye Jia , Wei Han , Yuan Cao , Aramys Miranda

This paper presents key principles and solutions to the challenges faced in designing a domain-specific conversational agent for the legal domain. It includes issues of scope, platform, architecture and preparation of input data. It…

Human-Computer Interaction · Computer Science 2021-09-02 Mudita Sharma , Tony Russell-Rose , Lina Barakat , Akitaka Matsuo

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted…

With the popularity of mobile devices, personalized speech recognizer becomes more realizable today and highly attractive. Each mobile device is primarily used by a single user, so it's possible to have a personalized recognizer well…

Computation and Language · Computer Science 2016-11-23 Bo-Hsiang Tseng , Hung-Yi Lee , Lin-Shan Lee

The availability of realistic simulated corpora is of key importance for the future progress of distant speech recognition technology. The reliability, flexibility and low computational cost of a data simulation process may ultimately allow…

Audio and Speech Processing · Electrical Eng. & Systems 2017-11-28 Mirco Ravanelli , Piergiorgio Svaizer , Maurizio Omologo

Disentangling uncorrelated information in speech utterances is a crucial research topic within speech community. Different speech-related tasks focus on extracting distinct speech representations while minimizing the affects of other…

Computation and Language · Computer Science 2023-09-26 Siqi Zheng , Luyao Cheng , Yafeng Chen , Hui Wang , Qian Chen

Domain-specific data is the crux of the successful transfer of machine learning systems from benchmarks to real life. In simple problems such as image classification, crowdsourcing has become one of the standard tools for cheap and…

Sound · Computer Science 2021-10-22 Nikita Pavlichenko , Ivan Stelmakh , Dmitry Ustalov

This paper introduces RyanSpeech, a new speech corpus for research on automated text-to-speech (TTS) systems. Publicly available TTS corpora are often noisy, recorded with multiple speakers, or lack quality male speech data. In order to…

Computation and Language · Computer Science 2021-06-17 Rohola Zandie , Mohammad H. Mahoor , Julia Madsen , Eshrat S. Emamian

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirable to design a…

Computation and Language · Computer Science 2024-08-27 Chien-yu Huang , Min-Han Shih , Ke-Han Lu , Chi-Yuan Hsiao , Hung-yi Lee

Conversational agents are gaining popularity with the increasing ubiquity of smart devices. However, training agents in a data driven manner is challenging due to a lack of suitable corpora. This paper presents a novel method for gathering…

Computation and Language · Computer Science 2018-09-20 Joachim Fainberg , Ben Krause , Mihai Dobre , Marco Damonte , Emmanuel Kahembwe , Daniel Duma , Bonnie Webber , Federico Fancellu

We introduce MyVoice, a crowdsourcing platform designed to collect Arabic speech to enhance dialectal speech technologies. This platform offers an opportunity to design large dialectal speech datasets; and makes them publicly available.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Yousseif Elshahawy , Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali

Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled…

Computation and Language · Computer Science 2022-03-08 Massa Baali , Wassim El-Hajj , Ahmed Ali

Audio captioning is a novel field of multi-modal translation and it is the task of creating a textual description of the content of an audio signal (e.g. "people talking in a big room"). The creation of a dataset for this task requires a…

Sound · Computer Science 2019-07-23 Samuel Lipping , Konstantinos Drossos , Tuomas Virtanen
‹ Prev 1 2 3 10 Next ›