English
Related papers

Related papers: CSS10: A Collection of Single Speaker Speech Datas…

200 papers

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend on explicit labels or reference audio, which limits both…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Yun Chen , Qi Chen , Zheqi Dai , Arshdeep Singh , Philip J. B. Jackson , Mark D. Plumbley

Speaker verification (SV) provides billions of voice-enabled devices with access control, and ensures the security of voice-driven technologies. As a type of biometrics, it is necessary that SV is unbiased, with consistent and reliable…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-14 Wiebke Toussaint Hutiri , Lauriane Gorce , Aaron Yi Ding

Goal: Numerous studies had successfully differentiated normal and abnormal voice samples. Nevertheless, further classification had rarely been attempted. This study proposes a novel approach, using continuous Mandarin speech instead of a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-23 Syu-Siang Wang , Chi-Te Wang , Chih-Chung Lai , Yu Tsao , Shih-Hau Fang

Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals,…

Sound · Computer Science 2026-01-28 Tianhua Li , Chenda Li , Wei Wang , Xin Zhou , Xihui Chen , Jianqing Gao , Yanmin Qian

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose…

Computation and Language · Computer Science 2025-11-11 Yuan Ge , Junxiang Zhang , Xiaoqian Liu , Bei Li , Xiangnan Ma , Chenglong Wang , Kaiyang Ye , Yangfan Du , Linfeng Zhang , Yuxin Huang , Tong Xiao , Zhengtao Yu , JingBo Zhu

MOS (Mean Opinion Score) is a subjective method used for the evaluation of a system's quality. Telecommunications (for voice and video), and speech synthesis systems (for generated speech) are a few of the many applications of the method.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-26 Bálint Gyires-Tóth , Csaba Zainkó

While many recent any-to-any voice conversion models succeed in transferring some target speech's style information to the converted speech, they still lack the ability to faithfully reproduce the speaking style of the target speaker. In…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-18 Hyungseob Lim , Kyungguen Byun , Sunkuk Moon , Erik Visser

Code-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phenomenon poses challenges for Automatic Speech Recognition…

Computation and Language · Computer Science 2025-06-03 Peng Xie , Xingyuan Liu , Tsz Wai Chan , Yequan Bie , Yangqiu Song , Yang Wang , Hao Chen , Kani Chen

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

Speech recognition and speech synthesis models are typically trained separately, each with its own set of learning objectives, training data, and model parameters, resulting in two distinct large networks. We propose a parameter-efficient…

Computation and Language · Computer Science 2024-10-25 Hawau Olamide Toyin , Hao Li , Hanan Aldarmaki

Synthetic data has become an important tool in the fine-tuning of language models to follow instructions and solve complex problems. Nevertheless, the majority of open data to date is often lacking multi-turn data and collected on closed…

Computation and Language · Computer Science 2024-07-29 Nathan Lambert , Hailey Schoelkopf , Aaron Gokaslan , Luca Soldaini , Valentina Pyatkin , Louis Castricato

Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide range of tasks.…

Computation and Language · Computer Science 2025-10-31 Pedro Corrêa , João Lima , Victor Moreno , Lucas Ueda , Paula Dornhofer Paro Costa

It is relatively easy to mine a large parallel corpus for any machine learning task, such as speech-to-text or speech-to-speech translation. Although these mined corpora are large in volume, their quality is questionable. This work shows…

Computation and Language · Computer Science 2024-02-06 Md Mahfuz Ibn Alam , Antonios Anastasopoulos

Models are increasing in size and complexity in the hunt for SOTA. But what if those 2\% increase in performance does not make a difference in a production use case? Maybe benefits from a smaller, faster model outweigh those slight…

Computation and Language · Computer Science 2022-04-12 Krzysztof Rajda , Łukasz Augustyniak , Piotr Gramacki , Marcin Gruza , Szymon Woźniak , Tomasz Kajdanowicz

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-22 Gaspard Michel , Elena V. Epure , Christophe Cerisara

The DISC project aims to (a) build an in-depth understanding of the state-of-the-art in spoken language dialogue systems (SLDSs) and components development and evaluation with the purpose of (b) developing a first best practice methodology…

Computation and Language · Computer Science 2007-05-23 Niels Ole Bernsen , Laila Dybkjaer , eds.

For real-life applications, it is crucial that end-to-end spoken language translation models perform well on continuous audio, without relying on human-supplied segmentation. For online spoken language translation, where models need to…

Computation and Language · Computer Science 2022-10-25 Chantal Amrhein , Barry Haddow

Although discrete speech tokens have exhibited strong potential for language model-based speech generation, their high bitrates and redundant timbre information restrict the development of such models. In this work, we propose LSCodec, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Yiwei Guo , Zhihan Li , Chenpeng Du , Hankun Wang , Xie Chen , Kai Yu

Many commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Shareef Babu Kalluri , Deepu Vijayasenan , Sriram Ganapathy , Ragesh Rajan M , Prashant Krishnan

Sequence-to-Sequence Text-to-Speech architectures that directly generate low level acoustic features from phonetic sequences are known to produce natural and expressive speech when provided with adequate amounts of training data. Such…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-26 Raul Fernandez , David Haws , Guy Lorberbom , Slava Shechtman , Alexander Sorin