中文
相关论文

相关论文: ToneUnit: A Speech Discretization Approach for Ton…

200 篇论文

Generating synthesised singing voice with models trained on speech data has many advantages due to the models' flexibility and controllability. However, since the information about the temporal relationship between segments and beats are…

声音 · 计算机科学 2021-09-07 Cong Zhang , Jian Zhu

Learning accent from crowd-sourced data is a feasible way to achieve a target speaker TTS system that can synthesize accent speech. To this end, there are two challenging problems to be solved. First, direct use of the poor acoustic quality…

声音 · 计算机科学 2022-11-01 Yongmao Zhang , Zhichao Wang , Peiji Yang , Hongshen Sun , Zhisheng Wang , Lei Xie

Recent advancements in speech-based topic segmentation have highlighted the potential of pretrained speech encoders to capture semantic representations directly from speech. Traditionally, topic segmentation has relied on a pipeline…

计算与语言 · 计算机科学 2024-09-11 Sakshi Deo Shukla , Pavel Denisov , Tugtekin Turan

This paper proposes a direct text to speech translation system using discrete acoustic units. This framework employs text in different source languages as input to generate speech in the target language without the need for text…

计算与语言 · 计算机科学 2023-09-15 Victoria Mingote , Pablo Gimeno , Luis Vicente , Sameer Khurana , Antoine Laurent , Jarod Duret

In recent years, the field of image generation has been revolutionized by the application of autoregressive transformers and DDPMs. These approaches model the process of image generation as a step-wise probabilistic processes and leverage…

声音 · 计算机科学 2023-05-25 James Betker

Vocoders received renewed attention as main components in statistical parametric text-to-speech (TTS) synthesis and speech transformation systems. Even though there are vocoding techniques give almost accepted synthesized speech, their high…

声音 · 计算机科学 2021-06-22 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Géza Németh

Synthesized speech is common today due to the prevalence of virtual assistants, easy-to-use tools for generating and modifying speech signals, and remote work practices. Synthesized speech can also be used for nefarious purposes, including…

声音 · 计算机科学 2022-05-05 Emily R. Bartusiak , Edward J. Delp

As a special machine translation task, dialect translation has two main characteristics: 1) lack of parallel training corpus; and 2) possessing similar grammar between two sides of the translation. In this paper, we investigate how to…

计算与语言 · 计算机科学 2022-10-20 Yu Wan , Baosong Yang , Derek F. Wong , Lidia S. Chao , Haihua Du , Ben C. H. Ao

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-synthesis paradigm.…

声音 · 计算机科学 2025-10-16 Xue Jiang , Xiulian Peng , Yuan Zhang , Yan Lu

Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discretizing continuous…

机器学习 · 计算机科学 2025-10-28 Luca Della Libera , Francesco Paissan , Cem Subakan , Mirco Ravanelli

Token-based language modeling is a prominent approach for speech generation, where tokens are obtained by quantizing features from self-supervised learning (SSL) models and extracting codes from neural speech codecs, generally referred to…

音频与语音处理 · 电气工程与系统科学 2025-06-30 Yang Yang , Yunpeng Li , George Sung , Shao-Fu Shih , Craig Dooley , Alessio Centazzo , Ramanan Rajeswaran

This paper propose to combine pretrained language models with the modular dialogue paradigm for open-domain dialogue modeling. Our method, semantic-enhanced finetuning, instantiates conversation understanding, planning, and response…

计算与语言 · 计算机科学 2022-05-25 Yinhe Zheng , Yida Wang , Pei Ke , Zhenyu Yang , Minlie Huang

How can speech-to-text translation (ST) perform as well as machine translation (MT)? The key point is to bridge the modality gap between speech and text so that useful MT techniques can be applied to ST. Recently, the approach of…

计算与语言 · 计算机科学 2023-05-22 Dong Zhang , Rong Ye , Tom Ko , Mingxuan Wang , Yaqian Zhou

This paper proposes a unimodal aggregation (UMA) based nonautoregressive model for both English and Mandarin speech recognition. The original UMA explicitly segments and aggregates acoustic frames (with unimodal weights that first…

计算与语言 · 计算机科学 2025-09-19 Ying Fang , Xiaofei Li

Recent advances in cross-lingual text-to-speech (TTS) made it possible to synthesize speech in a language foreign to a monolingual speaker. However, there is still a large gap between the pronunciation of generated cross-lingual speech and…

声音 · 计算机科学 2022-02-23 Jianhao Ye , Hongbin Zhou , Zhiba Su , Wendi He , Kaimeng Ren , Lin Li , Heng Lu

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modalities remains a…

音频与语音处理 · 电气工程与系统科学 2026-04-27 Chunyu Qiang , Xiaopeng Wang , Kang Yin , Yuzhe Liang , Yuxin Guo , Teng Ma , Ziyu Zhang , Tianrui Wang , Cheng Gong , Yushen Chen , Ruibo Fu , Chen Zhang , Longbiao Wang , Jianwu Dang

Recent advances in text-to-speech have significantly improved the expressiveness of synthesized speech. However, it is still challenging to generate speech with contextually appropriate and coherent speaking style for multi-sentence text in…

声音 · 计算机科学 2023-04-14 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Shiyin Kang , Helen Meng

Increasing amount of research has shed light on machine perception of audio events, most of which concerns detection and classification tasks. However, human-like perception of audio scenes involves not only detecting and classifying audio…

声音 · 计算机科学 2020-05-11 Mengyue Wu , Heinrich Dinkel , Kai Yu

Domain shift is a prominent problem in Deep Learning, causing a model pre-trained on a source dataset to suffer significant performance degradation on test datasets. This research aims to address the issue of audio classification under…

机器学习 · 计算机科学 2025-07-22 Weichuang Shao , Iman Yi Liao , Tomas Henrique Bode Maul , Tissa Chandesa

In recent years, Text-To-Speech (TTS) has been used as a data augmentation technique for speech recognition to help complement inadequacies in the training data. Correspondingly, we investigate the use of a multi-speaker TTS system to…

音频与语音处理 · 电气工程与系统科学 2020-11-25 Yiling Huang , Yutian Chen , Jason Pelecanos , Quan Wang