中文
相关论文

相关论文: Rich Prosody Diversity Modelling with Phone-level …

200 篇论文

This paper presents a self-supervised learning framework, named MGF, for general-purpose speech representation learning. In the design of MGF, speech hierarchy is taken into consideration. Specifically, we propose to use generative learning…

声音 · 计算机科学 2021-02-04 Yucheng Zhao , Dacheng Yin , Chong Luo , Zhiyuan Zhao , Chuanxin Tang , Wenjun Zeng , Zheng-Jun Zha

The Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter is an almost exact closed-form approximation to the Bayes-optimal multi-target tracking algorithm. Due to its optimality guarantees and ease of implementation, it has been…

信号处理 · 电气工程与系统科学 2025-05-20 Shiraz Khan , Yi-Chieh Sun , Inseok Hwang

A method for statistical parametric speech synthesis incorporating generative adversarial networks (GANs) is proposed. Although powerful deep neural networks (DNNs) techniques can be applied to artificially synthesize speech waveform, the…

声音 · 计算机科学 2017-09-26 Yuki Saito , Shinnosuke Takamichi , Hiroshi Saruwatari

Language models (LMs) for text data have been studied extensively for their usefulness in language generation and other downstream tasks. However, language modelling purely in the speech domain is still a relatively unexplored topic, with…

计算与语言 · 计算机科学 2021-11-02 Anurag Katakkar , Alan W Black

The evolutionary fitness landscape of biological molecules is extremely sparse and heterogeneous, with functional sequences forming isolated dense ``islands'' within a vast combinatorial space of largely non-functional variants. Protein…

In multi-target tracking (MTT), non-Gaussian measurement noise from sensors can diminish the performance of the Gaussian-assumed Gaussian mixture probability hypothesis density (GM-PHD) filter. In this paper, an approach that transforms the…

系统与控制 · 电气工程与系统科学 2023-09-18 Jiacheng He , Shan Zhong , Bei Peng , Gang Wang , Qizhen Wang

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. In this paper, a…

音频与语音处理 · 电气工程与系统科学 2021-12-24 Qicong Xie , Tao Li , Xinsheng Wang , Zhichao Wang , Lei Xie , Guoqiao Yu , Guanglu Wan

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily accessible data to clone a person's voice. (2) How to clone a…

音频与语音处理 · 电气工程与系统科学 2021-10-11 Dongyang Dai , Yuanzhe Chen , Li Chen , Ming Tu , Lu Liu , Rui Xia , Qiao Tian , Yuping Wang , Yuxuan Wang

Diffusion models have emerged as an expressive family of generative models rivaling GANs in sample quality and autoregressive models in likelihood scores. Standard diffusion models typically require hundreds of forward passes through the…

机器学习 · 计算机科学 2022-02-14 Daniel Watson , William Chan , Jonathan Ho , Mohammad Norouzi

Large-scale data collections in the wild, are invariably noisy. Thus developing data pruning strategies that remain robust even in the presence of corruption is critical in practice. In this work, we propose Geometric Median ($\gm$)…

机器学习 · 计算机科学 2025-01-20 Anish Acharya , Inderjit S Dhillon , Sujay Sanghavi

Generative models are increasingly able to produce remarkably high quality images and text. The community has developed numerous evaluation metrics for comparing generative models. However, these metrics do not effectively quantify data…

机器学习 · 计算机科学 2020-10-15 Liam Fowl , Micah Goldblum , Arjun Gupta , Amr Sharaf , Tom Goldstein

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

计算机视觉与模式识别 · 计算机科学 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for…

声音 · 计算机科学 2021-07-28 Shifeng Pan , Lei He

Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks…

音频与语音处理 · 电气工程与系统科学 2025-08-14 Eray Eren , Qingju Liu , Hyeongwoo Kim , Pablo Garrido , Abeer Alwan

This work addresses the challenge of making generative models suitable for resource-constrained environments like mobile wireless communication systems. We propose a generative model that integrates Autoregressive (AR) parameterization into…

信号处理 · 电气工程与系统科学 2026-05-19 Kathrin Klein , Benedikt Böck , Nurettin Turan , Wolfgang Utschick

State-of-the-art non-autoregressive text-to-speech (TTS) models based on FastSpeech 2 can efficiently synthesise high-fidelity and natural speech. For expressive speech datasets however, we observe characteristic audio distortions. We…

声音 · 计算机科学 2024-09-19 Fabian Kögel , Bac Nguyen , Fabien Cardinaux

We investigate multi-stage pretraining for prosody modeling in diffusion-based TTS. A speaker-conditioned dual-stream encoder is trained with masked language modeling followed by SigLIP-style cross-modal contrastive learning using…

We introduce Generalized Discrete Diffusion from Snapshots (GDDS), a unified framework for discrete diffusion modeling that supports arbitrary noising processes over large discrete state spaces. Our formulation encompasses all existing…

机器学习 · 统计学 2026-03-24 Oussama Zekri , Théo Uscidda , Nicolas Boullé , Anna Korba

For multichannel speech enhancement, this letter derives a robust maximum likelihood distortionless response beamformer by modeling speech sparse priors with a complex generalized Gaussian distribution, where we refer to as the CGGD-MLDR…

音频与语音处理 · 电气工程与系统科学 2021-02-22 Weixin Meng , Chengshi Zheng , Xiaodong Li

Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision-language inputs, multimodal distillation must preserve…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Jongoh Jeong , Hoyong Kwon , Minseok Kim , Kuk-Jin Yoon