中文
相关论文

相关论文: DINO-VITS: Data-Efficient Zero-Shot TTS with Self-…

200 篇论文

The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remains challenging to learn and are typically developed for task-specific…

机器人学 · 计算机科学 2025-02-04 Gaoyue Zhou , Hengkai Pan , Yann LeCun , Lerrel Pinto

Deep learning-based speech enhancement has shown unprecedented performance in recent years. The most popular mono speech enhancement frameworks are end-to-end networks mapping the noisy mixture into an estimate of the clean speech. With…

音频与语音处理 · 电气工程与系统科学 2022-02-02 Bahareh Tolooshams , Kazuhito Koishida

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen speakers have diverse…

音频与语音处理 · 电气工程与系统科学 2022-04-04 Yihan Wu , Xu Tan , Bohan Li , Lei He , Sheng Zhao , Ruihua Song , Tao Qin , Tie-Yan Liu

In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this…

计算机视觉与模式识别 · 计算机科学 2021-05-25 Mathilde Caron , Hugo Touvron , Ishan Misra , Hervé Jégou , Julien Mairal , Piotr Bojanowski , Armand Joulin

Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Mu Yang , John H. L. Hansen

Data-driven speech enhancement employing deep neural networks (DNNs) can provide state-of-the-art performance even in the presence of non-stationary noise. During the training process, most of the speech enhancement neural networks are…

音频与语音处理 · 电气工程与系统科学 2021-04-01 Ziyi Xu , Maximilian Strake , Tim Fingscheidt

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Takaaki Saeki , Soumi Maiti , Xinjian Li , Shinji Watanabe , Shinnosuke Takamichi , Hiroshi Saruwatari

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method consists of first measuring the spectral tilt of unlabeled…

音频与语音处理 · 电气工程与系统科学 2022-03-30 Tuomo Raitio , Petko Petkov , Jiangchuan Li , Muhammed Shifas , Andrea Davis , Yannis Stylianou

Zero-shot speaker adaptation aims to clone an unseen speaker's voice without any adaptation time and parameters. Previous researches usually use a speaker encoder to extract a global fixed speaker embedding from reference speech, and…

声音 · 计算机科学 2022-11-14 Yixuan Zhou , Changhe Song , Xiang Li , Luwen Zhang , Zhiyong Wu , Yanyao Bian , Dan Su , Helen Meng

Computed tomography (CT) has played a vital role in medical diagnosis, assessment, and therapy planning, etc. In clinical practice, concerns about the increase of X-ray radiation exposure attract more and more attention. To lower the X-ray…

图像与视频处理 · 电气工程与系统科学 2022-01-19 Zhicheng Zhang , Xiaokun Liang , Wei Zhao , Lei Xing

Since large number of high-quality remote sensing images are readily accessible, exploiting the corpus of images with less manual annotation draws increasing attention. Self-supervised models acquire general feature representations by…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Xinye Wanyan , Sachith Seneviratne , Shuchang Shen , Michael Kirley

Zero-shot voice conversion is a technique that alters the speaker identity of an input speech to match a target speaker using only a single reference utterance, without requiring additional training. Recent approaches extensively utilize…

声音 · 计算机科学 2025-09-11 Youngjun Sim , Jinsung Yoon , Wooyeol Jeong , Young-Joo Suh

In the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers' data is a common approach. However, in terms of ultimately achieved system performance for target speaker(s), the actual…

音频与语音处理 · 电气工程与系统科学 2021-10-11 Guangyan Zhang , Yichong Leng , Daxin Tan , Ying Qin , Kaitao Song , Xu Tan , Sheng Zhao , Tan Lee

Zero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model on large-scale video datasets, which struggle to generalize…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Weihuang Liu , Xi Shen , Haolun Li , Xiuli Bi , Bo Liu , Chi-Man Pun , Xiaodong Cun

Recently, deep neural network (DNN)-based speech enhancement (SE) systems have been used with great success. During training, such systems require clean speech data - ideally, in large quantity with a variety of acoustic conditions, many…

音频与语音处理 · 电气工程与系统科学 2021-05-27 Koichi Saito , Stefan Uhlich , Giorgio Fabbro , Yuki Mitsufuji

We propose a novel self-supervised image blind denoising approach in which two neural networks jointly predict the clean signal and infer the noise distribution. Assuming that the noisy observations are independent conditionally to the…

机器学习 · 计算机科学 2021-02-17 Jean Ollion , Charles Ollion , Elisabeth Gassiat , Luc Lehéricy , Sylvain Le Corff

To enhance the performance of end-to-end (E2E) speech recognition systems in noisy or low signal-to-noise ratio (SNR) conditions, this paper introduces NoisyD-CT, a novel tri-stage training framework built on the Conformer-Transducer…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Shuangyuan Chen , Shuang Wei , Dongxing Xu , Yanhua Long

Despite their exceptional performance in vision tasks, deep learning models often struggle when faced with domain shifts during testing. Test-Time Training (TTT) methods have recently gained popularity by their ability to enhance the…

Over the last few years, deep learning has grown in popularity for speaker verification, identification, and diarization. Inarguably, a significant part of this success is due to the demonstrated effectiveness of their speaker…

声音 · 计算机科学 2022-10-07 Yehoshua Dissen , Felix Kreuk , Joseph Keshet

Supervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-world degraded speech data that may better represent the…

音频与语音处理 · 电气工程与系统科学 2021-09-22 Yangyang Xia , Buye Xu , Anurag Kumar