中文
相关论文

相关论文: Pre-training with Synthetic Patterns for Audio

200 篇论文

Masked autoencoder (MAE), a simple and effective self-supervised learning framework based on the reconstruction of masked image regions, has recently achieved prominent success in a variety of vision tasks. Despite the emergence of…

机器学习 · 计算机科学 2023-06-09 Lingjing Kong , Martin Q. Ma , Guangyi Chen , Eric P. Xing , Yuejie Chi , Louis-Philippe Morency , Kun Zhang

Supervised training of speech recognition models requires access to transcribed audio data, which often is not possible due to confidentiality issues. Our approach to this problem is to generate synthetic audio from a text-only corpus using…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Yanis Perrin , Gilles Boulianne

Although end-to-end text-to-speech (TTS) models such as Tacotron have shown excellent results, they typically require a sizable set of high-quality <text, audio> pairs for training, which are expensive to collect. In this paper, we propose…

计算与语言 · 计算机科学 2018-08-31 Yu-An Chung , Yuxuan Wang , Wei-Ning Hsu , Yu Zhang , RJ Skerry-Ryan

We investigate applying audio manipulations using pretrained neural network-based autoencoders as an alternative to traditional signal processing methods, since the former may provide greater semantic or perceptual organization. To…

音频与语音处理 · 电气工程与系统科学 2023-04-11 Scott H. Hawley , Christian J. Steinmetz

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by…

机器学习 · 计算机科学 2023-04-06 Alexandros Haliassos , Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Maja Pantic

Constructing a dataset for replay spoofing detection requires a physical process of playing an utterance and re-recording it, presenting a challenge to the collection of large-scale datasets. In this study, we propose a self-supervised…

机器学习 · 计算机科学 2020-08-20 Hye-jin Shim , Hee-Soo Heo , Jee-weon Jung , Ha-Jin Yu

The advent of generative AI models has revolutionized digital content creation, yet it introduces challenges in maintaining copyright integrity due to generative parroting, where models mimic their training data too closely. Our research…

机器学习 · 计算机科学 2024-06-21 Saeid Asgari Taghanaki , Joseph Lambourne

We examine the text-free speech representations of raw audio obtained from a self-supervised learning (SSL) model by analyzing the synthesized speech using the SSL representations instead of conventional text representations. Since raw…

计算与语言 · 计算机科学 2024-12-05 Joonyong Park , Daisuke Saito , Nobuaki Minematsu

In this paper, we suggest a framework to make use of mutual information as a regularization criterion to train Auto-Encoders (AEs). In the proposed framework, AEs are regularized by minimization of the mutual information between input and…

机器学习 · 计算机科学 2017-08-08 Yan Zhang , Mete Ozay , Zhun Sun , Takayuki Okatani

We present a multimodal framework to learn general audio representations from videos. Existing contrastive audio representation learning methods mainly focus on using the audio modality alone during training. In this work, we show that…

声音 · 计算机科学 2021-04-29 Luyu Wang , Pauline Luc , Adria Recasens , Jean-Baptiste Alayrac , Aaron van den Oord

In this paper, we are interested in unsupervised (unknown noise) audio-visual speech enhancement based on variational autoencoders (VAEs), where the probability distribution of clean speech spectra is simulated using an encoder-decoder…

音频与语音处理 · 电气工程与系统科学 2021-03-10 Mostafa Sadeghi , Xavier Alameda-Pineda

Supervised learning for single-channel speech enhancement requires carefully labeled training examples where the noisy mixture is input into the network and the network is trained to produce an output close to the ideal target. To relax the…

音频与语音处理 · 电气工程与系统科学 2020-06-19 Yu-Che Wang , Shrikant Venkataramani , Paris Smaragdis

Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still rely on an RNN-based context network to model temporal…

声音 · 计算机科学 2024-08-20 Pengfei Cai , Yan Song , Kang Li , Haoyu Song , Ian McLoughlin

Speech representation learning has improved both speech understanding and speech synthesis tasks for single language. However, its ability in cross-lingual scenarios has not been explored. In this paper, we extend the pretraining method for…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Xiaoran Fan , Chao Pang , Tian Yuan , He Bai , Renjie Zheng , Pengfei Zhu , Shuohuan Wang , Junkun Chen , Zeyu Chen , Liang Huang , Yu Sun , Hua Wu

Deep neural networks have been applied to audio spectrograms for respiratory sound classification, but it remains challenging to achieve satisfactory performance due to the scarcity of available data. Moreover, domain mismatch may be…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Peidong Wei , Shiyu Miao , Lin Li

Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression systems and generation approaches like next-token prediction and…

声音 · 计算机科学 2025-09-10 Dimitrios Bralios , Paris Smaragdis , Jonah Casebeer

Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis systems only consider the information within each sentence…

声音 · 计算机科学 2023-05-19 Ya-Jie Zhang , Wei Song , Yanghao Yue , Zhengchen Zhang , Youzheng Wu , Xiaodong He

A significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs. Semi-supervised algorithms rely on labeled data to learn…

声音 · 计算机科学 2024-09-27 Pengfei Cai , Yan Song , Nan Jiang , Qing Gu , Ian McLoughlin

The NLP community has broadly focused on text-only approaches of cognitive state tasks, but audio can provide vital missing cues through prosody. We posit that text-to-speech models learn to track aspects of cognitive state in order to…

声音 · 计算机科学 2025-02-12 Adil Soubki , John Murzaku , Peter Zeng , Owen Rambow

This paper investigates a novel method for designing linear precoders with finite alphabet inputs based on autoencoders (AE) without the knowledge of the channel model. By model-free training of the autoencoder in a multiple-input…

信息论 · 计算机科学 2022-08-05 Chen Cao , Biqian Feng , Yongpeng Wu , Derrick Wing Kwan Ng , Wenjun Zhang