中文
相关论文

相关论文: STAR: Speech-to-Audio Generation via Representatio…

200 篇论文

Voice conversion (VC) transforms an utterance to sound like another person without changing the linguistic content. A recently proposed generative adversarial network-based VC method, StarGANv2-VC is very successful in generating…

音频与语音处理 · 电气工程与系统科学 2023-09-15 Arnab Das , Suhita Ghosh , Tim Polzehl , Sebastian Stober

This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised…

End-to-end models are an attractive new approach to spoken language understanding (SLU) in which the meaning of an utterance is inferred directly from the raw audio without employing the standard pipeline composed of a separately trained…

音频与语音处理 · 电气工程与系统科学 2019-10-22 Loren Lugosch , Brett Meyer , Derek Nowrouzezahrai , Mirco Ravanelli

The black-box nature of end-to-end speech translation (E2E ST) systems makes it difficult to understand how source language inputs are being mapped to the target language. To solve this problem, we would like to simultaneously generate…

计算与语言 · 计算机科学 2022-11-14 Motoi Omachi , Brian Yan , Siddharth Dalmia , Yuya Fujita , Shinji Watanabe

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

机器学习 · 计算机科学 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

In recent years, text-to-audio models have revolutionized the field of automatic audio generation. This paper investigates their application in generating synthetic datasets for training data-driven models. Specifically, this study analyzes…

音频与语音处理 · 电气工程与系统科学 2024-07-09 Francesca Ronchini , Luca Comanducci , Fabio Antonacci

Spoken language understanding system is traditionally designed as a pipeline of a number of components. First, the audio signal is processed by an automatic speech recognizer for transcription or n-best hypotheses. With the recognition…

计算与语言 · 计算机科学 2018-02-26 Dmitriy Serdyuk , Yongqiang Wang , Christian Fuegen , Anuj Kumar , Baiyang Liu , Yoshua Bengio

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

音频与语音处理 · 电气工程与系统科学 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by…

计算与语言 · 计算机科学 2025-01-10 Sanjana Sankar , Martin Lenglet , Gerard Bailly , Denis Beautemps , Thomas Hueber

In automatic speech translation (ST), traditional cascade approaches involving separate transcription and translation steps are giving ground to increasingly competitive and more robust direct solutions. In particular, by translating speech…

计算与语言 · 计算机科学 2020-12-10 Marco Gaido , Beatrice Savoldi , Luisa Bentivogli , Matteo Negri , Marco Turchi

Text-to-speech conversion has traditionally been performed either by concatenating short samples of speech or by using rule-based systems to convert a phonetic representation of speech into an acoustic representation, which is then…

神经与进化计算 · 计算机科学 2007-05-23 Orhan Karaali , Gerald Corrigan , Ira Gerson

End-to-end speech recognition systems have achieved competitive results compared to traditional systems. However, the complex transformations involved between layers given highly variable acoustic signals are hard to analyze. In this paper,…

计算与语言 · 计算机科学 2019-11-05 Chung-Yi Li , Pei-Chieh Yuan , Hung-Yi Lee

When building state-of-the-art speech translation models, the need for large computational resources is a significant obstacle due to the large training data size and complex models. The availability of pre-trained models is a promising…

计算与语言 · 计算机科学 2022-11-10 Zhaolin Li , Jan Niehues

Encoder-decoder based sequence-to-sequence models have demonstrated state-of-the-art results in end-to-end automatic speech recognition (ASR). Recently, the transformer architecture, which uses self-attention to model temporal context…

声音 · 计算机科学 2020-07-02 Niko Moritz , Takaaki Hori , Jonathan Le Roux

This paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by transfer learning from…

音频与语音处理 · 电气工程与系统科学 2019-08-14 Pavel Denisov , Ngoc Thang Vu

We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike previous approaches that rely on pretrained self-supervised…

声音 · 计算机科学 2025-06-05 Nadav Har-Tuv , Or Tal , Yossi Adi

Supervised training of speech recognition models requires access to transcribed audio data, which often is not possible due to confidentiality issues. Our approach to this problem is to generate synthetic audio from a text-only corpus using…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Yanis Perrin , Gilles Boulianne

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or reference image,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Foivos Paraperas Papantoniou , Stathis Galanakis , Rolandos Alexandros Potamias , Bernhard Kainz , Stefanos Zafeiriou

Brain-to-speech technology represents a fusion of interdisciplinary applications encompassing fields of artificial intelligence, brain-computer interfaces, and speech synthesis. Neural representation learning based intention decoding and…

人工智能 · 计算机科学 2024-02-28 Seo-Hyun Lee , Young-Eun Lee , Soowon Kim , Byung-Kwan Ko , Jun-Young Kim , Seong-Whan Lee

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

计算机视觉与模式识别 · 计算机科学 2019-10-03 Gaurav Mittal , Baoyuan Wang