English
Related papers

Related papers: Deep Speech Synthesis from Multimodal Articulatory…

200 papers

When the available data of a target speaker is insufficient to train a high quality speaker-dependent neural text-to-speech (TTS) system, we can combine data from multiple speakers and train a multi-speaker TTS model instead. Many studies…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-09 Hieu-Thi Luong , Xin Wang , Junichi Yamagishi , Nobuyuki Nishizawa

Multilingual end-to-end models have shown great improvement over monolingual systems. With the development of pre-training methods on speech, self-supervised multilingual speech representation learning like XLSR has shown success in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-08 Fenglin Ding , Genshun Wan , Pengcheng Li , Jia Pan , Cong Liu

Neural machine translation models have shown to achieve high quality when trained and fed with well structured and punctuated input texts. Unfortunately, the latter condition is not met in spoken language translation, where the input is…

Computation and Language · Computer Science 2019-10-24 Mattia Antonino Di Gangi , Robert Enyedi , Alessandra Brusadin , Marcello Federico

The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC)…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-04 Yilei Wu , Changyan Zheng , Xingyu Zhang , Yakun Zhang , Chengshi Zheng , Shuang Yang , Ye Yan , Erwei Yin

Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speech 3D gesture motion from text are a new and emerging field.…

Human-Computer Interaction · Computer Science 2024-05-01 Shivam Mehta , Anna Deichler , Jim O'Regan , Birger Moëll , Jonas Beskow , Gustav Eje Henter , Simon Alexanderson

The importance of modeling speech articulation for high-quality audiovisual (AV) speech synthesis is widely acknowledged. Nevertheless, while state-of-the-art, data-driven approaches to facial animation can make use of sophisticated motion…

Human-Computer Interaction · Computer Science 2012-09-25 Ingmar Steiner , Korin Richmond , Slim Ouni

In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level features from both text…

Computation and Language · Computer Science 2018-02-26 Yue Gu , Shuhong Chen , Ivan Marsic

Multimodal sentiment analysis is an important area for understanding the user's internal states. Deep learning methods were effective, but the problem of poor interpretability has gradually gained attention. Previous works have attempted to…

Computation and Language · Computer Science 2023-05-15 Sixia Li , Shogo Okada

Building a good speech recognition system usually requires large amounts of transcribed data, which is expensive to collect. To tackle this problem, many unsupervised pre-training methods have been proposed. Among these methods, Masked…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-24 Dongwei Jiang , Wubo Li , Ruixiong Zhang , Miao Cao , Ne Luo , Yang Han , Wei Zou , Xiangang Li

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-consuming,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-29 Tuomo Raitio , Javier Latorre , Andrea Davis , Tuuli Morrill , Ladan Golipour

Direct speech-to-text translation systems encounter an important drawback in data scarcity. A common solution consists on pretraining the encoder on automatic speech recognition, hence losing efficiency in the training process. In this…

Computation and Language · Computer Science 2024-09-27 Belen Alastruey , Gerard I. Gállego , Marta R. Costa-jussà

In end-to-end automatic speech recognition system, one of the difficulties for language expansion is the limited paired speech and text training data. In this paper, we propose a novel method to generate augmented samples with unpaired…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-01 Eric Sun , Jinyu Li , Jian Xue , Yifan Gong

This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tuned for a specific classification task. A classical issue with the fusion of many embeddings…

Computation and Language · Computer Science 2026-04-07 Nicolas Calbucura , Jose Guillen , Valentin Barriere

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Viet Anh Trinh , Sebastian Braun

Video-to-speech synthesis is the task of reconstructing the speech signal from a silent video of a speaker. Most established approaches to date involve a two-step process, whereby an intermediate representation from the video, such as a…

Sound · Computer Science 2024-10-28 Triantafyllos Kefalas , Yannis Panagakis , Maja Pantic

While neural text-to-speech systems perform remarkably well in high-resource scenarios, they cannot be applied to the majority of the over 6,000 spoken languages in the world due to a lack of appropriate training data. In this work, we use…

Computation and Language · Computer Science 2022-03-08 Florian Lux , Ngoc Thang Vu

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from real scenes often contains noise and generally needs to be…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Qiushi Zhu , Yu Gu , Rilin Chen , Chao Weng , Yuchen Hu , Lirong Dai , Jie Zhang

We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Julian Wechsler , Srikanth Raj Chetupalli , Wolfgang Mack , Emanuël A. P. Habets

Brain-to-speech technology represents a fusion of interdisciplinary applications encompassing fields of artificial intelligence, brain-computer interfaces, and speech synthesis. Neural representation learning based intention decoding and…

Artificial Intelligence · Computer Science 2024-02-28 Seo-Hyun Lee , Young-Eun Lee , Soowon Kim , Byung-Kwan Ko , Jun-Young Kim , Seong-Whan Lee

Spoken language understanding is typically based on pipeline architectures including speech recognition and natural language understanding steps. These components are optimized independently to allow usage of available data, but the overall…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-13 Pavel Denisov , Ngoc Thang Vu
‹ Prev 1 3 4 5 6 7 10 Next ›