English
Related papers

Related papers: Teaching Machines to Speak Using Articulatory Cont…

200 papers

Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce conversational behaviour that adapts dynamically to the context. Current spoken…

Computation and Language · Computer Science 2026-04-16 Maike Züfle , Ondrej Klejch , Nicholas Sanders , Jan Niehues , Alexandra Birch , Tsz Kin Lam

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality assessment needs a…

Computation and Language · Computer Science 2026-05-21 Vinicius Ribeiro , Yves Laprie

During speech, people spontaneously gesticulate, which plays a key role in conveying information. Similarly, realistic co-speech gestures are crucial to enable natural and smooth interactions with social agents. Current end-to-end co-speech…

Human-Computer Interaction · Computer Science 2021-01-15 Taras Kucherenko , Patrik Jonell , Sanne van Waveren , Gustav Eje Henter , Simon Alexanderson , Iolanda Leite , Hedvig Kjellström

Acoustics-to-word models are end-to-end speech recognizers that use words as targets without relying on pronunciation dictionaries or graphemes. These models are notoriously difficult to train due to the lack of linguistic knowledge. It is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-14 Hao Tang , James Glass

Recently, neural network based dialogue systems have become ubiquitous in our increasingly digitalized society. However, due to their inherent opaqueness, some recently raised concerns about using neural models are starting to be taken…

Computation and Language · Computer Science 2020-05-28 Haochen Liu , Zhiwei Wang , Tyler Derr , Jiliang Tang

Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Jiachen Lian , Alan W Black , Yijing Lu , Louis Goldstein , Shinji Watanabe , Gopala K. Anumanchipalli

In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Chun-Yi Kuan , Chih-Kai Yang , Wei-Ping Huang , Ke-Han Lu , Hung-yi Lee

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Tamás Gábor Csapó

We study open domain dialogue generation with dialogue acts designed to explain how people engage in social chat. To imitate human behavior, we propose managing the flow of human-machine interactions with the dialogue acts as policies. The…

Computation and Language · Computer Science 2018-07-23 Can Xu , Wei Wu , Yu Wu

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

Computer Vision and Pattern Recognition · Computer Science 2019-10-03 Gaurav Mittal , Baoyuan Wang

Speech sounds of spoken language are obtained by varying configuration of the articulators surrounding the vocal tract. They contain abundant information that can be utilized to better understand the underlying mechanism of human speech…

Image and Video Processing · Electrical Eng. & Systems 2021-06-17 Laxmi Pandey , Ahmed Sabbir Arif

Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech…

Sound · Computer Science 2025-04-28 Leena G Pillai , D. Muhammad Noorul Mubarak , Elizabeth Sherly

Infant speech perception and learning is modeled using Echo State Network classification and Reinforcement Learning. Ambient speech for the modeled infant learner is created using the speech synthesizer Vocaltractlab. An auditory system is…

Sound · Computer Science 2016-10-21 Philip Zurbuchen

Self-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun "probing" models like HuBERT,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Cheol Jun Cho , Abdelrahman Mohamed , Alan W Black , Gopala K. Anumanchipalli

Effective spoken dialog systems should facilitate natural interactions with quick and rhythmic timing, mirroring human communication patterns. To reduce response times, previous efforts have focused on minimizing the latency in automatic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Oswald Zink , Yosuke Higuchi , Carlos Mullov , Alexander Waibel , Tetsunori Kobayashi

Automatic optimization of spoken dialog management policies that are robust to environmental noise has long been the goal for both academia and industry. Approaches based on reinforcement learning have been proved to be effective. However,…

Human-Computer Interaction · Computer Science 2016-11-18 Hang Ren , Weiqun Xu , Yonghong Yan

Despite significant advancements in natural language generation, controlling language models to produce texts with desired attributes remains a formidable challenge. In this work, we introduce RSA-Control, a training-free controllable text…

Artificial Intelligence · Computer Science 2024-10-28 Yifan Wang , Vera Demberg

This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by…

Computation and Language · Computer Science 2025-01-10 Sanjana Sankar , Martin Lenglet , Gerard Bailly , Denis Beautemps , Thomas Hueber

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Saba Tabatabaee , Suzanne Boyce , Liran Oren , Mark Tiede , Carol Espy-Wilson