English
Related papers

Related papers: Gesture2Speech: How Far Can Hand Movements Shape E…

200 papers

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

This paper presents an accented text-to-speech (TTS) synthesis framework with limited training data. We study two aspects concerning accent rendering: phonetic (phoneme difference) and prosodic (pitch pattern and phoneme duration)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-09 Xuehao Zhou , Mingyang Zhang , Yi Zhou , Zhizheng Wu , Haizhou Li

Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven methods, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jiahui Chen , Yang Huan , Runhua Shi , Chanfan Ding , Xiaoqi Mo , Siyu Xiong , Yinong He

Modern neural TTS systems are capable of generating natural and expressive speech when provided with sufficient amounts of training data. Such systems can be equipped with prosody-control functionality, allowing for more direct shaping of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-21 Slava Shechtman , Raul Fernandez

We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME), a new mesh-level…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Haiyang Liu , Zihao Zhu , Giorgio Becherini , Yichen Peng , Mingyang Su , You Zhou , Xuefei Zhe , Naoya Iwamoto , Bo Zheng , Michael J. Black

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive…

Sound · Computer Science 2022-01-19 Yi Lei , Shan Yang , Xinsheng Wang , Lei Xie

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to learn prosody…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Jakub Swiatkowski , Duo Wang , Mikolaj Babianski , Patrick Lumban Tobing , Ravichander Vipperla , Vincent Pollet

Speech is one of the most common forms of communication in humans. Speech commands are essential parts of multimodal controlling of prosthetic hands. In the past decades, researchers used automatic speech recognition systems for controlling…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-30 Mohsen Jafarzadeh , Yonas Tadesse

With the rapid advancement in deep generative models, recent neural Text-To-Speech(TTS) models have succeeded in synthesizing human-like speech. There have been some efforts to generate speech with various prosody beyond monotonous prosody…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-24 Seongho Joo , Hyukhun Koh , Kyomin Jung

Intonations play an important role in delivering the intention of a speaker. However, current end-to-end TTS systems often fail to model proper intonations. To alleviate this problem, we propose a novel, intuitive method to synthesize…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-08 Jihwan Lee , Joun Yeop Lee , Heejin Choi , Seongkyu Mun , Sangjun Park , Jae-Sung Bae , Chanwoo Kim

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a…

Sound · Computer Science 2025-05-29 Long-Khanh Pham , Thanh V. T. Tran , Minh-Tan Pham , Van Nguyen

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

While generative methods have progressed rapidly in recent years, generating expressive prosody for an utterance remains a challenging task in text-to-speech synthesis. This is particularly true for systems that model prosody explicitly…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-02 Paul Mayer , Florian Lux , Alejandro Pérez-González-de-Martos , Angelina Elizarova , Lindsey Vanderlyn , Dirk Väth , Ngoc Thang Vu

During speech, people spontaneously gesticulate, which plays a key role in conveying information. Similarly, realistic co-speech gestures are crucial to enable natural and smooth interactions with social agents. Current end-to-end co-speech…

Human-Computer Interaction · Computer Science 2021-01-15 Taras Kucherenko , Patrik Jonell , Sanne van Waveren , Gustav Eje Henter , Simon Alexanderson , Iolanda Leite , Hedvig Kjellström

This paper presents a novel framework for automatic speech-driven gesture generation, applicable to human-agent interaction including both virtual agents and robots. Specifically, we extend recent deep-learning-based, data-driven methods…

Human-Computer Interaction · Computer Science 2019-06-12 Taras Kucherenko , Dai Hasegawa , Gustav Eje Henter , Naoshi Kaneko , Hedvig Kjellström

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Anna Deichler , Jim O'Regan , Teo Guichoux , David Johansson , Jonas Beskow

Neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the synthetic speech often represents the average prosodic style of the database instead of having more versatile prosodic…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-24 Tuomo Raitio , Jiangchuan Li , Shreyas Seshadri

We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history. Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and…

Body language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the…

Robotics · Computer Science 2022-10-14 Hitoshi Teshima , Naoki Wake , Diego Thomas , Yuta Nakashima , Hiroshi Kawasaki , Katsushi Ikeuchi

This paper presents a simple yet effective method to achieve prosody transfer from a reference speech signal to synthesized speech. The main idea is to incorporate well-known acoustic correlates of prosody such as pitch and loudness…

Sound · Computer Science 2020-05-19 Siddharth Gururani , Kilol Gupta , Dhaval Shah , Zahra Shakeri , Jervis Pinto
‹ Prev 1 4 5 6 7 8 10 Next ›