中文
相关论文

相关论文: SPEAK WITH YOUR HANDS Using Continuous Hand Gestur…

200 篇论文

I present "SnakeSynth," a web-based lightweight audio synthesizer that combines audio generated by a deep generative model and real-time continuous two-dimensional (2D) input to create and control variable-length generative sounds through…

人机交互 · 计算机科学 2023-07-13 Eric Easthope

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

计算机视觉与模式识别 · 计算机科学 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

The hand gestures are one of the typical methods used in sign language. It is very difficult for the hearing-impaired people to communicate with the world. This project presents a solution that will not only automatically recognize the hand…

计算机视觉与模式识别 · 计算机科学 2018-11-30 K. Manikandan , Ayush Patidar , Pallav Walia , Aneek Barman Roy

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality assessment needs a…

计算与语言 · 计算机科学 2026-05-21 Vinicius Ribeiro , Yves Laprie

When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker…

计算机视觉与模式识别 · 计算机科学 2022-06-29 Christen Millerdurai , Lotfy Abdel Khaliq , Timon Ulrich

Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuanced articulatory…

音频与语音处理 · 电气工程与系统科学 2024-06-24 Zehua Kcriss Li , Meiying Melissa Chen , Yi Zhong , Pinxin Liu , Zhiyao Duan

During speech, people spontaneously gesticulate, which plays a key role in conveying information. Similarly, realistic co-speech gestures are crucial to enable natural and smooth interactions with social agents. Current end-to-end co-speech…

Articulatory features can provide interpretable and flexible controls for the synthesis of human vocalizations by allowing the user to directly modify parameters like vocal strain or lip position. To make this manipulation through…

声音 · 计算机科学 2023-07-11 David Südholt , Mateo Cámara , Zhiyuan Xu , Joshua D. Reiss

We present a non-supervised approach to optimize and evaluate the synthesis of non-speech audio effects from a speech production model. We use the Pink Trombone synthesizer as a case study of a simplified production model of the vocal tract…

音频与语音处理 · 电气工程与系统科学 2023-09-27 Mateo Cámara , Zhiyuan Xu , Yisu Zong , José Luis Blanco , Joshua D. Reiss

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations.…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Yihao Zhi , Xiaodong Cun , Xuelin Chen , Xi Shen , Wen Guo , Shaoli Huang , Shenghua Gao

Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline. This can lead to…

人机交互 · 计算机科学 2021-08-27 Siyang Wang , Simon Alexanderson , Joakim Gustafson , Jonas Beskow , Gustav Eje Henter , Éva Székely

Speech is one of the most common forms of communication in humans. Speech commands are essential parts of multimodal controlling of prosthetic hands. In the past decades, researchers used automatic speech recognition systems for controlling…

音频与语音处理 · 电气工程与系统科学 2020-11-30 Mohsen Jafarzadeh , Yonas Tadesse

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repetitive beat gestures…

人机交互 · 计算机科学 2025-10-21 Hendric Voss , Stefan Kopp

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

音频与语音处理 · 电气工程与系统科学 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotional speech synthesis often needs manual labels or reference…

声音 · 计算机科学 2020-11-18 Yi Lei , Shan Yang , Lei Xie

We present a novel neural encoder system for acoustic-to-articulatory inversion. We leverage the Pink Trombone voice synthesizer that reveals articulatory parameters (e.g tongue position and vocal cord configuration). Our system is designed…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Mateo Cámara , Fernando Marcos , José Luis Blanco

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interlocutor which is known as entrainment. To produce human-like…

音频与语音处理 · 电气工程与系统科学 2021-06-22 Jian Cong , Shan Yang , Na Hu , Guangzhi Li , Lei Xie , Dan Su

The automated synthesis of high-quality 3D gestures from speech is of significant value in virtual humans and gaming. Previous methods focus on synthesizing gestures that are synchronized with speech rhythm, yet they frequently overlook the…

人机交互 · 计算机科学 2024-09-24 Qingrong Cheng , Xu Li , Xinghui Fu , Fei Xia , Zhongqian Sun

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

音频与语音处理 · 电气工程与系统科学 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has made progress by using acoustic and semantic information as input and adopting classify method to…

声音 · 计算机科学 2024-04-16 Fan Zhang , Naye Ji , Fuxing Gao , Siyuan Zhao , Zhaohan Wang , Shunman Li