中文
相关论文

相关论文: Tongue pressure recordings during speech using com…

200 篇论文

Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1). Accented TTS synthesis is challenging as L2 is different from L1 in both in terms of phonetic rendering and…

声音 · 计算机科学 2022-09-23 Rui Liu , Berrak Sisman , Guanglai Gao , Haizhou Li

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

计算与语言 · 计算机科学 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture. To allow for supervised training, a teacher-student approach is employed: the teacher computes the target…

音频与语音处理 · 电气工程与系统科学 2023-09-20 Tobias Cord-Landwehr , Christoph Boeddeker , Cătălin Zorilă , Rama Doddipatla , Reinhold Haeb-Umbach

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

计算机视觉与模式识别 · 计算机科学 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

Humans involuntarily tend to infer parts of the conversation from lip movements when the speech is absent or corrupted by external noise. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate natural…

计算机视觉与模式识别 · 计算机科学 2020-05-19 K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

The noise of a device under test (DUT) is measured simultaneously with two instruments, each of which contributes its own background. The average cross power spectral density converges to the DUT power spectral density. This method enables…

仪器与探测器 · 物理学 2010-03-02 Enrico Rubiola , Francois Vernotte

Collapsible tubes can be employed to study the sound generation mechanism in the human respiratory system. The goals of this work are (a) to determine the airflow characteristics connected to three different collapse states of a…

流体动力学 · 物理学 2024-05-21 Marco Laudato , Elias Zea , Elias Sundström , Susann Boij , Mihai Mihaescu

The goal of this contribution is to use a parametric speech synthesis system for reducing background noise and other interferences from recorded speech signals. In a first step, Hidden Markov Models of the synthesis system are trained. Two…

声音 · 计算机科学 2017-07-06 Daniel Dzibela , Armin Sehr

The speech signal is a consummate example of time-series data. The acoustics of the signal change over time, sometimes dramatically. Yet, the most common type of comparison we perform in phonetics is between instantaneous acoustic…

音频与语音处理 · 电气工程与系统科学 2023-04-18 Matthew C. Kelley

The analysis of speech measures in individuals with amyotrophic lateral sclerosis (ALS) can provide essential information for early diagnosis and tracking disease progression. However, current methods for extracting speech and pause…

声音 · 计算机科学 2022-08-24 Saeid Alavi Naeini , Leif Simmatis , Yana Yunusova , Babak Taati

The aim of this project was to develop and implement an English language Text-to-Speech synthesis system. This involved a study of mechanisms of human speech production, a review of techniques in speech synthesis, and analysis of tests used…

声音 · 计算机科学 2017-09-25 David Ferris

Sound production due to turbulence is widely shown to be an important phenomenon involved in a.o. fricatives, singing, whispering and speech pathologies. In spite of its relevance turbulent flow is not considered in classical physical…

经典物理 · 物理学 2009-07-24 Xavier Grandchamp , Annemie Van Hirtum , Xavier Pelorson

The development of novel texture adaptations for the management of swallowing disorders could be accelerated if reliable in vitro tests were made available. This study addresses some of the limitations of swallowing in vitro models, by…

医学物理 · 物理学 2020-03-04 Marco Marconati , Silvia Pani , Jan Engmann , Adam Burbidge , Marco Ramaioli

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression,…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Baiqin Wang , Xiangyu Zhu , Fan Shen , Hao Xu , Zhen Lei

Speech data has rich acoustic and paralinguistic information with important cues for understanding a speaker's tone, emotion, and intent, yet traditional large language models such as BERT do not incorporate this information. There has been…

计算与语言 · 计算机科学 2023-11-14 Fatema Hasan , Yulong Li , James Foulds , Shimei Pan , Bishwaranjan Bhattacharjee

Prosody plays a vital role in verbal communication. Acoustic cues of prosody have been examined extensively. However, prosodic characteristics are not only perceived auditorily, but also visually based on head and facial movements. The…

计算与语言 · 计算机科学 2022-09-14 Hartmut Meister , Isa Samira Winter , Moritz Waeachtler , Pascale Sandmann , Khaled Abdellatif

Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and…

图形学 · 计算机科学 2025-04-01 Lee Chae-Yeon , Oh Hyun-Bin , Han EunGi , Kim Sung-Bin , Suekyeong Nam , Tae-Hyun Oh

Hand interactions are increasingly used as the primary input modality in immersive environments, but they are not always feasible due to situational impairments, motor limitations, and environmental constraints. Speech interfaces have been…

人机交互 · 计算机科学 2025-07-25 Chen Liang , Yuxuan Liu , Martez Mott , Anhong Guo

The human ear canal couples the external sound field to the eardrum and the solid parts of the middle ear. Therefore, knowledge of the acoustic impedance of the human ear is widely used in the industry to develop audio devices such as…

医学物理 · 物理学 2018-11-09 Søren Jønsson , Andreas Schuhmacher , Henrik Ingerslev Jørgensen

Many hearables contain an in-ear microphone, which may be used to capture the own voice of its user. However, due to the hearable occluding the ear canal, the in-ear microphone mostly records body-conducted speech, typically suffering from…

音频与语音处理 · 电气工程与系统科学 2024-09-09 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo