中文
相关论文

相关论文: Deep-Learning-Based Audio-Visual Speech Enhancemen…

200 篇论文

Humans tend to change their way of speaking when they are immersed in a noisy environment, a reflex known as Lombard effect. Current speech enhancement systems based on deep learning do not usually take into account this change in the…

音频与语音处理 · 电气工程与系统科学 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan , Sigurdur Sigurdsson , Jesper Jensen

Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all of them ignore the impact of the Lombard effect, i.e., the…

音频与语音处理 · 电气工程与系统科学 2019-07-10 Pingchuan Ma , Stavros Petridis , Maja Pantic

Speakers usually adjust their way of talking in noisy environments involuntarily for effective communication. This adaptation is known as the Lombard effect. Although speech accompanying the Lombard effect can improve the intelligibility of…

音频与语音处理 · 电气工程与系统科学 2019-04-10 Yi Zhao , Atsushi Ando , Shinji Takaki , Junichi Yamagishi , Satoshi Kobashikawa

Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and the Lombard effect is…

声音 · 计算机科学 2025-07-15 Dominika Woszczyk , Manuel Sam Ribeiro , Thomas Merritt , Daniel Korzekwa

The Lombard effect refers to individuals' unconscious modulation of vocal effort in response to variations in the ambient noise levels, intending to enhance speech intelligibility. The impact of different decibel levels and types of…

声音 · 计算机科学 2023-09-15 Qingmu Liu , Yuhong Yang , Baifeng Li , Hongyang Chen , Weiping Tu , Song Lin

The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for…

声音 · 计算机科学 2026-01-21 Seymanur Akti , Alexander Waibel

It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user. In this paper, we present and compare various approaches for…

音频与语音处理 · 电气工程与系统科学 2021-01-15 Qiong Hu , Tobias Bleisch , Petko Petkov , Tuomo Raitio , Erik Marchi , Varun Lakshminarasimhan

This study investigates the Lombard effect, where individuals adapt their speech in noisy environments. We introduce an enhanced Mandarin Lombard grid (EMALG) corpus with meaningful sentences , enhancing the Mandarin Lombard grid (MALG)…

声音 · 计算机科学 2024-01-10 Baifeng Li , Qingmu Liu , Yuhong Yang , Hongyang Chen , Weiping Tu , Song Lin

This study explores how sentence types affect the Lombard effect and intelligibility enhancement, focusing on comparisons between natural and grid sentences. Using the Lombard Chinese-TIMIT (LCT) corpus and the Enhanced MAndarin Lombard…

声音 · 计算机科学 2024-07-10 Hongyang Chen , Yuhong Yang , Zhongyuan Wang , Weiping Tu , Haojun Ai , Song Lin

For a better understanding of the mechanisms underlying speech perception and the contribution of different signal features, computational models of speech recognition have a long tradition in hearing research. Due to the diverse range of…

音频与语音处理 · 电气工程与系统科学 2022-04-15 Maximilian Karl Scharf , Sabine Hochmuth , Lena L. N. Wong , Birger Kollmeier , Anna Warzybok

In this study, we present a new experiment in order to study the Lombard effect in telepresence robotics. In this experiment, one person talks with a robot controled remotely by someone in a different room. The remote pilot (R) is immersed…

机器人学 · 计算机科学 2023-05-31 Ambre Davat , Gang Feng , Véronique Aubergé

Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yihuan Huang , Jun Xue , Liu Jiajun , Daixian Li , Tong Zhang , Zhuolin Yi , Yanzhen Ren , Kai Li

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the possibility of…

音频与语音处理 · 电气工程与系统科学 2025-11-18 Mark R. Saddler , Andrew Francl , Jenelle Feather , Kaizhi Qian , Yang Zhang , Josh H. McDermott

Deep learning technology has been widely applied to speech enhancement. While testing the effectiveness of various network structures, researchers are also exploring the improvement of the loss function used in network training. Although…

音频与语音处理 · 电气工程与系统科学 2023-04-25 Tianrui Wang , Weibin Zhu

Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment…

声音 · 计算机科学 2023-01-27 Jianwei Zhang , Julie Liss , Suren Jayasuriya , Visar Berisha

Eliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition that stills remains an important challenge. Data-driven supervised approaches, including ones based on…

Data augmentation is conventionally used to inject robustness in Speaker Verification systems. Several recently organized challenges focus on handling novel acoustic environments. Deep learning based speech enhancement is a modern solution…

音频与语音处理 · 电气工程与系统科学 2020-04-29 Saurabh Kataria , Phani Sankar Nidadavolu , Jesús Villalba , Najim Dehak

Listening in noisy environments can be difficult even for individuals with a normal hearing thresholds. The speech signal can be masked by noise, which may lead to word misperceptions on the side of the listener, and overall difficulty to…

计算与语言 · 计算机科学 2021-07-20 Anupama Chingacham , Vera Demberg , Dietrich Klakow

Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of…

音频与语音处理 · 电气工程与系统科学 2024-10-31 Nils L. Westhausen , Hendrik Kayser , Theresa Jansen , Bernd T. Meyer
‹ 上一页 1 2 3 10 下一页 ›