中文
相关论文

相关论文: DurIAN: Duration Informed Attention Network For Mu…

200 篇论文

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-syncing, being…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Georgios Milis , Panagiotis P. Filntisis , Anastasios Roussos , Petros Maragos

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast, generative Speech…

音频与语音处理 · 电气工程与系统科学 2025-10-22 Bunlong Lay , Rostislav Makarov , Simon Welker , Maris Hillemann , Timo Gerkmann

Recently, GAN based speech synthesis methods, such as MelGAN, have become very popular. Compared to conventional autoregressive based methods, parallel structures based generators make waveform generation process fast and stable. However,…

音频与语音处理 · 电气工程与系统科学 2020-11-25 Qiao Tian , Yi Chen , Zewang Zhang , Heng Lu , Linghui Chen , Lei Xie , Shan Liu

Ensuring intelligible speech communication for hearing assistive devices in low-latency scenarios presents significant challenges in terms of speech enhancement, coding and transmission. In this paper, we propose novel solutions for…

音频与语音处理 · 电气工程与系统科学 2024-05-01 Mohammad Bokaei , Jesper Jensen , Simon Doclo , Jan Østergaard

Attention based neural TTS is elegant speech synthesis pipeline and has shown a powerful ability to generate natural speech. However, it is still not robust enough to meet the stability requirements for industrial products. Besides, it…

音频与语音处理 · 电气工程与系统科学 2020-11-03 Qiao Tian , Zewang Zhang , Chao Liu , Heng Lu , Linghui Chen , Bin Wei , Pujiang He , Shan Liu

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved…

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local…

声音 · 计算机科学 2022-02-16 Tianchi Liu , Rohan Kumar Das , Kong Aik Lee , Haizhou Li

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While…

声音 · 计算机科学 2026-02-13 Chung-Soo Ahn , Rajib Rana , Sunil Sivadas , Carlos Busso , Jagath C. Rajapakse

Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to relying on previous statistical parametric speech synthesis…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Ammar Abbas , Thomas Merritt , Alexis Moinet , Sri Karlapati , Ewa Muszynska , Simon Slangen , Elia Gatti , Thomas Drugman

This paper describes a practical dual-process speech enhancement system that adapts environment-sensitive frame-online beamforming (front-end) with help from environment-free block-online source separation (back-end). To use minimum…

音频与语音处理 · 电气工程与系统科学 2022-07-25 Aditya Arie Nugraha , Kouhei Sekiguchi , Mathieu Fontaine , Yoshiaki Bando , Kazuyoshi Yoshii

Deep dilated temporal convolutional networks (TCN) have been proved to be very effective in sequence modeling. In this paper we propose several improvements of TCN for end-to-end approach to monaural speech separation, which consists of 1)…

声音 · 计算机科学 2023-06-27 Liwen Zhang , Ziqiang Shi , Jiqing Han , Anyan Shi , Ding Ma

Talking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diffusion-based talking head generation, almost all methods rely…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Hanbo Cheng , Limin Lin , Chenyu Liu , Pengcheng Xia , Pengfei Hu , Jiefeng Ma , Jun Du , Jia Pan

In this paper, we propose a multi-channel network for simultaneous speech dereverberation, enhancement and separation (DESNet). To enable gradient propagation and joint optimization, we adopt the attentional selection mechanism of the…

声音 · 计算机科学 2020-11-17 Yihui Fu , Jian Wu , Yanxin Hu , Mengtao Xing , Lei Xie

The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on…

At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current…

音频与语音处理 · 电气工程与系统科学 2024-11-19 Sheng Yan , Cunhang fan , Hongyu Zhang , Xiaoke Yang , Jianhua Tao , Zhao Lv

Different from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the…

计算与语言 · 计算机科学 2019-10-25 Zheng Lian , Jianhua Tao , Bin Liu , Jian Huang

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature…

多媒体 · 计算机科学 2023-03-07 Zhongweiyang Xu , Xulin Fan , Mark Hasegawa-Johnson

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

计算与语言 · 计算机科学 2023-05-15 Fei Tao , Carlos Busso