中文
相关论文

相关论文: DiffV2S: Diffusion-based Video-to-Speech Synthesis…

200 篇论文

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker…

声音 · 计算机科学 2025-03-10 Yifan Liu , Yu Fang , Zhouhan Lin

The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Joanna Hong , Minsu Kim , Yong Man Ro

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining…

音频与语音处理 · 电气工程与系统科学 2023-10-31 Suyeon Lee , Chaeyoung Jung , Youngjoon Jang , Jaehun Kim , Joon Son Chung

Building a high-quality singing corpus for a person who is not good at singing is non-trivial, thus making it challenging to create a singing voice synthesizer for this person. Learn2Sing is dedicated to synthesizing the singing voice of a…

声音 · 计算机科学 2022-05-27 Heyang Xue , Xinsheng Wang , Yongmao Zhang , Lei Xie , Pengcheng Zhu , Mengxiao Bi

Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Runwu Shi , Kai Li , Chang Li , Jiang Wang , Sihan Tan , Kazuhiro Nakadai

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are collected in…

音频与语音处理 · 电气工程与系统科学 2024-11-05 Ju-Chieh Chou , Chung-Ming Chien , Karen Livescu

The field of Singing Voice Synthesis (SVS) has seen significant advancements in recent years due to the rapid progress of diffusion-based approaches. However, capturing vocal style, genre-specific pitch inflections, and language-dependent…

声音 · 计算机科学 2025-12-01 Sandipan Dhar , Mayank Gupta , Preeti Rao

Video-to-speech synthesis involves reconstructing the speech signal of a speaker from a silent video. The implicit assumption of this task is that the sound signal is either missing or contains a high amount of noise/corruption such that it…

声音 · 计算机科学 2024-10-28 Triantafyllos Kefalas , Yannis Panagakis , Maja Pantic

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple…

音频与语音处理 · 电气工程与系统科学 2021-05-21 Dan Oneata , Adriana Stan , Horia Cucu

Generating high-quality and person-generic visual dubbing remains a challenge. Recent innovation has seen the advent of a two-stage paradigm, decoupling the rendering and lip synchronization process facilitated by intermediate…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Tao Liu , Chenpeng Du , Shuai Fan , Feilong Chen , Kai Yu

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous singing acoustic…

音频与语音处理 · 电气工程与系统科学 2022-03-23 Jinglin Liu , Chengxi Li , Yi Ren , Feiyang Chen , Zhou Zhao

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped…

音频与语音处理 · 电气工程与系统科学 2025-01-03 Shota Horiguchi , Takafumi Moriya , Atsushi Ando , Takanori Ashihara , Hiroshi Sato , Naohiro Tawara , Marc Delcroix

In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) framework designed to generate natural and intelligible speech directly from silent talking face videos. While recent V2S systems have shown promising results on constrained…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Jeongsoo Choi , Ji-Hoon Kim , Jinyu Li , Joon Son Chung , Shujie Liu

Enabling Visual Semantic Models to effectively handle multi-view description matching has been a longstanding challenge. Existing methods typically learn a set of embeddings to find the optimal match for each view's text and compute…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Yang Liu , Wentao Feng , Zhuoyao Liu , Shudong Huang , Jiancheng Lv

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embeddings obtained from…

音频与语音处理 · 电气工程与系统科学 2023-06-05 Julius Richter , Simone Frintrop , Timo Gerkmann

Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an autoregressive teacher…

音频与语音处理 · 电气工程与系统科学 2022-08-09 Yi Ren , Chenxu Hu , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , Tie-Yan Liu

Speech-driven 3D facial animation synthesis has been a challenging task both in industry and research. Recent methods mostly focus on deterministic deep learning methods meaning that given a speech input, the output is always the same.…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Stefan Stan , Kazi Injamamul Haque , Zerrin Yumak

We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently…

Significant progress has been made in speaker dependent Lip-to-Speech synthesis, which aims to generate speech from silent videos of talking faces. Current state-of-the-art approaches primarily employ non-autoregressive sequence-to-sequence…

声音 · 计算机科学 2023-07-06 Neha Sahipjohn , Neil Shah , Vishal Tambrahalli , Vineet Gandhi

Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods are slow due to numerous denoising steps and costly…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Beijia Lu , Ziyi Chen , Jing Xiao , Jun-Yan Zhu
‹ 上一页 1 2 3 10 下一页 ›