中文
相关论文

相关论文: Sequence-to-Sequence Multi-Modal Speech In-Paintin…

200 篇论文

The process of reconstructing missing parts of speech audio from context is called speech in-painting. Human perception of speech is inherently multi-modal, involving both audio and visual (AV) cues. In this paper, we introduce and study a…

多媒体 · 计算机科学 2024-06-04 Mahsa Kadkhodaei Elyaderani , Shahram Shirani

In this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliable audio context and uncorrupted visual information. Recent…

音频与语音处理 · 电气工程与系统科学 2021-02-04 Giovanni Morrone , Daniel Michelsanti , Zheng-Hua Tan , Jesper Jensen

Speech inpainting consists in reconstructing corrupted or missing speech segments using surrounding context, a process that closely resembles the pretext tasks in Self-Supervised Learning (SSL) for speech encoders. This study investigates…

声音 · 计算机科学 2025-12-09 Ihab Asaad , Maxime Jacquelin , Olivier Perrotin , Laurent Girin , Thomas Hueber

Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, \ie synthesizing missing audio segments that correspond to their accompanying…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Hang Zhou , Ziwei Liu , Xudong Xu , Ping Luo , Xiaogang Wang

Long (> 200 ms) audio inpainting, to recover a long missing part in an audio segment, could be widely applied to audio editing tasks and transmission loss recovery. It is a very challenging problem due to the high dimensional, complex and…

声音 · 计算机科学 2019-11-18 Ya-Liang Chang , Kuan-Ying Lee , Po-Yu Wu , Hung-yi Lee , Winston Hsu

Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal…

声音 · 计算机科学 2023-06-02 Juan F. Montesinos , Daniel Michelsanti , Gloria Haro , Zheng-Hua Tan , Jesper Jensen

Video-to-speech synthesis is the task of reconstructing the speech signal from a silent video of a speaker. Most established approaches to date involve a two-step process, whereby an intermediate representation from the video, such as a…

声音 · 计算机科学 2024-10-28 Triantafyllos Kefalas , Yannis Panagakis , Maja Pantic

This paper presents a method of sequence-to-sequence (seq2seq) voice conversion using non-parallel training data. In this method, disentangled linguistic and speaker representations are extracted from acoustic features, and voice conversion…

音频与语音处理 · 电气工程与系统科学 2020-01-14 Jing-Xuan Zhang , Zhen-Hua Ling , Li-Rong Dai

Recent research in AI is focusing towards generating narrative stories about visual scenes. It has the potential to achieve more human-like understanding than just basic description generation of images- in-sequence. In this work, we…

人工智能 · 计算机科学 2018-09-25 Marko Smilevski , Ilija Lalkovski , Gjorgji Madjarov

We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input. We demonstrate that the model performs speech inpainting with the appropriate content, while maintaining…

声音 · 计算机科学 2022-03-31 Zalán Borsos , Matt Sharifi , Marco Tagliasacchi

Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introduce the first approach that applies discrete diffusion over…

声音 · 计算机科学 2026-02-18 Tali Dror , Iftach Shoham , Moshe Buchris , Oren Gal , Haim Permuter , Gilad Katz , Eliya Nachmani

Learning a new language involves constantly comparing speech productions with reference productions from the environment. Early in speech acquisition, children make articulatory adjustments to match their caregivers' speech. Grownup…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Talia Ben-Simon , Felix Kreuk , Faten Awwad , Jacob T. Cohen , Joseph Keshet

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and acoustic scene…

音频与语音处理 · 电气工程与系统科学 2020-07-14 Abhinav Shukla , Stavros Petridis , Maja Pantic

Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture. Promising techniques have been proposed for this task, but earlier works have…

声音 · 计算机科学 2018-05-16 Hiroshi Seki , Takaaki Hori , Shinji Watanabe , Jonathan Le Roux , John R. Hershey

Audio inpainting aims to reconstruct missing segments in corrupted recordings. Most of existing methods produce plausible reconstructions when the gap lengths are short, but struggle to reconstruct gaps larger than about 100 ms. This paper…

音频与语音处理 · 电气工程与系统科学 2025-01-13 Eloi Moliner , Vesa Välimäki

Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video inpainting. To…

音频与语音处理 · 电气工程与系统科学 2023-10-12 Kyuyeon Kim , Junsik Jung , Woo Jae Kim , Sung-Eui Yoon

A sequence-to-sequence model is a neural network module for mapping two sequences of different lengths. The sequence-to-sequence model has three core modules: encoder, decoder, and attention. Attention is the bridge that connects the…

计算与语言 · 计算机科学 2018-07-24 Andros Tjandra , Sakriani Sakti , Satoshi Nakamura

Encoder-decoder models have become an effective approach for sequence learning tasks like machine translation, image captioning and speech recognition, but have yet to show competitive results for handwritten text recognition. To this end,…

计算机视觉与模式识别 · 计算机科学 2019-07-16 Johannes Michael , Roger Labahn , Tobias Grüning , Jochen Zöllner

Speech 'in-the-wild' is a handicap for speaker recognition systems due to the variability induced by real-life conditions, such as environmental noise and the emotional state of the speaker. Taking advantage of the principles of…

音频与语音处理 · 电气工程与系统科学 2022-05-17 Esther Rituerto-González , Carmen Peláez-Moreno

Image inpainting approaches have achieved significant progress with the help of deep neural networks. However, existing approaches mainly focus on leveraging the priori distribution learned by neural networks to produce a single inpainting…

计算机视觉与模式识别 · 计算机科学 2022-01-27 Wangbo Yu , Jinhao Du , Ruixin Liu , Yixuan Li , Yuesheng zhu
‹ 上一页 1 2 3 10 下一页 ›