中文
相关论文

相关论文: Multi-view Temporal Alignment for Non-parallel Art…

200 篇论文

Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural…

音频与语音处理 · 电气工程与系统科学 2019-05-17 Ahmed Hussen Abdelaziz , Barry-John Theobald , Justin Binder , Gabriele Fanelli , Paul Dixon , Nicholas Apostoloff , Thibaut Weise , Sachin Kajareker

End-to-end (E2E) automatic speech recognition (ASR) systems directly map acoustics to words using a unified model. Previous works mostly focus on E2E training a single model which integrates acoustic and language model into a whole.…

计算与语言 · 计算机科学 2018-03-06 Zhehuai Chen , Qi Liu , Hao Li , Kai Yu

We propose a novel training algorithm for a multi-speaker neural text-to-speech (TTS) model based on multi-task adversarial training. A conventional generative adversarial network (GAN)-based training algorithm significantly improves the…

声音 · 计算机科学 2022-09-27 Yusuke Nakai , Yuki Saito , Kenta Udagawa , Hiroshi Saruwatari

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal…

声音 · 计算机科学 2024-09-16 Zhiqi Huang , Dan Luo , Jun Wang , Huan Liao , Zhiheng Li , Zhiyong Wu

This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike…

计算与语言 · 计算机科学 2020-07-28 Srikanth Ronanki , Oliver Watts , Simon King , Gustav Eje Henter

We propose to model parallel streams of data, such as overlapped speech, using shuffles. Specifically, this paper shows how the shuffle product and partial order finite-state automata (FSAs) can be used for alignment and speaker-attributed…

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its capacity to modulate the distribution of silence and overlap…

音频与语音处理 · 电气工程与系统科学 2023-10-20 Tae Jin Park , He Huang , Coleman Hooper , Nithin Koluguri , Kunal Dhawan , Ante Jukic , Jagadeesh Balam , Boris Ginsburg

We present an end-to-end text-to-speech (TTS) synthesis system that generates audio and synchronized tongue motion directly from text. This is achieved by adapting a 3D model of the tongue surface to an articulatory dataset and training a…

人机交互 · 计算机科学 2018-04-17 Ingmar Steiner , Sébastien Le Maguer , Alexander Hewer

Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source methods mainly rely on either dual-tower designs with posterior alignment or fully…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Longbin Ji , Guan Wang , Xuan Wei , Chenye Yang , Xiangrui Liu , Zhenyu Zhang , Shuohuan Wang , Yu Sun , Jingzhou He

This paper introduces Parallel Tacotron 2, a non-autoregressive neural text-to-speech model with a fully differentiable duration model which does not require supervised duration signals. The duration model is based on a novel attention…

声音 · 计算机科学 2021-08-31 Isaac Elias , Heiga Zen , Jonathan Shen , Yu Zhang , Ye Jia , RJ Skerry-Ryan , Yonghui Wu

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Davide Berghi , Philip J. B. Jackson

We present a learning-based approach for generating binaural audio from mono audio using multi-task learning. Our formulation leverages additional information from two related tasks: the binaural audio generation task and the flipped audio…

声音 · 计算机科学 2021-09-03 Sijia Li , Shiguang Liu , Dinesh Manocha

The aim of this research is to refine knowledge transfer on audio-image temporal agreement for audio-text cross retrieval. To address the limited availability of paired non-speech audio-text data, learning methods for transferring the…

音频与语音处理 · 电气工程与系统科学 2024-03-19 Shunsuke Tsubaki , Daisuke Niizumi , Daiki Takeuchi , Yasunori Ohishi , Noboru Harada , Keisuke Imoto

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding…

声音 · 计算机科学 2021-10-12 Cheng Gong , Longbiao Wang , Zhenhua Ling , Ju Zhang , Jianwu Dang

We study the problem of syncing the lip movement in a video with the audio stream. Our solution finds an optimal alignment using a dual-domain recurrent neural network that is trained on synthetic data we generate by dropping and…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Yoav Shalev , Lior Wolf

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

计算机视觉与模式识别 · 计算机科学 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Understanding how humans express and synchronize emotions across multiple communication channels particularly facial expressions and speech has significant implications for emotion recognition systems and human computer interaction.…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Von Ralph Dane Marquez Herbuela , Yukie Nagai

We propose a step-by-step video-to-audio (V2A) generation method for finer controllability over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach aims to comprehensively capture…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefined conditions,…

声音 · 计算机科学 2020-12-01 Peng Zhang , Jiaming Xu , Jing shi , Yunzhe Hao , Bo Xu