中文
相关论文

相关论文: Audio-Driven Dubbing for User Generated Contents v…

200 篇论文

Automatic dubbing (AD) is among the machine translation (MT) use cases where translations should match a given length to allow for synchronicity between source and target speech. For neural MT, generating translations of length close to the…

计算与语言 · 计算机科学 2022-02-18 Surafel M. Lakew , Yogesh Virkar , Prashant Mathur , Marcello Federico

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interlocutor which is known as entrainment. To produce human-like…

音频与语音处理 · 电气工程与系统科学 2021-06-22 Jian Cong , Shan Yang , Na Hu , Guangzhi Li , Lei Xie , Dan Su

The development of Audio Description (AD) has been a pivotal step forward in making video content more accessible and inclusive. Traditionally, AD production has demanded a considerable amount of skilled labor, while existing automated…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Peng Chu , Jiang Wang , Andre Abrantes

Cross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language. This is a challenging task, and resulting speech can suffer from reduced naturalness, accented speech, and/or loss…

声音 · 计算机科学 2022-04-04 Marcel de Korte , Jaebok Kim , Aki Kunikoshi , Adaeze Adigwe , Esther Klabbers

End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with significantly lower latency. However, these systems remain…

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Hui Fu , Zeqing Wang , Ke Gong , Keze Wang , Tianshui Chen , Haojie Li , Haifeng Zeng , Wenxiong Kang

In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attention due to its wide…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Jeongsoo Choi , Ji-Hoon Kim , Kim Sung-Bin , Tae-Hyun Oh , Joon Son Chung

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With video feature, text, character bank and context information as…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Hanlin Wang , Zhan Tong , Kecheng Zheng , Yujun Shen , Limin Wang

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Xi Liu , Ying Guo , Cheng Zhen , Tong Li , Yingying Ao , Pengfei Yan

In response to the rising prominence of the Metaverse, omnidirectional videos (ODVs) have garnered notable interest, gradually shifting from professional-generated content (PGC) to user-generated content (UGC). However, the study of…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Fei Zhao , Da Pan , Zelu Qi , Ping Shi

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos,…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Lin Zhang , Zefan Cai , Yufan Zhou , Shentong Mo , Jinhong Lin , Cheng-En Wu , Yibing Wei , Yijing Zhang , Ruiyi Zhang , Wen Xiao , Tong Sun , Junjie Hu , Pedro Morgado

In real-world applications, the sample distribution at the inference stage often differs from the one at the training stage, causing performance degradation of trained deep models. The research on domain generalization (DG) aims to develop…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Jiao Zhang , Jian Xu , Xu-Yao Zhang , Cheng-Lin Liu

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing…

Speech enhancement is crucial for ubiquitous human-computer interaction. Recently, ultrasound-based acoustic sensing has emerged as an attractive choice for speech enhancement because of its superior ubiquity and performance. However, due…

声音 · 计算机科学 2025-05-20 Luca Jiang-Tao Yu , Running Zhao , Sijie Ji , Edith C. H. Ngai , Chenshu Wu

Can a generative model be trained to produce images from a specific domain, guided by a text prompt only, without seeing any image? In other words: can an image generator be trained "blindly"? Leveraging the semantic power of large scale…

计算机视觉与模式识别 · 计算机科学 2021-12-17 Rinon Gal , Or Patashnik , Haggai Maron , Gal Chechik , Daniel Cohen-Or

Goal-oriented dialog systems enable users to complete specific goals like requesting information about a movie or booking a ticket. Typically the dialog system pipeline contains multiple ML models, including natural language understanding,…

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Yazhou Xing , Yingqing He , Zeyue Tian , Xintao Wang , Qifeng Chen

Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). To solve the task, we propose a universal…

计算与语言 · 计算机科学 2020-02-04 Zekang Li , Zongjia Li , Jinchao Zhang , Yang Feng , Cheng Niu , Jie Zhou

Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial…