中文
相关论文

相关论文: TalkVid: A Large-Scale Diversified Dataset for Aud…

200 篇论文

The advancement of AI systems for mental health support is hindered by limited access to therapeutic conversation data, particularly for trauma treatment. We present Thousand Voices of Trauma, a synthetic benchmark dataset of 3,000 therapy…

计算机与社会 · 计算机科学 2025-05-19 Suhas BN , Andrew M. Sherrill , Rosa I. Arriaga , Chris W. Wiese , Saeed Abdullah

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Xinhan Di , Kristin Qi , Pengqian Yu

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Weixia Zhang , Chengguang Zhu , Jingnan Gao , Yichao Yan , Guangtao Zhai , Xiaokang Yang

Talking Head Generation (THG) has emerged as a transformative technology in computer vision, enabling the synthesis of realistic human faces synchronized with image, audio, text, or video inputs. This paper provides a comprehensive review…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Vineet Kumar Rakesh , Soumya Mazumdar , Research Pratim Maity , Sarbajit Pal , Amitabha Das , Tapas Samanta

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Wentao Hu , Shunkai Li , Ziqiao Peng , Haoxian Zhang , Fan Shi , Xiaoqiang Liu , Pengfei Wan , Di Zhang , Hui Tian

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zeyu Zhu , Weijia Wu , Mike Zheng Shou

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Yiming Ju , Jijin Hu , Zhengxiong Luo , Haoge Deng , hanyu Zhao , Li Du , Chengwei Wu , Donglin Hao , Xinlong Wang , Tengfei Pan

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

This paper presents the Multi-Language Audio Anti-Spoofing Dataset (MLAAD), version 10: a dataset of synthetic audio to train and evaluate audio deepfake detection models. It features 175 Text-to-Speech (TTS) models, comprising a total of…

In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead…

图形学 · 计算机科学 2022-12-09 Zhentao Yu , Zixin Yin , Deyu Zhou , Duomin Wang , Finn Wong , Baoyuan Wang

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However,…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Xinqi Xiong , Prakrut Patel , Qingyuan Fan , Amisha Wadhwa , Sarathy Selvam , Xiao Guo , Luchao Qi , Xiaoming Liu , Roni Sengupta

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that…

图形学 · 计算机科学 2026-01-28 Radek Daněček , Carolin Schmitt , Senya Polikovsky , Michael J. Black

High-quality Text-to-Speech (TTS) model training requires extensive and diverse text and speech data. It is challenging to procure such data from real sources due to issues of domain specificity, licensing, and scalability. Large language…

计算与语言 · 计算机科学 2025-10-03 Karan Dua , Puneet Mittal , Ranjeet Gupta , Hitesh Laxmichand Patel

Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchmarks remain largely designed for human-recorded videos or…

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Deng , Youxin Pang , Xiaochen Zhao , Chao Xu , Lizhen Wang , Hongjiang Xiao , Shi Yan , Hongwen Zhang , Yebin Liu

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

音频与语音处理 · 电气工程与系统科学 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited controllability, diversity in deepfake methods, and licensing…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Yongyi Zang , Jiatong Shi , You Zhang , Ryuichi Yamamoto , Jionghao Han , Yuxun Tang , Shengyuan Xu , Wenxiao Zhao , Jing Guo , Tomoki Toda , Zhiyao Duan

Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Sijing Wu , Yunhao Li , Weitian Zhang , Jun Jia , Yucheng Zhu , Yichao Yan , Guangtao Zhai , Xiaokang Yang