English
Related papers

Related papers: TalkVid: A Large-Scale Diversified Dataset for Aud…

200 papers

Recent efforts in Spoken Dialogue Modeling aim to synthesize spoken dialogue without the need for direct transcription, thereby preserving the wealth of non-textual information inherent in speech. However, this approach faces a challenge…

Computation and Language · Computer Science 2024-07-03 Yu-Kuan Fu , Cheng-Kuang Lee , Hsiu-Hsuan Wang , Hung-yi Lee

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

Recent advancements in deep learning and computer vision have led to a surge of interest in generating realistic talking heads. This paper presents a comprehensive survey of state-of-the-art methods for talking head generation. We…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Shreyank N Gowda , Dheeraj Pandey , Shashank Narayana Gowda

Lecture slide presentations, a sequence of pages that contain text and figures accompanied by speech, are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and…

Artificial Intelligence · Computer Science 2022-08-18 Dong Won Lee , Chaitanya Ahuja , Paul Pu Liang , Sanika Natu , Louis-Philippe Morency

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world.…

Sound · Computer Science 2026-02-04 Chengyuan Ma , Jiawei Jin , Ruijie Xiong , Chunxiang Jin , Canxiang Yan , Wenming Yang

Talk2AI is a large-scale longitudinal dataset of 3,080 conversations (totaling 30,800 turns) between human participants and Large Language Models (LLMs), designed to support research on persuasion, opinion change, and human-AI interaction.…

Human-Computer Interaction · Computer Science 2026-04-07 Alexis Carrillo , Enrique Taietta , Ali Aghazadeh Ardebili , Giuseppe Alessandro Veltri , Massimo Stella

Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Yifeng Ma , Suzhen Wang , Yu Ding , Bowen Ma , Tangjie Lv , Changjie Fan , Zhipeng Hu , Zhidong Deng , Xin Yu

We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN…

Computer Vision and Pattern Recognition · Computer Science 2022-09-12 Mohammed M. Alghamdi , He Wang , Andrew J. Bulpitt , David C. Hogg

Spurred by recent advances in Large Language Models (LLMs), virtual assistants are poised to take a leap forward in terms of their dialogue capabilities. Yet a major bottleneck to achieving genuinely transformative task-oriented dialogue…

Computation and Language · Computer Science 2024-05-06 Joe Stacey , Jianpeng Cheng , John Torr , Tristan Guigue , Joris Driesen , Alexandru Coca , Mark Gaynor , Anders Johannsen

The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present…

Sound · Computer Science 2025-05-15 Yicheng Gu , Chaoren Wang , Junan Zhang , Xueyao Zhang , Zihao Fang , Haorui He , Zhizheng Wu

Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-03 Kratika Bhagtani , Amit Kumar Singh Yadav , Paolo Bestagini , Edward J. Delp

This paper introduces SpoofCeleb, a dataset designed for Speech Deepfake Detection (SDD) and Spoofing-robust Automatic Speaker Verification (SASV), utilizing source data from real-world conditions and spoofing attacks generated by…

Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviate this issue, with…

Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-20 Matthew Le , Apoorv Vyas , Bowen Shi , Brian Karrer , Leda Sari , Rashel Moritz , Mary Williamson , Vimal Manohar , Yossi Adi , Jay Mahadeokar , Wei-Ning Hsu

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Mattia Soldan , Alejandro Pardo , Juan León Alcázar , Fabian Caba Heilbron , Chen Zhao , Silvio Giancola , Bernard Ghanem

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural…

‹ Prev 1 3 4 5 6 7 10 Next ›