中文
相关论文

相关论文: A Speech-to-Video Synthesis Approach Using Spatio-…

200 篇论文

Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Rishi Jain , Bohan Yu , Peter Wu , Tejas Prabhune , Gopala Anumanchipalli

Text to video generation has emerged as a critical frontier in generative artificial intelligence, yet existing approaches struggle with maintaining temporal consistency, compositional understanding, and fine grained control over visual…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Piyushkumar Patel

Understanding the underlying relationship between tongue and oropharyngeal muscle deformation seen in tagged-MRI and intelligible speech plays an important role in advancing speech motor control theories and treatment of speech…

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

计算与语言 · 计算机科学 2025-08-19 Shumin Que , Anton Ragni

Recent advances in synthetic imaging open up opportunities for obtaining additional data in the field of surgical imaging. This data can provide reliable supplements supporting surgical applications and decision-making through computer…

图像与视频处理 · 电气工程与系统科学 2023-12-07 Simeon Allmendinger , Patrick Hemmer , Moritz Queisner , Igor Sauer , Leopold Müller , Johannes Jakubik , Michael Vössing , Niklas Kühl

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Diffusion-based video generation models have made significant strides, producing outputs with improved visual fidelity, temporal coherence, and user control. These advancements hold great promise for improving surgical education by enabling…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Joseph Cho , Samuel Schmidgall , Cyril Zakka , Mrudang Mathur , Dhamanpreet Kaur , Rohan Shad , William Hiesinger

Text-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdles in (a) accurately…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Hyeonho Jeong , Geon Yeong Park , Jong Chul Ye

Recent advances in text-to-video generation have demonstrated the utility of powerful diffusion models. Nevertheless, the problem is not trivial when shaping diffusion models to animate static image (i.e., image-to-video generation). The…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Zhongwei Zhang , Fuchen Long , Yingwei Pan , Zhaofan Qiu , Ting Yao , Yang Cao , Tao Mei

Recent advances in the diffusion models have significantly improved text-to-image generation. However, generating videos from text is a more challenging task than generating images from text, due to the much larger dataset and higher…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Taegyeong Lee , Soyeong Kwon , Taehwan Kim

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Ruobing Zheng , Zhou Zhu , Bo Song , Changjiang Ji

Generating a coherent sequence of images that tells a visual story, using text-to-image diffusion models, often faces the critical challenge of maintaining subject consistency across all story scenes. Existing approaches, which typically…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Gopalji Gaur , Mohammadreza Zolfaghari , Thomas Brox

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

Streaming speech-to-avatar synthesis creates real-time animations for a virtual character from audio data. Accurate avatar representations of speech are important for the visualization of sound in linguistics, phonetics, and phonology,…

声音 · 计算机科学 2023-10-26 Tejas S. Prabhune , Peter Wu , Bohan Yu , Gopala K. Anumanchipalli

Deep generative models have demonstrated the ability to create realistic audiovisual content, sometimes driven by domains of different nature. However, smooth temporal dynamics in video generation is a challenging problem. This work focuses…

声音 · 计算机科学 2024-06-25 Rafael Redondo

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Shuai Yang , Yifan Zhou , Ziwei Liu , Chen Change Loy

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee