中文
相关论文

相关论文: ShoulderShot: Generating Over-the-Shoulder Dialogu…

200 篇论文

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate…

Although large language models(LLMs) show amazing capabilities, among various exciting applications discovered for LLMs fall short in other low-resource languages. Besides, most existing methods depend on large-scale dialogue corpora and…

计算与语言 · 计算机科学 2024-08-19 Yongkang Liu , Feng Shi , Daling Wang , Yifei Zhang , Hinrich Schütze

Current video generation models excel at short clips but fail to produce cohesive multi-shot narratives due to disjointed visual dynamics and fractured storylines. Existing solutions either rely on extensive manual scripting/editing or…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Mingzhe Zheng , Yongqi Xu , Haojian Huang , Xuran Ma , Yexin Liu , Wenjie Shu , Yatian Pang , Feilong Tang , Qifeng Chen , Harry Yang , Ser-Nam Lim

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Hong Chen , Xin Wang , Yipeng Zhang , Yuwei Zhou , Zeyang Zhang , Siao Tang , Wenwu Zhu

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

The one-shot talking-head generation learns to synthesize a talking-head video with one source portrait image under the driving of same or different identity video. Usually these methods require plane-based pixel transformations via Jacobin…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Luchuan Song , Pinxin Liu , Guojun Yin , Chenliang Xu

We consider the challenging problem of audio to animated video generation. We propose a novel method OneShotAu2AV to generate an animated video of arbitrary length using an audio clip and a single unseen image of a person as an input. The…

计算机视觉与模式识别 · 计算机科学 2021-02-22 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall , Mujtaba Hasan , Pranshu Agarwal , Dipankar Sarkar

Shot transitions play a pivotal role in multi-shot video generation, as they determine the overall narrative expression and the directorial design of visual storytelling. However, recent progress has primarily focused on low-level visual…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Xiaoxue Wu , Xinyuan Chen , Yaohui Wang , Yu Qiao

In order to better simulate the real human conversation process, models need to generate dialogue utterances based on not only preceding textual contexts but also visual contexts. However, with the development of multi-modal dialogue…

计算与语言 · 计算机科学 2021-09-29 Shuhe Wang , Yuxian Meng , Xiaoya Li , Xiaofei Sun , Rongbin Ouyang , Jiwei Li

Recent years have witnessed remarkable advances in audio-driven talking head generation. However, existing approaches predominantly focus on single-character scenarios. While some methods can create separate conversation videos between two…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Yubo Huang , Weiqiang Wang , Sirui Zhao , Tong Xu , Lin Liu , Enhong Chen

Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds…

Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven methods, we focus on…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Jiahui Chen , Yang Huan , Runhua Shi , Chanfan Ding , Xiaoqi Mo , Siyu Xiong , Yinong He

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Faraz Waseem , Muhammad Shahzad

Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue context. Due to the lack of a large-scale dataset specifically…

人工智能 · 计算机科学 2024-08-13 Hee Suk Yoon , Eunseop Yoon , Joshua Tian Jin Tee , Kang Zhang , Yu-Jung Heo , Du-Seong Chang , Chang D. Yoo

Gestures that accompany speech are an essential part of natural and efficient embodied human communication. The automatic generation of such co-speech gestures is a long-standing problem in computer animation and is considered an enabling…

图形学 · 计算机科学 2023-04-11 Simbarashe Nyatsanga , Taras Kucherenko , Chaitanya Ahuja , Gustav Eje Henter , Michael Neff

The advent of stereoscopic videos has opened new horizons in multimedia, particularly in extended reality (XR) and virtual reality (VR) applications, where immersive content captivates audiences across various platforms. Despite its growing…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Qiao Jin , Xiaodong Chen , Wu Liu , Tao Mei , Yongdong Zhang

Recent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Canyu Zhao , Mingyu Liu , Wen Wang , Weihua Chen , Fan Wang , Hao Chen , Bo Zhang , Chunhua Shen

Recent advancements in conversational systems have significantly enhanced human-machine interactions across various domains. However, training these systems is challenging due to the scarcity of specialized dialogue data. Traditionally,…

计算与语言 · 计算机科学 2026-05-29 Heydar Soudani , Roxana Petcu , Evangelos Kanoulas , Faegheh Hasibi

Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the generated videos. This work presents Moonshot, a new video…

计算机视觉与模式识别 · 计算机科学 2024-01-04 David Junhao Zhang , Dongxu Li , Hung Le , Mike Zheng Shou , Caiming Xiong , Doyen Sahoo