English
Related papers

Related papers: AutoAD III: The Prequel -- Back to the Pixels

200 papers

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully,…

Computer Vision and Pattern Recognition · Computer Science 2019-05-10 Huda Alamri , Vincent Cartillier , Abhishek Das , Jue Wang , Anoop Cherian , Irfan Essa , Dhruv Batra , Tim K. Marks , Chiori Hori , Peter Anderson , Stefan Lee , Devi Parikh

Text-to-audio generation models (TAG) have achieved significant advances in generating audio conditioned on text descriptions. However, a critical challenge lies in the lack of transparency regarding how each textual input impacts the…

Sound · Computer Science 2025-10-20 Hyunju Kang , Geonhee Han , Yoonjae Jeong , Hogun Park

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing…

Multimedia · Computer Science 2025-07-15 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Guangtao Zhai

The ability to map descriptions of scenes to 3D geometric representations has many applications in areas such as art, education, and robotics. However, prior work on the text to 3D scene generation task has used manually specified object…

Computation and Language · Computer Science 2015-06-08 Angel Chang , Will Monroe , Manolis Savva , Christopher Potts , Christopher D. Manning

Recent video action recognition methods have shown excellent performance by adapting large-scale pre-trained language-image models to the video domain. However, language models contain rich common sense priors - the scene contexts that…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Xiaodan Hu , Chuhang Zou , Suchen Wang , Jaechul Kim , Narendra Ahuja

We introduce Audio-Agent, a multimodal framework for audio generation, editing and composition based on text or video inputs. Conventional approaches for text-to-audio (TTA) tasks often make single-pass inferences from text descriptions.…

Sound · Computer Science 2025-01-15 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

3D content creation plays a vital role in various applications, such as gaming, robotics simulation, and virtual reality. However, the process is labor-intensive and time-consuming, requiring skilled designers to invest considerable effort…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Chenhan Jiang

Achieving expressive 3D motion reconstruction and automatic generation for isolated sign words can be challenging, due to the lack of real-world 3D sign-word data, the complex nuances of signing motions, and the cross-modal understanding of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Lu Dong , Lipisha Chaudhary , Fei Xu , Xiao Wang , Mason Lary , Ifeoma Nwogu

Currently, digital avatars can be created manually using human images as reference. Systems such as Bitmoji are excellent producers of detailed avatar designs, with hundreds of choices for customization. A supervised learning model could be…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 An Ngo , Daniel Phelps , Derrick Lai , Thanyared Wong , Lucas Mathias , Anish Shivamurthy , Mustafa Ajmal , Minghao Liu , James Davis

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Shulei Ji , Songruoyao Wu , Zihao Wang , Shuyu Li , Kejun Zhang

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Recently, interactive digital human video generation has attracted widespread attention and achieved remarkable progress. However, building such a practical system that can interact with diverse input signals in real time remains…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Ming Chen , Liyuan Cui , Wenyuan Zhang , Haoxian Zhang , Yan Zhou , Xiaohan Li , Songlin Tang , Jiwen Liu , Borui Liao , Hejia Chen , Xiaoqiang Liu , Pengfei Wan

Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting…

Computer Vision and Pattern Recognition · Computer Science 2015-01-13 Anna Rohrbach , Marcus Rohrbach , Niket Tandon , Bernt Schiele

End-to-end differentiable learning for autonomous driving (AD) has recently become a prominent paradigm. One main bottleneck lies in its voracious appetite for high-quality labeled data e.g. 3D bounding boxes and semantic segmentation,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Han Lu , Xiaosong Jia , Yichen Xie , Wenlong Liao , Xiaokang Yang , Junchi Yan

Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtain. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Naila Murray , Antonio Manuel López

3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Sijing Wu , Yunhao Li , Yichao Yan , Huiyu Duan , Ziwei Liu , Guangtao Zhai

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Ricong Huang , Peiwen Lai , Yipeng Qin , Guanbin Li

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified…

‹ Prev 1 8 9 10 Next ›