中文
相关论文

相关论文: AutoAD III: The Prequel -- Back to the Pixels

200 篇论文

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Mengxue Hu , Yunfeng Diao , Changtao Miao , Zhiqing Guo , Jianshu Li , Zhe Li , Joey Tianyi Zhou

Given an arbitrary audio clip, audio-driven 3D facial animation aims to generate lifelike lip motions and facial expressions for a 3D head. Existing methods typically rely on training their models using limited public 3D datasets that…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Liying Lu , Tianke Zhang , Yunfei Liu , Xuangeng Chu , Yu Li

Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation…

计算机视觉与模式识别 · 计算机科学 2021-08-20 Yudong Guo , Keyu Chen , Sen Liang , Yong-Jin Liu , Hujun Bao , Juyong Zhang

End-to-end architectures in autonomous driving (AD) face a significant challenge in interpretability, impeding human-AI trust. Human-friendly natural language has been explored for tasks such as driving explanation and 3D captioning.…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Kairui Ding , Boyuan Chen , Yuchen Su , Huan-ang Gao , Bu Jin , Chonghao Sima , Wuqiang Zhang , Xiaohui Li , Paul Barsch , Hongyang Li , Hao Zhao

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Weixia Zhang , Chengguang Zhu , Jingnan Gao , Yichao Yan , Guangtao Zhai , Xiaokang Yang

Learning how to generate descriptions of images or videos received major interest both in the Computer Vision and Natural Language Processing communities. While a few works have proposed to learn a grounding during the generation process in…

计算机视觉与模式识别 · 计算机科学 2017-04-06 Anna Rohrbach , Marcus Rohrbach , Siyu Tang , Seong Joon Oh , Bernt Schiele

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature of the audio-visual features involved. Recent works have…

声音 · 计算机科学 2024-09-17 Mingjing Yi , Ming Li

Automated audio captioning is a cross-modal translation task that aims to generate natural language descriptions for given audio clips. This task has received increasing attention with the release of freely available datasets in recent…

音频与语音处理 · 电气工程与系统科学 2022-09-28 Xinhao Mei , Xubo Liu , Mark D. Plumbley , Wenwu Wang

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Xuyang Shen , Dong Li , Jinxing Zhou , Zhen Qin , Bowen He , Xiaodong Han , Aixuan Li , Yuchao Dai , Lingpeng Kong , Meng Wang , Yu Qiao , Yiran Zhong

In recent years, text-to-audio models have revolutionized the field of automatic audio generation. This paper investigates their application in generating synthetic datasets for training data-driven models. Specifically, this study analyzes…

音频与语音处理 · 电气工程与系统科学 2024-07-09 Francesca Ronchini , Luca Comanducci , Fabio Antonacci

Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. Unlike standard video captioning, it involves not only describing key visual details but also inferring plots that unfold…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Zihao Yue , Yepeng Zhang , Ziheng Wang , Qin Jin

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Video Multimethod Assessment Fusion (VMAF) [1], [2], [3] is a popular tool in the industry for measuring coded video quality. In this study, we propose an auditory-inspired frontend in existing VMAF for creating videos of reference and…

音频与语音处理 · 电气工程与系统科学 2023-08-08 Arijit Biswas , Harald Mundt

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. This begs the…

声音 · 计算机科学 2023-05-23 Guy Yariv , Itai Gat , Lior Wolf , Yossi Adi , Idan Schwartz

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

计算机视觉与模式识别 · 计算机科学 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

The analysis, processing, and extraction of meaningful information from sounds all around us is the subject of the broader area of audio analytics. Audio captioning is a recent addition to the domain of audio analytics, a cross-modal…

音频与语音处理 · 电气工程与系统科学 2023-05-04 Sandeep Kothinti , Dimitra Emmanouilidou

Recent text-to-video models have enabled the generation of high-resolution driving scenes from natural language prompts. These AI-generated driving videos (AIGVs) offer a low-cost, scalable alternative to real or simulator data for…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Xinhao Xiang , Abhijeet Rastogi , Jiawei Zhang

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…