中文
相关论文

相关论文: FoleyBench: A Benchmark For Video-to-Audio Models

200 篇论文

The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-shot paradigms, lacking the comprehensive story assets and…

多媒体 · 计算机科学 2026-03-02 Haoyuan Shi , Yunxin Li , Nanhao Deng , Zhenran Xu , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

In this research, we present an interface based on Variational Autoencoders trained with a wide range of natural sounds for the innovative creation of Foley effects. The model can transfer new sound features to prerecorded audio or…

音频与语音处理 · 电气工程与系统科学 2023-10-25 Mateo Cámara , José Luis Blanco

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval,…

声音 · 计算机科学 2026-01-30 Hui Wang , Jinghua Zhao , Junyang Cheng , Cheng Liu , Yuhang Jia , Haoqin Sun , Jiaming Zhou , Yong Qin

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

声音 · 计算机科学 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Existing VLM frameworks predominantly assess…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Pavan Kumar Anasosalu Vasu , Cem Koc , Fartash Faghri , Chun-Liang Li , Bo Feng , Zhengfeng Lai , Meng Cao , Oncel Tuzel , Hadi Pouransari

Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Hongbo Liu , Jingwen He , Yi Jin , Dian Zheng , Yuhao Dong , Fan Zhang , Ziqi Huang , Yinan He , Yangguang Li , Weichao Chen , Yu Qiao , Wanli Ouyang , Shengjie Zhao , Ziwei Liu

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have…

声音 · 计算机科学 2025-09-29 Zeyu Xie , Yaoyun Zhang , Xuenan Xu , Yongkang Yin , Chenxing Li , Mengyue Wu , Yuexian Zou

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

计算与语言 · 计算机科学 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

Generative models have driven significant progress in a variety of AI tasks, including text-to-video generation, where models like Video LDM and Stable Video Diffusion can produce realistic, movie-level videos from textual instructions.…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Xuyang Guo , Zekai Huang , Jiayan Huo , Yingyu Liang , Zhenmei Shi , Zhao Song , Jiahao Zhang

Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some…

Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses two core evaluation challenges:…

Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as images or audio.…

计算与语言 · 计算机科学 2025-11-13 Yiming Gao , Bin Wang , Chengwei Wei , Shuo Sun , AiTi Aw

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehension, necessitating…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Junming Lin , Zheng Fang , Chi Chen , Zihao Wan , Fuwen Luo , Peng Li , Yang Liu , Maosong Sun

Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their…

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

图形学 · 计算机科学 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tasks, hindering the…

As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speaking, how they sound, and where the conversation takes place…

声音 · 计算机科学 2026-04-21 Yuxiang Wang , Hongyu Liu , Yijiang Xu , Qinke Ni , Li Wang , Wan Lin , Kunyu Feng , Dekun Chen , Xu Tan , Lei Wang , Jie Shi , Zhizheng Wu

Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Lei Zhu , Xing Cai , Yingjie Chen , Yiheng Li , Binxin Yang , Hao Liu , Jie Chen , Chen Li , Jing LYu

"Foley" refers to sound effects that are added to multimedia during post-production to enhance its perceived acoustic properties, e.g., by simulating the sounds of footsteps, ambient environmental sounds, or visible objects on the screen.…

声音 · 计算机科学 2022-07-25 Keunwoo Choi , Sangshin Oh , Minsung Kang , Brian McFee
‹ 上一页 1 8 9 10 下一页 ›