中文
相关论文

相关论文: PhyAVBench: A Challenging Audio Physics-Sensitivit…

200 篇论文

As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organization, and intelligent assistants. However, most existing benchmarks concentrate on…

人工智能 · 计算机科学 2026-05-07 Honglei Zhang , Yuting Chen , Chenpeng Hu , Siyue Zhang , Yilei Shi

Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result, their adoption has increased in many computer vision…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Elena Izzo , Luca Parolari , Davide Vezzaro , Lamberto Ballan

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu

Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Guangzhao Li , Kangrui Cen , Baixuan Zhao , Yi Xin , Siqi Luo , Guangtao Zhai , Lei Zhang , Xiaohong Liu

We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Guy Yariv , Yuval Kirstain , Amit Zohar , Shelly Sheynin , Yaniv Taigman , Yossi Adi , Sagie Benaim , Adam Polyak

Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely designed for realistic videos-struggle to evaluate animation-style generation with its…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Leyi Wu , Pengjun Fang , Kai Sun , Yazhou Xing , Yinwei Wu , Songsong Wang , Ziqi Huang , Dan Zhou , Yingqing He , Ying-Cong Chen , Qifeng Chen

Evaluating the quality of videos generated from text-to-video (T2V) models is important if they are to produce plausible outputs that convince a viewer of their authenticity. We examine some of the metrics used in this area and highlight…

计算机视觉与模式识别 · 计算机科学 2023-09-18 Iya Chivileva , Philip Lynch , Tomas E. Ward , Alan F. Smeaton

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional…

声音 · 计算机科学 2025-05-16 Detao Bai , Zhiheng Ma , Xihan Wei , Liefeng Bo

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Yiqing Shen , Chenjia Li , Chenxiao Fan , Mathias Unberath

Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from limited video duration and narrow evaluation scopes, making it…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Yangyang Zhong , Ji Qi , Yuan Yao , Pengxin Luo , Yunfeng Yan , Donglian Qi , Zhiyuan Liu , Tat-Seng Chua

We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Jinxing Zhou , Xuyang Shen , Jianyuan Wang , Jiayi Zhang , Weixuan Sun , Jing Zhang , Stan Birchfield , Dan Guo , Lingpeng Kong , Meng Wang , Yiran Zhong

Unlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches, leveraging transformer…

声音 · 计算机科学 2025-02-24 Jia Li , Wenjie Zhao , Ziru Huang , Yunhui Guo , Yapeng Tian

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and…

多媒体 · 计算机科学 2024-10-01 Kun Su , Xiulong Liu , Eli Shlizerman

Text-to-Visualization (Text2VIS) enables users to create visualizations from natural language queries, making data insights more accessible. However, Text2VIS faces challenges in interpreting ambiguous queries, as users often express their…

计算与语言 · 计算机科学 2026-01-06 Tianqi Luo , Chuhan Huang , Leixian Shen , Boyan Li , Shuyu Shen , Wei Zeng , Nan Tang , Yuyu Luo

Video generation models are increasingly used as world simulators for storytelling, simulation, and embodied AI. As these models advance, a key question arises: do generated videos obey the physical laws of the real world? Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Qin Zhang , Peiyu Jing , Hong-Xing Yu , Fangqiang Ding , Fan Nie , Weimin Wang , Yilun Du , James Zou , Jiajun Wu , Bing Shuai

Text-to-video (T2V) generation has gained significant attention due to its wide applications to video generation, editing, enhancement and translation, \etc. However, high-quality (HQ) video synthesis is extremely challenging because of the…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Tao Yang , Yangming Shi , Yunwen Huang , Feng Chen , Yin Zheng , Lei Zhang

Recent breakthroughs in Vision-Language (V&L) joint research have achieved remarkable results in various text-driven tasks. High-quality Text-to-video (T2V), a task that has been long considered mission-impossible, was proven feasible with…

人工智能 · 计算机科学 2022-11-28 Yuxing Qiu , Feng Gao , Minchen Li , Govind Thattai , Yin Yang , Chenfanfu Jiang