中文
相关论文

相关论文: MSG Score: Automated Video Verification for Reliab…

200 篇论文

Motivated by the success of coarse-grained or fine-grained contrast in text-video retrieval, there emerge multi-grained contrastive learning methods which focus on the integration of contrasts with different granularity. However, due to the…

信息检索 · 计算机科学 2025-04-08 Xiaolun Jing , Genke Yang , Jian Chu

Retrieval-Augmented Generation (RAG) has demonstrated remarkable success in enhancing Large Language Models (LLMs) through external knowledge integration, yet its application has primarily focused on textual content, leaving the rich domain…

信息检索 · 计算机科学 2025-02-04 Xubin Ren , Lingrui Xu , Long Xia , Shuaiqiang Wang , Dawei Yin , Chao Huang

Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous…

计算与语言 · 计算机科学 2025-06-13 Tian Lan , Yang-Hao Zhou , Zi-Ao Ma , Fanshu Sun , Rui-Qing Sun , Junyu Luo , Rong-Cheng Tu , Heyan Huang , Chen Xu , Zhijing Wu , Xian-Ling Mao

Parallel test-time scaling, which generates multiple candidate solutions for a single problem, is a powerful technique for improving large language model performance. However, it is hindered by two key bottlenecks: accurately selecting the…

密码学与安全 · 计算机科学 2026-03-05 Yegon Kim , Seungyoo Lee , Chaeyun Jang , Hyungi Lee , Juho Lee

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Recent advances in text-to-video generation, particularly with autoregressive models, have enabled the synthesis of high-quality videos depicting individual scenes. However, extending these models to generate long, cross-scene videos…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Xueji Fang , Liyuan Ma , Zhiyang Chen , Mingyuan Zhou , Guo-jun Qi

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Najmeh Sadoughi , Xinyu Li , Avijit Vajpayee , David Fan , Bing Shuai , Hector Santos-Villalobos , Vimal Bhat , Rohith MV

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

Synthesizing consistent and photorealistic 3D scenes is an open problem in computer vision. Video diffusion models generate impressive videos but cannot directly synthesize 3D representations, i.e., lack 3D consistency in the generated…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Katja Schwarz , Norman Mueller , Peter Kontschieder

We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (using vision language models). Recently, MVGs have been the…

图形学 · 计算机科学 2025-07-02 Xianghui Xie , Chuhang Zou , Meher Gitika Karumuri , Jan Eric Lenssen , Gerard Pons-Moll

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Tanzila Rahman , Hsin-Ying Lee , Jian Ren , Sergey Tulyakov , Shweta Mahajan , Leonid Sigal

Effective content moderation is essential for video platforms to safeguard user experience and uphold community standards. While traditional video classification models effectively handle well-defined moderation tasks, they struggle with…

机器学习 · 计算机科学 2025-07-24 Zixuan Wang , Jinghao Shi , Hanzhong Liang , Xiang Shen , Vera Wen , Zhiqian Chen , Yifan Wu , Zhixin Zhang , Hongyu Xiong

Dynamic Scene Graph Generation (DSGG) for videos is a challenging task in computer vision. While existing approaches often focus on sophisticated architectural design and solely use recall during evaluation, we take a closer look at their…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xuanming Cui , Jaiminkumar Ashokbhai Bhoi , Chionh Wei Peng , Adriel Kuek , Ser Nam Lim

Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Yaosi Hu , Chong Luo , Zhenzhong Chen

Goal-oriented generative script learning aims to generate subsequent steps to reach a particular goal, which is an essential task to assist robots or humans in performing stereotypical activities. An important aspect of this process is the…

计算与语言 · 计算机科学 2025-06-11 Qingyun Wang , Manling Li , Hou Pong Chan , Lifu Huang , Julia Hockenmaier , Girish Chowdhary , Heng Ji

Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive computational cost of processing dense frame sequences. Prevailing solutions, which select a keyframe subset,…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Shaoguang Wang , Weiyu Guo , Ziyang Chen , Xuming Hu , Hui Xiong

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Sridhar S , Nithin A , Shakeel Rifath , Vasantha Raj

Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some…