English
Related papers

Related papers: VideoScore: Building Automatic Metrics to Simulate…

200 papers

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of textual instructions. This is a challenging problem as it…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Tomáš Souček , Prajwal Gatti , Michael Wray , Ivan Laptev , Dima Damen , Josef Sivic

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond…

Machine learning-based video codecs have made significant progress in the past few years. A critical area in the development of ML-based video codecs is an accurate evaluation metric that does not require an expensive and slow subjective…

Image and Video Processing · Electrical Eng. & Systems 2023-09-06 Abrar Majeedi , Babak Naderi , Yasaman Hosseinkashi , Juhee Cho , Ruben Alvarez Martinez , Ross Cutler

Automatic video summarization is still an unsolved problem due to several challenges. We take steps towards making automatic video summarization more realistic by addressing them. Firstly, the currently available datasets either have very…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Vishal Kaushal , Suraj Kothawade , Rishabh Iyer , Ganesh Ramakrishnan

There is growing interest in generating skeleton-based human motions from natural language descriptions. While most efforts have focused on developing better neural architectures for this task, there has been no significant work on…

Computation and Language · Computer Science 2023-09-20 Jordan Voas , Yili Wang , Qixing Huang , Raymond Mooney

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yucheng Suo , Fan Ma , Linchao Zhu , Tianyi Wang , Fengyun Rao , Yi Yang

We present a comprehensive solution to learn and improve text-to-image models from human preference feedback. To begin with, we build ImageReward -- the first general-purpose text-to-image human preference reward model -- to effectively…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiazheng Xu , Xiao Liu , Yuchen Wu , Yuxuan Tong , Qinkai Li , Ming Ding , Jie Tang , Yuxiao Dong

Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, given the difficulty of evaluating free-form answers, largely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Shaden Shaar , Bradon Thymes , Sirawut Chaixanien , Claire Cardie , Bharath Hariharan

Video summarization has been extensively studied in the past decades. However, user-generated video summarization is much less explored since there lack large-scale video datasets within which human-generated video summaries are…

Computation and Language · Computer Science 2019-04-15 Zhuo Lei , Chao Zhang , Qian Zhang , Guoping Qiu

Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet remains difficult due to the lack of suitable data and deployable editing models.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yiren Song , Cheng Liu , Yuxin Jiang , Mike Zheng Shou

We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifications, enabling…

Graphics · Computer Science 2025-12-02 Haoyi Duan , Hong-Xing Yu , Sirui Chen , Li Fei-Fei , Jiajun Wu

Humans share a strong tendency to memorize/forget some of the visual information they encounter. This paper focuses on providing computational models for the prediction of the intrinsic memorability of visual content. To address this new…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Romain Cohendet , Claire-Hélène Demarty , Ngoc Q. K. Duong , Martin Engilberge

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy benchmarks target manipulation-based forgeries, and recent synthetic-video benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Roberto Leotta , Salvatore Alfio Sambataro , Claudio Vittorio Ragaglia , Mirko Casu , Yuri Petralia , Francesco Guarnera , Luca Guarnera , Sebastiano Battiato

Automatic video summarization is still an unsolved problem due to several challenges. The currently available datasets either have very short videos or have few long videos of only a particular type. We introduce a new benchmarking video…

Computer Vision and Pattern Recognition · Computer Science 2021-01-27 Vishal Kaushal , Suraj Kothawade , Anshul Tomar , Rishabh Iyer , Ganesh Ramakrishnan

The task of automated code review has recently gained a lot of attention from the machine learning community. However, current review comment evaluation metrics rely on comparisons with a human-written reference for a given code change…

Software Engineering · Computer Science 2025-03-18 Atharva Naik , Marcus Alenius , Daniel Fried , Carolyn Rose

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential, a key research…

Robotics · Computer Science 2025-12-12 James Ni , Zekai Wang , Wei Lin , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik , Roei Herzig

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Keming Wu , Yijing Cui , Wenhan Xue , Qijie Wang , Xuan Luo , Zhiyuan Feng , Zuhao Yang , Sudong Wang , Sicong Jiang , Haowei Zhu , Zihan Wang , Ping Nie , Wenhu Chen , Bin Wang

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

Sound · Computer Science 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di