English
Related papers

Related papers: OSCBench: Benchmarking Object State Change in Text…

200 papers

Generative AI models, particularly Text-to-Video (T2V) systems, offer a promising avenue for transforming science education by automating the creation of engaging and intuitive visual explanations. In this work, we take a first step toward…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Megha Mariam K. M , Aditya Arun , Zakaria Laskar , C. V. Jawahar

The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making the existing benchmarks increasingly insufficient to evaluate state-of-the-art T2V models.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Zeqing Wang , Xinyu Wei , Bairui Li , Zhen Guo , Jinrui Zhang , Hongyang Wei , Keze Wang , Lei Zhang

Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit impressive capabilities, crafting visually stunning videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jay Zhangjie Wu , Guian Fang , Haoning Wu , Xintao Wang , Yixiao Ge , Xiaodong Cun , David Junhao Zhang , Jia-Wei Liu , Yuchao Gu , Rui Zhao , Weisi Lin , Wynne Hsu , Ying Shan , Mike Zheng Shou

Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuheng Chen , Teng Hu , Jiangning Zhang , Zhucun Xue , Ran Yi , Lizhuang Ma

Existing text-to-video (T2V) evaluation benchmarks, such as VBench and EvalCrafter, suffer from two limitations. (i) While the emphasis is on subject-centric prompts or static camera scenes, camera motion essential for producing cinematic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Nithin C. Babu , Aniruddha Mahapatra , Harsh Rangwani , Rajiv Soundararajan , Kuldeep Kulkarni

Text-to-3D (T23D) generation has emerged as a crucial visual generation task, aiming at synthesizing 3D content from textual descriptions. Studies of this task are currently shifting from per-scene T23D, which requires optimization of the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Xiao Cai , Sitong Su , Jingkuan Song , Pengpeng Zeng , Ji Zhang , Qinhong Du , Mengqi Li , Heng Tao Shen , Lianli Gao

A plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Qihe Pan , Zhen Zhao , Zicheng Wang , Sifan Long , Yiming Wu , Wei Ji , Haoran Liang , Ronghua Liang

Most of these text-to-video (T2V) generative models often produce single-scene video clips that depict an entity performing a particular action (e.g., 'a red panda climbing a tree'). However, it is pertinent to generate multi-scene videos…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Hritik Bansal , Yonatan Bitton , Michal Yarom , Idan Szpektor , Aditya Grover , Kai-Wei Chang

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shenghai Yuan , Xianyi He , Yufan Deng , Yang Ye , Jinfa Huang , Bin Lin , Jiebo Luo , Li Yuan

The burgeoning field of Artificial Intelligence Generated Content (AIGC) is witnessing rapid advancements, particularly in video generation. This paper introduces AIGCBench, a pioneering comprehensive and scalable benchmark designed to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Fanda Fan , Chunjie Luo , Wanling Gao , Jianfeng Zhan

Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Yang Du , Zhuoran Lin , Kaiqiang Song , Biao Wang , Zhicheng Zheng , Tiezheng Ge , Bo Zheng , Qin Jin

In recent years, large-scale models have achieved significant advancements, accompanied by the emergence of numerous high-quality benchmarks for evaluating various aspects of their comprehension abilities. However, most existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Kangning Li , Zheyang Jia , Anyu Ying

The capability of intelligent models to extrapolate and comprehend changes in object states is a crucial yet demanding aspect of AI research, particularly through the lens of human interaction in real-world settings. This task involves…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Nguyen Nguyen , Jing Bi , Ali Vosoughi , Yapeng Tian , Pooyan Fazli , Chenliang Xu

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the…

Sound · Computer Science 2024-03-14 Shentong Mo , Jing Shi , Yapeng Tian

Querying generative AI models, e.g., large language models (LLMs), has become a prevalent method for information acquisition. However, existing query-answer datasets primarily focus on textual responses, making it challenging to address…

Artificial Intelligence · Computer Science 2025-06-03 Shuting Wang , Yunqi Liu , Zixin Yang , Ning Hu , Zhicheng Dou , Chenyan Xiong

We propose a novel text-to-video (T2V) generation benchmark, ChronoMagic-Bench, to evaluate the temporal and metamorphic capabilities of the T2V models (e.g. Sora and Lumiere) in time-lapse video generation. In contrast to existing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Shenghai Yuan , Jinfa Huang , Yongqi Xu , Yaoyang Liu , Shaofeng Zhang , Yujun Shi , Ruijie Zhu , Xinhua Cheng , Jiebo Luo , Li Yuan

Current text-to-image generative models struggle to accurately represent object states (e.g., "a table without a bottle," "an empty tumbler"). In this work, we first design a fully-automatic pipeline to generate high-quality synthetic data…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Tianle Chen , Chaitanya Chakka , Deepti Ghadiyaram

Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daili Hua , Xizhi Wang , Bohan Zeng , Xinyi Huang , Hao Liang , Junbo Niu , Xinlong Chen , Quanqing Xu , Wentao Zhang

Following recipes while cooking is an important but difficult task for visually impaired individuals. We developed OSCAR (Object Status Context Awareness for Recipes), a novel approach that provides recipe progress tracking and…

Human-Computer Interaction · Computer Science 2025-03-11 Franklin Mingzhe Li , Kaitlyn Ng , Bin Zhu , Patrick Carrington

Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-driven methods cannot…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Fangyi Zhu , Jenq-Neng Hwang , Zhanyu Ma , Guang Chen , Jun Guo