English
Related papers

Related papers: TMD-Bench: A Multi-Level Evaluation Paradigm for M…

200 papers

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities independently and do not examine whether…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Weixing Wang , Liudvikas Zekas , Anton Hackl , Constantin Alexander Auga , Parisa Shahabinejad , Jona Otholt , Antonio Rueda-Toicen , Gerard de Melo

We propose a novel text-to-video (T2V) generation benchmark, ChronoMagic-Bench, to evaluate the temporal and metamorphic capabilities of the T2V models (e.g. Sora and Lumiere) in time-lapse video generation. In contrast to existing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Shenghai Yuan , Jinfa Huang , Yongqi Xu , Yaoyang Liu , Shaofeng Zhang , Yujun Shi , Ruijie Zhu , Xinhua Cheng , Jiebo Luo , Li Yuan

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos…

Multimedia · Computer Science 2024-09-12 Yan-Bo Lin , Yu Tian , Linjie Yang , Gedas Bertasius , Heng Wang

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu

Existing text-to-video (T2V) evaluation benchmarks, such as VBench and EvalCrafter, suffer from two limitations. (i) While the emphasis is on subject-centric prompts or static camera scenes, camera motion essential for producing cinematic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Nithin C. Babu , Aniruddha Mahapatra , Harsh Rangwani , Rajiv Soundararajan , Kuldeep Kulkarni

Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to calibrate semantic consistency. However, their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Juanxi Tian , Siyuan Li , Conghui He , Lijun Wu , Cheng Tan

Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textual semantics. However, most existing works study these two problems separately, without…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jidong Kuang , Hongsong Wang , Jie Gui

Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in other modalities, their…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Baisen Wang , Le Zhuo , Zhaokai Wang , Chenxi Bao , Wu Chengjing , Xuecheng Nie , Jiao Dai , Jizhong Han , Yue Liao , Si Liu

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu

Dance serves as a powerful medium for expressing human emotions, but the lifelike generation of dance is still a considerable challenge. Recently, diffusion models have showcased remarkable generative abilities across various domains. They…

Sound · Computer Science 2024-06-25 Canyu Zhang , Youbao Tang , Ning Zhang , Ruei-Sung Lin , Mei Han , Jing Xiao , Song Wang

Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing across visual and textual modalities. However, developing a unified framework for UMMs…

Artificial Intelligence · Computer Science 2026-05-21 Yinyi Luo , Wenwen Wang , Hayes Bai , Hongyu Zhu , Hao Chen , Pan He , Marios Savvides , Sharon Li , Jindong Wang

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Hongxiang Li , Yaowei Li , Bin Lin , Yuwei Niu , Yuhang Yang , Xiaoshuang Huang , Jiayin Cai , Xiaolong Jiang , Yao Hu , Long Chen

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Youcan Xu , Zhen Wang , Jiaxin Shi , Kexin Li , Feifei Shao , Jun Xiao , Yi Yang , Jun Yu , Long Chen

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Yang Luo , Xuanlei Zhao , Baijiong Lin , Lingting Zhu , Liyao Tang , Yuqi Liu , Ying-Cong Chen , Shengju Qian , Xin Wang , Yang You

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Shangkun Sun , Xiaoyu Liang , Songlin Fan , Wenxu Gao , Wei Gao

Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popularity of the technology. This progress has created a growing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Zelu Qi , Ping Shi , Shuqi Wang , Chaoyang Zhang , Fei Zhao , Zefeng Ying , Da Pan , Xi Yang , Zheqi He , Teng Dai

Machine-generated music (MGM) has become a groundbreaking innovation with wide-ranging applications, such as music therapy, personalised editing, and creative inspiration within the music industry. However, the unregulated proliferation of…

Sound · Computer Science 2026-04-30 Yupei Li , Qiyang Sun , Hanqian Li , Lucia Specia , Björn W. Schuller

In pop music, accompaniments are usually played by multiple instruments (tracks) such as drum, bass, string and guitar, and can make a song more expressive and contagious by arranging together with its melody. Previous works usually…

Sound · Computer Science 2020-08-19 Yi Ren , Jinzheng He , Xu Tan , Tao Qin , Zhou Zhao , Tie-Yan Liu