English
Related papers

Related papers: MagicAnime: A Hierarchically Annotated, Multimodal…

200 papers

Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically evaluated on a small selection of public datasets, a limited…

Machine Learning · Computer Science 2026-05-07 Leyan Xue , Changqing Zhang , Kecheng Xue , Xiaohong Liu , Guangyu Wang , Zongbo Han

Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Hailong Yan , Shice Liu , Tao Wang , Xiangtao Zhang , Yijie Zhong , Jinwei Chen , Le Zhang , Bo Li

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

Recent advances in AIGC (Artificial Intelligence Generated Content) models have enabled significant progress in image and video generation. However, users still struggle to obtain content that aligns with their preferences due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Zitong Xu , Dake Shen , Yaosong Du , Kexiang Hao , Jinghan Huang , Xiande Huang

Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Sankalp Sinha , Mohammad Sadil Khan , Muhammad Usama , Shino Sam , Didier Stricker , Sk Aziz Ali , Muhammad Zeshan Afzal

Imitation learning from a large set of human demonstrations has proved to be an effective paradigm for building capable robot agents. However, the demonstrations can be extremely costly and time-consuming to collect. We introduce MimicGen,…

Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input granularity and motion accuracy. Existing methods using…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 He Feng , Yongjia Ma , Donglin Di , Lei Fan , Tonghua Su

Monitoring animal behavior can facilitate conservation efforts by providing key insights into wildlife health, population status, and ecosystem function. Automatic recognition of animals and their behaviors is critical for capitalizing on…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Jun Chen , Ming Hu , Darren J. Coker , Michael L. Berumen , Blair Costelloe , Sara Beery , Anna Rohrbach , Mohamed Elhoseiny

Text-to-video generative models convert textual prompts into dynamic visual content, offering wide-ranging applications in film production, gaming, and education. However, their real-world performance often falls short of user expectations.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Wenhao Wang , Yi Yang

We present AniME, a director-oriented multi-agent system for automated long-form anime production, covering the full workflow from a story to the final video. The director agent keeps a global memory for the whole workflow, and coordinates…

Artificial Intelligence · Computer Science 2025-10-13 Lisai Zhang , Baohan Xu , Siqian Yang , Mingyu Yin , Jing Liu , Chao Xu , Siqi Wang , Yidi Wu , Yuxin Hong , Zihao Zhang , Yanzhang Liang , Yudong Jiang

Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Sijing Wu , Yunhao Li , Weitian Zhang , Jun Jia , Yucheng Zhu , Yichao Yan , Guangtao Zhai , Xiaokang Yang

The generative capabilities of Large Language Models (LLMs) are rapidly expanding from static code to dynamic, interactive visual artifacts. This progress is bottlenecked by a critical evaluation gap: established benchmarks focus on…

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first…

Computation and Language · Computer Science 2025-09-23 Fei Zhao , Chengqiang Lu , Yufan Shen , Qimeng Wang , Yicheng Qian , Haoxin Zhang , Yan Gao , Yi Wu , Yao Hu , Zhen Wu , Shangyu Xing , Xinyu Dai

Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Kai Zhang , Lingbo Mo , Wenhu Chen , Huan Sun , Yu Su

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Xiaofeng Wang , Kang Zhao , Feng Liu , Jiayu Wang , Guosheng Zhao , Xiaoyi Bao , Zheng Zhu , Yingya Zhang , Xingang Wang

Generating large-scale multi-character interactions is a challenging and important task in character animation. Multi-character interactions involve not only natural interactive motions but also characters coordinated with each other for…

Graphics · Computer Science 2025-05-21 Ziyi Chang , He Wang , George Alex Koulieris , Hubert P. H. Shum

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we…

Artificial Intelligence · Computer Science 2025-05-26 Jihan Yao , Yushi Hu , Yujie Yi , Bin Han , Shangbin Feng , Guang Yang , Bingbing Wen , Ranjay Krishna , Lucy Lu Wang , Yulia Tsvetkov , Noah A. Smith , Banghua Zhu

We introduce a novel framework for 3D human avatar generation and personalization, leveraging text prompts to enhance user engagement and customization. Central to our approach are key innovations aimed at overcoming the challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Armand Comas-Massagué , Di Qiu , Menglei Chai , Marcel Bühler , Amit Raj , Ruiqi Gao , Qiangeng Xu , Mark Matthews , Paulo Gotardo , Octavia Camps , Sergio Orts-Escolano , Thabo Beeler

Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a…

Human-Computer Interaction · Computer Science 2025-07-09 Pegah Salehi , Sajad Amouei Sheshkal , Vajira Thambawita , Michael A. Riegler , Pål Halvorsen

Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Xinhang Gao , Junlin Guan , Shuhan Luo , Wenzhuo Li , Guanghuan Tan , Jiacheng Wang