中文
相关论文

相关论文: Customized Visual Storytelling with Unified Multim…

200 篇论文

We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual content generation. While existing models excel at generating…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Zhengjian Yao , Yongzhi Li , Xinyuan Gao , Quan Chen , Peng Jiang , Yanye Lu

Chapter generation becomes practical technique for online videos nowadays. The chapter breakpoints enable users to quickly find the parts they want and get the summative annotations. However, there is no public method and dataset for this…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Xiao Cao , Zitan Chen , Canyu Le , Lei Meng

Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, is still challenging. Object customization, using reference…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Lingjie Kong , Kai Wu , Xiaobin Hu , Wenhui Han , Jinlong Peng , Chengming Xu , Donghao Luo , Mengtian Li , Jiangning Zhang , Chengjie Wang , Yanwei Fu

Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We argue that these two…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Jiaxu Zhang , Tianshu Hu , Yuan Zhang , Zenan Li , Linjie Luo , Guosheng Lin , Xin Chen

Large Multimodal Models (LMMs), or Vision-Language Models (VLMs), have shown impressive capabilities in a wide range of visual tasks. However, they often struggle with fine-grained visual reasoning, failing to identify domain-specific…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Yucheng Shi , Quanzheng Li , Jin Sun , Xiang Li , Ninghao Liu

As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models into true personal assistants hinges on visual…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Zihui Xue , Ami Baid , Sangho Kim , Mi Luo , Kristen Grauman

Medical report generation from imaging data remains a challenging task in clinical practice. While large language models (LLMs) show great promise in addressing this challenge, their effective integration with medical imaging data still…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Chunlei Li , Jingyang Hou , Yilei Shi , Jingliang Hu , Xiao Xiang Zhu , Lichao Mou

Analyzing literature involves tracking interactions between characters, locations, and themes. Visualization has the potential to facilitate the mapping and analysis of these complex relationships, but capturing structured information from…

人机交互 · 计算机科学 2025-08-12 Catherine Yeh , Tara Menon , Robin Singh Arya , Helen He , Moira Weigel , Fernanda Viégas , Martin Wattenberg

Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detailed, temporally grounded scripts remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Junfu Pu , Yuxin Chen , Teng Wang , Ying Shan

Vision-Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic representations over time. Previous methods resolved…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Tommaso Galliena , Stefano Rosa , Tommaso Apicella , Pietro Morerio , Alessio Del Bue , Lorenzo Natale

Language-modeling--based approaches to story plot generation attempt to construct a plot by sampling from a language model (LM) to predict the next character, word, or sentence to add to the story. LM techniques lack the ability to receive…

计算与语言 · 计算机科学 2023-01-19 Pradyumna Tambwekar , Murtaza Dhuliawala , Lara J. Martin , Animesh Mehta , Brent Harrison , Mark O. Riedl

Characters are important in narratives. They move the plot forward, create emotional connections, and embody the story's themes. Visual storytelling methods focus more on the plot and events relating to it, without building the narrative…

计算与语言 · 计算机科学 2025-03-04 Danyang Liu , Mirella Lapata , Frank Keller

Ensuring the functional correctness and safety of autonomous vehicles is a major challenge for the automotive industry. However, exhaustive physical test drives are not feasible, as billions of driven kilometers would be required to obtain…

软件工程 · 计算机科学 2021-02-09 Barbara Schuett , Thilo Braun , Stefan Otten , Eric Sax

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Daoan Zhang , Wenlin Yao , Xiaoyang Wang , Yebowen Hu , Jiebo Luo , Dong Yu

In order to offer a customized script tool and inspire professional scriptwriters, we present VScript. It is a controllable pipeline that generates complete scripts, including dialogues and scene descriptions, as well as presents visually…

计算与语言 · 计算机科学 2022-11-24 Ziwei Ji , Yan Xu , I-Tsun Cheng , Samuel Cahyawijaya , Rita Frieske , Etsuko Ishii , Min Zeng , Andrea Madotto , Pascale Fung

Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length…

多智能体系统 · 计算机科学 2025-10-03 Haoyuan Shi , Yunxin Li , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to…

人工智能 · 计算机科学 2023-11-23 Jiayang Wu , Wensheng Gan , Zefeng Chen , Shicheng Wan , Philip S. Yu

Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Zhifei Yang , Guangyao Zhai , Keyang Lu , YuYang Yin , Chao Zhang , Zhen Xiao , Jieyi Long , Nassir Navab , Yikai Wang

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in…

人机交互 · 计算机科学 2025-04-21 Zhen Wen , Luoxuan Weng , Yinghao Tang , Runjin Zhang , Yuxin Liu , Bo Pan , Minfeng Zhu , Wei Chen