中文

LLM 作为艺术总监(LaDi):利用 LLM 改进文本到媒体生成器

计算与语言 2023-11-08 v1 人工智能 计算机视觉与模式识别

摘要

近期文本到图像生成的进展通过自动生成高质量、上下文感知的图像与视频,革新了艺术与电影等诸多领域。然而,这些技术的实用性常受限于文本提示在引导生成器产出艺术连贯且与主题相关图像方面的不足。本文描述了可使大语言模型(Large Language Models, LLMs)充当艺术总监以增强图像与视频生成的技术。我们描述了为此构建的统一系统“LaDi”。我们探讨了 LaDi 如何整合多种技术以增强文本到图像生成器(T2Is)与文本到视频生成器(T2Vs)的能力,重点包括约束解码、智能提示、微调与检索。LaDi 及这些技术目前正应用于 Plai Labs 所开发的应用与平台中。

关键词

引用

@article{arxiv.2311.03716,
  title  = {LLM as an Art Director (LaDi): Using LLMs to improve Text-to-Media Generators},
  author = {Allen Roush and Emil Zakirov and Artemiy Shirokov and Polina Lunina and Jack Gane and Alexander Duffy and Charlie Basil and Aber Whitcomb and Jim Benedetto and Chris DeWolfe},
  journal= {arXiv preprint arXiv:2311.03716},
  year   = {2023}
}

备注

12 pages, System Demonstration/Industry paper. Preprint