English

Gaud\'i: Conversational Interactions with Deep Representations to Generate Image Collections

Artificial Intelligence 2021-12-09 v1 Information Retrieval Machine Learning

Abstract

Based on recent advances in realistic language modeling (GPT-3) and cross-modal representations (CLIP), Gaud\'i was developed to help designers search for inspirational images using natural language. In the early stages of the design process, with the goal of eliciting a client's preferred creative direction, designers will typically create thematic collections of inspirational images called "mood-boards". Creating a mood-board involves sequential image searches which are currently performed using keywords or images. Gaud\'i transforms this process into a conversation where the user is gradually detailing the mood-board's theme. This representation allows our AI to generate new search queries from scratch, straight from a project briefing, following a theme hypothesized by GPT-3. Compared to previous computational approaches to mood-board creation, to the best of our knowledge, ours is the first attempt to represent mood-boards as the stories that designers tell when presenting a creative direction to a client.

Keywords

Cite

@article{arxiv.2112.04404,
  title  = {Gaud\'i: Conversational Interactions with Deep Representations to Generate Image Collections},
  author = {Victor S. Bursztyn and Jennifer Healey and Vishwa Vinay},
  journal= {arXiv preprint arXiv:2112.04404},
  year   = {2021}
}

Comments

Accepted at the NeurIPS 2021 Workshop on Machine Learning for Creativity and Design

R2 v1 2026-06-24T08:09:21.142Z