中文
相关论文

相关论文: Gaud\'i: Conversational Interactions with Deep Rep…

200 篇论文

Technological progress has persistently shaped the dynamics of human-machine interactions in task execution. In response to the advancements in Generative AI, this paper outlines a detailed study plan that investigates various human-AI…

人机交互 · 计算机科学 2024-02-13 Zijian Ding

Development of multimodal interactive systems is hindered by the lack of rich, multimodal (text, images) conversational data, which is needed in large quantities for LLMs. Previous approaches augment textual dialogues with retrieved images,…

Emotion is vital to information and message processing, playing a key role in attitude formation. Consequently, creating a mood that evokes an emotional response is essential to any compelling piece of outreach communication. Many…

人机交互 · 计算机科学 2024-03-20 Samia Menon , Sitong Wang , Lydia Chilton

AI illustrator aims to automatically design visually appealing images for books to provoke rich thoughts and emotions. To achieve this goal, we propose a framework for translating raw descriptions with complex semantics into semantically…

计算机视觉与模式识别 · 计算机科学 2022-09-09 Yiyang Ma , Huan Yang , Bei Liu , Jianlong Fu , Jiaying Liu

Generative artificial intelligence (GenAI) can rapidly produce large and diverse volumes of content. This lends to it a quality of creativity which can be empowering in the early stages of design. In seeking to understand how creative ways…

人机交互 · 计算机科学 2024-03-20 Gionnieve Lim , Simon T. Perrault

GPT-Vision has impressed us on a range of vision-language tasks, but it comes with the familiar new challenge: we have little idea of its capabilities and limitations. In our study, we formalize a process that many have instinctively been…

计算与语言 · 计算机科学 2023-11-06 Alyssa Hwang , Andrew Head , Chris Callison-Burch

Current research has explored how Generative AI can support the brainstorming process for content creators, but a gap remains in exploring support-tools for the pre-writing process. Specifically, our research is focused on supporting users…

人机交互 · 计算机科学 2024-06-19 Grace Li , Tao Long , Lydia B. Chilton

Text-to-image generative models have demonstrated remarkable capabilities in generating high-quality images based on textual prompts. However, crafting prompts that accurately capture the user's creative intent remains challenging. It often…

人机交互 · 计算机科学 2023-04-20 Stephen Brade , Bryan Wang , Mauricio Sousa , Sageev Oore , Tovi Grossman

Text-to-image generation models have seen considerable advancement, catering to the increasing interest in personalized image creation. Current customization techniques often necessitate users to provide multiple images (typically 3-5) for…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Linhao Zhong , Yan Hong , Wentao Chen , Binglin Zhou , Yiyi Zhang , Jianfu Zhang , Liqing Zhang

In this work, we introduce Mozualization, a music generation and editing tool that creates multi-style embedded music by integrating diverse inputs, such as keywords, images, and sound clips (e.g., segments from various pieces of music or…

人机交互 · 计算机科学 2025-04-22 Wanfang Xu , Lixiang Zhao , Haiwen Song , Xinheng Song , Zhaolin Lu , Yu Liu , Min Chen , Eng Gee Lim , Lingyun Yu

A method for generating narratives by analyzing single images or image sequences is presented, inspired by the time immemorial tradition of Narrative Art. The proposed method explores the multimodal capabilities of GPT-4o to interpret…

计算与语言 · 计算机科学 2024-08-22 Edirlei Soares de Lima , Marco A. Casanova , Antonio L. Furtado

A big part of achieving Artificial General Intelligence(AGI) is to build a machine that can see and listen like humans. Much work has focused on designing models for image classification, video classification, object detection, pose…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Ruotian Luo

Generative AI has increasingly been used for artistic creation, but little work has explored how it shapes the experiential meaning of practice. We consider how generative AI might transform the embodied and tangible process of instant…

人机交互 · 计算机科学 2026-05-05 Michael Yin , Angela Chiang , Robert Xiao

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that…

声音 · 计算机科学 2024-09-13 Tanisha Hisariya , Huan Zhang , Jinhua Liang

Generative AI techniques like those that synthesize images from text (text-to-image models) offer new possibilities for creatively imagining new ideas. We investigate the capabilities of these models to help communities engage in…

人机交互 · 计算机科学 2022-06-22 Ziv Epstein , Hope Schroeder , Dava Newman

The human brain is naturally equipped to comprehend and interpret visual information rapidly. When confronted with complex problems or concepts, we use flowcharts, sketches, and diagrams to aid our thought process. Leveraging this inherent…

计算机视觉与模式识别 · 计算机科学 2023-11-17 Fanxu Meng , Haotong Yang , Yiding Wang , Muhan Zhang

As technologies become more and more pervasive, there is a need for considering the affective dimension of interaction with computer systems to make them more human-like. Current demands for this matter include accurate emotion recognition,…

人机交互 · 计算机科学 2018-06-13 Barbara Giżycka , Grzegorz J. Nalepa , Paweł Jemioło

Visual metaphors are powerful rhetorical devices used to persuade or communicate creative ideas through images. Similar to linguistic metaphors, they convey meaning implicitly through symbolism and juxtaposition of the symbols. We propose a…

This work investigates the integration of generative visual aids in human-robot task communication. We developed GenComUI, a system powered by large language models that dynamically generates contextual visual aids (such as map annotations,…

人机交互 · 计算机科学 2025-02-18 Yate Ge , Meiying Li , Xipeng Huang , Yuanda Hu , Qi Wang , Xiaohua Sun , Weiwei Guo

Automatic image captioning has recently approached human-level performance due to the latest advances in computer vision and natural language understanding. However, most of the current models can only generate plain factual descriptions…

计算机视觉与模式识别 · 计算机科学 2018-01-31 Quanzeng You , Hailin Jin , Jiebo Luo