中文
相关论文

相关论文: TaleForge: Interactive Multimodal System for Perso…

200 篇论文

With the advancement of natural language generation (NLG) technologies, creative story generation systems have gained increasing attention. However, current systems often fail to accurately translate user intent into satisfactory story…

计算与语言 · 计算机科学 2025-12-03 Yunchao Wang , Guodao Sun , Zihang Fu , Zhehao Liu , Kaixing Du , Haidong Gao , Ronghua Liang

Personalized interaction is highly valued by parents in their story-reading activities with children. While AI-empowered story-reading tools have been increasingly used, their abilities to support personalized interaction with children are…

人机交互 · 计算机科学 2025-03-28 Jiaju Chen , Minglong Tang , Yuxuan Lu , Bingsheng Yao , Elissa Fan , Xiaojuan Ma , Ying Xu , Dakuo Wang , Yuling Sun , Liang He

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inaccurate dialogue…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Ayan Banerjee , Josep Llados , Umapada Pal , Anjan Dutta

Automated plot generation for games enhances the player's experience by providing rich and immersive narrative experience that adapts to the player's actions. Traditional approaches adopt a symbolic narrative planning method which limits…

人机交互 · 计算机科学 2024-11-05 Yi Wang , Qian Zhou , David Ledo

The emergence of large language models (LLMs) has revolutionized the capabilities of text comprehension and generation. Multi-modal generation attracts great attention from both the industry and academia, but there is little work on…

信息检索 · 计算机科学 2024-04-16 Xiaoteng Shen , Rui Zhang , Xiaoyan Zhao , Jieming Zhu , Xi Xiao

Storytelling is an open-ended task that entails creative thinking and requires a constant flow of ideas. Natural language generation (NLG) for storytelling is especially challenging because it requires the generated text to follow an…

计算与语言 · 计算机科学 2021-09-17 Eden Bensaid , Mauro Martino , Benjamin Hoover , Hendrik Strobelt

With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods mostly rely on predefined scripts, lacking dynamism and…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xiaolu Hou , Bing Ma , Jiaxiang Cheng , Xuhua Ren , Kai Yu , Wenyue Li , Tianxiang Zheng , Qinglin Lu

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, most approaches rely…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Wei-Hua Li , Cheng Sun , Chu-Song Chen

In this paper, we introduce NarrativePlay, a novel system that allows users to role-play a fictional character and interact with other characters in narratives such as novels in an immersive environment. We leverage Large Language Models…

计算与语言 · 计算机科学 2023-10-04 Runcong Zhao , Wenjia Zhang , Jiazheng Li , Lixing Zhu , Yanran Li , Yulan He , Lin Gui

Product designers often begin their design process with handcrafted personas. While personas are intended to ground design decisions in consumer preferences, they often fall short in practice by remaining abstract, expensive to produce, and…

Visual storytelling is an emerging field that combines images and narratives to create engaging and contextually rich stories. Despite its potential, generating coherent and emotionally resonant visual stories remains challenging due to the…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Xiaochuan Lin , Xiangyong Chen

We introduce AvatarForge, a framework for generating animatable 3D human avatars from text or image inputs using AI-driven procedural generation. While diffusion-based methods have made strides in general 3D object generation, they struggle…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

Collecting human-chatbot dialogues typically demands substantial manual effort and is time-consuming, which limits and poses challenges for research on conversational AI. In this work, we propose DialogueForge - a framework for generating…

计算与语言 · 计算机科学 2025-07-22 Ruizhe Zhu , Hao Zhu , Yaxuan Li , Syang Zhou , Shijing Cai , Malgorzata Lazuka , Elliott Ash

Simulation is an invaluable tool for developing and evaluating controllers for self-driving cars. Current simulation frameworks are driven by highly-specialist domain specific languages, and so a natural language interface would greatly…

人工智能 · 计算机科学 2023-10-27 Antonio Valerio Miceli-Barone , Alex Lascarides , Craig Innes

This study explores the effectiveness of Large Language Models (LLMs) in creating personalized "mirror stories" that reflect and resonate with individual readers' identities, addressing the significant lack of diversity in literature. We…

计算与语言 · 计算机科学 2024-09-25 Sarfaroz Yunusov , Hamza Sidat , Ali Emami

Every individual carries a unique and personal life story shaped by their memories and experiences. However, these memories are often scattered and difficult to organize into a coherent narrative, a challenge that defines the task of…

人机交互 · 计算机科学 2025-09-30 Shayan Talaei , Meijin Li , Kanu Grover , James Kent Hippler , Diyi Yang , Amin Saberi

Storyboarding is an established method for designing user experiences. Generative AI can support this process by helping designers quickly create visual narratives. However, existing tools only focus on accurate text-to-image generation.…

人机交互 · 计算机科学 2024-07-11 Zhaohui Liang , Xiaoyu Zhang , Kevin Ma , Zhao Liu , Xipei Ren , Kosa Goucher-Lambert , Can Liu

This study presents RoleCraft-GLM, an innovative framework aimed at enhancing personalized role-playing with Large Language Models (LLMs). RoleCraft-GLM addresses the key issue of lacking personalized interactions in conversational AI, and…

计算与语言 · 计算机科学 2024-04-05 Meiling Tao , Xuechen Liang , Tianyu Shi , Lei Yu , Yiting Xie

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating…

人机交互 · 计算机科学 2025-09-19 Taesoo Kim , Yongsik Jo , Hyunmin Song , Taehwan Kim

Large Language Models (LLMs) excel in handling general knowledge tasks, yet they struggle with user-specific personalization, such as understanding individual emotions, writing styles, and preferences. Personalized Large Language Models…

人工智能 · 计算机科学 2025-09-23 Jiahong Liu , Zexuan Qiu , Zhongyang Li , Quanyu Dai , Wenhao Yu , Jieming Zhu , Minda Hu , Menglin Yang , Tat-Seng Chua , Irwin King
‹ 上一页 1 2 3 10 下一页 ›