English
Related papers

Related papers: MUSE: A Multi-agent Framework for Unconstrained St…

200 papers

Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present…

Machine Learning · Computer Science 2026-03-04 Zhongxi Wang , Yueqian Lin , Jingyang Zhang , Hai Helen Li , Yiran Chen

Writing compelling fiction is a multifaceted process combining elements such as crafting a plot, developing interesting characters, and using evocative language. While large language models (LLMs) show promise for story writing, they…

Computation and Language · Computer Science 2025-03-17 Fantine Huot , Reinald Kim Amplayo , Jennimaria Palomaki , Alice Shoshana Jakobovits , Elizabeth Clark , Mirella Lapata

Metacognition, defined as the awareness and regulation of one's cognitive processes, is central to human adaptability in unknown situations. In contrast, current autonomous agents often struggle in novel environments due to their limited…

Machine Learning · Computer Science 2025-11-18 Rodolfo Valiente , Praveen K. Pilly

Large language model (LLM) agents rely on reusable skills to solve complex tasks. However, existing skill creation approaches treat skills as isolated and static artifacts, limiting their reusability, reliability, and long-term improvement.…

Artificial Intelligence · Computer Science 2026-05-27 Huawei Lin , Peng Li , Jie Song , Fuxin Jiang , Tieying Zhang

Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design. Existing benchmarks focus primarily on generating single-part CAD models and evaluate them…

Artificial Intelligence · Computer Science 2026-05-28 Xiaoyu Dong , Zhi Li , Xiao-Ming Wu

Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Kiymet Akdemir , Tahira Kazimi , Pinar Yanardag

Images evoke emotions that profoundly influence perception, often prioritized over content. Current Image Emotional Synthesis (IES) approaches artificially separate generation and editing tasks, creating inefficiencies and limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Yingjie Xia , Xi Wang , Jinglei Shi , Vicky Kalogeiton , Jian Yang

Digital storytelling, essential in entertainment, education, and marketing, faces challenges in production scalability and flexibility. The StoryAgent framework, introduced in this paper, utilizes Large Language Models and generative tools…

Computation and Language · Computer Science 2024-06-24 Samuel S. Sohn , Danrui Li , Sen Zhang , Che-Jui Chang , Mubbasir Kapadia

Story visualization is the transformation of narrative elements into image sequences. While existing research has primarily focused on visual contextual coherence, the deeper narrative essence of stories often remains overlooked. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Seungkwon Kim , GyuTae Park , Sangyeon Kim , Seung-Hun Nam

Accurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Mihee Lee , Samuel S. Sohn , Seonghyeon Moon , Sejong Yoon , Mubbasir Kapadia , Vladimir Pavlovic

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining roles, and…

This work addresses the problem of sensing the world: how to learn a multimodal representation of a reinforcement learning agent's environment that allows the execution of tasks under incomplete perceptual conditions. To address such…

Machine Learning · Computer Science 2022-02-01 Miguel Vasco , Hang Yin , Francisco S. Melo , Ana Paiva

Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on predicting isolated mental state labels, lacking structured annotations of complex…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Xiaoyu Yuan , Niklas Heikkala , Tiina Törmänen , Hanna Järvenoja , Guoying Zhao , Haoyu Chen

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, most approaches rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Wei-Hua Li , Cheng Sun , Chu-Song Chen

Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an open challenge, where prior works mainly decompose this task into separate stages, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bo Gao , Chang Liu , Yuyang Miao , Siyuan Ma , Ser-Nam Lim

In this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Donggeun Lim , Jinseok Bae , Inwoo Hwang , Seungmin Lee , Hwanhee Lee , Young Min Kim

Long-horizon tasks that require sustained reasoning and multiple tool interactions remain challenging for LLM agents: small errors compound across steps, and even state-of-the-art models often hallucinate or lose coherence. We identify…

Artificial Intelligence · Computer Science 2025-10-13 Guangya Wan , Mingyang Ling , Xiaoqi Ren , Rujun Han , Sheng Li , Zizhao Zhang

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits expressivity. Recent end-to-end methods reduce latency and…

Generating coherent and communicative visual sequences, such as image sequences and videos, remains a significant challenge for current multimodal systems. Despite advances in visual quality and the integration of world knowledge, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Chutian Meng , Fan Ma , Chi Zhang , Jiaxu Miao , Yi Yang , Yueting Zhuang

Modern language agents must operate over long-horizon, multi-turn histories, yet deploying such agents with Small Language Models (SLMs) remains fundamentally difficult. Full-context prompting causes context overflow, flat retrieval exposes…

Multiagent Systems · Computer Science 2026-05-06 Jiayi Chen , Yingcong Li , Guiling Wang