中文
相关论文

相关论文: CANVAS: Continuity-Aware Narratives via Visual Age…

200 篇论文

Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Shuowei Li , Yuming Zhao , Parth Bhalerao , Oana Ignat

Text-to-image generation has advanced rapidly, but existing models still struggle with faithfully composing multiple objects and preserving their attributes in complex scenes. We propose coDrawAgents, an interactive multi-agent dialogue…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Chunhan Li , Qifeng Wu , Jia-Hui Pan , Ka-Hei Hui , Jingyu Hu , Yuming Jiang , Bin Sheng , Xihui Liu , Wenjuan Gong , Zhengzhe Liu

Embodied agents are expected to perform object navigation in dynamic, open-world environments. However, existing approaches typically rely on static trajectories and a fixed set of object categories during training, overlooking the…

机器人学 · 计算机科学 2026-04-07 Ming-Ming Yu , Fei Zhu , Wenzhuo Liu , Yirong Yang , Qunbo Wang , Wenjun Wu , Jing Liu

Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpainting, or interpolation, treating spatio-temporal control as a set of isolated problems. We…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Minghong Cai , Qiulin Wang , Zongli Ye , Wenze Liu , Quande Liu , Weicai Ye , Xintao Wang , Pengfei Wan , Kun Gai , Xiangyu Yue

Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates, evaluating them, and using feedback to guide future search. However, existing methods are…

The ability for computational agents to reason about the high-level content of real world scene images is important for many applications. Existing attempts at addressing the problem of complex scene understanding lack representational…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Zachary A. Daniels , Dimitris N. Metaxas

Scene graphs -- objects as nodes and visual relationships as edges -- describe the whereabouts and interactions of the things and stuff in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all…

计算机视觉与模式识别 · 计算机科学 2019-08-12 Long Chen , Hanwang Zhang , Jun Xiao , Xiangnan He , Shiliang Pu , Shih-Fu Chang

Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle complex queries requiring multi-step reasoning, agentic VRAG systems interleave reasoning with…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Yucheng Shen , Jiulong Wu , Jizhou Huang , Dawei Yin , Lingyong Yan , Min Cao

Deep reinforcement learning has been applied successfully to solve various real-world problems and the number of its applications in the multi-agent settings has been increasing. Multi-agent learning distinctly poses significant challenges…

机器学习 · 计算机科学 2021-02-24 Ngoc Duy Nguyen , Thanh Thi Nguyen , Doug Creighton , Saeid Nahavandi

Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Xinhang Gao , Junlin Guan , Shuhan Luo , Wenzhuo Li , Guanghuan Tan , Jiacheng Wang

Despite the substantial progress in recent years, the image captioning techniques are still far from being perfect.Sentences produced by existing methods, e.g. those based on RNNs, are often overly rigid and lacking in variability. This…

计算机视觉与模式识别 · 计算机科学 2017-08-14 Bo Dai , Sanja Fidler , Raquel Urtasun , Dahua Lin

Story rewriting aims to adapt existing narratives to diverse reader preferences while preserving plot consistency and narrative coherence. Unlike conventional work on style transfer, we argue that effective story rewriting demands…

计算与语言 · 计算机科学 2026-05-28 Hanwen Cui , Yuting Mei , Yuhang Fu , Dingyi Yang , Qin Jin

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Han Zhang , Wanting Jiang , Tomasz Kornuta , Tian Zheng , Vidya Murali

Storyboarding is an established method for designing user experiences. Generative AI can support this process by helping designers quickly create visual narratives. However, existing tools only focus on accurate text-to-image generation.…

人机交互 · 计算机科学 2024-07-11 Zhaohui Liang , Xiaoyu Zhang , Kevin Ma , Zhao Liu , Xipei Ren , Kosa Goucher-Lambert , Can Liu

Visual reasoning -- the ability to interpret the visual world -- is crucial for embodied agents that operate within three-dimensional scenes. Progress in AI has led to vision and language models capable of answering questions from images.…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Damiano Marsili , Rohun Agrawal , Yisong Yue , Georgia Gkioxari

Video production workflows offer a rich and demanding arena for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this…

密码学与安全 · 计算机科学 2026-05-28 Zongheng Cao , Yi Zheng , Rui Song , Xinyu Hu

Visual localization remains challenging in dynamic environments where fluctuating lighting, adverse weather, and moving objects disrupt appearance cues. Despite advances in feature representation, current absolute pose regression methods…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Zhongtao Tian , Wenhao Huang , Zhidong Chen , Xiao Wei Sun

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inaccurate dialogue…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Ayan Banerjee , Josep Llados , Umapada Pal , Anjan Dutta

Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with coherent scene…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Hanwen Shen , Jiajie Lu , Yupeng Cao , Xiaonan Yang

Motion forecasting for agents in autonomous driving is highly challenging due to the numerous possibilities for each agent's next action and their complex interactions in space and time. In real applications, motion forecasting takes place…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Nan Song , Bozhou Zhang , Xiatian Zhu , Li Zhang