中文
相关论文

相关论文: VisionCreator: A Native Visual-Generation Agentic …

200 篇论文

Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocation capabilities. Yet, existing agents remain at the mercy of underlying black-box image…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Junyan Ye , Jun He , Zilong Huang , Dongzhi Jiang , Xuan Yang , Rui Chen , Weijia Li

Large language models (LLMs) have recently enabled coding agents capable of generating, executing, and revising visualization code. However, existing models often fail in practical workflows due to limited language coverage, unreliable…

Recent agentic systems demonstrate that large language models can generate scientific visualizations from natural language. However, reliability remains a major limitation: systems may execute invalid operations, introduce subtle but…

人机交互 · 计算机科学 2026-03-27 Nathaniel Gorski , Shusen Liu , Bei Wang

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

The proliferation of generative AI has transformed creative workflows, yet current systems face critical challenges in controllability and content protection. We propose a novel multi-agent framework that addresses both limitations through…

多智能体系统 · 计算机科学 2026-01-21 Haris Khan , Sadia Asif

Agentic tasks, which require multi-step problem solving with autonomy, tool use, and adaptive reasoning, are becoming increasingly central to the advancement of NLP and AI. However, existing instruction data lacks tool interaction, and…

Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We present Visual Foresight Planning (ForeAct), a general and…

机器人学 · 计算机科学 2026-02-16 Zhuoyang Zhang , Shang Yang , Qinghao Hu , Luke J. Huang , James Hou , Yufei Sun , Yao Lu , Song Han

The automated generation of interactive 3D cities is a critical challenge with broad applications in autonomous driving, virtual reality, and embodied intelligence. While recent advances in generative models and procedural techniques have…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Zishan Liu , Zecong Tang , RuoCheng Wu , Xinzhe Zheng , Jingyu Hu , Ka-Hei Hui , Haoran Xie , Bo Dai , Zhengzhe Liu

A key challenge in training Vision-Language Model (VLM) agents, compared to Language Model (LLM) agents, lies in the shift from textual states to complex visual observations. This transition introduces partial observability and demands…

Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for rare entities and long-tail facts. Most existing retrieval-augmented generation (RAG) methods…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Zhuohong Chen , Zhenxian Wu , Yunyao Yu , Hangrui Xu , Zirui Liao , Zhifang Liu , Xiangwen Deng , Pen Jiao , Haoqian Wang

Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Qinglin Zeng , Kaitong Cai , Ruiqi Chen , Qinhan Lv , Keze Wang

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Mingde Yao , Zhiyuan You , King-Man Tam , Menglu Wang , Tianfan Xue

In the accelerating era of human-instructed visual content creation, diffusion models have demonstrated remarkable generative potential. Yet their deployment is constrained by a dual bottleneck: semantic ambiguity in diverse prompts and the…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Jie Qin , Jie Wu , Weifeng Chen , Yueming Lyu

Common knowledge indicates that the process of constructing image datasets usually depends on the time-intensive and inefficient method of manual collection and annotation. Large models offer a solution via data generation. Nonetheless,…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Haoran Sun , Haoyu Bian , Shaoning Zeng , Yunbo Rao , Xu Xu , Lin Mei , Jianping Gou

Creating lifelike virtual agents capable of interacting with their environments is a longstanding goal in computer graphics. This paper addresses the challenge of generating natural head rotations, a critical aspect of believable agent…

图形学 · 计算机科学 2025-02-17 Juyeong Hwang , Seong-Eun Hong , Hyeongyeop Kang

Scaling Vision-Language-Action (VLA) models requires massive datasets that are both semantically coherent and physically feasible. However, existing scene generation methods often lack context-awareness, making it difficult to synthesize…

机器人学 · 计算机科学 2026-04-13 Yaru Liu , Ao-bo Wang , Nanyang Ye

The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to…

Considering the importance of building a good Visual Dialog (VD) Questioner, many researchers study the topic under a Q-Bot-A-Bot image-guessing game setting, where the Questioner needs to raise a series of questions to collect information…

计算与语言 · 计算机科学 2021-09-07 Duo Zheng , Zipeng Xu , Fandong Meng , Xiaojie Wang , Jiaan Wang , Jie Zhou

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

人工智能 · 计算机科学 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Chi-Pin Huang , Yueh-Hua Wu , Min-Hung Chen , Yu-Chiang Frank Wang , Fu-En Yang