中文
相关论文

相关论文: From Pixels to Paths: A Multi-Agent Framework for …

200 篇论文

Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorrect results. Existing benchmarks either rely on image-centric or subjective metrics…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Tong Zhang , Honglin Lin , Zhou Liu , Chong Chen , Wentao Zhang

Recent text-to-image (T2I) models have demonstrated impressive capabilities in photorealistic synthesis and instruction following. However, their reliability in knowledge-intensive settings remains largely unexplored. Unlike natural image…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Ran Zhao , Sheng Jin , Size Wu , Kang Liao , Zerui Gong , Zujin Guo , Yang Xiao , Wei Li

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Yulin Pan , Xiangteng He , Chaojie Mao , Zhen Han , Zeyinzi Jiang , Jingfeng Zhang , Yu Liu

In computational pathology, understanding and generation have evolved along disparate paths: advanced understanding models already exhibit diagnostic-level competence, whereas generative models largely simulate pixels. Progress remains…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Minghao Han , Yichen Liu , Yizhou Liu , Zizhi Chen , Jingqun Tang , Xuecheng Wu , Dingkang Yang , Lihua Zhang

Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layout editing ability…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Alessandro Fontanella , Petru-Daniel Tudosiu , Yongxin Yang , Shifeng Zhang , Sarah Parisot

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

In many science papers, "Figure 1" serves as the primary visual summary of the core research idea. These figures are visually simple yet conceptually rich, often requiring significant effort and iteration by human authors to get right,…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Yaohan Guan , Pristina Wang , Najim Dehak , Alan Yuille , Jieneng Chen , Daniel Khashabi

Visual design instructors often provide multi-modal feedback, mixing annotations with text. Prior theory emphasizes the importance of actionable feedback, where "actionability" lies on a spectrum--from surfacing relevant design concepts to…

人机交互 · 计算机科学 2026-03-06 Mingyi Li , Mengyi Chen , Sarah Luo , Yining Cao , Haijun Xia , Maitraye Das , Steven P. Dow , Jane L. E

High-quality scientific illustrations are essential for communicating complex scientific and technical concepts, yet existing automated systems remain limited in editability, stylistic controllability, and efficiency. We present…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Zhen Lin , Qiujie Xie , Minjun Zhu , Shichen Li , Qiyao Sun , Enhao Gu , Yiran Ding , Ke Sun , Fang Guo , Panzhong Lu , Zhiyuan Ning , Yixuan Weng , Yue Zhang

The rapid growth of scientific literature demands robust tools for automated survey-generation. However, current large language model (LLM)-based methods often lack in-depth analysis, structural coherence, and reliable citations. To address…

人工智能 · 计算机科学 2025-07-22 Xiaofeng Shi , Qian Kou , Yuduo Li , Ning Tang , Jinxin Xie , Longbin Yu , Songjing Wang , Hua Zhou

In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language instructions. However, an important yet underexplored research…

Scientific diagrams are vital tools for communicating structured knowledge across disciplines. However, they are often published as static raster images, losing symbolic semantics and limiting reuse. While Multimodal Large Language Models…

人工智能 · 计算机科学 2025-10-14 Zhiqing Cui , Jiahao Yuan , Hanqing Wang , Yanshu Li , Chenxu Du , Zhenglong Ding

Witnessed by the recent advancements on leveraging LLM for coding and multimodal understanding, we present WebGen-V, a new benchmark and framework for instruction-to-HTML generation that enhances both data quality and evaluation…

人工智能 · 计算机科学 2025-10-20 Kuang-Da Wang , Zhao Wang , Yotaro Shimose , Wei-Yao Wang , Shingo Takamatsu

Recent advances in image generation have achieved remarkable visual quality, while a fundamental challenge remains: Can image generation be controlled at the element level, enabling intuitive modifications such as adjusting shapes, altering…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Lanqing Guo , Xi Liu , Yufei Wang , Zhihao Li , Siyu Huang

Existing text-guided image editing methods primarily rely on end-to-end pixel-level inpainting paradigm. Despite its success in simple scenarios, this paradigm still significantly struggles with compositional editing tasks that require…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Jinghan Yu , Junhao Xiao , Chenyu Zhu , Jiaming Li , Jia Li , HanMing Deng , Xirui Wang , Guoli Jia , Jianjun Li , Xiang Bai , Bowen Zhou , Zhiyuan Ma

The exponential growth of scientific literature in PDF format necessitates advanced tools for efficient and accurate document understanding, summarization, and content optimization. Traditional methods fall short in handling complex layouts…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Kun Qian , Wenjie Li , Tianyu Sun , Wenhong Wang , Wenhan Luo

Scientists often explore and analyze large-scale scientific simulation data by leveraging two- and three-dimensional visualizations. The data and tasks can be complex and therefore best supported using myriad display technologies, from…

人机交互 · 计算机科学 2024-04-30 Thomas Marrinan , Madeleine Moeller , Alina Kanayinkal , Victor A. Mateevitsi , Michael E. Papka

Autonomous science agents built on large language models (LLMs) are increasingly used to generate hypotheses, design experiments, and produce reports. However, prior work mainly targets open-ended scientific problems with subjective outputs…

计算与语言 · 计算机科学 2026-03-24 Tianshu Zhang , Huan Sun

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning,…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Feng Han , Yibin Wang , Chenglin Li , Zheming Liang , Dianyi Wang , Yang Jiao , Zhipeng Wei , Chao Gong , Cheng Jin , Jingjing Chen , Jiaqi Wang

Real-world design tasks - such as picture book creation, film storyboard development using character sets, photo retouching, visual effects, and font transfer - are highly diverse and complex, requiring deep interpretation and extraction of…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Chen Liang , Lianghua Huang , Jingwu Fang , Huanzhang Dou , Wei Wang , Zhi-Fan Wu , Yupeng Shi , Junge Zhang , Xin Zhao , Yu Liu