中文
相关论文

相关论文: From Pixels to Paths: A Multi-Agent Framework for …

200 篇论文

Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a…

计算与语言 · 计算机科学 2025-08-14 Lingjie Jiang , Shaohan Huang , Xun Wu , Yixia Li , Dongdong Zhang , Furu Wei

We present Zero-Painter, a novel training-free framework for layout-conditional text-to-image synthesis that facilitates the creation of detailed and controlled imagery from textual prompts. Our method utilizes object masks and individual…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Marianna Ohanyan , Hayk Manukyan , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Generating engaging, accurate short-form videos from scientific papers is challenging due to content complexity and the gap between expert authors and readers. Existing end-to-end methods often suffer from factual inaccuracies and visual…

计算与语言 · 计算机科学 2025-04-29 Jong Inn Park , Maanas Taneja , Qianwen Wang , Dongyeop Kang

Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactive medical image analysis, their application in dermatology…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yize Liu , Siyuan Yan , Ming Hu , Lie Ju , Xieji Li , Feilong Tang , Wei Feng , Zongyuan Ge

Exploratory analysis of high-dimensional data relies on embedding the data into a low-dimensional space (typically 2D or 3D), based on which visualization plot is produced to uncover meaningful structures and to communicate geometric and…

人机交互 · 计算机科学 2026-04-23 Burak Susam , Tingting Mu

Caption quality has emerged as a critical bottleneck in training high-quality text-to-image (T2I) and text-to-video (T2V) generative models. While visual language models (VLMs) are commonly deployed to generate captions from visual data,…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Varun Ananth , Baqiao Liu , Haoran Cai

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xin Gu , Ming Li , Libo Zhang , Fan Chen , Longyin Wen , Tiejian Luo , Sijie Zhu

Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding,…

计算机视觉与模式识别 · 计算机科学 2025-10-15 A. Alfarano , L. Venturoli , D. Negueruela del Castillo

We present VizGenie, a self-improving, agentic framework that advances scientific visualization through large language model (LLM) by orchestrating of a collection of domain-specific and dynamically generated modules. Users initially access…

We propose a weakly-supervised approach for conditional image generation of complex scenes where a user has fine control over objects appearing in the scene. We exploit sparse semantic maps to control object shapes and classes, as well as…

计算机视觉与模式识别 · 计算机科学 2020-11-23 Dario Pavllo , Aurelien Lucchi , Thomas Hofmann

Computational thematic analysis is rapidly emerging as a method of using large text corpora to understand the lived experience of people across the continuum of health care: patients, practitioners, and everyone in between. However, many…

人机交互 · 计算机科学 2024-12-20 Luka Ugaya Mazza , Plinio Morita , James R. Wallace

Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to serialize this intent into text, a bottleneck that compresses…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Yaofang Liu , Kangning Cui , Meng Chu , Zhaoqing Li , Suiyun Zhang , Jean-Michel Morel , Xiaodong Cun , Haoxuan Che , Rui Liu , Raymond H. Chan

Converting process sketches into executable simulation models remains a major bottleneck in process systems engineering, requiring substantial manual effort and simulator-specific expertise. Recent advances in generative AI have improved…

软件工程 · 计算机科学 2026-03-27 Abdullah Bahamdan , Emma Pajak , John D. Hedengren , Antonio del Rio Chanona

Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing document agents mainly transform papers into static artifacts such as summaries,…

计算与语言 · 计算机科学 2026-03-31 Dasen Dai , Biao Wu , Meng Fang , Wenhao Wang

This paper proposes an image-to-painting translation method that generates vivid and realistic painting artworks with controllable styles. Different from previous image-to-image translation methods that formulate the translation as…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Zhengxia Zou , Tianyang Shi , Shuang Qiu , Yi Yuan , Zhenwei Shi

Visual compatibility is critical for fashion analysis, yet is missing in existing fashion image synthesis systems. In this paper, we propose to explicitly model visual compatibility through fashion image inpainting. To this end, we present…

计算机视觉与模式识别 · 计算机科学 2019-04-25 Xintong Han , Zuxuan Wu , Weilin Huang , Matthew R. Scott , Larry S. Davis

To bridge the gap between artists and non-specialists, we present a unified framework, Neural-Polyptych, to facilitate the creation of expansive, high-resolution paintings by seamlessly incorporating interactive hand-drawn sketches with…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Yiming Zhao , Dewen Guo , Zhouhui Lian , Yue Gao , Jianhong Han , Jie Feng , Guoping Wang , Bingfeng Zhou , Sheng Li

Multimodal reasoning models often produce fluent answers supported by seemingly coherent rationales. Existing benchmarks evaluate only final-answer correctness. They do not support atomic visual entailment verification of intermediate…

人工智能 · 计算机科学 2026-03-25 Saleem Ahmed , Srirangaraj Setlur , Venu Govindaraju

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Kangcheng Zhou , Jun Jiang , Qing Zhang , Shuang Zheng , Qingli Li , Shugong Xu

Hand-drawn sketches are a natural and efficient medium for capturing and conveying ideas. Despite significant advancements in controllable natural image generation, translating freehand sketches into structured, machine-readable diagrams…

人工智能 · 计算机科学 2025-08-05 Cheng Tan , Qi Chen , Jingxuan Wei , Gaowei Wu , Zhangyang Gao , Siyuan Li , Bihui Yu , Ruifeng Guo , Stan Z. Li