中文
相关论文

相关论文: APEX: Academic Poster Editing Agentic Expert

200 篇论文

As generative models achieve unprecedented visual quality, the gold standard for image evaluation remains traditional feature-distribution metrics (e.g., FID). However, these metrics are provably hindered by the closed-vocabulary bottleneck…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Caterina Gallegati , Monica Bianchini , Franco Scarselli , Vittorio Murino , Barbara Toniella Corradini

Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation or…

人工智能 · 计算机科学 2026-04-01 Yinuo Liu , Zi Qian , Heng Zhou , Jiahao Zhang , Yajie Zhang , Zhihang Li , Mengyu Zhou , Erchao Zhao , Xiaoxi Jiang , Guanjun Jiang

Interactive documents help readers engage with complex ideas through dynamic visualization, interactive animations, and exploratory interfaces. However, creating such documents remains costly, as it requires both domain expertise and web…

人机交互 · 计算机科学 2026-03-31 Yinghao Tang , Yupeng Xie , Yingchaojie Feng , Tingfeng Lan , Jiale Lao , Yue Cheng , Wei Chen

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Yang Ye , Xianyi He , Zongjian Li , Bin Lin , Shenghai Yuan , Zhiyuan Yan , Bohan Hou , Li Yuan

Text-to-image generative models have achieved remarkable visual quality but still struggle with compositionality$-$accurately capturing object relationships, attribute bindings, and fine-grained details in prompts. A key limitation is that…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Arman Zarei , Jiacheng Pan , Matthew Gwilliam , Soheil Feizi , Zhenheng Yang

The paper-to-video task converts a research paper into a structured video abstract, distilling key concepts, methods, and conclusions into an accessible, well-organized format. While state-of-the-art video generation models demonstrate…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Jingwei Liu , Ling Yang , Hao Luo , Fan Wang , Hongyan Li , Mengdi Wang

Integration of artificial intelligent (AI) agents in higher education is transforming teaching, learning and administrative processes. Although existing AI agents effectively support individual tasks, their implementation remains fragmented…

As more and more academic papers are being submitted to conferences and journals, evaluating all these papers by professionals is time-consuming and can cause inequality due to the personal factors of the reviewers. In this paper, in order…

计算与语言 · 计算机科学 2018-05-11 Pengcheng Yang , Xu Sun , Wei Li , Shuming Ma

Facial expression image editing requires fine-grained control to strictly preserve human identity and background while precisely manipulating expression. However, existing editing benchmarks primarily focus on general scenarios, lacking…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Fengjian Xue , Xuecheng Wu , Heli Sun , Yunyun Shi , Shi Chen , Liangyu Fu , Jinheng Xie , Dingkang Yang , Hao Wang , Junxiao Xue , Liang He

We present iPoster, an interactive layout generation framework that empowers users to guide content-aware poster layout design by specifying flexible constraints. iPoster enables users to specify partial intentions within the intention…

人机交互 · 计算机科学 2026-04-01 Xudong Zhou , Jinyuan Liang , Qiuyi Guo , Guozheng Li

Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm drains computational and annotation resources, inflates…

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Xiangyu Zhao , Peiyuan Zhang , Kexian Tang , Xiaorong Zhu , Hao Li , Wenhao Chai , Zicheng Zhang , Renqiu Xia , Guangtao Zhai , Junchi Yan , Hua Yang , Xue Yang , Haodong Duan

Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models is notably absent. In this paper, we first introduce the…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Jian Ma , Xujie Zhu , Zihao Pan , Qirong Peng , Xu Guo , Chen Chen , Haonan Lu

Autonomous agents are moving beyond simple retrieval tasks to become economic actors that invoke APIs, sequence workflows, and make real-time decisions. As this shift accelerates, API providers need request-level monetization with…

密码学与安全 · 计算机科学 2026-04-03 Mohd Safwan Uddin , Mohammed Mouzam , Mohammed Imran , Syed Badar Uddin Faizan

AI agents have drawn increasing attention mostly on their ability to perceive environments, understand tasks, and autonomously achieve goals. To advance research on AI agents in mobile scenarios, we introduce the Android Multi-annotation…

人机交互 · 计算机科学 2025-08-15 Yuxiang Chai , Siyuan Huang , Yazhe Niu , Han Xiao , Liang Liu , Dingyu Zhang , Shuai Ren , Hongsheng Li

Scientific discoveries must be communicated clearly to realize their full potential. Without effective communication, even the most groundbreaking findings risk being overlooked or misunderstood. The primary way scientists communicate their…

We present ActuBench, a multi-agent LLM pipeline for the automated generation and evaluation of advanced actuarial assessment items aligned with the International Actuarial Association (IAA) Education Syllabus. The pipeline separates four…

人工智能 · 计算机科学 2026-04-23 Jan-Philipp Schmidt

Despite the remarkable capabilities of text-to-image (T2I) generation models, real-world applications often demand fine-grained, iterative image editing that existing methods struggle to provide. Key challenges include granular instruction…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Zihan Liang , Jiahao Sun , Haoran Ma

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Juntong Wang , Jiarui Wang , Huiyu Duan , Jiaxiang Kang , Guangtao Zhai , Xiongkuo Min

Content-aware visual-textual presentation layout aims at arranging spatial space on the given canvas for pre-defined elements, including text, logo, and underlay, which is a key to automatic template-free creative graphic design. In…

计算机视觉与模式识别 · 计算机科学 2023-03-29 HsiaoYuan Hsu , Xiangteng He , Yuxin Peng , Hao Kong , Qing Zhang