中文
相关论文

相关论文: TechImage-Bench: Rubric-Based Evaluation for Techn…

200 篇论文

Deep Research Systems (DRS) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. However, how to rigorously evaluate these systems remains under-explored. Existing deep-research…

计算与语言 · 计算机科学 2026-02-02 Ruizhe Li , Mingxuan Du , Benfeng Xu , Chiwei Zhu , Xiaorui Wang , Zhendong Mao

With the rapid advancement of large multimodal models (LMMs), recent text-to-image (T2I) models can generate high-quality images and demonstrate great alignment to short prompts. However, they still struggle to effectively understand and…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Juntong Wang , Huiyu Duan , Jiarui Wang , Ziheng Jia , Guangtao Zhai , Xiongkuo Min

Reward modeling lies at the core of reinforcement learning from human feedback (RLHF), yet most existing reward models rely on scalar or pairwise judgments that fail to capture the multifaceted nature of human preferences. Recent studies…

计算与语言 · 计算机科学 2026-02-04 Tianci Liu , Ran Xu , Tony Yu , Ilgee Hong , Carl Yang , Tuo Zhao , Haoyu Wang

Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubrics authored by domain experts. While effective, rubric authoring is time-consuming and…

计算与语言 · 计算机科学 2026-04-15 Xin Xia , Nejla Yuruk , Yun Wang , Xiaoming Zhai

Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions, and content moderation. In these settings, models are typically trained on historical…

机器学习 · 计算机科学 2026-05-21 Calvin Isley , Johann D. Gaebler , Sharad Goel

The rapid advancement of generative AI has revolutionized image creation, enabling high-quality synthesis from text prompts while raising critical challenges for media authenticity. We present Ai-GenBench, a novel benchmark designed to…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Lorenzo Pellegrini , Davide Cozzolino , Serafino Pandolfini , Davide Maltoni , Matteo Ferrara , Luisa Verdoliva , Marco Prati , Marco Ramilli

Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Ruoxuan Zhang , Jidong Gao , Bin Wen , Hongxia Xie , Chenming Zhang , Hong-Han Shuai , Wen-Huang Cheng

Recently, diffusion-based deep generative models (e.g., Stable Diffusion) have shown impressive results in text-to-image synthesis. However, current text-to-image models often require multiple passes of prompt engineering by humans in order…

计算与语言 · 计算机科学 2023-11-14 Tingfeng Cao , Chengyu Wang , Bingyan Liu , Ziheng Wu , Jinhui Zhu , Jun Huang

The rapid progress of text-to-image diffusion models raises significant concerns regarding the unauthorized reproduction of trademarked content. While prior work targets general concepts (e.g., styles, celebrities), it fails to address…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Dawid Malarz , Filip Manjak , Maciej Zięba , Przemysław Spurek , Artur Kasymov

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of figures. Unlike…

计算与语言 · 计算机科学 2026-03-13 Michiko Yoshitake , Yuta Suzuki , Ryo Igarashi , Yoshitaka Ushiku , Keisuke Nagato

DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,800 evaluation instances across six programming languages and six task categories derived from real…

Current text-to-image generative models struggle to accurately represent object states (e.g., "a table without a bottle," "an empty tumbler"). In this work, we first design a fully-automatic pipeline to generate high-quality synthetic data…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Tianle Chen , Chaitanya Chakka , Deepti Ghadiyaram

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Junjie Yang , Yuhao Yan , Gang Wu , Yuxuan Wang , Ruoyu Liang , Xinjie Jiang , Xiang Wan , Fenglei Fan , Yongquan Zhang , Feiwei Qin , Changmiao Wang

As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical models has become a critical gap. Computational mechanics,…

机器学习 · 计算机科学 2025-12-25 Saeed Mohammadzadeh , Erfan Hamdi , Joel Shor , Emma Lejeune

Evaluating and comparing text-to-image models is a challenging problem. Significant advances in the field have recently been made, piquing interest of various industrial sectors. As a consequence, a gold standard in the field should cover a…

计算机视觉与模式识别 · 计算机科学 2022-12-16 Federico A. Galatolo , Mario G. C. A. Cimino , Edoardo Cogotti

Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structure of the feedback signal: unstructured critique helps inconsistently across models, and…

机器学习 · 计算机科学 2026-05-19 Will LeVine , Brendan Evers , Sam Saltwick , Abhay Venkatesh

The rapid development and reduced barriers to entry for Text-to-Image (T2I) models have raised concerns about the biases in their outputs, but existing research lacks a holistic definition and evaluation framework of biases, limiting the…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Hanjun Luo , Ziye Deng , Ruizhe Chen , Zuozhu Liu

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both methodologies are…

计算与语言 · 计算机科学 2026-05-26 Russell Yang , Ruishi Chen , Pierce Kelaita , Riya Ranjan , Sibo Ma , Charles Dickens , Matthew Guillod , Megan Ma , Julian Nyarko

Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols either emphasize…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Tianyang Han , Junhao Su , Junjie Hu , Peizhen Yang , Hengyu Shi , Junfeng Luo , Jialin Gao

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. However, existing frameworks typically treat rubrics only as…

计算与语言 · 计算机科学 2026-05-11 Jiachen Yu , Zhihao Xu , Junjie Wang , Yujiu Yang