中文
相关论文

相关论文: Auditing Sabotage Bench: A Benchmark for Detecting…

200 篇论文

Legal judgments may contain errors due to the complexity of case circumstances and the abstract nature of legal concepts, while existing appellate review mechanisms face efficiency pressures from a surge in case volumes. Although current…

计算与语言 · 计算机科学 2026-02-02 Yifei Li , Richong Zhang , Wanyu Tu , Zhijie Nie , Haokun Luo , Chuantao Yin , Pengchong Li

The rapid progress in Automated Program Repair (APR) has been fueled by advances in AI, particularly large language models (LLMs) and agent-based systems. SWE-Bench is a benchmark designed to evaluate repair systems using real issues mined…

软件工程 · 计算机科学 2026-02-05 Matias Martinez , Xavier Franch

Modern coding scaffolds turn LLMs into capable software agents, but their ability to follow scaffold-specified instructions remains under-examined, especially when constraints are heterogeneous and persist across interactions. To fill this…

Large Language Model (LLM) agents increasingly act through external tools, making their safety contingent on tool-call workflows rather than text generation alone. While recent benchmarks evaluate agents across diverse environments and risk…

软件工程 · 计算机科学 2026-03-20 Xuan Chen , Lu Yan , Ruqi Zhang , Xiangyu Zhang

Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by…

LLM-based software engineering assistants fail not only by producing incorrect outputs, but also by allocating trust to the wrong artifact when code, documentation, and tests disagree. Existing evaluations focus mainly on downstream…

软件工程 · 计算机科学 2026-04-07 Noshin Ulfat , Ahsanul Ameen Sabit , Soneya Binta Hossain

This study introduces an evaluation benchmark for middle school algebra to be used in artificial intelligence(AI) based educational platforms. The goal is to support the design of AI systems that can enhance learner conceptual understanding…

人机交互 · 计算机科学 2024-12-06 Otero Nancy , Druga Stefania , Lan Andrew

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehensive industry-standard benchmark for assessing AI-product…

计算机与社会 · 计算机科学 2025-04-22 Shaona Ghosh , Heather Frase , Adina Williams , Sarah Luger , Paul Röttger , Fazl Barez , Sean McGregor , Kenneth Fricklas , Mala Kumar , Quentin Feuillade--Montixi , Kurt Bollacker , Felix Friedrich , Ryan Tsang , Bertie Vidgen , Alicia Parrish , Chris Knotz , Eleonora Presani , Jonathan Bennion , Marisa Ferrara Boston , Mike Kuniavsky , Wiebke Hutiri , James Ezick , Malek Ben Salem , Rajat Sahay , Sujata Goswami , Usman Gohar , Ben Huang , Supheakmungkol Sarin , Elie Alhajjar , Canyu Chen , Roman Eng , Kashyap Ramanandula Manjusha , Virendra Mehta , Eileen Long , Murali Emani , Natan Vidra , Benjamin Rukundo , Abolfazl Shahbazi , Kongtao Chen , Rajat Ghosh , Vithursan Thangarasa , Pierre Peigné , Abhinav Singh , Max Bartolo , Satyapriya Krishna , Mubashara Akhtar , Rafael Gold , Cody Coleman , Luis Oala , Vassil Tashev , Joseph Marvin Imperial , Amy Russ , Sasidhar Kunapuli , Nicolas Miailhe , Julien Delaunay , Bhaktipriya Radharapu , Rajat Shinde , Tuesday , Debojyoti Dutta , Declan Grabb , Ananya Gangavarapu , Saurav Sahay , Agasthya Gangavarapu , Patrick Schramowski , Stephen Singam , Tom David , Xudong Han , Priyanka Mary Mammen , Tarunima Prabhakar , Venelin Kovatchev , Rebecca Weiss , Ahmed Ahmed , Kelvin N. Manyeki , Sandeep Madireddy , Foutse Khomh , Fedor Zhdanov , Joachim Baumann , Nina Vasan , Xianjun Yang , Carlos Mougn , Jibin Rajan Varghese , Hussain Chinoy , Seshakrishna Jitendar , Manil Maskey , Claire V. Hardgrove , Tianhao Li , Aakash Gupta , Emil Joswin , Yifan Mai , Shachi H Kumar , Cigdem Patlak , Kevin Lu , Vincent Alessi , Sree Bhargavi Balija , Chenhe Gu , Robert Sullivan , James Gealy , Matt Lavrisa , James Goel , Peter Mattson , Percy Liang , Joaquin Vanschoren

Existing approaches to monitoring AI agents rely on supervised evaluation: human-written rules or LLM-based judges that check for known failure modes. However, novel misbehaviors may fall outside predefined categories entirely and LLM-based…

人工智能 · 计算机科学 2026-04-14 Ziqian Zhong , Shashwat Saxena , Aditi Raghunathan

Recent advances in frontier large language models have enabled code review agents that operate in open-ended, reasoning-intensive settings. However, the lack of standardized benchmarks and granular evaluation protocols makes it difficult to…

软件工程 · 计算机科学 2026-03-13 Kristen Pereira , Neelabh Sinha , Rajat Ghosh , Debojyoti Dutta

We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was run. The motivation came from a familiar frustration: two papers will report results on…

机器学习 · 计算机科学 2026-05-21 Mahdi Naser Moghadasi , Faezeh Ghaderi

We introduce MacroBench, a code-first benchmark that evaluates whether LLMs can synthesize reusable browser-automation programs (macros) from natural-language goals by reading HTML/DOM and emitting Selenium. MacroBench instantiates seven…

软件工程 · 计算机科学 2025-10-10 Hyunjun Kim , Sejong Kim

Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation contracts: artifacts whose later use is constrained by the external system that produced…

软件工程 · 计算机科学 2026-05-19 Jicheng Wang , Yifeng He , Zili Wang , Hanwen Xing , Arkaprava De , Hao Chen

Large language models can generate fluent peer reviews, yet their assessments often lack sufficient critical rigor when substantive issues are subtle and distributed across a paper. In this paper, we introduce PaperAudit-Bench, which…

计算与语言 · 计算机科学 2026-01-29 Songjun Tu , Yiwen Ma , Jiahao Lin , Qichao Zhang , Xiangyuan Lan , Junfeng. Li , Nan Xu , Linjing Li , Dongbin Zhao

As AI models progress beyond simple chatbots into more complex workflows, we draw ever closer to the event horizon beyond which AI systems will be utilized in autonomous, self-maintaining feedback loops. Any autonomous AI system will depend…

人工智能 · 计算机科学 2026-03-06 Benjamin Feuer , Lucas Rosenblatt , Oussama Elachqar

As Large Language Models (LLMs) advance, Machine-Generated Texts (MGTs) have become increasingly fluent, high-quality, and informative. Existing wide-range MGT detectors are designed to identify MGTs to prevent the spread of plagiarism and…

密码学与安全 · 计算机科学 2025-03-14 Jingyi Zheng , Junfeng Wang , Zhen Sun , Wenhan Dong , Yule Liu , Xinlei He

Benchmark Data Contamination (BDC)-the inclusion of benchmark testing samples in the training set-has raised increasing concerns in Large Language Model (LLM) evaluation, leading to falsely inflated performance estimates and undermining…

人工智能 · 计算机科学 2025-03-21 Yifan Sun , Han Wang , Dongbai Li , Gang Wang , Huan Zhang

Large Language Models (LLMs) have been suggested for use in automated vulnerability repair, but benchmarks showing they can consistently identify security-related bugs are lacking. We thus develop SecLLMHolmes, a fully automated evaluation…

密码学与安全 · 计算机科学 2024-07-25 Saad Ullah , Mingji Han , Saurabh Pujar , Hammond Pearce , Ayse Coskun , Gianluca Stringhini

Language models have improved by orders of magnitude with the recent emergence of Transformer-based Large Language Models (LLMs). LLMs have demonstrated their ability to generate natural code that is highly similar to code written by…

软件工程 · 计算机科学 2024-04-24 Aidan Z. H. Yang , Sophia Kolak , Vincent J. Hellendoorn , Ruben Martins , Claire Le Goues

The success of Large Language Models (LLMs) relies heavily on the huge amount of pre-training data learned in the pre-training phase. The opacity of the pre-training process and the training data causes the results of many benchmark tests…

计算与语言 · 计算机科学 2025-03-03 Shiwen Ni , Xiangtao Kong , Chengming Li , Xiping Hu , Ruifeng Xu , Jia Zhu , Min Yang
‹ 上一页 1 8 9 10 下一页 ›