中文
相关论文

相关论文: Accurate Failure Prediction in Agents Does Not Imp…

200 篇论文

Multi-agent LLM workflows -- systems composed of multiple role-specific LLM calls -- often outperform single-prompt baselines, but they remain difficult to debug and refine. Failures can originate from subtle errors in intermediate outputs…

计算与语言 · 计算机科学 2026-05-19 Kazuki Kawamura , Satoshi Waki , Kei Tateno

Verifying LLM-generated systems code is hard: bugs are prevalent, formal specifications are missing, and safety contracts are encoded implicitly at call sites rather than enforced at function boundaries. We propose agentic model checking, a…

软件工程 · 计算机科学 2026-05-21 Youcheng Sun , Jiawen Liu , Daniel Kroening , Jason Xue

Social media platforms mediate how billions form opinions and engage with public discourse. As autonomous AI agents increasingly participate in these spaces, understanding their behavioral fidelity becomes critical for platform governance…

计算与语言 · 计算机科学 2026-04-23 Ljubisa Bojic , Alexander Felfernig , Bojana Dinic , Velibor Ilic , Achim Rettinger , Vera Mevorah , Damian Trilling

When evaluating the performance of clinical machine learning models, one must consider the deployment population. When the population of patients with observed labels is only a subset of the deployment population (label selection), standard…

机器学习 · 计算机科学 2022-09-20 Conor K. Corbin , Michael Baiocchi , Jonathan H. Chen

Large Language Models are increasingly proposed as cognitive components for robotic systems, yet their opaque decision processes make it difficult to explain success or failure in closed-loop embodied tasks. Following an empirical AI…

人工智能 · 计算机科学 2026-05-20 Oussama Zenkri , Oliver Brock

[Context:] The acceptance of candidate patches in automated program repair has been typically based on testing oracles. Testing requires typically a costly process of building the application while ML models can be used to quickly classify…

软件工程 · 计算机科学 2025-04-10 Maria Camporese , Fabio Massacci

Large Language Models (LLMs) are increasingly being used to simulate human-like decision making in agent-based financial market models (ABMs). As models become more powerful and accessible, researchers can now incorporate individual LLM…

机器学习 · 计算机科学 2025-01-29 Alicia Vidler , Toby Walsh

The enhanced capabilities of LLM-based agents come with an emergency for model planning and tool-use abilities. Attributing to helpful-harmless trade-off from LLM alignment, agents typically also inherit the flaw of "over-refusal", which is…

计算与语言 · 计算机科学 2026-02-05 Xinyue Wang , Yuanhe Zhang , Zhengshuo Gong , Haoran Gao , Fanyu Meng , Zhenhong Zhou , Li Sun , Yang Liu , Sen Su

LLM agents increasingly perform end-to-end ML engineering tasks where success is judged by a single scalar test metric. This creates a structural vulnerability: an agent can increase the reported score by compromising the evaluation…

人工智能 · 计算机科学 2026-03-13 Yonas Atinafu , Robin Cohen

LLM prompting is widely used for naturally stated tasks, yet it is unreliable it may succeed on a few test cases but fail at deployment time. We study performance prediction: given a program, either symbolic (e.g. Python) or a prompt…

机器学习 · 计算机科学 2026-05-22 Chengqi Zheng , Keya Hu , Shuzhi Liu , Tao Wu , Kevin Ellis , Yewen Pu

The prevalent deployment of Large Language Model agents such as OpenClaw unlocks potential in real-world applications, while amplifying safety concerns. Among these concerns, the self-replication risk of LLM agents driven by objective…

人工智能 · 计算机科学 2026-04-02 Boxuan Zhang , Yi Yu , Jiaxuan Guo , Jing Shao

LLM-assisted defect discovery has a precision crisis: plausible-but-wrong reports overwhelm maintainers and degrade credibility for real findings. We present Refute-or-Promote, an inference-time reliability pattern combining Stratified…

密码学与安全 · 计算机科学 2026-04-22 Abhinav Agarwal

Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Prior work on adversarial robustness in NLP has emphasized…

机器学习 · 计算机科学 2025-10-16 Ivan Dubrovsky , Anastasia Orlova , Illarion Iov , Nina Gubina , Irena Gureeva , Alexey Zaytsev

Large Language Model (LLM) inference systems present significant challenges in statistical performance characterization due to dynamic workload variations, diverse hardware architectures, and complex interactions between model size, batch…

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models…

人工智能 · 计算机科学 2026-01-15 Logan Ritchie , Sushant Mehta , Nick Heiner , Mason Yu , Edwin Chen

LLM-based agents show promise for automating penetration testing, yet reported performance varies widely across systems and benchmarks. We analyze 28 LLM-based penetration testing systems and evaluate five representative implementations…

密码学与安全 · 计算机科学 2026-02-20 Gelei Deng , Yi Liu , Yuekang Li , Ruozhao Yang , Xiaofei Xie , Jie Zhang , Han Qiu , Tianwei Zhang

End-to-end LLM trading agents have moved quickly from research curiosity to a small ecosystem of named systems, including FinCon, FinMem, TradingAgents, FinAgent, QuantAgent, and FLAG-Trader. Several of these report headline Sharpe ratios…

计算工程、金融与科学 · 计算机科学 2026-05-19 Yuxuan Ye , Jun Han , Ao Hu , Juncheng Bu , Yiyi Chen , Liangjian Wen , Danilo Mandic , Danny Dongning Sun , Xu Yinghui , Zenglin Xu

Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that circumvent their safety guardrails. Existing evaluation frameworks typically report binary…

密码学与安全 · 计算机科学 2026-05-14 Zvi Topol

Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise validity. Despite rapid uptake of Large Language Models (LLMs), there is limited evidence on…

软件工程 · 计算机科学 2026-05-01 Gilberto Sussumu Hida , Danilo Monteiro Ribeiro , Erika Yahata

The transition of Large Language Models (LLMs) from exploratory tools to active "silicon subjects" in social science lacks extensive validation of operational validity. This study introduces Conditioned Comment Prediction (CCP), a task in…

计算与语言 · 计算机科学 2026-03-27 Nils Schwager , Simon Münker , Alistair Plum , Achim Rettinger