English
Related papers

Related papers: Agentic Confidence Calibration

200 papers

Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to facilitate AI deployment. However, hallucinations generated by LLMs-where outputs are inconsistent with…

Machine Learning · Computer Science 2025-07-23 Siyuan Liu , Wenjing Liu , Zhiwei Xu , Xin Wang , Bo Chen , Tao Li

We introduce the notion of heterogeneous calibration that applies a post-hoc model-agnostic transformation to model outputs for improving AUC performance on binary classification tasks. We consider overconfident models, whose performance is…

Machine Learning · Statistics 2022-02-11 David Durfee , Aman Gupta , Kinjal Basu

Current agentic AI benchmarks predominantly evaluate task completion accuracy, while overlooking critical enterprise requirements such as cost-efficiency, reliability, and operational stability. Through systematic analysis of 12 main…

Artificial Intelligence · Computer Science 2025-11-19 Sushant Mehta

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models…

Artificial Intelligence · Computer Science 2026-01-15 Logan Ritchie , Sushant Mehta , Nick Heiner , Mason Yu , Edwin Chen

Large language model (LLM)-based agents have shown promise in tackling complex tasks by interacting dynamically with the environment. Existing work primarily focuses on behavior cloning from expert demonstrations or preference learning…

Machine Learning · Computer Science 2025-05-30 Hanlin Wang , Jian Wang , Chak Tou Leong , Wenjie Li

AI is moving from domain-specific autonomy in closed, predictable settings to large-language-model-driven agents that plan and act in open, cross-organizational environments. As a result, the cybersecurity risk landscape is changing in…

Cryptography and Security · Computer Science 2026-02-03 Alsharif Abuadbba , Nazatul Sultan , Surya Nepal , Sanjay Jha

The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, existing static benchmarks are ill-equipped to address the dynamic nature of AI risks and…

Artificial Intelligence · Computer Science 2026-05-15 Yixu Wang , Xin Wang , Yang Yao , Xinyuan Li , Xibang Yang , Yan Teng , Xingjun Ma , Yingchun Wang

Recent proliferation of powerful AI systems has created a strong need for capabilities that help users to calibrate trust in those systems. As AI systems grow in scale, information required to evaluate their trustworthiness becomes less…

Human-Computer Interaction · Computer Science 2025-11-20 Scott T Steinmetz , Asmeret Naugle , Paul Schutte , Matt Sweitzer , Alex Washburne , Lisa Linville , Daniel Krofcheck , Michal Kucer , Samuel Myren

Large language models (LLMs) often suffer from hallucinations, posing significant challenges for real-world applications. Confidence calibration, as an effective indicator of hallucination, is thus essential to enhance the trustworthiness…

Computation and Language · Computer Science 2025-11-21 Caiqi Zhang , Ruihan Yang , Zhisong Zhang , Xinting Huang , Sen Yang , Dong Yu , Nigel Collier

As machine learning techniques become widely adopted in new domains, especially in safety-critical systems such as autonomous vehicles, it is crucial to provide accurate output uncertainty estimation. As a result, many approaches have been…

Machine Learning · Computer Science 2021-12-28 Sooyong Jang , Radoslav Ivanov , Insup Lee , James Weimer

The rapid advancement of AI, including Large Language Models, has propelled autonomous agents forward, accelerating the human-agent teaming (HAT) paradigm to leverage complementary strengths. However, HAT research remains fragmented, often…

Human-Computer Interaction · Computer Science 2025-04-16 Mengyao Wang , Jiayun Wu , Shuai Ma , Nuo Li , Peng Zhang , Ning Gu , Tun Lu

A fundamental limitation of current LLM-based AI agents is their inability to build differentiated trust among each other at the onset of an agent-to-agent dialogue. However, autonomous and interoperable trust establishment becomes…

Cryptography and Security · Computer Science 2026-05-28 Sandro Rodriguez Garzon , Awid Vaziry , Enis Mert Kuzu , Dennis Enrique Gehrmann , Buse Varkan , Alexander Gaballa , Axel Küpper

Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data…

Artificial Intelligence · Computer Science 2026-05-07 Chenglin Yang

The wide-spread adoption of representation learning technologies in clinical decision making strongly emphasizes the need for characterizing model reliability and enabling rigorous introspection of model behavior. While the former need is…

Machine Learning · Computer Science 2020-05-01 Jayaraman J. Thiagarajan , Prasanna Sattigeri , Deepta Rajan , Bindya Venkatesh

Warning: This paper contains content that may be inappropriate or offensive. AI agents have gained significant recent attention due to their autonomous tool usage capabilities and their integration in various real-world applications. This…

Artificial Intelligence · Computer Science 2025-06-24 Ninareh Mehrabi , Tharindu Kumarage , Kai-Wei Chang , Aram Galstyan , Rahul Gupta

Today, AI is being increasingly used to help human experts make decisions in high-stakes scenarios. In these scenarios, full automation is often undesirable, not only due to the significance of the outcome, but also because human experts…

Artificial Intelligence · Computer Science 2020-01-08 Yunfeng Zhang , Q. Vera Liao , Rachel K. E. Bellamy

The consensus strategies used in collaborative multi-agent systems (MAS) face notable challenges related to adaptability, scalability, and convergence certainties. These approaches, including structured workflows, debate models, and…

Multiagent Systems · Computer Science 2025-11-25 Rathin Chandra Shit , Sharmila Subudhi

Artificial intelligence is undergoing a structural transformation marked by the rise of agentic systems capable of open-ended action trajectories, generative representations and outputs, and evolving objectives. These properties introduce…

Artificial Intelligence · Computer Science 2026-03-06 Bowen Lou , Tian Lu , T. S. Raghu , Yingjie Zhang

Agentic artificial intelligence (AI) is an AI paradigm that can perceive the environment, reason over observations, and execute actions to achieve specific goals. Task-oriented communication supports agentic AI by transmitting only the…

Image and Video Processing · Electrical Eng. & Systems 2026-02-16 Sin-Yu Huang , Lele Wang , Vincent W. S. Wong

Agentic AI systems present both significant opportunities and novel risks due to their capacity for autonomous action, encompassing tasks such as code execution, internet interaction, and file modification. This poses considerable…

Artificial Intelligence · Computer Science 2025-12-30 Shaun Khoo , Jessica Foo , Roy Ka-Wei Lee