中文
相关论文

相关论文: Certainty-Validity: A Diagnostic Framework for Dis…

200 篇论文

Large Language and Vision-Language Models (LLMs/VLMs) are increasingly used in safety-critical applications, yet their opaque decision-making complicates risk assessment and reliability. Uncertainty quantification (UQ) helps assess…

Discovering potential failures of an autonomous system is important prior to deployment. Falsification-based methods are often used to assess the safety of such systems, but the cost of running many accurate simulation can be high. The…

机器人学 · 计算机科学 2023-10-03 Marc R. Schlichting , Nina V. Boord , Anthony L. Corso , Mykel J. Kochenderfer

Large language models (LLMs) often produce confident but incorrect answers in settings where abstention would be safer. Standard evaluation protocols, however, require a response and do not account for how confidence should guide decisions…

计算与语言 · 计算机科学 2026-04-06 Sean Wu , Fredrik K. Gustafsson , Edward Phillips , Boyan Gao , Anshul Thakur , David A. Clifton

Most deep anomaly detection models are based on learning normality from datasets due to the difficulty of defining abnormality by its diverse and inconsistent nature. Therefore, it has been a common practice to learn normality under the…

机器学习 · 计算机科学 2023-09-19 Minkyung Kim , Jongmin Yu , Junsik Kim , Tae-Hyun Oh , Jun Kyun Choi

Debiased machine learning is a meta algorithm based on bias correction and sample splitting to calculate confidence intervals for functionals, i.e. scalar summaries, of machine learning algorithms. For example, an analyst may desire the…

机器学习 · 统计学 2022-10-25 Victor Chernozhukov , Whitney K. Newey , Rahul Singh

Data-driven visual odometry (VO) is a critical subroutine for autonomous edge robotics, and recent progress in the field has produced highly accurate point predictions in complex environments. However, emerging autonomous edge robotics…

计算机视觉与模式识别 · 计算机科学 2023-03-07 Alex C. Stutts , Danilo Erricolo , Theja Tulabandhula , Amit Ranjan Trivedi

Deep neural networks, despite their high accuracy, often exhibit poor confidence calibration, limiting their reliability in high-stakes applications. Current ad-hoc confidence calibration methods attempt to fix this during training but face…

机器学习 · 计算机科学 2026-04-15 Sandra Gómez-Gálvez , Tobias Olenyi , Gillian Dobbie , Katerina Taškova

Robust reinforcement learning methods typically focus on suppressing unreliable experiences or corrupted rewards, but they lack the ability to reason about the reliability of their own learning process. As a result, such methods often…

机器学习 · 计算机科学 2026-03-24 Zhipeng Zhang , Xiongfei Su , Kai Li

Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relationships. This vulnerability leads to causal confusion, where models exploit dataset biases as…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Jiacheng Tang , Zhiyuan Zhou , Zhuolin He , Jia Zhang , Kai Zhang , Jian Pu

Model deficiency that results from incomplete training data is a form of structural blindness that leads to costly errors, oftentimes with high confidence. During the training of classification tasks, underrepresented class-conditional…

机器学习 · 计算机科学 2021-02-09 Bruno Abrahao , Zheng Wang , Haider Ahmed , Yuchen Zhu

In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonly referred to as calibration, is essential for risk-aware decision-making. In regression a…

机器学习 · 计算机科学 2026-04-23 Jelke Wibbeke , Nico Schönfisch , Sebastian Rohjans , Andreas Rauh

While deep learning offers tremendous promise for scientific and medical imaging, any failures and hallucinations (predictions that do not coincide with reality) are hard to pinpoint and can have serious downstream consequences. Uncertainty…

图像与视频处理 · 电气工程与系统科学 2026-05-26 Cassandra Tong Ye , Shamus Li , Tyler King , Kristina Monakhova

Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what the evidence supports. We study this failure mode as overcommitment control and…

计算与语言 · 计算机科学 2026-05-19 Tianyi Huang , Samuel Xu , Jason Tansong Dang , Samuel Yan , Kimberley Yin

Continuous perception, the ability to integrate visual observations over time in a continuous stream fashion, is essential for robust real-world understanding, yet remains largely untested in current multimodal models. We introduce…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Zeyu Wang , Zhenzhen Weng , Serena Yeung-Levy

Weighted Majority Voting (WMV) is a well-known optimal decision rule for collective decision making, given the probability of sources to provide accurate information (trustworthiness). However, in reality, the trustworthiness is not a known…

人工智能 · 计算机科学 2024-07-02 Shaojie Bai , Dongxia Wang , Tim Muller , Peng Cheng , Jiming Chen

Prediction credibility measures, in the form of confidence intervals or probability distributions, are fundamental in statistics and machine learning to characterize model robustness, detect out-of-distribution samples (outliers), and…

机器学习 · 计算机科学 2020-11-26 Luiz F. O. Chamon , Santiago Paternain , Alejandro Ribeiro

A fundamental challenge in robust visual-inertial odometry (VIO) is to dynamically assess the reliability of sensor measurements. This assessment is crucial for properly weighting the contribution of each measurement to the state estimate.…

机器人学 · 计算机科学 2025-10-03 Seungwon Choi , Donggyu Park , Seo-Yeon Hwang , Tae-Wan Kim

The incompleteness of positive labels and the presence of many unlabelled instances are common problems in binary classification applications such as in review helpfulness classification. Various studies from the classification literature…

信息检索 · 计算机科学 2020-08-17 Xi Wang , Iadh Ounis , Craig Macdonald

The focus in deep learning research has been mostly to push the limits of prediction accuracy. However, this was often achieved at the cost of increased complexity, raising concerns about the interpretability and the reliability of deep…

计算机视觉与模式识别 · 计算机科学 2020-06-08 Abdelrahman Eldesokey , Michael Felsberg , Karl Holmquist , Mikael Persson

A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) use the same…

计算机视觉与模式识别 · 计算机科学 2020-12-21 Robert Geirhos , Kristof Meding , Felix A. Wichmann