English
Related papers

Related papers: PROXIMA: A Reliability Scoring Framework for Proxy…

200 papers

Verbal confidence elicitation is widely used to extract uncertainty estimates from LLMs. We tested whether seven instruction-tuned open-weight models (3-9B parameters, four families) produce verbalised confidence that meets minimal validity…

Computation and Language · Computer Science 2026-04-27 Jon-Paul Cacioli

Current evaluation of mathematical reasoning in language models relies primarily on answer accuracy, potentially masking fundamental failures in logical computation. We introduce a diagnostic framework that distinguishes genuine…

Computation and Language · Computer Science 2025-12-02 Subramanyam Sahoo , Vinija Jain , Saanidhya Vats , Siddharth Mohapatra , Rui Min , Aman Chadha , Divya Chaudhary

Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on…

Computation and Language · Computer Science 2026-03-20 Trishita Dhara , Siddhesh Sheth

To provide rigorous uncertainty quantification for online learning models, we develop a framework for constructing uncertainty sets that provably control risk -- such as coverage of confidence intervals, false negative rate, or F1 score --…

Machine Learning · Computer Science 2023-01-30 Shai Feldman , Liran Ringel , Stephen Bates , Yaniv Romano

Mobile devices that connect to the Internet via cellular networks are rapidly becoming the primary medium for accessing Web content. Cellular service providers (CSPs) commonly deploy Web proxies and other middleboxes for security,…

Networking and Internet Architecture · Computer Science 2015-11-17 Huijing Zhang , David Choffnes

Within process mining, a relevant activity is conformance checking. Such activity consists of establishing the extent to which actual executions of a process conform the expected behavior of a reference model. Current techniques focus on…

Artificial Intelligence · Computer Science 2022-01-25 Andrea Burattin

Sequential recommender systems have achieved steady gains in offline accuracy, yet it remains unclear how close current models are to the intrinsic accuracy limit imposed by the data. A reliable, model-agnostic estimate of this ceiling…

Information Retrieval · Computer Science 2026-04-15 En Xu , Jingtao Ding , Yong Li

Aligning Large Language Models (LLMs) with high-stakes medical standards remains a significant challenge, primarily due to the dissonance between coarse-grained preference signals and the complex, multi-dimensional nature of clinical…

Artificial Intelligence · Computer Science 2026-04-10 He Geng , Yangmin Huang , Lixian Lai , Qianyun Du , Hui Chu , Zhiyang He , Jiaxue Hu , Xiaodong Tao

Online conformal prediction has demonstrated its capability to construct a prediction set for each incoming data point that covers the true label with a predetermined probability. To cope with potential distribution shift, multi-model…

Machine Learning · Computer Science 2025-10-14 Erfan Hajihashemi , Yanning Shen

While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-trained language…

Machine Learning · Computer Science 2026-04-02 Cai Zhou , Zekai Wang , Menghua Wu , Qianyu Julie Zhu , Flora C. Shi , Chenyu Wang , Ashia Wilson , Tommi Jaakkola , Stephen Bates

Foundation model reliability assessment typically requires thousands of evaluation examples, making it computationally expensive and time-consuming for real-world deployment. We introduce microprobe, a novel approach that achieves…

Artificial Intelligence · Computer Science 2025-12-25 Aayam Bansal , Ishaan Gangwani

Online reviews enable consumers to engage with companies and provide important feedback. Due to the complexity of the high-dimensional text, these reviews are often simplified as a single numerical score, e.g., ratings or sentiment scores.…

Machine Learning · Computer Science 2022-01-04 Lu Cheng , Ruocheng Guo , Huan Liu

Estimation of crossed random effects models commonly requires computational costs that grow faster than linearly in the sample size $N$, often as fast as $\Omega(N^{3/2})$, making them unsuitable for large data sets. For non-Gaussian…

Methodology · Statistics 2025-05-01 Ruggero Bellio , Swarnadip Ghosh , Art B. Owen , Cristiano Varin

Contextual online decision-making problems with constraints appear in a wide range of real-world applications, such as adaptive experimental design under safety constraints, personalized recommendation with resource limits, and dynamic…

Machine Learning · Statistics 2025-05-23 Haichen Hu , David Simchi-Levi , Navid Azizan

Post-hoc explanations provide transparency and are essential for guiding model optimization, such as prompt engineering and data sanitation. However, applying model-agnostic techniques to Large Language Models (LLMs) is hindered by…

Machine Learning · Computer Science 2026-04-13 Junhao Liu , Haonan Yu , Zhenyu Yan , Xin Zhang

Large language models (LLMs) are increasingly used to support question answering and decision-making in high-stakes, domain-specific settings such as natural hazard response and infrastructure planning, where effective answers must convey…

Computation and Language · Computer Science 2026-02-11 Homaira Huda Shomee , Rochana Chaturvedi , Yangxinyu Xie , Tanwi Mallick

Recent works proposed test-time alignment methods that rely on a small aligned model as a proxy that guides the generation of a larger base (unaligned) model. The implicit reward approach skews the large model distribution, whereas the…

Computation and Language · Computer Science 2026-04-21 Ayoub Hammal , Pierre Zweigenbaum , Caio Corro

Consistency under paraphrase, the property that semantically equivalent prompts yield identical predictions, is increasingly used as a proxy for reliability when deploying medical vision-language models (VLMs). We show this proxy is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Binesh Sadanandan , Vahid Behzadan

Many online experiments exhibit dependence between users and items. For example, in online advertising, observations that have a user or an ad in common are likely to be associated. Because of this, even in experiments involving millions of…

Methodology · Statistics 2017-10-26 Eytan Bakshy , Dean Eckles

Trajectory prediction, the task of forecasting future agent behavior from past data, is central to safe and efficient autonomous driving. A diverse set of methods (e.g., rule-based or learned with different architectures and datasets) have…

Robotics · Computer Science 2025-02-21 Alex Tong , Apoorva Sharma , Sushant Veer , Marco Pavone , Heng Yang