English
Related papers

Related papers: PROXIMA: A Reliability Scoring Framework for Proxy…

200 papers

Reliability is an essential measure of how closely observed scores represent latent scores (reflecting constructs), assuming some latent variable measurement model. We present a general theoretical framework of reliability, placing emphasis…

Methodology · Statistics 2024-10-29 Yang Liu , Jolynn Pek , Alberto Maydeu-Olivares

Despite their remarkable success, large language models (LLMs) have shown limited ability on safety-critical code tasks such as vulnerability detection. Typically, static analysis (SA) tools, like CodeQL, CodeGuru Security, etc., are used…

Cryptography and Security · Computer Science 2025-09-15 Ira Ceka , Feitong Qiao , Anik Dey , Aastha Valecha , Gail Kaiser , Baishakhi Ray

A standard assumption for causal inference about the joint effects of time-varying treatment is that one has measured sufficient covariates to ensure that within covariate strata, subjects are exchangeable across observed treatment values,…

Methodology · Statistics 2022-08-04 Andrew Ying , Wang Miao , Xu Shi , Eric J. Tchetgen Tchetgen

Modern media firms require automated and efficient methods to identify content that is most engaging and appealing to users. Leveraging a large-scale dataset from Upworthy (a news publisher), which includes 17,681 headline A/B tests, we…

Machine Learning · Computer Science 2024-11-27 Zikun Ye , Hema Yoganarasimhan , Yufeng Zheng

Promptable segmentation models (e.g., the Segment Anything Models) enable generalizable, zero-shot segmentation across diverse domains. Although predictions are deterministic for a fixed image-prompt pair, the robustness of these models to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Elodie Germani , Krystel Nyangoh-Timoh , Pierre Jannin , John S H Baxter

Assessing response quality to instructions in language models is vital but challenging due to the complexity of human language across different contexts. This complexity often results in ambiguous or inconsistent interpretations, making…

Online propaganda detection pipelines expose measurable privacy risks at multiple stages including data collection, feature extraction, and model inference. We conduct a structured analysis of $162$ peer-reviewed studies and formalize the…

Cryptography and Security · Computer Science 2026-04-21 Dhiman Goswami , Al Nahian Bin Emran , Md Hasan Ullah Sadi , Sanchari Das

Online experimentation, also known as A/B testing, is the gold standard for measuring product impacts and making business decisions in the tech industry. The validity and utility of experiments, however, hinge on unbiasedness and sufficient…

Applications · Statistics 2020-12-17 Min Liu , Jialiang Mao , Kang Kang

Conformal prediction is a framework for uncertainty quantification that constructs prediction sets for previously unseen data, guaranteeing coverage of the true label with a specified probability. However, the efficiency of these prediction…

Machine Learning · Computer Science 2026-01-06 Erfan Hajihashemi , Yanning Shen

Local Interpretable Model-Agnostic Explanations (LIME) is a popular method to perform interpretability of any kind of Machine Learning (ML) model. It explains one ML prediction at a time, by learning a simple linear model around the…

Machine Learning · Computer Science 2022-02-09 Giorgio Visani , Enrico Bagli , Federico Chesani

Uncertainty quantification is crucial in safety-critical systems, where decisions must be made under uncertainty. In particular, we consider the problem of online uncertainty quantification, where data points arrive sequentially. Online…

Machine Learning · Computer Science 2026-04-21 Junyoung Yang , Kyungmin Kim , Sangdon Park

Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Benchmarks for these agents must both reliably compare models and yield on-policy training…

Leaderboard scores on public benchmarks have been steadily rising and converging, with many frontier language models now separated by only marginal differences. However, these scores often fail to match users' day to day experience, because…

Artificial Intelligence · Computer Science 2026-02-05 Yiliang Song , Hongjun An , Jiangong Xiao , Haofei Zhao , Jiawei Shao , Xuelong Li

A critical challenge in recommender systems is to establish reliable relationships between offline and online metrics that predict real-world performance. Motivated by recent advances in Pareto front approximation, we introduce a pragmatic…

Information Retrieval · Computer Science 2025-07-15 Timo Wilm , Philipp Normann

Information retrieval systems, such as online marketplaces, news feeds, and search engines, are ubiquitous in today's digital society. They facilitate information discovery by ranking retrieved items on predicted relevance, i.e. likelihood…

Econometrics · Economics 2022-05-16 Rina Friedberg , Karthik Rajkumar , Jialiang Mao , Qian Yao , YinYin Yu , Min Liu

System prompt configuration can make the difference between near-total phishing blindness and near-perfect detection in LLM email agents. We present PhishNChips, a study of 11 models under 10 prompt strategies, showing that prompt-model…

Cryptography and Security · Computer Science 2026-03-27 Ron Litvak

In many situations it is either impossible or impractical to develop and evaluate agents entirely on the target domain on which they will be deployed. This is particularly true in robotics, where doing experiments on hardware is much more…

Machine Learning · Computer Science 2021-10-08 Anthony Courchesne , Andrea Censi , Liam Paull

Confidence calibration is central to providing accurate and interpretable uncertainty estimates, especially under safety-critical scenarios. However, we find that existing calibration algorithms often overlook the issue of *proximity bias*,…

Machine Learning · Computer Science 2024-03-19 Miao Xiong , Ailin Deng , Pang Wei Koh , Jiaying Wu , Shen Li , Jianqing Xu , Bryan Hooi

It is often desirable for Large Language Models (LLMs) to capture multiple objectives when providing a response. In document-grounded response generation, for example, agent responses are expected to be relevant to a user's query while also…

Computation and Language · Computer Science 2024-03-05 Keshav Ramji , Young-Suk Lee , Ramón Fernandez Astudillo , Md Arafat Sultan , Tahira Naseem , Asim Munawar , Radu Florian , Salim Roukos

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a…

Econometrics · Economics 2025-01-14 Ke Sun , Linglong Kong , Hongtu Zhu , Chengchun Shi
‹ Prev 1 4 5 6 7 8 10 Next ›