English
Related papers

Related papers: PROXIMA: A Reliability Scoring Framework for Proxy…

200 papers

Estimating uncertainty for AI agents in real-world multi-turn tool-using interaction with humans is difficult because failures are often triggered by sparse critical episodes (e.g., looping, incoherent tool use, or user-agent…

Artificial Intelligence · Computer Science 2026-02-13 Sina Tayebati , Divake Kumar , Nastaran Darabi , Davide Ettori , Ranganath Krishnan , Amit Ranjan Trivedi

Purchase data from retail chains provide proxy measures of private household expenditure on items that are the most troublesome to collect in the traditional expenditure survey. Due to the sheer amount of proxy data, the bias due to…

Econometrics · Economics 2019-06-27 Li-Chun Zhang

Sensitivity analysis is widely used to assess the robustness of causal conclusions in observational studies, yet its interaction with the structure of measured covariates is often overlooked. When latent confounders cannot be directly…

Methodology · Statistics 2026-02-17 Abhinandan Dalal , Iris Horng , Yang Feng , Dylan S. Small

Flexible estimation of heterogeneous treatment effects lies at the heart of many statistical challenges, such as personalized medicine and optimal resource allocation. In this paper, we develop a general class of two-step algorithms for…

Machine Learning · Statistics 2020-08-07 Xinkun Nie , Stefan Wager

A/B testing has become the cornerstone of decision-making in online markets, guiding how platforms launch new features, optimize pricing strategies, and improve user experience. In practice, we typically employ the pairwise $t$-test to…

Machine Learning · Statistics 2025-10-29 Junpeng Gong , Chunkai Wang , Hao Li , Jinyong Ma , Haoxuan Li , Xu He

Structured extraction with LLMs fails in production not because models lack understanding, but because output formatting is unreliable across models and prompts. A prompt that returns clean JSON on GPT-4 may produce fenced, prose-wrapped,…

Machine Learning · Computer Science 2026-01-13 Varun Kotte

Deploying large language model (LLM)-driven conversational agents in enterprise settings requires prompts that are simultaneously correct at launch and resilient to the non-deterministic behavioral drift that characterizes production LLM…

Artificial Intelligence · Computer Science 2026-05-18 Keshava Chaitanya , Jahnavi Gundakaram

A common concern when a policymaker draws causal inferences from and makes decisions based on observational data is that the measured covariates are insufficiently rich to account for all sources of confounding, i.e., the standard no…

Methodology · Statistics 2023-10-25 Tao Shen , Yifan Cui

The growing reliance on artificial intelligence in safety- and security-critical applications is raising concerns about the robustness of neural networks to erroneous or adversarial input. Certification is a methodology for ensuring model…

Machine Learning · Computer Science 2026-05-01 Anton Björklund , Mykola Zaitsev , Paolo Morettin , Marta Kwiatkowska

We consider the problem of indirect comparison, where a treatment arm of interest is absent by design in one randomized controlled trial but available in the other. The former is the target trial, and the latter is the source trial. The…

Methodology · Statistics 2025-06-06 Zehao Su , Helene C. W. Rytgaard , Henrik Ravn , Frank Eriksson

We study ratio metrics in A/B testing at the presence of correlation among observations coming from the same user and provides practical guidance especially when two metrics contradict each other. We propose new estimating methods to…

Applications · Statistics 2020-07-24 Keyu Nie , Yinfei Kong , Ted Tao Yuan , Pauline Berry Burke

We propose an interpretable AI-assisted reliability diagnostic framework for parameterized root-finding schemes based on kNN-LLE proxy stability profiling and multi-horizon early prediction. The approach augments a numerical solver with a…

Numerical Analysis · Mathematics 2026-03-19 Bruno Carpentieri , Andrei Velichko , Mudassir Shams , Paola Lecca

Remote and webcam-based eye tracking in multi-line reading suffers from various noise factors and layout ambiguity, precisely where real-time reading support needs reliable, per-fixation line assignment. Prior work largely addresses this…

Neurons and Cognition · Quantitative Biology 2026-05-04 Franziska Kaltenberger , Wei-Ling Chen , Enkeleda Thaqi , Enkelejda Kasneci

Reasoning with LLMs increasingly unfolds inside a broader verification loop. Internally, systems use cheap checks, such as self-consistency or proxy rewards, which we call weak verification. Externally, users inspect outputs and steer the…

Machine Learning · Computer Science 2026-02-20 Shayan Kiyani , Sima Noorani , George Pappas , Hamed Hassani

Mechanistic simulations typically assume fixed ontologies: variables, causal relationships, and resolution policies are static. This assumption fails when the true causal structure is contested or unidentifiable-as in antimicrobial…

Computational Physics · Physics 2026-04-02 Kinson Vernet

Large language models (LLMs) are increasingly used for automated tutoring, but their reliability in structured symbolic domains remains unclear. We study step-level feedback for propositional logic proofs, which require precise symbolic…

Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding…

Computation and Language · Computer Science 2025-07-02 Alexander Hoyle , Lorena Calvo-Bartolomé , Jordan Boyd-Graber , Philip Resnik

In industry, online randomized controlled experiment (a.k.a. A/B experiment) is a standard approach to measure the impact of a causal change. These experiments have small treatment effect to reduce the potential blast radius. As a result,…

Econometrics · Economics 2025-05-29 Tanmoy Das , Dohyeon Lee , Arnab Sinha

We present ProbReach, a tool for verifying probabilistic reachability for stochastic hybrid systems, i.e., computing the probability that the system reaches an unsafe region of the state space. In particular, ProbReach will compute an…

Logic in Computer Science · Computer Science 2015-03-06 Fedor Shmarov , Paolo Zuliani

A/B testing, or online experiment is a standard business strategy to compare a new product with an old one in pharmaceutical, technological, and traditional industries. Major challenges arise in online experiments of two-sided marketplace…

Machine Learning · Computer Science 2022-11-04 Chengchun Shi , Xiaoyu Wang , Shikai Luo , Hongtu Zhu , Jieping Ye , Rui Song