English
Related papers

Related papers: PROXIMA: A Reliability Scoring Framework for Proxy…

200 papers

Many digital platforms offer advertisers experimentation tools like Meta's Lift and A/B tests to optimize their ad campaigns. Lift tests compare outcomes between users eligible to see ads versus users in a no-ad control group. In contrast,…

General Economics · Economics 2025-09-01 Gordon Burtch , Robert Moakler , Brett R. Gordon , Poppy Zhang , Shawndra Hill

Conformal prediction provides a principled framework for uncertainty quantification with finite-sample coverage guarantees. While recent work has extended conformal prediction to online and sequential settings, existing methods typically…

Machine Learning · Statistics 2026-05-14 Eduardo Ochoa Rivera , Ambuj Tewari

We consider the problem of distance metric learning (DML), where the task is to learn an effective similarity measure between images. We revisit ProxyNCA and incorporate several enhancements. We find that low temperature scaling is a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Eu Wern Teh , Terrance DeVries , Graham W. Taylor

Companies offering web services routinely run randomized online experiments to estimate the causal impact associated with the adoption of new features and policies on key performance metrics of interest. These experiments are used to…

Methodology · Statistics 2023-07-13 Lorenzo Masoero , Doug Hains , James McQueen

Existing metric learning losses can be categorized into two classes: pair-based and proxy-based losses. The former class can leverage fine-grained semantic relations between data points, but slows convergence in general due to its high…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Sungyeon Kim , Dongwon Kim , Minsu Cho , Suha Kwak

Evaluating the robustness of LLMs to adversarial attacks is crucial for safe deployment, yet current red-teaming methods are often prohibitively expensive. We compare the ability of fast proxy metrics to predict the real-world robustness of…

Cryptography and Security · Computer Science 2025-02-18 Tim Beyer , Jan Schuchardt , Leo Schwinn , Stephan Günnemann

Existing high-dimensional statistical methods are largely established for analyzing individual-level data. In this work, we study estimation and inference for high-dimensional linear models where we only observe "proxy data", which include…

Methodology · Statistics 2022-01-12 Sai Li , T. Tony Cai , Hongzhe Li

During early stages of CPU design, benchmarks can only run on simulators to evaluate CPU performance. However, most big data benchmarks are too huge at code size scale, which causes them to be unable to finish running on simulators at an…

Performance · Computer Science 2023-09-20 Yikang Yang , Lei Wang , Jianfeng Zhan

For some kernel matrices, low-rank approximations can be quickly obtained via analytic techniques. One important class of analytic methods that has received attention in recent years is based on the use of proxy points. Accuracy analysis…

Numerical Analysis · Mathematics 2026-05-26 Mikhail Lepilov , Jianlin Xia

Evaluating the reliability of noisy quantum circuits is essential for implementing quantum algorithms on noisy quantum devices. However, current quantum hardware exhibits diverse noise mechanisms whose compounded effects make accurate and…

Quantum Physics · Physics 2026-02-23 Jindi Wu , Tianjie Hu , Qun Li

Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real world web use consists of long-horizon, multi-site workflows. Common web navigation tasks,…

Machine Learning · Computer Science 2026-04-29 Lawrence Keunho Jang , Jing Yu Koh , Daniel Fried , Ruslan Salakhutdinov

Recommender systems are central to online services, enabling users to navigate through massive amounts of content across various domains. However, their evaluation remains challenging due to the disconnect between offline metrics and online…

Information Retrieval · Computer Science 2026-04-14 Nicolas Bougie , Gian Maria Marconi , Xiaotong Ye , Narimasa Watanabe

While there exists a large amount of literature on the general challenges of and best practices for trustworthy online A/B testing, there are limited studies on sample size estimation, which plays a crucial role in trustworthy and efficient…

Methodology · Statistics 2023-08-21 Jing Zhou , Jiannan Lu , Anas Shallah

As the modern microservice architecture for cloud applications grows in popularity, cloud services are becoming increasingly complex and more vulnerable to misconfiguration and software bugs. Traditional approaches rely on expert input to…

Software Engineering · Computer Science 2026-05-21 Rohan Kumar , Jason Li , Zongshun Zhang , Syed Mohammad Qasim , Gianluca Stringhini , Ayse K. Coskun

Both in academic and industry-based research, online evaluation methods are seen as the golden standard for interactive applications like recommendation systems. Naturally, the reason for this is that we can directly measure utility metrics…

Information Retrieval · Computer Science 2022-09-20 Imad Aouali , Amine Benhalloum , Martin Bompaire , Benjamin Heymann , Olivier Jeunen , David Rohde , Otmane Sakhi , Flavian Vasile

Online controlled experiments, or A/B tests, are large-scale randomized trials in digital environments. This paper investigates the estimands of the difference-in-means estimator in these experiments, focusing on scenarios with repeated…

Methodology · Statistics 2024-11-12 Sebastian Ankargren , Mattias Frånberg , Mårten Schultzberg

We study offline model-based optimization to maximize a black-box objective function with a static dataset of designs and scores. These designs encompass a variety of domains, including materials, robots and DNA sequences. A common approach…

Computational Engineering, Finance, and Science · Computer Science 2023-10-11 Can Chen , Christopher Beckham , Zixuan Liu , Xue Liu , Christopher Pal

Online A/B testing plays a critical role in the high-tech industry to guide product development and accelerate innovation. It performs a null hypothesis statistical test to determine which variant is better. However, a typical A/B test…

Methodology · Statistics 2021-09-03 Miao Yu , Wenbin Lu , Rui Song

Who should we prioritize for treatment when causal effects cannot be estimated? In practice, organizations often rely on predictive proxies: ads are targeted using purchase probabilities, and retention incentives are allocated using…

Machine Learning · Statistics 2025-10-15 Carlos Fernández-Loría , Jorge Loría

Autonomous satellite servicing missions must execute close-range rendezvous under stringent safety and operational constraints while remaining computationally tractable for onboard use and robust to uncertainty in sensing, actuation, and…

Robotics · Computer Science 2026-02-16 Minduli C. Wijayatunga , Julian Guinane , Nathan D. Wallace , Xiaofeng Wu