English
Related papers

Related papers: Stylish Risk-Limiting Audits in Practice

200 papers

Testing autonomous robotic systems, such as self-driving cars and unmanned aerial vehicles, is challenging due to their interaction with highly unpredictable environments. A common practice is to first conduct simulation-based testing,…

Neural and Evolutionary Computing · Computer Science 2025-03-27 Dmytro Humeniuk , Foutse Khomh

Randomized smoothing (RS) has successfully been used to improve the robustness of predictions for deep neural networks (DNNs) by adding random noise to create multiple variations of an input, followed by deciding the consensus. To…

Machine Learning · Computer Science 2024-04-29 Emmanouil Seferis , Stefanos Kollias , Chih-Hong Cheng

We transform the randomness of LLMs into precise assurances using an actuator at the API interface that applies a user-defined risk constraint in finite samples via Conformal Risk Control (CRC). This label-free and model-agnostic actuator…

Methodology · Statistics 2025-09-30 Lingyou Pang , Lei Huang , Jianyu Lin , Tianyu Wang , Alexander Aue , Carey E. Priebe

This paper explores the use of clustering methods and machine learning algorithms, including Natural Language Processing (NLP), to identify and classify problems identified in credit risk models through textual information contained in…

Machine Learning · Computer Science 2023-06-05 Szymon Lis , Mariusz Kubkowski , Olimpia Borkowska , Dobromił Serwa , Jarosław Kurpanik

Accurate detection of errors in large language models (LLM) responses is central to the success of scalable oversight, or providing effective supervision to superhuman intelligence. Yet, self-diagnosis is often unreliable on complex tasks…

Machine Learning · Computer Science 2025-10-27 Yongqiang Chen , Gang Niu , James Cheng , Bo Han , Masashi Sugiyama

Scam detection remains a critical challenge in cybersecurity as adversaries craft messages that evade automated filters. We propose a Hierarchical Scam Detection System (HSDS) that combines a lightweight multi-model voting front end with a…

Cryptography and Security · Computer Science 2025-11-11 Chen-Wei Chang , Shailik Sarkar , Hossein Salemi , Hyungmin Kim , Shutonu Mitra , Hemant Purohit , Fengxiu Zhang , Michin Hong , Jin-Hee Cho , Chang-Tien Lu

In safety-critical robotic tasks, potential failures must be reduced, and multiple constraints must be met, such as avoiding collisions, limiting energy consumption, and maintaining balance. Thus, applying safe reinforcement learning (RL)…

Machine Learning · Computer Science 2023-12-27 Dohyeong Kim , Kyungjae Lee , Songhwai Oh

Online reinforcement learning (RL) algorithms are often difficult to deploy in complex human-facing applications as they may learn slowly and have poor early performance. To address this, we introduce a practical algorithm for incorporating…

Artificial Intelligence · Computer Science 2022-01-03 Tong Mu , Georgios Theocharous , David Arbour , Emma Brunskill

This paper presents a systematic analysis of biases in open-source Large Language Models (LLMs), across gender, religion, and race. Our study evaluates bias in smaller-scale Llama and Gemma models using the SALT ($\textbf{S}$ocial…

Computation and Language · Computer Science 2025-02-19 Samee Arif , Zohaib Khan , Maaidah Kaleem , Suhaib Rashid , Agha Ali Raza , Awais Athar

Risk Limiting Dispatch (RLD) was proposed recently as a mechanism that utilizes information and market recourse to reduce reserve capacity requirements, emissions and achieve other system operator objectives. It induces a set of simple…

Optimization and Control · Mathematics 2012-12-04 Junjie Qin , Han-I Su , Ram Rajagopal

The explosive growth of AI research has driven paper submissions at flagship AI conferences to unprecedented levels, necessitating many venues in 2025 (e.g., CVPR, ICCV, KDD, AAAI, IJCAI, WSDM) to enforce strict per-author submission limits…

Data Structures and Algorithms · Computer Science 2025-06-26 Xiaoyu Li , Zhao Song , Jiahao Zhang

In large-scale open-source projects, hundreds of pull requests land daily, each a potential source of regressions. Diff risk scoring (DRS) estimates how likely an individual code change is to introduce a defect. This score can help…

Software Engineering · Computer Science 2026-01-23 Ali Sayedsalehi , Peter C. Rigby , Audris Mockus

Large language models (LLMs) are increasingly used for high-stakes decisions, yet their susceptibility to spurious features remains poorly characterized. We introduce ICE-Guard, a framework applying intervention consistency testing to…

Computation and Language · Computer Science 2026-03-20 Abhinaba Basu , Pavan Chakraborty

Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalated to a more expensive expert. Existing cascades delegate based on probe uncertainty,…

Machine Learning · Computer Science 2026-04-17 Edoardo Pona , Milad Kazemi , Mehran Hosseini , Yali Du , David Watson , Osvaldo Simeone , Nicola Paoletti

Owing to the merits of early safety and reliability guarantee, autonomous driving virtual testing has recently gains increasing attention compared with closed-loop testing in real scenarios. Although the availability and quality of…

Robotics · Computer Science 2021-07-01 Pengliang Ji , Li Ruan , Yunzhi Xue , Limin Xiao , Qian Dong

The spread of election misinformation and harmful political content conveys misleading narratives and poses a serious threat to democratic integrity. Detecting harmful content at early stages is essential for understanding and potentially…

Human-Computer Interaction · Computer Science 2026-02-24 Qile Wang , Prerana Khatiwada , Carolina Coimbra Vieira , Benjamin E. Bagozzi , Kenneth E. Barner , Matthew Louis Mauriello

The standard voting methods in the United States, plurality and ranked choice (or instant runoff) voting, are susceptible to significant voting failures. These flaws include Condorcet and majority failures as well as monotonicity and…

General Economics · Economics 2025-01-10 N. Bradley Fox , Benjamin Bruyns

Data quality is a critical factor in the effectiveness of machine learning models. Label errors, present even in widely used benchmarks, introduce noise into training data and reduce model generalization. In this work, we conduct a…

Computation and Language · Computer Science 2026-05-29 Egor Shevchenko , Elena Bruches

Deep learning models achieve state-of-the-art performance across domains but face scalability challenges in real-time or resource-constrained scenarios. To address this, we propose Correlation of Loss Differences (CLD), a simple and…

Machine Learning · Computer Science 2025-11-20 Manish Nagaraj , Deepak Ravikumar , Kaushik Roy

Many voter-verifiable, coercion-resistant schemes have been proposed, but even the most carefully designed systems necessarily leak information via the announced result. In corner cases, this may be problematic. For example, if all the…

Cryptography and Security · Computer Science 2019-08-15 Wojciech Jamroga , Peter B. Roenne , Peter Y. A. Ryan , Philip B. Stark
‹ Prev 1 4 5 6 7 8 10 Next ›