中文
相关论文

相关论文: Refunded but Rewarded: The Double Dip Attack on Ca…

200 篇论文

Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intended task. We systematically study this phenomenon in coding tasks using an environment-manipulation…

机器学习 · 计算机科学 2026-04-03 Rui Wu , Ruixiang Tang

The rapid integration of Artificial Intelligence (AI)-based systems offers benefits for various domains of the economy and society but simultaneously raises concerns due to emerging scandals. These scandals have led to the increasing…

计算机与社会 · 计算机科学 2024-11-28 L. H. Nguyen , S. Lins , G. Du , A. Sunyaev

Points-based rewards programs are a prevalent way to incentivize customer loyalty; in these programs, customers who make repeated purchases from a seller accumulate points, working toward eventual redemption of a free reward. These programs…

机器学习 · 计算机科学 2025-06-05 Chamsi Hssaine , Yichun Hu , Ciara Pike-Burke

Artificial currencies have grown in popularity in many real-world resource allocation settings, gaining traction in government benefits programs like food assistance and transit benefits programs. However, such programs are susceptible to…

系统与控制 · 电气工程与系统科学 2024-02-27 Devansh Jalota , Matthew Tsao , Marco Pavone

Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of completed responses,…

计算与语言 · 计算机科学 2026-03-05 Patrick Wilhelm , Thorsten Wittkopp , Odej Kao

Most reinforcement learning algorithms implicitly assume strong synchrony. We present novel attacks targeting Q-learning that exploit a vulnerability entailed by this assumption by delaying the reward signal for a limited time period. We…

机器学习 · 计算机科学 2022-09-09 Anindya Sarkar , Jiarui Feng , Yevgeniy Vorobeychik , Christopher Gill , Ning Zhang

Credit card fraud causes significant financial losses and frequently occurs as fraud attack, defined as short-term sequence of fraudulent transactions associated with high transaction rates and amounts, business areas historically tied to…

最优化与控制 · 数学 2025-03-27 Alexander Stotsky

Reinforcement Learning from Verifiable Rewards (RLVR) has recently shown that large language models (LLMs) can develop their own reasoning without direct supervision. However, applications in the medical domain, specifically for question…

机器学习 · 计算机科学 2025-09-22 Mirza Farhan Bin Tarek , Rahmatollah Beheshti

Language models are capable of iteratively improving their outputs based on natural language feedback, thus enabling in-context optimization of user preference. In place of human users, a second language model can be used as an evaluator,…

计算与语言 · 计算机科学 2024-07-08 Jane Pan , He He , Samuel R. Bowman , Shi Feng

Nowadays, rating systems play a crucial role in the attraction of customers for different services. However, as it is difficult to detect a fake rating, attackers can potentially impact the rating's aggregated score unfairly. This malicious…

计算机科学与博弈论 · 计算机科学 2022-08-05 Iman Vakilinia , Peyman Faizian , Mohammad Mahdi Khalili

Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified…

机器学习 · 计算机科学 2022-02-15 Alexander Pan , Kush Bhatia , Jacob Steinhardt

The large variety of digital payment choices available to consumers today has been a key driver of e-commerce transactions in the past decade. Unfortunately, this has also given rise to cybercriminals and fraudsters who are constantly…

机器学习 · 计算机科学 2021-12-09 Siddharth Vimal , Kanishka Kayathwal , Hardik Wadhwa , Gaurav Dhama

Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinforcement learning updates in language models and argue that…

机器学习 · 计算机科学 2026-05-26 Wenlong Deng , Jiaji Huang , Kaan Ozkara , Yushu Li , Christos Thrampoulidis , Xiaoxiao Li , Youngsuk Park

In this paper, we navigate the intricate domain of reviewer rewards in open-access academic publishing, leveraging the precision of mathematics and the strategic acumen of game theory. We conceptualize the prevailing voucher-based reviewer…

人工智能 · 计算机科学 2023-05-23 Minhyeok Lee

The Denial of Wallet (DoW) attack poses a unique and growing threat to serverless architectures that rely on Function-as-a-Service (FaaS) models, exploiting the cost structure of pay-as-you-go billing to financially burden application…

密码学与安全 · 计算机科学 2025-08-28 Mark Dorsett , Scott Mann , Jabed Chowdhury , Abdun Mahmood

This study critically examines the methodological rigor in credit card fraud detection research, revealing how fundamental evaluation flaws can overshadow algorithmic sophistication. Through deliberate experimentation with improper…

机器学习 · 计算机科学 2025-11-11 Khizar Hayat , Baptiste Magnier

Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task. Precisely measuring reward hacking occurrence remains challenging because true task rewards are often…

机器学习 · 计算机科学 2026-04-21 Muhammad Khalifa , Zohaib Khan , Omer Tafveez , Hao Peng , Lu Wang

External reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving. However, these systems are prone to reward hacking, where…

机器学习 · 计算机科学 2025-08-07 Ruike Song , Zeen Song , Huijie Guo , Wenwen Qiang

Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a…

人工智能 · 计算机科学 2026-05-13 Anas Mahmoud , MohammadHossein Rezaei , Zihao Wang , Anisha Gunjal , Bing Liu , Yunzhong He

Reward hacking--where agents exploit flaws in imperfect reward functions rather than performing tasks as intended--poses risks for AI alignment. Reward hacking has been observed in real training runs, with coding agents learning to…

人工智能 · 计算机科学 2025-08-26 Mia Taylor , James Chua , Jan Betley , Johannes Treutlein , Owain Evans
‹ 上一页 1 2 3 10 下一页 ›