English
Related papers

Related papers: Generative Adversarial Post-Training Mitigates Rew…

200 papers

Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an \emph{online} manner, meaning simultaneously with…

Algorithmic music composition is a way of composing musical pieces with minimal to no human intervention. While recurrent neural networks are traditionally applied to many sequence-to-sequence prediction tasks, including successful…

Machine Learning · Computer Science 2022-11-03 Moseli Mots'oehli , Anna Sergeevna Bosman , Johan Pieter De Villiers

Recent advances in generative artificial intelligence (AI) have created models capable of high-quality musical content generation. However, little consideration is given to how to use these models for real-time or cooperative jamming…

Human-Computer Interaction · Computer Science 2025-03-03 Alexander Scarlatos , Yusong Wu , Ian Simon , Adam Roberts , Tim Cooijmans , Natasha Jaques , Cassie Tarakajian , Cheng-Zhi Anna Huang

Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of completed responses,…

Computation and Language · Computer Science 2026-03-05 Patrick Wilhelm , Thorsten Wittkopp , Odej Kao

Reinforcement learning has been widely successful in producing agents capable of playing games at a human level. However, this requires complex reward engineering, and the agent's resulting policy is often unpredictable. Going beyond…

Machine Learning · Computer Science 2023-08-16 William Ahlberg , Alessandro Sestini , Konrad Tollmar , Linus Gisslén

Live performances of music are always charming, with the unpredictability of improvisation due to the dynamic between musicians and interactions with the audience. Jazz improvisation is a particularly noteworthy example for further…

Physics and Society · Physics 2024-03-07 Vedant Tapiavala , Joshua Piesner , Sourjyamoy Barman , Feng Fu

Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores while violating human intent. Existing mitigations rely on…

Artificial Intelligence · Computer Science 2026-02-03 Mohammad Beigi , Ming Jin , Junshan Zhang , Qifan Wang , Lifu Huang

Reinforcement learning (RL) has become a standard approach for post-training large language models and, more recently, for improving image generation models, which uses reward functions to enhance generation quality and human preference…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Yunqi Hong , Kuei-Chun Kao , Hengguang Zhou , Cho-Jui Hsieh

This paper studies the problem of mitigating reactive jamming, where a jammer adopts a dynamic policy of selecting channels and sensing thresholds to detect and jam ongoing transmissions. The transmitter-receiver pair learns to avoid…

Machine Learning · Computer Science 2025-10-03 Yalin E. Sagduyu , Tugba Erpek , Kemal Davaslioglu , Sastry Kompella

Most of the current anti-jamming algorithms for wireless communications only consider how to avoid jamming attacks, but ignore that the communication waveform or frequency action may be obtained by the jammers. Although existing…

Information Theory · Computer Science 2020-12-24 Yifan Wang , Xin Liu , Mei Wang , Yu Yu

Reward modeling has emerged as a promising approach for the scalable alignment of language models. However, contemporary reward models (RMs) often lack robustness, awarding high rewards to low-quality, out-of-distribution (OOD) samples.…

There are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems. In this setting, an online user is the environment; neither the reward function nor the environment dynamics are clearly…

Machine Learning · Computer Science 2020-01-03 Xinshi Chen , Shuang Li , Hui Li , Shaohua Jiang , Yuan Qi , Le Song

Reward hacking--where agents exploit flaws in imperfect reward functions rather than performing tasks as intended--poses risks for AI alignment. Reward hacking has been observed in real training runs, with coding agents learning to…

Artificial Intelligence · Computer Science 2025-08-26 Mia Taylor , James Chua , Jan Betley , Johannes Treutlein , Owain Evans

Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodological hurdles. Current data collection paradigms, reliant on…

Machine Learning · Computer Science 2026-03-17 Jiesong Lian , Ruizhe Zhong , Zixiang Zhou , Xiaoyue Mi , Long Hu , Yuan Zhou , Qinglin Lu , Yixue Hao , Junchi Yan

Aligning AI systems with human preferences typically suffers from the infamous reward hacking problem, where optimization of an imperfect reward model leads to undesired behaviors. In this paper, we investigate reward hacking in offline…

Machine Learning · Computer Science 2024-12-13 Paria Rashidinejad , Yuandong Tian

This paper presents a deep reinforcement learning algorithm for online accompaniment generation, with potential for real-time interactive human-machine duet improvisation. Different from offline music generation and harmonization, online…

Machine Learning · Computer Science 2020-02-11 Nan Jiang , Sheng Jin , Zhiyao Duan , Changshui Zhang

Large language models (LLMs) with explicit reasoning capabilities excel at mathematical reasoning yet still commit process errors, such as incorrect calculations, brittle logic, and superficially plausible but invalid steps. In this paper,…

Artificial Intelligence · Computer Science 2026-03-26 Qihao Liu , Luoxin Ye , Wufei Ma , Yu-Cheng Chou , Alan Yuille

As the applications of deep reinforcement learning (DRL) in wireless communications grow, sensitivity of DRL based wireless communication strategies against adversarial attacks has started to draw increasing attention. In order to address…

Signal Processing · Electrical Eng. & Systems 2022-09-13 Feng Wang , Chen Zhong , M. Cenk Gursoy , Senem Velipasalar

Reinforcement Learning (RL) methods have emerged as a popular choice for training an efficient and effective dialogue policy. However, these methods suffer from sparse and unstable reward signals returned by a user simulator only when a…

Artificial Intelligence · Computer Science 2020-09-18 Ziming Li , Sungjin Lee , Baolin Peng , Jinchao Li , Julia Kiseleva , Maarten de Rijke , Shahin Shayandeh , Jianfeng Gao

The integration of Generative AI models into AI-native network systems offers a transformative path toward achieving autonomous and adaptive control. However, the application of such models to continuous control tasks is impeded by…

Artificial Intelligence · Computer Science 2026-03-12 Yuanhao Li , Haozhe Wang , Geyong Min , Nektarios Georgalas , Wang Miao
‹ Prev 1 2 3 10 Next ›