English
Related papers

Related papers: Group Relative Policy Optimization for Robust Blin…

200 papers

Group Relative Policy Optimization (GRPO) has proven effective in RLVR by using outcome-based rewards. While fine-grained dense rewards can theoretically improve performance, we reveal that under practical sampling budgets, Monte Carlo…

Machine Learning · Computer Science 2026-04-13 Fengwei Teng , Jinyi Bai , Xinhao Yao , Demi Ruohan Wang , Jiahao Zhao , Zhijiang Guo

The Group Relative Policy Optimization (GRPO), a reinforcement learning method used to fine-tune large language models (LLMs), has proved its effectiveness in practical applications such as DeepSeek-R1. It raises a question whether GRPO can…

Machine Learning · Computer Science 2025-11-20 Yanchen Xu , Ziheng Jiao , Hongyuan Zhang , Xuelong Li

Over-the-air (OTA) federated learning (FL) effectively utilizes communication bandwidth, yet it is vulnerable to errors during analog aggregation. While removing users with unfavorable channel conditions can mitigate these errors, it also…

Signal Processing · Electrical Eng. & Systems 2025-03-04 Yang Zhao , Minrui Xu , Ping Wang , Dusit Niyato

Group Relative Policy Optimization (GRPO) is a promising policy-based approach for Large Language Model alignment, yet its performance is often limited by training instability and suboptimal convergence. In this paper, we identify and…

Machine Learning · Computer Science 2025-12-12 Marco Simoni , Aleksandar Fontana , Giulio Rossolini , Andrea Saracino , Paolo Mori

We introduce a novel received signal strength intensity (RSSI)-based positioning method using fluid antenna systems (FAS), leveraging their inherent channel correlation properties to improve location accuracy. By enabling a single antenna…

Signal Processing · Electrical Eng. & Systems 2025-03-04 Wenzhi Liu , Zhisheng Rong , Xiayue Liu , Yufei Jiang , Xu Zhu

In this paper, we investigate cell-free massive MIMO (CF-mMIMO) systems in which access points (APs) are equipped with fluid antennas (FAs) and develop a comprehensive framework for channel estimation, antenna port selection, and uplink…

Information Theory · Computer Science 2025-12-30 Maryam Olyaee , Giovanni Interdonato , Stefano Buzzi

Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different…

Computation and Language · Computer Science 2024-05-31 Shyam Sundhar Ramesh , Yifan Hu , Iason Chaimalas , Viraj Mehta , Pier Giuseppe Sessa , Haitham Bou Ammar , Ilija Bogunovic

Traditional single-input single-output (SISO) systems face fundamental limitations in achieving accurate three-dimensional (3D) localization due to limited spatial degrees of freedom (DoF) and the adverse impact of multipath propagation.…

Signal Processing · Electrical Eng. & Systems 2025-09-17 Hua Chen , Tao Gong , Tuo Wu , Maged Elkashlan , Baiyang Liu , Chan-Byoung Chae , Kin-Fai Tong , Kai-Kit Wong

Reinforcement learning with verifiable rewards (RLVR) has become a practical route to improve large language model reasoning, and Group Relative Policy Optimization (GRPO) is a widely used optimizer in this setting. However, RLVR training…

Machine Learning · Computer Science 2026-05-14 Tue Le , Linh Ngo Van , Trung Le

This letter investigates the reconfigurable intelligent surface (RIS)-assisted multiple-input single-output (MISO) wireless system, where both half-duplex (HD) and full-duplex (FD) operating modes are considered together, for the first time…

Signal Processing · Electrical Eng. & Systems 2021-10-12 Alice Faisal , Ibrahim Al-Nahhal , Octavia A. Dobre , Telex M. N. Ngatched

Reparameterization Policy Gradient (RPG) has emerged as a powerful paradigm for model-based reinforcement learning, enabling high sample efficiency by backpropagating gradients through differentiable dynamics. However, prior RPG approaches…

Machine Learning · Computer Science 2026-02-04 Hai Zhong , Zhuoran Li , Xun Wang , Longbo Huang

The policy represented by the deep neural network can overfit the spurious features in observations, which hamper a reinforcement learning agent from learning effective policy. This issue becomes severe in high-dimensional state, where the…

Machine Learning · Computer Science 2023-05-01 Md Masudur Rahman , Yexiang Xue

Movable antenna (MA) technology offers a flexible approach to enhancing wireless channel conditions by adjusting antenna positions within a designated region. While most existing works focus on narrowband MA systems, this paper investigates…

Signal Processing · Electrical Eng. & Systems 2025-10-03 Ruixi Feng , Weidong Mei , Lele Lu , Xin Wei , Zhi Chen , Zhen Gao , Boyu Ning

Movable antenna (MA) technology has emerged as a promising solution for reconfiguring wireless channel conditions through local antenna movement within confined regions. Unlike previous works assuming perfect channel state information…

Information Theory · Computer Science 2025-05-13 Haifeng Ma , Weidong Mei , Xin Wei , Boyu Ning , Zhi Chen

We consider a multi-user (MU) full-duplex (FD) multiple-input multiple-output (MIMO) communication system, in which the base station transceiver is equipped with transmit and receive position reconfigurable antennas (PRAs) to mitigate both…

Signal Processing · Electrical Eng. & Systems 2025-11-06 Chengjie Zhao , Yuanzhe Gong , Tho Le-Ngoc

In this letter, we investigate the fundamental limits of localization in fluid antenna systems (FAS) utilizing a Fisher-information-theoretic framework. We develop a unified model to quantify the localization information extractable from…

Signal Processing · Electrical Eng. & Systems 2025-12-17 Abdelhamid Salem , Kai-Kit Wong , Hyundong Shin , Yangyang Zhang

Rate splitting multiple access (RSMA) is regarded as a crucial and powerful physical layer (PHY) paradigm for next-generation communication systems. Particularly, users employ successive interference cancellation (SIC) to decode part of the…

Information Theory · Computer Science 2024-11-15 Cixiao Zhang , Size Peng , Yin Xu , Qingqing Wu , Xiaowu Ou , Xinghao Guo , Dazhi He , Wenjun Zhang

Reinforcement learning (RL) has proven effective in strengthening the reasoning capabilities of large language models (LLMs). A widely adopted method, Group Relative Policy Optimization (GRPO), has shown strong empirical results in training…

Machine Learning · Computer Science 2026-03-11 Peter Chen , Xiaopeng Li , Ziniu Li , Xi Chen , Tianyi Lin

Effective frequency control in power grids has become increasingly important with the increasing demand for renewable energy sources. Here, we propose a novel strategy for resolving this challenge using graph convolutional proximal policy…

The initial alignment provides an accurate attitude for SINS (strapdown inertial navigation system). By further estimating the IMU's bias and misalignment angle, the recursive Bayesian filter is accurate. However, the prior heading error…

Robotics · Computer Science 2023-06-07 Hanwen Zhou , Xiufen Ye