English
Related papers

Related papers: Group Relative Policy Optimization for Robust Blin…

200 papers

This paper addresses the challenge of co-channel interference and intentional jamming in low-altitude air-ground communications. Since conventional fixed-position antenna (FPA) systems lack spatial adaptability to dynamically balance signal…

Signal Processing · Electrical Eng. & Systems 2025-09-03 Yifan Guo , Junshan Luo , Fanggang Wang , Haiyang Ding , Shilian Wang , Zhenhai Xu

Flow-matching policies have emerged as a powerful paradigm for generalist robotics. These models are trained to imitate an action chunk, conditioned on sensor observations and textual instructions. Often, training demonstrations are…

Machine Learning · Computer Science 2025-07-22 Samuel Pfrommer , Yixiao Huang , Somayeh Sojoudi

Federated learning (FL) is an emerging machine learning paradigm with immense potential to support advanced services and applications in future industries. However, when deployed over wireless communication systems, FL suffers from…

Signal Processing · Electrical Eng. & Systems 2025-03-06 Sangjun Park , Hyowoon Seo

Group Relative Policy Optimization (GRPO), which is widely adopted by R1-like reasoning models, has advanced mathematical reasoning. Nevertheless, GRPO faces challenges in reward sparsity, verbosity, and inadequate focus on problem…

Computation and Language · Computer Science 2025-09-23 Jixiao Zhang , Chunsheng Zuo

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. Within these frameworks, the policy update relies on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Jing Wang , Jiajun Liang , Jie Liu , Henglin Liu , Gongye Liu , Jun Zheng , Wanyuan Pang , Ao Ma , Zhenyu Xie , Xintao Wang , Meng Wang , Pengfei Wan , Xiaodan Liang

Large Language Models (LLMs) have shown promise in solving complex mathematical problems, yet they still fall short of producing accurate and consistent solutions. Reinforcement Learning (RL) is a framework for aligning these models with…

Artificial Intelligence · Computer Science 2026-02-10 Ali Hatamizadeh , Shrimai Prabhumoye , Igor Gitman , Ximing Lu , Seungju Han , Wei Ping , Yejin Choi , Jan Kautz

Stacked intelligent metasurfaces (SIMs) have recently emerged as a powerful wave-domain technology that enables multi-stage manipulation of electromagnetic signals through multilayer programmable architectures. While SIMs offer…

Networking and Internet Architecture · Computer Science 2026-05-29 Le-Hung Hoang , Quang-Trung Luu , Dinh Thai Hoang , Diep N. Nguyen , Van-Dinh Nguyen

Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improving the reasoning capabilities of large language models…

Machine Learning · Computer Science 2026-05-21 Xixiang He , Qiyao Sun , Ao Cheng , Xingming Li , Xuanyu Ji , Hailun Lu , Runke Huang , Qingyong Hu

Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become…

Machine Learning · Computer Science 2026-05-18 Youngjae Cho , Jongsuk Kim , Ji-Hoon Kim

Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous input structures across their generation steps, which creates…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Shikun Sun , Liao Qu , Huichao Zhang , Yiheng Liu , Yangyang Song , Xian Li , Xu Wang , Yi Jiang , Daniel K. Du , Xinglong Wu , Jia Jia

This paper investigates the performance of a multi-reconfigurable intelligent surface (RIS)-assisted fluid antenna system (FAS). In this system, a single-antenna transmitter communicates with a receiver equipped with a planar FAS through…

Signal Processing · Electrical Eng. & Systems 2025-04-29 Mahmoud Aldababsa , Taissir Y. Elganimi , Mahmoud A. Albreem , Saeed Abdallah

Reinforcement learning (RL) is widely used for post-training large language models (LLMs) in code editing, where group-relative methods, such as GRPO, are popular due to their critic-free and normalized advantage estimation. However, in…

Machine Learning · Computer Science 2026-01-09 Jianqing Zhang , Zhezheng Hao , Wei Xia , Hande Dong , Hong Wang , Chenxing Wei , Yuyan Zhou , Yubin Qi , Qiang Lin , Jian Cao

The concept of fluid reconfigurable intelligent surface (FRIS) upgrades the conventional reconfigurable intelligent surface (RIS) paradigm by empowering its reflecting elements with positioning reconfigurability. This letter aims to…

Information Theory · Computer Science 2025-11-21 Xusheng Zhu , Kai-Kit Wong , Boyi Tang , Wen Chen , Chan-Byoung Chae

This paper investigates a reconfigurable intelligent surface (RIS)-aided multi-user multiple-input multiple-output (MIMO) system by considering only the statistical channel state information (CSI) at the base station (BS). We aim to…

Information Theory · Computer Science 2021-12-23 Huan Zhang , Shaodan Ma , Zheng Shi , Xin Zhao , Guanghua Yang

Proximal constraints are fundamental to the stability of the Large Language Model reinforcement learning. While the canonical clipping mechanism in PPO serves as an efficient surrogate for trust regions, we identify a critical bottleneck:…

Machine Learning · Computer Science 2026-03-06 Yuan Li , Bo Wang , Yufei Gao , Yuqian Yao , Xinyuan Wang , Zhangyue Yin , Xipeng Qiu

The paper presents a joint beamforming algorithm using statistical channel state information (S-CSI) for reconfigurable intelligent surfaces (RIS) for multiuser MISO wireless communications. We used S-CSI, which is a long-term average of…

Signal Processing · Electrical Eng. & Systems 2022-09-21 Mahdi Eskandari , Huiling Zhu , Arman Shojaeifard , Jiangzhou Wang

Fluid antenna systems (FAS) have emerged as a revolutionary technology offering enhanced spatial diversity within a compact form factor. Concurrently, unmanned aerial vehicles (UAVs) are integral to future networks, necessitating channel…

Information Theory · Computer Science 2025-11-24 Xusheng Zhu , Kai-Kit Wong , Qingqing Wu , Hyundong Shin , Yangyang Zhang

The revolutionary convergence of fluid antenna systems (FAS) and reconfigurable intelligent surfaces (RIS) creates unprecedented opportunities for secure wireless communications, yet the practical implications of hardware impairments on…

Signal Processing · Electrical Eng. & Systems 2026-03-10 Tuo Wu , Jianchao Zheng , Xiazhi Lai , Maged Elkashlan , Hyundong Shin , Naofal Al-Dhahir

In this paper, we investigate the performance of a fluid antenna relay (FAR)-assisted downlink communication system utilizing non-orthogonal multiple access (NOMA). The FAR, which integrates a fluid antenna system (FAS), is equipped on an…

Signal Processing · Electrical Eng. & Systems 2026-03-27 Ruopeng Xu , Songling Zhang , Zhaohui Yang , Yixuan Chen , Mingzhe Chen , Zhaoyang Zhang , Kai-Kit Wong

In this paper, an reconfigurable intelligent surface (RIS)-aided millimeter wave (mmWave) non-orthogonal multiple access (NOMA) system is considered. In particular, we consider an RIS-aided mmWave-NOMA downlink system with a hybrid…

Signal Processing · Electrical Eng. & Systems 2020-10-20 Yue Xiu , Jun Zhao , Wei Sun , Marco Di Renzo , Guan Gui , Zhongpei Zhang , Ning Wei