English
Related papers

Related papers: Group Relative Policy Optimization for Robust Blin…

200 papers

Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentally limited by the ambiguity of textual visual correspondence:…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ruiying Liu , Yuanzhi Liang , Haibin Huang , Tianshu Yu , Chi Zhang

Proximal Policy Optimization (PPO) dominates reinforcement learning and LLM alignment but relies on a "hard clipping" mechanism that discards valuable gradients. Conversely, unconstrained methods like SPO expose the optimization to…

Artificial Intelligence · Computer Science 2026-05-07 Yiheng Zhang , Yiming Wang , Kaiyan Zhao , Zhenglin Wan , Jiayu Chen , Leong Hou U

The emerging technology of fluid antenna systems (FASs) represents a promising next-generation reconfigurable antenna solution, capable of exploiting the full spatial diversity within a predefined space by finely reconfiguring the positions…

Information Theory · Computer Science 2026-05-26 Jiangsheng Huangfu , Zhengyu Song , Tianwei Hou , Anna Li , Yuanwei Liu , Arumugam Nallanathan

Alignment methodologies have emerged as a critical pathway for enhancing language model alignment capabilities. While SFT (supervised fine-tuning) accelerates convergence through direct token-level loss intervention, its efficacy is…

Movable antenna (MA) and intelligent reflecting surface (IRS) are considered promising technologies for the next-generation wireless communication systems due to their shared channel reconfiguration capabilities. This, however, raises a…

Information Theory · Computer Science 2025-07-04 Xin Wei , Weidong Mei , Qingqing Wu , Qiaoran Jia , Boyu Ning , Zhi Chen , Jun Fang

Emerging communication networks are envisioned to support massive wireless connectivity of heterogeneous devices with sporadic traffic and diverse requirements in terms of latency, reliability, and bandwidth. Providing multiple access to an…

Information Theory · Computer Science 2023-04-19 Sajad Daei , Marios Kountouris

The Fluid Antenna System (FAS) overcomes the spatial degree-of-freedom limitations of conventional static antenna arrays in wireless communications.This capability critically depends on acquiring full Channel State Information across all…

Signal Processing · Electrical Eng. & Systems 2025-08-05 Xuehui Dong , Kai Wan , Shuangyang Li , Robert Caiming Qiu , Giuseppe Caire

This paper investigates a design framework for sparse fluid antenna systems (FAS) enabling high-performance direction-of-arrival (DOA) estimation, particularly in challenging millimeter-wave (mmWave) environments. By ingeniously harnessing…

Signal Processing · Electrical Eng. & Systems 2025-08-15 He Xu , Tuo Wu , Ye Tian , Ming Jin , Wei Liu , Qinghua Guo , Maged Elkashlan , Matthew C. Valenti , Chan-Byoung Chae , Kin-Fai Tong , Kai-Kit Wong

In this letter, we investigate the discrete phase shift design of the intelligent reflecting surface (IRS) in a time division duplexing (TDD) multi-user multiple input multiple output (MIMO) system.We modify the design of deep reinforcement…

Information Theory · Computer Science 2023-07-31 Fengyu Zhao , Wen Chen , Ziwei Liu , Jun Li , Qingqing Wu

Group Relative Policy Optimization (GRPO) was introduced and used recently for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean + variance calibration of these rewards induces a weighted contrastive loss…

Machine Learning · Computer Science 2025-10-22 Youssef Mroueh

Reconfigurable-antenna systems have received increasing attention for their ability to adapt wireless channels. However, existing architectures exhibit scenario-dependent limitations: fluid antennas provide strong diversity gains in…

Signal Processing · Electrical Eng. & Systems 2026-05-12 Xiao Lin , Yizhe Zhao , Xiangyang Wang , Halvin Yang , Bingxin Zhang

Movable antenna (MA) has shown significant potential for improving the performance of integrated sensing and communication (ISAC) systems. In this paper, we model an MA-aided ISAC system operating in a communication full-duplex mono-static…

Information Theory · Computer Science 2025-05-22 Size Peng , Yin Xu , Guanli Yi , Cixiao Zhang , Dazhi He , Wenjun Zhang

We present Group Orthogonalized Policy Optimization (GOPO), a new alignment algorithm for large language models derived from the geometry of Hilbert function spaces. Instead of optimizing on the probability simplex and inheriting the…

Machine Learning · Computer Science 2026-02-26 Wang Zixian

Future sixth-generation (6G) networks require high spectral efficiency (SE), massive connectivity, and stringent reliability under imperfect channel state information at the transmitter. Rate-splitting multiple access (RSMA) addresses part…

Signal Processing · Electrical Eng. & Systems 2026-04-15 Jinyuan Liu , Yong Liang Guan , Hong Niu , Qian Zhang , Mérouane Debbah , Hyundong Shin , Bruno Clerckx

Fluid antenna system (FAS) is emerging as a key technology for enhancing spatial flexibility and sensing accuracy in future wireless systems. This paper investigates an unmanned aerial vehicle (UAV)-enabled FAS for multi-target wireless…

Information Theory · Computer Science 2025-09-29 Xuhui Zhang , Wenchao Liu , Chunjie Wang , Jinke Ren , Huijun Xing , Shuqiang Wang , Yanyan Shen

Stacked intelligent metasurfaces (SIMs) and fluid antenna systems (FAS) are emerging technologies for wave-domain and spatial signal manipulation, respectively.This letter proposes a novel joint SIM-FAS communication model in which…

Information Theory · Computer Science 2026-05-22 Anastasios Papazafeiropoulos

Powered by position-flexible antennas, the emerging fluid antenna system (FAS) technology is postulated as a key enabler for massive connectivity in 6G networks. The free movement of antenna elements enables the opportunistic minimization…

Signal Processing · Electrical Eng. & Systems 2024-07-24 Pablo Ramirez-Espinosa , David Morales-Jimenez , Kai-Kit Wong

The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning models (LRMs). In this work, we analyze the GRPO objective…

Machine Learning · Computer Science 2026-01-07 Gang Li , Ming Lin , Tomer Galanti , Zhengzhong Tu , Tianbao Yang

Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction…

Machine Learning · Computer Science 2026-03-05 Chengxuan Lu , Zhenquan Zhang , Shukuan Wang , Qunzhi Lin , Baigui Sun , Yang Liu

This correspondence investigates the novel fluid antenna system (FAS) technology, combining with reconfigurable intelligent surface (RIS) for wireless communications, where a base station (BS) communicates with a FAS-enabled user with the…

Signal Processing · Electrical Eng. & Systems 2024-08-27 Junteng Yao , Jianchao Zheng , Tuo Wu , Ming Jin , Chau Yuen , Kai-Kit Wong , Fumiyuki Adachi