English
Related papers

Related papers: Group Relative Policy Optimization for Robust Blin…

200 papers

Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14B parametered model typically demands hundreds of GPU days…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Xiaoxuan He , Siming Fu , Zeyue Xue , Weijie Wang , Ruizhe He , Yuming Li , Dacheng Yin , Shuai Dong , Haoyang Huang , Hongfa Wang , Nan Duan , Bohan Zhuang

Reinforcement learning (RL) plays a central role in large language model (LLM) post-training. Among existing approaches, Group Relative Policy Optimization (GRPO) is widely used, especially for RL with verifiable rewards (RLVR) fine-tuning.…

The Fluid Antenna System (FAS), which enables flexible Multiple-Input Multiple-Output (MIMO) communications, introduces new spatial degrees of freedom for next-generation wireless networks. Unlike traditional MIMO, FAS involves joint port…

Information Theory · Computer Science 2025-06-18 Chao Wang , Kai-Kit Wong , Zan Li , Liang Jin , Chan-Byoung Chae

Group Relative Policy Optimisation (GRPO) enhances large language models by estimating advantages across a group of sampled trajectories. However, mapping these trajectory-level advantages to policy updates requires aggregating token-level…

Reinforcement learning (RL) plays an increasingly important role in enhancing the reasoning capabilities of large language models (LLMs), yet stable and performant policy optimization remains challenging. Token-level importance ratios often…

Machine Learning · Computer Science 2025-12-02 Chang Gao , Chujie Zheng , Xiong-Hui Chen , Kai Dang , Shixuan Liu , Bowen Yu , An Yang , Shuai Bai , Jingren Zhou , Junyang Lin

The integration of electromagnetic metasurfaces into wireless communications enables intelligent control of the propagation environment. Recently, flexible intelligent metasurfaces (FIMs) have evolved beyond conventional reconfigurable…

Signal Processing · Electrical Eng. & Systems 2026-01-23 Ling He , Vaibhav Kumar , Anastasios Papazafeiropoulos , Miaowen Wen , Le-Nam Tran , Marwa Chafii

The integrated sensing and communication (ISAC) technology has been extensively researched to enhance communication rates and radar sensing capabilities. Additionally, a new technology known as fluid antenna system (FAS) has recently been…

Information Theory · Computer Science 2025-11-14 Yuqi Ye , Li You , Hao Xu , Ahmed Elzanaty , Kai-Kit Wong , Xiqi Gao

This paper focuses on optimal beamforming to maximize the mean signal-to-noise ratio (SNR) for a reconfigurable intelligent surface (RIS)-aided MISO downlink system under correlated Rician fading. The beamforming problem becomes non-convex…

Information Theory · Computer Science 2023-07-04 Kali Krishna Kota , M. S. S. Manasa , Praful D. Mankar , Harpreet S. Dhillon

Interference alignment (IA) is a widely recognized approach for mitigating inter-cell interference in multi-user multiple-input multiple-output (MIMO) networks. Despite its effectiveness, practical deployment remains constrained by two…

Signal Processing · Electrical Eng. & Systems 2026-04-30 Samitha Gunarathne , Eslam Eldeeb , Nurul Huda Mahmood , Italo Atzeni

Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-based optimization frameworks (e.g., GRPO) face a critical limitation: the rapid decay…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Sujie Hu , Chubin Chen , Jiashu Zhu , Jiahong Wu , Xiangxiang Chu , Xiu Li

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing the reasoning capabilities of Large Language Models (LLMs). However, dominant approaches like Group Relative Policy Optimization (GRPO) face critical…

Machine Learning · Computer Science 2026-02-24 Kevin Han , Yuhang Zhou , Mingze Gao , Gedi Zhou , Serena Li , Abhishek Kumar , Xiangjun Fan , Weiwei Li , Lizhu Zhang

Visual-Language-Action (VLA) models have demonstrated strong cross-scenario generalization capabilities in various robotic tasks through large-scale pre-training and task-specific fine-tuning. However, their training paradigm mainly relies…

Robotics · Computer Science 2025-09-30 Zengjue Chen , Runliang Niu , He Kong , Qi Wang , Qianli Xing , Zipei Fan

We revisit policy-gradient optimization for Large Language Models (LLMs) from a single-stream perspective. Prevailing group-based methods like GRPO reduce variance with on-the-fly baselines but suffer from critical flaws: frequent…

Machine Learning · Computer Science 2025-09-24 Zhongwen Xu , Zihan Ding

In this letter, we propose to employ movable antenna (MA) to enhance covert communications with noise uncertainty, where the confidential data is transmitted from an MA-aided access point (AP) to multiple users with a warden attempting to…

Signal Processing · Electrical Eng. & Systems 2025-04-29 Haobin Mao , Xiangyu Pi , Lipeng Zhu , Zhenyu Xiao , Xiang-Gen Xia , Rui Zhang

In order to explore how blind interference alignment (BIA) schemes may take advantage of side-information in computation tasks, we study the degrees of freedom (DoF) of a $K$ user wireless network setting that arises in full-duplex wireless…

Information Theory · Computer Science 2025-09-26 Yuxiang Lu , Syed A. Jafar

Configuring intelligent surface (IS) or passive antenna array without any channel knowledge, namely blind beamforming, is a frontier research topic in the wireless communication field. Existing methods in the previous literature for blind…

Information Theory · Computer Science 2024-09-25 Wenhai Lai , Wenyu Wang , Fan Xu , Xin Li , Shaobo Niu , Kaiming Shen

Fluid antennas (FAs) is a promising technology for introducing flexibility and reconfigurability in wireless networks. Recent research efforts have highlighted the potential gains that can be achieved in comparison to conventional antennas.…

Signal Processing · Electrical Eng. & Systems 2023-11-03 Constantinos Psomas , Peter J. Smith , Himal A. Suraweera , Ioannis Krikidis

Group Relative Policy Optimization (GRPO) has emerged as an effective method for training reasoning models. While it computes advantages based on group mean, GRPO treats each output as an independent sample during the optimization and…

Artificial Intelligence · Computer Science 2026-03-16 Yu Li , Tian Lan , Zhengling Qi

This paper addresses the fairness issue within fluid antenna system (FAS)-assisted non-orthogonal multiple access (NOMA) and orthogonal multiple access (OMA) systems, where a single fixed-antenna base station (BS) transmits…

Signal Processing · Electrical Eng. & Systems 2024-03-04 Junteng Yao , Liaoshi Zhou , Tuo Wu , Ming Jin , Cunhua Pan , Maged Elkashlan , Kai-Kit Wong

In wireless communication systems, mmWave beam tracking is a critical task that affects both sensing and communications, as it is related to the knowledge of the wireless channel. We consider a setup in which a Base Station (BS) needs to…

Signal Processing · Electrical Eng. & Systems 2023-06-28 Cristian J. Vaca-Rubio , Carles Navarro Manchón , Ramoni Adeogun , Petar Popovski
‹ Prev 1 4 5 6 7 8 10 Next ›