English
Related papers

Related papers: Group Relative Policy Optimization for Robust Blin…

200 papers

Group Relative Policy Optimization (GRPO) has shown promise in discrete action spaces by eliminating value function dependencies through group-based advantage estimation. However, its application to continuous control remains unexplored,…

Robotics · Computer Science 2025-07-29 Rajat Khanda , Mohammad Baqar , Sambuddha Chakrabarti , Satyasaran Changdar

In-band full-duplex (IBFD) systems are expected to double the spectral efficiency compared to half-duplex systems, provided that loopback self-interference (SI) can be effectively suppressed. The inherent interference mitigation…

Information Theory · Computer Science 2025-06-09 Hanjiang Hong , Kai-Kit Wong , Hao Xu , Yiyan Wu , Sai Xu , Chan-Byoung Chae , Baiyang Liu , Kin-Fai Tong

Policy gradient reinforcement learning techniques enable an agent to directly learn an optimal action policy through the interactions with the environment. Nevertheless, despite its advantages, it sometimes suffers from slow convergence…

Information Theory · Computer Science 2020-08-05 Mohammad G. Khoshkholgh , Halim Yanikomeroglu

In this paper, we investigate the problem of resource allocation for fluid antenna relay (FAR) system with antenna location optimization. In the considered model, each user transmits information to a base station (BS) with help of FAR. The…

Information Theory · Computer Science 2024-06-28 Ruopeng Xu , Yixuan Chen , Jiawen Kang , Minrui Xu , Zhaohui Yang , Chongwen Huang , Dusit Niyato

Group-Relative Policy Optimization (GRPO) is a key technique for training large reasoning models, yet it suffers from a critical vulnerability: the \emph{Think-Answer Mismatch}, where noisy reward signals corrupt the learning process. This…

Machine Learning · Computer Science 2025-08-11 Si Shen , Peijun Shen , Wenhua Zhao , Danhao Zhu

Fluid antenna system (FAS), which continuously repositions a single physical element across a deployment region $[0, D]$, breaks this limit by freeing antenna positions from the discrete grid entirely. This paper establishes the theoretical…

Signal Processing · Electrical Eng. & Systems 2026-05-20 Tuo Wu , Jie Tang , Ye Tian , Cheng Zeng , Matthew C. Valenti , Hing Cheung So

Vision-Language-Action (VLA) models such as OpenVLA, Octo, and $\pi_0$ have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fundamentally constrained by the quality and coverage of…

Machine Learning · Computer Science 2025-10-14 Mingyang Lyu , Yinqian Sun , Erliang Lin , Huangrui Li , Ruolin Chen , Feifei Zhao , Yi Zeng

Most existing integrated sensing and communication (ISAC) studies focus on enabling a base station (BS) to support sensing and communication over shared resources through advanced waveform design and power allocation. In contrast, the…

Signal Processing · Electrical Eng. & Systems 2026-05-25 Noor Waqar , Kai-Kit Wong , Chan-Byoung Chae , Ross Murch

Fluid antenna system (FAS)/movable antenna (MA) has emerged as a promising technology to fully exploit the spatial degrees of freedom (DoFs). In this paper, we propose a new rotatable antenna (RA) model, as a simplified implementation of…

Information Theory · Computer Science 2025-02-28 Qingjie Wu , Beixiong Zheng , Tiantian Ma , Rui Zhang

Stacked intelligent metasurfaces (SIMs) have emerged as a disruptive technology for future wireless networks. To investigate their capabilities, we study the sum rate maximization problem in an SIM-based multiuser (MU) multiple-input…

Signal Processing · Electrical Eng. & Systems 2025-08-22 Eduard E. Bahingayi , Shuying Lin , Murat Uysal , Marco Di Renzo , Le-Nam Tran

Latent reasoning offers a more efficient alternative to explicit reasoning by compressing intermediate reasoning into continuous representations and substantially shortening reasoning chains. However, existing latent reasoning methods…

Machine Learning · Computer Science 2026-05-01 Jingcheng Deng , Zihao Wei , Liang Pang , Junhong Wu , Shicheng Xu , Zenghao Duan , Huawei Shen

Group-based reinforcement learning methods, like Group Relative Policy Optimization (GRPO), are widely used nowadays to post-train large language models. Despite their empirical success, they exhibit structural mismatches between reward…

Machine Learning · Computer Science 2026-01-09 Aleksandar Fontana , Marco Simoni , Giulio Rossolini , Andrea Saracino , Paolo Mori

Fluid antenna enables position reconfigurability that gives transceiver access to a high-resolution spatial signal and the ability to avoid interference through the ups and downs of fading channels. Previous studies investigated this fluid…

Signal Processing · Electrical Eng. & Systems 2025-04-30 Tianyu Han , Yongxu Zhu , Kai-Kit Wong , Gan Zheng , Hyundong Shin

Group Relative Policy Optimization (GRPO) has recently emerged as an effective approach for improving the reasoning capabilities of large language models through online multi-objective reinforcement learning. While personalization on…

Machine Learning · Computer Science 2026-02-03 Ziyao Wang , Daeun Jung , Yexiao He , Guoheng Sun , Zheyu Shen , Myungjin Lee , Ang Li

The fluid antenna system (FAS) refers to a family of reconfigurable antenna technologies that provide substantial spatial gains within a compact, predefined small space, thereby offering extensive degrees of freedom in the physical layer…

Information Theory · Computer Science 2025-12-18 Zhentian Zhang , Jian Dang , David Morales-Jimenez , Hao Jiang , Zaichen Zhang , Christos Masouros , Chan-Byoung Chae

Reconfigurable intelligent Surfaces (RIS) and half-duplex decoded and forwarded (DF) relays can collaborate to optimize wireless signal propagation in communication systems. Users typically have different rate demands and are clustered into…

Information Theory · Computer Science 2025-06-04 Huijun Tang , Jieling Zhang , Zhidong Zhao , Huaming Wu , Hongjian Sun , Pengfei Jiao

The advantage function is a central concept in RL that helps reduce variance in policy gradient estimates. For language modeling, Group Relative Policy Optimization (GRPO) was proposed to use the within-group sample mean as a baseline for…

Machine Learning · Computer Science 2026-04-23 Hu Wang , Congbo Ma , Ian Reid , Mohammad Yaqub

An emerging fluid antenna system (FAS) brings a new dimension, i.e., the antenna positions, to deal with the deep fading, but simultaneously introduces challenges related to the transmit design. This paper proposes an ``unsupervised…

Signal Processing · Electrical Eng. & Systems 2025-02-07 Changpeng He , Yang Lu , Wei Chen , Bo Ai , Kai-Kit Wong , Dusit Niyato

This paper advocates a fluid antenna system (FAS)-assisted long-range communication (LoRa-FAS) for Internet-of-Things (IoT) applications. \textcolor{blue}{In the proposed system, FAS provides spatial diversity gains for LoRa, eliminating…

Signal Processing · Electrical Eng. & Systems 2025-07-08 Gaoze Mu , Yanzhao Hou , Kai-Kit Wong , Mingjie Chen , Qimei Cui , Xiaofeng Tao , Ping Zhang

Fluid antenna system (FAS) represents the concept of treating antenna as a reconfigurable physical-layer resource to broaden system design and network optimization and inspire next-generation reconfigurable antennas. FAS can unleash new…

Signal Processing · Electrical Eng. & Systems 2025-12-23 Tuo Wu , Kai-Kit Wong , Jie Tang , Junteng Yao , Baiyang Liu , Kin-Fai Tong , Chan-Byoung Chae , Matthew C. Valenti , Kwai-Man Luk
‹ Prev 1 3 4 5 6 7 10 Next ›