English
Related papers

Related papers: Muti-Agent Proximal Policy Optimization For Data F…

200 papers

Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorithm. Despite its ubiquity, we argue that the core ratio…

Machine Learning · Computer Science 2026-05-27 Penghui Qi , Xiangxin Zhou , Zichen Liu , Tianyu Pang , Chao Du , Min Lin , Wee Sun Lee

In this paper, multi-unmanned aerial vehicle (UAV) enabled mobile edge computing (MEC), i.e., UAVE is studied, where several UAVs are deployed as flying MEC platform to provide computing resource to ground user equipments (UEs). Compared to…

Networking and Internet Architecture · Computer Science 2019-04-18 Liang Wang , Peiqiu Huang , Kezhi Wang , Guopeng Zhang , Lei Zhang , Nauman Aslam , Kun Yang

We discuss surveillance with multiple unmanned aerial vehicles (UAV) that minimize idleness (the time between consecutive visits of sensing locations) and constrain latency (the time between capturing data at a sensing location and its…

Robotics · Computer Science 2021-01-12 Jürgen Scherer , Bernhard Rinner

Unmanned aerial vehicles (UAVs) have emerged as the potential aerial base stations (BSs) to improve terrestrial communications. However, the limited onboard energy and antenna power of a UAV restrict its communication range and transmission…

Neural and Evolutionary Computing · Computer Science 2025-02-11 Geng Sun , Jian Xiao , Jiahui Li , Jiacheng Wang , Jiawen Kang , Dusit Niyato , Shiwen Mao

Unmanned Aerial Vehicles (UAVs) play an increasingly critical role in Intelligence, Surveillance, and Reconnaissance (ISR) missions such as border patrolling and criminal detection, thanks to their ability to access remote areas and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-16 Niloufar Mehrabi , Sayed Pedram Haeri Boroujeni , Jenna Hofseth , Abolfazl Razi , Long Cheng , Manveen Kaur , James Martin , Rahul Amin

This paper proposes a Shared Backbone Proximal Policy Optimization (Shared Backbone PPO) algorithm. By sharing the base module between the Actor and Critic networks, the algorithm achieves efficient training and improved performance. The…

Artificial Intelligence · Computer Science 2026-05-19 Z. Jiang

Unmanned aerial vehicles serving as aerial base stations can rapidly restore connectivity after disasters, yet abrupt changes in user mobility and traffic demands shift the quality of service trade-offs and induce strong non-stationarity.…

Multiagent Systems · Computer Science 2026-04-13 Wen Qiu , Zhiqiang He , Wei Zhao , Hiroshi Masui

Protecting endangered wildlife from illegal poaching presents a critical challenge, particularly in vast and partially observable environments where real-time response is essential. This paper introduces a novel Expectation-Maximization…

Machine Learning · Computer Science 2025-10-13 Mazyar Taghavi , Rahman Farnoosh

In this paper, we propose a multi-unmanned aerial vehicle (UAV)-assisted integrated sensing, communication, and computation network. Specifically, the treble-functional UAVs are capable of offering communication and edge computing services…

Information Theory · Computer Science 2024-10-08 Sicong Peng , Bin Li , Lei Liu , Zesong Fei , Dusit Niyato

Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RLVR convergence and raises the performance ceiling, yet…

Machine Learning · Computer Science 2026-04-23 Chuanyu Qin , Chenxu Yang , Qingyi Si , Naibin Gu , Dingyu Yao , Zheng Lin , Peng Fu , Nan Duan , Jiaqi Wang

Federated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the…

Machine Learning · Computer Science 2024-05-30 Sheng Yue , Zerui Qin , Xingyuan Hua , Yongheng Deng , Ju Ren

Software-Defined Networking (SDN) is increasingly adopted to secure Internet-of-Things (IoT) networks due to its centralized control and programmable forwarding. However, SDN-IoT defense is inherently a closed-loop control problem in which…

Cryptography and Security · Computer Science 2026-04-02 Saeid Jamshidi , Negar Shahabi , Foutse Khomh , Carol Fung , Mohammad Hamdaqa

We consider the problem of using multiple agents to harvest data from a collection of sensor nodes (targets) scattered across a two-dimensional environment. These targets transmit their data to the agents that move in the space above them,…

Systems and Control · Electrical Eng. & Systems 2025-08-25 Shili Wu , Yancheng Zhu , Aniruddha Datta , Sean B. Andersson

We discuss the problem of decentralized multi-agent reinforcement learning (MARL) in this work. In our setting, the global state, action, and reward are assumed to be fully observable, while the local policy is protected as privacy by each…

Multiagent Systems · Computer Science 2021-11-02 Kuo Li , Qing-Shan Jia

We present a proximal policy optimization (PPO) agent trained through curriculum learning (CL) principles and meticulous reward engineering to optimize a real-world high-throughput waste sorting facility. Our work addresses the challenge of…

Machine Learning · Computer Science 2024-07-24 Abhijeet Pendyala , Asma Atamna , Tobias Glasmachers

The proximal policy optimization (PPO) algorithm stands as one of the most prosperous methods in the field of reinforcement learning (RL). Despite its success, the theoretical understanding of PPO remains deficient. Specifically, it is…

Machine Learning · Computer Science 2023-06-09 Han Zhong , Tong Zhang

This paper explores the potential of aerial reconfigurable intelligent surfaces (ARIS) to enhance coordinated multi-point non-orthogonal multiple access (CoMP-NOMA) networks. We consider a system model where a UAV-mounted RIS assists in…

Signal Processing · Electrical Eng. & Systems 2024-11-05 Muhammad Umer , Muhammad Ahmed Mohsin , Aamir Mahmood , Kapal Dev , Haejoon Jung , Mikael Gidlund , Syed Ali Hassan

Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcome-based supervision to strengthen internal LLM reasoning,…

Artificial Intelligence · Computer Science 2026-05-29 Siyao Song , Cong Ma , Zhihao Cheng , Shiye Lei , Minghao Li , Ying Zeng , Huaixiao Tou , Kai Jia

Unmanned aerial vehicles (UAVs) are promising for providing communication services due to their advantages in cost and mobility, especially in the context of the emerging Metaverse and Internet of Things (IoT). This paper considers a…

Machine Learning · Computer Science 2023-05-31 Peiyuan Si , Liangxin Qian , Jun Zhao , Kwok-Yan Lam

Obstacle avoidance for small unmanned aircraft is vital for the safety of future urban air mobility (UAM) and Unmanned Aircraft System (UAS) Traffic Management (UTM). There are many techniques for real-time robust drone guidance, but many…

Robotics · Computer Science 2021-11-16 Jueming Hu , Xuxi Yang , Weichang Wang , Peng Wei , Lei Ying , Yongming Liu
‹ Prev 1 8 9 10 Next ›