中文
相关论文

相关论文: H-TD2: Hybrid Temporal Difference Learning for Ada…

200 篇论文

The on-demand ride-hailing industry has experienced rapid growth, transforming transportation norms worldwide. Despite improvements in efficiency over traditional taxi services, significant challenges remain, including drivers' strategic…

计算机科学与博弈论 · 计算机科学 2025-05-12 Yunpeng Li , Antonis Dimakis , Costas A. Courcoubetis

This paper introduces an adaptive model-free deep reinforcement approach that can recognize and adapt to the diurnal patterns in the ride-sharing environment with car-pooling. Deep Reinforcement Learning (RL) suffers from catastrophic…

人工智能 · 计算机科学 2021-06-15 Marina Haliem , Vaneet Aggarwal , Bharat Bhargava

Accurate Travel Time Estimation (TTE) is critical for ride-hailing platforms, where errors directly impact user experience and operational efficiency. While existing production systems excel at holistic route-level dependency modeling, they…

机器学习 · 计算机科学 2026-01-07 Wenzhao Jiang , Jindong Han , Ruiqian Han , Hao Liu

Spatial-temporal prediction is a fundamental problem for constructing smart city, which is useful for tasks such as traffic control, taxi dispatching, and environmental policy making. Due to data collection mechanism, it is common to see…

机器学习 · 计算机科学 2020-08-25 Huaxiu Yao , Yiding Liu , Ying Wei , Xianfeng Tang , Zhenhui Li

Balancing passenger demand and vehicle availability is crucial for ensuring the sustainability and effectiveness of urban transportation systems. To address this challenge, we propose a novel hierarchical strategy for the efficient…

系统与控制 · 电气工程与系统科学 2024-06-17 Pengbo Zhu , Giancarlo Ferrari-Trecate , Nikolas Geroliminis

We initiate the study of federated reinforcement learning under environmental heterogeneity by considering a policy evaluation problem. Our setup involves $N$ agents interacting with environments that share the same state and action space…

机器学习 · 计算机科学 2024-07-02 Han Wang , Aritra Mitra , Hamed Hassani , George J. Pappas , James Anderson

In a multi-agent system, an agent's optimal policy will typically depend on the policies chosen by others. Therefore, a key issue in multi-agent systems research is that of predicting the behaviours of others, and responding promptly to…

多智能体系统 · 计算机科学 2019-10-22 Dongge Han , Wendelin Boehmer , Michael Wooldridge , Alex Rogers

In this paper, we propose a two-timescale delay-optimal base station Discontinuous Transmission (BS-DTX) control and user scheduling for downlink coordinated MIMO systems with energy harvesting capability. To reduce the complexity and…

系统与控制 · 计算机科学 2015-06-04 Ying Cui , Vincent K. N. Lau , Yueping Wu

We derive a learning framework to generate routing/pickup policies for a fleet of autonomous vehicles tasked with servicing stochastically appearing requests on a city map. We focus on policies that 1) give rise to coordination amongst the…

多智能体系统 · 计算机科学 2023-07-07 Daniel Garces , Sushmita Bhattacharya , Stephanie Gil , Dimitri Bertsekas

Taxi demand prediction is an important building block to enabling intelligent transportation systems in a smart city. An accurate prediction model can help the city pre-allocate resources to meet travel demand and to reduce empty taxis on…

机器学习 · 计算机科学 2018-02-28 Huaxiu Yao , Fei Wu , Jintao Ke , Xianfeng Tang , Yitian Jia , Siyu Lu , Pinghua Gong , Jieping Ye , Zhenhui Li

Rapid urbanization has led to a surge of customizable mobility demand in urban areas, which makes on-demand services increasingly popular. On-demand services are flexible while reducing the need for private cars, thus mitigating congestion…

最优化与控制 · 数学 2025-09-03 Xinling Li , Daniele Gammelli , Alex Wallar , Jinhua Zhao , Gioele Zardini

Temporal difference (TD) learning is a fundamental algorithm for estimating value functions in reinforcement learning. Recent finite-time analyses of TD with linear function approximation quantify its theoretical convergence rate. However,…

机器学习 · 计算机科学 2026-03-04 Yunxiang Li , Mark Schmidt , Reza Babanezhad , Sharan Vaswani

In experimenting with off-policy temporal difference (TD) methods in hierarchical reinforcement learning (HRL) systems, we have observed unwanted on-policy learning under reproducible conditions. Here we present modifications to several TD…

机器学习 · 计算机科学 2015-03-19 Mitchell Keith Bloch

We study the policy evaluation problem in multi-agent reinforcement learning, modeled by a Markov decision process. In this problem, the agents operate in a common environment under a fixed control policy, working together to discover the…

最优化与控制 · 数学 2020-01-13 Thinh T. Doan , Siva Theja Maguluri , Justin Romberg

In matching markets such as kidney exchanges and freight exchanges, delayed matching has been shown to improve overall market efficiency. The benefits of delay are highly sensitive to participants' sojourn times and departure behavior, and…

机器学习 · 计算机科学 2026-02-27 Ruiqi Zhou , Donghao Zhu , Houcai Shen

We introduce an improved algorithm for the dynamic taxi sharing problem, i.e. a dispatcher that schedules a fleet of shared taxis as it is used by services like UberXShare and Lyft Shared. We speed up the basic online algorithm that looks…

数据结构与算法 · 计算机科学 2023-11-06 Moritz Laupichler , Peter Sanders

A key problem in location-based modeling and forecasting lies in identifying suitable spatial and temporal resolutions. In particular, judicious spatial partitioning can play a significant role in enhancing the performance of location-based…

机器学习 · 计算机科学 2018-12-11 Neema Davis , Gaurav Raina , Krishna Jagannathan

Temporal-difference (TD) learning is highly effective at controlling and evaluating an agent's long-term outcomes. Most approaches in this paradigm implement a semi-gradient update to boost the learning speed, which consists of ignoring the…

The goal of this paper is to study a distributed version of the gradient temporal-difference (GTD) learning algorithm for a class of multi-agent Markov decision processes (MDPs). The temporal-difference (TD) learning is a reinforcement…

最优化与控制 · 数学 2020-04-29 Donghwan Lee , Jianghai Hu

This paper revisits the temporal difference (TD) learning algorithm for the policy evaluation tasks in reinforcement learning. Typically, the performance of TD(0) and TD($\lambda$) is very sensitive to the choice of stepsizes. Oftentimes,…

最优化与控制 · 数学 2021-10-12 Tao Sun , Han Shen , Tianyi Chen , Dongsheng Li