English
Related papers

Related papers: Behavior-Induced Mirror-Prox Temporal-Difference L…

200 papers

Trajectory prediction is essential for autonomous driving, enabling vehicles to anticipate the motion of surrounding agents to support safe planning. However, most existing predictors assume fixed-length histories and suffer substantial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Mingyu Fan , Yi Liu , Hao Zhou , Deheng Qian , Mohammad Haziq Khan , Matthias Raetsch

Learning high-performance control policies that remain consistent with expert behavior is a fundamental challenge in robotics. Reinforcement learning can discover high-performing strategies but often departs from desirable human behavior,…

Robotics · Computer Science 2026-04-06 Siwei Ju , Jan Tauberschmidt , Oleg Arenz , Peter van Vliet , Jan Peters

We proposed an iterate scheme for solving convex-concave saddle-point problems associated with general convex-concave functions. We demonstrated that when our iterate scheme is applied to a special class of convex-concave functions, which…

Optimization and Control · Mathematics 2023-11-01 Hui Ouyang

The inherently diverse and uncertain nature of trajectories presents a formidable challenge in accurately modeling them. Motion prediction systems must effectively learn spatial and temporal information from the past to forecast the future…

Robotics · Computer Science 2023-11-28 Pranav Singh Chib , Pravendra Singh

A transformed primal-dual (TPD) flow is developed for a class of nonlinear smooth saddle point system. The flow for the dual variable contains a Schur complement which is strongly convex. Exponential stability of the saddle point is…

Optimization and Control · Mathematics 2023-02-03 Long Chen , Jingrong Wei

Owing to their connection with generative adversarial networks (GANs), saddle-point problems have recently attracted considerable interest in machine learning and beyond. By necessity, most theoretical guarantees revolve around…

Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constraints can be formulated as a primal-dual optimization problem,…

Machine Learning · Computer Science 2026-02-26 Yining Li , Peizhong Ju , Ness Shroff

This paper investigates the problem of online prediction learning, where learning proceeds continuously as the agent interacts with an environment. The predictions made by the agent are contingent on a particular way of behaving,…

Machine Learning · Computer Science 2018-11-08 Sina Ghiassian , Andrew Patterson , Martha White , Richard S. Sutton , Adam White

Deep models for Multivariate Time Series (MTS) forecasting have recently demonstrated significant success. Channel-dependent models capture complex dependencies that channel-independent models cannot capture. However, the number of channels…

Machine Learning · Computer Science 2024-08-09 Xin Zhou , Weiqing Wang , Wray Buntine , Shilin Qu , Abishek Sriramulu , Weicong Tan , Christoph Bergmeir

We consider stochastic strongly-convex-strongly-concave (SCSC) saddle point (SP) problems which frequently arise in applications ranging from distributionally robust learning to game theory and fairness in machine learning. We focus on the…

Optimization and Control · Mathematics 2023-07-17 Yassine Laguel , Necdet Serhat Aybat , Mert Gürbüzbalaban

One of the main obstacles to broad application of reinforcement learning methods is the parameter sensitivity of our core learning algorithms. In many large-scale applications, online computation and function approximation represent key…

Artificial Intelligence · Computer Science 2016-10-25 Martha White , Adam White

Robot learning in high-dimensional control settings, such as humanoid locomotion, presents persistent challenges for reinforcement learning (RL) algorithms due to unstable dynamics, complex contact interactions, and sensitivity to…

Robotics · Computer Science 2025-05-21 Khang Nguyen , Khai Nguyen , An T. Le , Jan Peters , Manfred Huber , Ngo Anh Vien , Minh Nhat Vu

Given a dataset on actions and resulting long-term rewards, a direct estimation approach fits value functions that minimize prediction error on the training data. Temporal difference learning (TD) methods instead fit value functions by…

Machine Learning · Computer Science 2024-02-15 David Cheikhi , Daniel Russo

Saddle point problems arise from many wireless applications, and primal-dual iterative algorithms are widely applied to find the saddle points. In the existing literature, the convergence results of such algorithms are established assuming…

Information Theory · Computer Science 2017-04-26 Junting Chen , Vincent K. N. Lau

This paper studies equality-constrained minimization problems through the lens of feedback control. We introduce a unified control-theoretic framework by showing that a PID feedback law acting on the dual variable induces the PID…

Optimization and Control · Mathematics 2026-04-13 Veronica Centorrino , Rawan Hoteit , Efe C. Balta , John Lygeros

We propose a unified framework to study policy evaluation (PE) and the associated temporal difference (TD) methods for reinforcement learning in continuous time and space. We show that PE is equivalent to maintaining the martingale…

Machine Learning · Computer Science 2022-02-02 Yanwei Jia , Xun Yu Zhou

We propose two variants of the Primal Dual Hybrid Gradient (PDHG) algorithm for saddle point problems with block decomposable duals, hereafter called Multi-Timescale PDHG (MT-PDHG) and its accelerated variant (AMT-PDHG). Through novel…

Optimization and Control · Mathematics 2026-04-03 Junhui Zhang , Patrick Jaillet

Human behavior has the nature of indeterminacy, which requires the pedestrian trajectory prediction system to model the multi-modality of future motion states. Unlike existing stochastic trajectory prediction methods which usually use a…

Computer Vision and Pattern Recognition · Computer Science 2022-03-28 Tianpei Gu , Guangyi Chen , Junlong Li , Chunze Lin , Yongming Rao , Jie Zhou , Jiwen Lu

We consider a change-point test based on the Hill estimator to test for structural changes in the tail index of Long Memory Stochastic Volatility time series. In order to determine the asymptotic distribution of the corresponding test…

Statistics Theory · Mathematics 2020-06-05 Annika Betken , Davide Giraudo , Rafał Kulik

Off-policy reinforcement learning has many applications including: learning from demonstration, learning multiple goal seeking policies in parallel, and representing predictive knowledge. Recently there has been an proliferation of new…

Machine Learning · Computer Science 2016-04-01 Adam White , Martha White
‹ Prev 1 4 5 6 7 8 10 Next ›