English
Related papers

Related papers: Drain-Vortex Optimization: A Population-Based Meta…

200 papers

Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect supervision. Existing robust DPO methods often address this…

Machine Learning · Computer Science 2026-05-27 Jilong Liu , Yonghui Yang , Pengyang Shao , Wenjian Tao , Hao Zhan , Haokai Ma , Wei Qin , Richang Hong

We study the problem of Online Convex Optimization (OCO) with memory, which allows loss functions to depend on past decisions and thus captures temporal effects of learning problems. In this paper, we introduce dynamic policy regret as the…

Machine Learning · Computer Science 2023-08-16 Peng Zhao , Yu-Hu Yan , Yu-Xiang Wang , Zhi-Hua Zhou

Trajectory optimization is a fundamental stochastic optimal control problem. This paper deals with a trajectory optimization approach for dynamical systems subject to measurement noise that can be fitted into linear time-varying stochastic…

Systems and Control · Electrical Eng. & Systems 2021-08-24 Prakash Mallick , Zhiyong Chen

Bayesian optimization (BO) has shown impressive results in a variety of applications within low-to-moderate dimensional Euclidean spaces. However, extending BO to high-dimensional settings remains a significant challenge. We address this…

Machine Learning · Statistics 2024-03-11 Shouri Hu , Jiawei Li , Zhibo Cai

Direct Preference Optimization (DPO) has been successfully used to align large language models (LLMs) according to human preferences, and more recently it has also been applied to improving the quality of text-to-image diffusion models.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Shivanshu Shekhar , Shreyas Singh , Tong Zhang

A new stochastic control problem of a dam-reservoir system installed in a river is analyzed both mathematically and numerically. Water balance dynamics of the reservoir are piece-wise deterministic and are driven by a stochastic…

Systems and Control · Electrical Eng. & Systems 2020-05-04 H. Yoshioka , Y. Yoshioka

Many model-based Visual Odometry (VO) algorithms have been proposed in the past decade, often restricted to the type of camera optics, or the underlying motion manifold observed. We envision robots to be able to learn and perform these…

Robotics · Computer Science 2017-05-30 Sudeep Pillai , John J. Leonard

This paper focuses on the active flow control of a computational fluid dynamics simulation over a range of Reynolds numbers using deep reinforcement learning (DRL). More precisely, the proximal policy optimization (PPO) method is used to…

Fluid Dynamics · Physics 2020-06-24 Hongwei Tang , Jean Rabault , Alexander Kuhnle , Yan Wang , Tongguang Wang

In offline reinforcement learning, value overestimation caused by out-of-distribution (OOD) actions significantly limits policy performance. Recently, diffusion models have been leveraged for their strong distribution-matching capabilities,…

Machine Learning · Computer Science 2025-11-13 Yunchang Ma , Tenglong Liu , Yixing Lan , Xin Yin , Changxin Zhang , Xinglong Zhang , Xin Xu

With the uptake of intelligent data-driven applications, edge computing infrastructures necessitate a new generation of admission control algorithms to maximize system performance under limited and highly heterogeneous resources. In this…

Networking and Internet Architecture · Computer Science 2024-07-01 A. Fox , F. De Pellegrini , F. Faticanti , E. Altman , F. Bronzino

Visual-inertial odometry (VIO) is the most common approach for estimating the state of autonomous micro aerial vehicles using only onboard sensors. Existing methods improve VIO performance by including a dynamics model in the estimation…

Robotics · Computer Science 2023-06-29 Giovanni Cioffi , Leonard Bauersfeld , Davide Scaramuzza

Flow-based models are powerful tools for designing probabilistic models with tractable density. This paper introduces Convex Potential Flows (CP-Flow), a natural and efficient parameterization of invertible models inspired by the optimal…

Machine Learning · Computer Science 2021-02-25 Chin-Wei Huang , Ricky T. Q. Chen , Christos Tsirigotis , Aaron Courville

We introduce Velocity-Regularized Adam (VRAdam), a physics-inspired optimizer for training deep neural networks that draws on ideas from quartic terms for kinetic energy with its stabilizing effects on various system dynamics. Previous…

Machine Learning · Computer Science 2026-05-13 Pranav Vaidhyanathan , Lucas Schorling , Natalia Ares , Michael A. Osborne

Regional flow duration curves (FDCs) often reflect streamflow influenced by human activities. We propose a new machine learning algorithm to predict naturalized FDCs at human influenced sites and multiple catchment scales. Separate Meta…

Geophysics · Physics 2025-01-03 Michael J. Friedel , Dave Stewart , Xiao Feng Lu , Pete Stevenson , Helen Manly , Tom Dyer

The ever increasing penetration of Renewable Energy Resources (RESs) in power distribution networks has brought, among others, the challenge of maintaining the grid voltages within the secure region. Employing droop voltage regulators on…

Optimization and Control · Mathematics 2022-03-18 H. Sekhavatmanesh , G. Ferrari-Trecate , S. Mastellone

In recent years, Wasserstein Distributionally Robust Optimization (DRO) has garnered substantial interest for its efficacy in data-driven decision-making under distributional uncertainty. However, limited research has explored the…

Machine Learning · Computer Science 2025-10-01 Ahmad-Reza Ehyaei , Golnoosh Farnadi , Samira Samadi

Preference optimization has become a central paradigm for aligning large language models with human feedback. Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback by directly optimizing pairwise…

Machine Learning · Computer Science 2026-05-05 Inoussa Mouiche

Bandit Convex Optimization (BCO) is a fundamental framework for modeling sequential decision-making with partial information, where the only feedback available to the player is the one-point or two-point function values. In this paper, we…

Machine Learning · Computer Science 2020-07-07 Peng Zhao , Guanghui Wang , Lijun Zhang , Zhi-Hua Zhou

Dynamic multiobjective optimization problems (DMOPs) feature time-varying objectives, which cause the Pareto optimal solution (POS) set to drift over time and make it difficult to maintain both convergence and diversity under limited…

Neural and Evolutionary Computing · Computer Science 2026-03-31 Jian Guan , Huolong Wu , Zhenzhong Wang , Gary G. Yen , Min Jiang

Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF). However, neither RLHF nor DPO take into account the fact that learning certain…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Florinel-Alin Croitoru , Vlad Hondru , Radu Tudor Ionescu , Nicu Sebe , Mubarak Shah