中文
相关论文

相关论文: Two-Stage Submodular Optimization of Dynamic Therm…

200 篇论文

Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), as the widely employed policy based reinforcement learning (RL) methods, are prone to converge to a sub-optimal solution as they limit the policy representation…

机器学习 · 计算机科学 2020-06-16 Jun Song , Chaoyue Zhao

In this paper, we propose several new stochastic second-order algorithms for policy optimization that only require gradient and Hessian-vector product in each iteration, making them computationally efficient and comparable to policy…

最优化与控制 · 数学 2023-01-31 Jinsong Liu , Chenghan Xie , Qi Deng , Dongdong Ge , Yinyu Ye

The proliferation of large-scale AI and data-intensive applications has driven the development of Computing Power Networks (CPN). It is a key paradigm for delivering ubiquitous, on-demand computational services with high efficiency.…

网络与互联网体系结构 · 计算机科学 2026-02-04 Haoxiang Luo , Kun Yang , Qi Huang , Marco Aiello , Schahram Dustdar

This paper proposes a neural stochastic optimization method for efficiently solving the two-stage stochastic unit commitment (2S-SUC) problem under high-dimensional uncertainty scenarios. The proposed method approximates the second-stage…

系统与控制 · 电气工程与系统科学 2026-04-16 Zhentong Shao , Jingtao Qin , Nanpeng Yu

We study two-stage distributionally robust optimization (DRO) problems with decision-dependent information discovery (DDID) wherein (a portion of) the uncertain parameters are revealed only if an (often costly) investment is made in the…

最优化与控制 · 数学 2025-10-07 Qing Jin , Angelos Georghiou , Phebe Vayanos , Grani A. Hanasusanto

Next-day delivery logistics services are redefining the industry by increasingly focusing on customer service. A challenge each logistics service provider faces is to jointly optimize time window assignment and vehicle routing for such…

最优化与控制 · 数学 2025-01-27 Sifa Celik , Layla Martin , Albert H. Schrotenboer , Tom Van Woensel

Transmission system operators employ reserves to deal with unexpected variations of demand and generation to guarantee the security of supply. The French transmission system operator RTE dynamically sizes the required margins using a…

最优化与控制 · 数学 2024-05-14 Jonathan Dumas

Dynamic Threshold Optimization (DTO) adaptively "compresses" the decision space (DS) in a global search and optimization problem by bounding the objective function from below. This approach is different from "shrinking" DS by reducing…

其他计算机科学 · 计算机科学 2012-06-07 Richard A. Formato

In environments with delayed observation, state augmentation by including actions within the delay window is adopted to retrieve Markovian property to enable reinforcement learning (RL). However, state-of-the-art (SOTA) RL techniques with…

机器学习 · 计算机科学 2024-10-23 Qingyuan Wu , Simon Sinong Zhan , Yixuan Wang , Yuhui Wang , Chung-Wei Lin , Chen Lv , Qi Zhu , Chao Huang

Time-distributed Optimization (TDO) is an approach for reducing the computational burden of Model Predictive Control (MPC). When using TDO, optimization iterations are distributed over time by maintaining a running solution estimate and…

最优化与控制 · 数学 2021-02-25 Dominic Liao-McPherson , Terrence Skibik , Jordan Leung , Ilya Kolmanovsky , Marco M. Nicotra

Tensor ring (TR) decomposition is a simple but effective tensor network for analyzing and interpreting latent patterns of tensors. In this work, we propose a doubly randomized optimization framework for computing TR decomposition. It can be…

数值分析 · 数学 2023-03-30 Yajie Yu , Hanyu Li , Jingchun Zhou

Reliability-based design optimization (RBDO) is traditionally formulated as a nested optimization and reliability problem. Although surrogate models are generally employed to improve efficiency, the approach remains computationally…

统计计算 · 统计学 2026-04-08 M. Moustapha , B. Sudret

This paper presents a multi-objective stochastic optimization method for tuning of the controller parameters of Refrigeration Systems based on Vapour Compression. Stochastic Multi Parameter Divergence Optimization (SMDO) algorithm is…

系统与控制 · 计算机科学 2018-06-05 Abdullah Ates , Jie Yuan , Sina Dehghan , Yang Zhao , Celaleddin Yeroglu , YangQuan Chen

A robust-to-dynamics optimization (RDO) problem is an optimization problem specified by two pieces of input: (i) a mathematical program (an objective function $f:\mathbb{R}^n\rightarrow\mathbb{R}$ and a feasible set…

最优化与控制 · 数学 2023-11-27 Amir Ali Ahmadi , Oktay Gunluk

Reliability-based design optimization (RBDO) is a methodology for designing systems and components under the consideration of probabilistic uncertainty. In practical engineering, the number of input data is often limited, which can damage…

最优化与控制 · 数学 2026-05-27 Takumi Fujiyama , Yoshihiro Kanno

Counterfactual learning to rank (CLTR) can be risky and, in various circumstances, can produce sub-optimal models that hurt performance when deployed. Safe CLTR was introduced to mitigate these risks when using inverse propensity scoring to…

机器学习 · 计算机科学 2024-08-08 Shashank Gupta , Harrie Oosterhuis , Maarten de Rijke

We introduce a two-level trust-region method (TLTR) for solving unconstrained nonlinear optimization problems. Our method uses a composite iteration step, which is based on two distinct search directions. The first search direction is…

数值分析 · 数学 2024-09-10 Andrea Angino , Alena Kopaničáková , Rolf Krause

In this paper, we investigate a novel reconfigurable distributed antennas and reflecting surface (RDARS) aided multi-user massive MIMO system with imperfect CSI and propose a practical two-timescale (TTS) transceiver design to reduce the…

信息论 · 计算机科学 2023-12-15 Chengzhi Ma , Jintao Wang , Xi Yang , Guanghua Yang , Wei Zhang , Shaodan Ma

There hardly exists a general solver that is efficient for scheduling problems due to their diversity and complexity. In this study, we develop a two-stage framework, in which reinforcement learning (RL) and traditional operations research…

人工智能 · 计算机科学 2021-03-11 Yongming He , Guohua Wu , Yingwu Chen , Witold Pedrycz

An adaptive delay-tolerant distributed space-time coding (DSTC) scheme with feedback is proposed for two-hop cooperative multiple-input multiple-output (MIMO) networks using an amplify-and-forward strategy and opportunistic relaying…

信息论 · 计算机科学 2014-12-16 T. Peng , R. C. de Lamare