中文
相关论文

相关论文: Relaxed constant positive linear dependence constr…

200 篇论文

During the last years, asymptotic (or sequential) constraint qualifications, which postulate upper semicontinuity of certain set-valued mappings and provide a natural companion of asymptotic stationarity conditions, have been shown to be…

最优化与控制 · 数学 2023-02-10 Matúš Benko , Patrick Mehlitz

This paper investigates the expected excess risk of in-context learning (ICL) for multiclass classification. We formalize each task as a sequence of labeled examples followed by a query input; a pretrained model then estimates the query's…

机器学习 · 统计学 2025-09-03 Chenrui Liu , Falong Tan , Chuanlong Xie , Yicheng Zeng , Lixing Zhu

Reinforcement Learning (RL) serves as a versatile framework for sequential decision-making, finding applications across diverse domains such as robotics, autonomous driving, recommendation systems, supply chain optimization, biology,…

机器学习 · 计算机科学 2024-08-26 Vaneet Aggarwal , Washim Uddin Mondal , Qinbo Bai

The paper concerns optimization problems with general equality and inequality constraints and with constraints expressed by a convex set. In order to solve these problems, the general constraints are treated by an exact penalty functions…

最优化与控制 · 数学 2026-05-26 Bogdan K. Jastrzębski , Radosław Pytlak

Bilevel optimization (BLO) offers a principled framework for hierarchical decision-making and has been widely applied in machine learning tasks such as hyperparameter optimization and meta-learning. While existing BLO methods are mostly…

最优化与控制 · 数学 2025-10-20 Zhuo Chen , Xinjian Xu , Shihui Ying , Tieyong Zeng

The notion of the relaxed Robust Control Lyapunov Function (relaxed RCLF) is introduced and is exploited for the design of robust feedback stabilizers for nonlinear systems. Particularly, it is shown for systems with input constraints that…

最优化与控制 · 数学 2008-10-07 Iasson Karafyllis , Costas Kravaris , Nicolas Kalogerakis

Policy gradient methods usually rely on entropy regularization to prevent premature convergence. However, maximizing entropy indiscriminately pushes the policy towards a uniform distribution, often overriding the reward signal if not…

机器学习 · 计算机科学 2026-03-06 Luca Serfilippi , Giorgio Franceschelli , Antonio Corradi , Mirco Musolesi

Traditional Evidence Deep Learning (EDL) methods rely on static hyperparameter for uncertainty calibration, limiting their adaptability in dynamic data distributions, which results in poor calibration and generalization in high-risk…

机器学习 · 计算机科学 2025-10-13 Zhen Yang , Yansong Ma , Lei Chen

Double Machine Learning is often justified by nuisance-rate conditions, yet finite-sample reliability also depends on the conditioning of the orthogonal-score Jacobian. This conditioning is typically assumed rather than tracked. When…

统计方法学 · 统计学 2026-01-08 Gabriel Saco

Our contribution in this paper is two folded. We consider first the case of linear programming with real coefficients and give a method which allows the computation of a new upper bound on the distance from the origin to a feasible point.…

最优化与控制 · 数学 2020-10-30 Beniamin Costandin , Marius Costandin , Petru Dobra

We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, based on an arbitrary prior or description language. Grunwald…

机器学习 · 统计学 2025-03-12 Xiaohan Zhu , Nathan Srebro

The concept of qualification for spectral regularization methods for inverse ill-posed problems is strongly associated to the optimal order of convergence of the regularization error. In this article, the definition of qualification is…

数值分析 · 数学 2010-08-31 Terry Herdman , Ruben D. Spies , Karina G. Temperini

For general quadratically-constrained quadratic programming (QCQP), we propose a parabolic relaxation described with convex quadratic constraints. An interesting property of the parabolic relaxation is that the original non-convex feasible…

最优化与控制 · 数学 2022-08-09 Ramtin Madani , Mersedeh Ashraphijuo , Mohsen Kheirandishfard , Alper Atamturk

A class of valued constraint satisfaction problems (VCSPs) is characterised by a valued constraint language, a fixed set of cost functions on a finite domain. An instance of the problem is specified by a sum of cost functions from the…

计算复杂性 · 计算机科学 2015-03-20 Vladimir Kolmogorov

Constrained optimization is popularly seen in reinforcement learning for addressing complex control tasks. From the perspective of dynamic system, iteratively solving a constrained optimization problem can be framed as the temporal…

机器学习 · 计算机科学 2025-01-28 Tianqi Zhang , Puzhen Yuan , Guojian Zhan , Ziyu Lin , Yao Lyu , Zhenzhi Qin , Jingliang Duan , Liping Zhang , Shengbo Eben Li

Reinforcement learning (RL) has revolutionized decision-making across a wide range of domains over the past few decades. Yet, deploying RL policies in real-world scenarios presents the crucial challenge of ensuring safety. Traditional safe…

系统与控制 · 电气工程与系统科学 2024-03-26 Lunet Yifru , Ali Baheri

Certified robustness is a critical property for deploying neural networks (NN) in safety-critical applications. A principle approach to achieving such guarantees is to constrain the global Lipschitz constant of the network. However,…

机器学习 · 计算机科学 2025-07-01 Zain ul Abdeen , Vassilis Kekatos , Ming Jin

Modern engineering systems, such as autonomous vehicles, flexible robotics, and intelligent aerospace platforms, require controllers that are robust to uncertainties, adaptive to environmental changes, and safety-aware under real-time…

机器人学 · 计算机科学 2025-12-16 Patrick Kostelac , Xuerui Wang , Anahita Jamshidnejad

Many practical optimization problems lack strong convexity. Fortunately, recent studies have revealed that first-order algorithms also enjoy linear convergences under various weaker regularity conditions. While the relationship among…

最优化与控制 · 数学 2026-02-05 Feng-Yi Liao , Lijun Ding , Yang Zheng

Group Relative Policy Optimization (GRPO) has emerged as an effective method for training reasoning models. While it computes advantages based on group mean, GRPO treats each output as an independent sample during the optimization and…

人工智能 · 计算机科学 2026-03-16 Yu Li , Tian Lan , Zhengling Qi