English
Related papers

Related papers: DAGs with No Fears: A Closer Look at Continuous Op…

200 papers

Bayesian networks are probabilistic graphical models often used in big data analytics. The problem of exact structure learning is to find a network structure that is optimal under certain scoring criteria. The problem is known to be NP-hard…

Artificial Intelligence · Computer Science 2017-03-22 Subhadeep Karan , Jaroslaw Zola

We propose an input convex neural network (ICNN)-based self-supervised learning framework to solve continuous constrained optimization problems. By integrating the augmented Lagrangian method (ALM) with the constraint correction mechanism,…

Optimization and Control · Mathematics 2025-05-08 Kang Liu , Wei Peng , Jianchen Hu

We show how to treat systematic uncertainties using Bayesian deep networks for regression. First, we analyze how these networks separately trace statistical and systematic uncertainties on the momenta of boosted top quarks forming fat jets.…

High Energy Physics - Phenomenology · Physics 2020-12-23 Gregor Kasieczka , Michel Luchmann , Florian Otterpohl , Tilman Plehn

We develop a novel switching dynamics that converges to the Karush-Kuhn-Tucker (KKT) point of a nonlinear optimisation problem. This new approach is particularly notable for its lower dimensionality compared to conventional primal-dual…

Optimization and Control · Mathematics 2026-02-03 Joel Ferguson , Saeed Ahmed , Juan E. Machado , Michele Cucuzzella , Jacquelien M. A. Scherpen

Attention operators have been applied on both 1-D data like texts and higher-order data such as images and videos. Use of attention operators on high-order data requires flattening of the spatial or spatial-temporal dimensions into a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Hongyang Gao , Zhengyang Wang , Shuiwang Ji

Learning a Bayesian networks with bounded treewidth is important for reducing the complexity of the inferences. We present a novel anytime algorithm (k-MAX) method for this task, which scales up to thousands of variables. Through extensive…

Artificial Intelligence · Computer Science 2018-02-08 Mauro Scanagatta , Giorgio Corani , Marco Zaffalon , Jaemin Yoo , U Kang

We develop refined Karush-Kuhn-Tucker (KKT) and Fritz-John (FJ)-type optimality conditions for nonsmooth, nonconvex mathematical pro\-gra\-mming problems. We pay special attention in the case that the functional constraint belongs to a…

Optimization and Control · Mathematics 2026-04-15 Felipe Lara , Alberto Ramos

When the objective function is not locally Lipschitz, constraint qualifications are no longer sufficient for Karush-Kuhn-Tucker (KKT) conditions to hold at a local minimizer, let alone ensuring an exact penalization. In this paper, we…

Optimization and Control · Mathematics 2017-01-17 Lei Guo , Jane J. Ye

This expository paper contains a concise introduction to some significant works concerning the Karush-Kuhn-Tucker condition, a necessary condition for a solution in local optimality in problems with equality and inequality constraints. The…

Optimization and Control · Mathematics 2020-06-08 Zhuoyu Xiao

Second-order optimization has been shown to accelerate the training of deep neural networks in many applications, often yielding faster progress per iteration on the training loss compared to first-order optimizers. However, the…

Machine Learning · Computer Science 2024-11-14 Davide Buffelli , Jamie McGowan , Wangkun Xu , Alexandru Cioba , Da-shan Shiu , Guillaume Hennequin , Alberto Bernacchia

It is well known that solving a (non-convex) quadratic program is NP-hard. We show that the problem remains hard even if we are only looking for a Karush-Kuhn-Tucker (KKT) point, instead of a global optimum. Namely, we prove that computing…

Computational Complexity · Computer Science 2025-07-30 John Fearnley , Paul W. Goldberg , Alexandros Hollender , Rahul Savani

A discrete Bayesian network is a directed acyclic graph (DAG) consisting of categorical variables. Two popular approaches for DBN modeling include classification and nonparametric methods. However, both methods often require a large number…

Methodology · Statistics 2026-04-29 Alexander Dombowsky , David B. Dunson

Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the Neural Tangent Kernel (NTK). This analysis leads to global…

Machine Learning · Statistics 2020-04-28 Colin Wei , Jason D. Lee , Qiang Liu , Tengyu Ma

Bayesian network structure learning is the notoriously difficult problem of discovering a Bayesian network that optimally represents a given set of training data. In this paper we study the computational worst-case complexity of exact…

Artificial Intelligence · Computer Science 2014-02-05 Sebastian Ordyniak , Stefan Szeider

This paper focuses on the minimization of a sum of a twice continuously differentiable function $f$ and a nonsmooth convex function. An inexact regularized proximal Newton method is proposed by an approximation to the Hessian of $f$…

Optimization and Control · Mathematics 2023-11-09 Ruyu Liu , Shaohua Pan , Yuqia Wu , Xiaoqi Yang

We extend the well-known BFGS quasi-Newton method and its memory-limited variant LBFGS to the optimization of nonsmooth convex objectives. This is done in a rigorous fashion by generalizing three components of BFGS to subdifferentials: the…

Machine Learning · Statistics 2010-11-30 Jin Yu , S. V. N. Vishwanathan , Simon Guenter , Nicol N. Schraudolph

Recurrent Neural Networks (RNNs) are widely used to model sequential data in a wide range of areas, such as natural language processing, speech recognition, machine translation, and time series analysis. In this paper, we model the training…

Optimization and Control · Mathematics 2024-08-20 Yue Wang , Chao Zhang , Xiaojun Chen

Machine learning algorithms frequently require careful tuning of model hyperparameters, regularization terms, and optimization parameters. Unfortunately, this tuning is often a "black art" that requires expert experience, unwritten rules of…

Machine Learning · Statistics 2012-08-30 Jasper Snoek , Hugo Larochelle , Ryan P. Adams

Convolutional Neural Networks experience catastrophic forgetting when optimized on a sequence of learning problems: as they meet the objective of the current training examples, their performance on previous tasks drops drastically. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Davide Abati , Jakub Tomczak , Tijmen Blankevoort , Simone Calderara , Rita Cucchiara , Babak Ehteshami Bejnordi

The implicit bias of neural networks has been extensively studied in recent years. Lyu and Li [2019] showed that in homogeneous networks trained with the exponential or the logistic loss, gradient flow converges to a KKT point of the max…

Machine Learning · Computer Science 2022-10-04 Gal Vardi , Ohad Shamir , Nathan Srebro