English
Related papers

Related papers: Linear Bounds between Contraction Coefficients for…

200 papers

In this article, we investigate posterior convergence in nonparametric regression models where the unknown regression function is modeled by some appropriate stochastic process. In this regard, we consider two setups. The first setup is…

Statistics Theory · Mathematics 2020-05-04 Debashis Chatterjee , Sourabh Bhattacharya

In this paper we explore the information-theoretic aspects of interference alignment and its relation to channel state information (CSI). For the $K-$user interference channel using different changing patterns between different users, we…

Information Theory · Computer Science 2015-12-08 Milad Johnny , Mohammad Reza Aref

This paper studies the basic question of whether a given channel $V$ can be dominated (in the precise sense of being more noisy) by a $q$-ary symmetric channel. The concept of "less noisy" relation between channels originated in network…

Information Theory · Computer Science 2018-12-04 Anuran Makur , Yury Polyanskiy

We introduce optimization methods for convolutional neural networks that can be used to improve existing gradient-based optimization in terms of generalization error. The method requires only simple processing of existing stochastic…

Machine Learning · Computer Science 2020-08-26 Dong Lao , Peihao Zhu , Peter Wonka , Ganesh Sundaramoorthi

The classical hypercontractive inequality for the noise operator on the discrete cube plays a crucial role in many of the fundamental results in the Analysis of Boolean functions, such as the KKL (Kahn-Kalai-Linial) theorem, Friedgut's…

Combinatorics · Mathematics 2019-06-14 Peter Keevash , Noam Lifshitz , Eoin Long , Dor Minzer

Infinitesimal contraction analysis provides exponential convergence rates between arbitrary pairs of trajectories of a system by studying the system's linearization. An essentially equivalent viewpoint arises through stability analysis of a…

Systems and Control · Electrical Eng. & Systems 2025-08-11 Akash Harapanahalli , Samuel Coogan

Classifiers that are linear in their parameters, and trained by optimizing a convex loss function, have predictable behavior with respect to changes in the training data, initial conditions, and optimization. Such desirable properties are…

Machine Learning · Computer Science 2020-12-22 Alessandro Achille , Aditya Golatkar , Avinash Ravichandran , Marzia Polito , Stefano Soatto

We provide a stochastic interpretation of non-commutative Dirichlet forms in the context of quantum filtering. For stochastic processes motivated by quantum optics experiments, we derive an optimal finite time deviation bound expressed in…

Quantum Physics · Physics 2022-08-10 Tristan Benoist , Lisa Hänggli , Cambyse Rouzé

Diffusion models have achieved great success in generating high-dimensional samples across various applications. While the theoretical guarantees for continuous-state diffusion models have been extensively studied, the convergence analysis…

Machine Learning · Computer Science 2025-04-15 Zikun Zhang , Zixiang Chen , Quanquan Gu

Recent work has highlighted several advantages of enforcing orthogonality in the weight layers of deep networks, such as maintaining the stability of activations, preserving gradient norms, and enhancing adversarial robustness by enforcing…

Machine Learning · Computer Science 2021-04-19 Asher Trockman , J. Zico Kolter

This paper is concerned with a class of nonmonotone descent methods for minimizing a proper lower semicontinuous KL function $\Phi$, which generates a sequence satisfying a nonmonotone decrease condition and a relative error tolerance.…

Optimization and Control · Mathematics 2022-07-19 Yitian Qian , Shaohua Pan

We interpret likelihood-based test functions from a geometric perspective where the Kullback-Leibler (KL) divergence is adopted to quantify the distance from a distribution to another. Such a test function can be seen as a sub-Gaussian…

Information Theory · Computer Science 2021-01-05 Yan Wang

Deep nonlinear models pose a challenge for fitting parameters due to lack of knowledge of the hidden layer and the potentially non-affine relation of the initial and observed layers. In the present work we investigate the use of information…

Optimization and Control · Mathematics 2016-12-20 Jacob S. Hunter , Nathan O. Hodas

A recent line of work has focused on the use of low-density generator matrix (LDGM) codes for lossy source coding. In this paper, wedevelop a generic technique for deriving lower bounds on the rate-distortion functions of binary linear…

Information Theory · Computer Science 2008-08-18 A. G. Dimakis , M. J. Wainwright , K. Ramchandran

Training convolutional neural networks (CNNs) with a strict 1-Lipschitz constraint under the $l_{2}$ norm is useful for adversarial robustness, interpretable gradients and stable training. 1-Lipschitz CNNs are usually designed by enforcing…

Machine Learning · Computer Science 2022-11-17 Sahil Singla , Soheil Feizi

Information-theoretic measures such as the entropy, cross-entropy and the Kullback-Leibler divergence between two mixture models is a core primitive in many signal processing tasks. Since the Kullback-Leibler divergence of mixtures provably…

Machine Learning · Computer Science 2017-02-01 Frank Nielsen , Ke Sun

The highly non-linear nature of deep neural networks causes them to be susceptible to adversarial examples and have unstable gradients which hinders interpretability. However, existing methods to solve these issues, such as adversarial…

Machine Learning · Computer Science 2023-01-11 Suraj Srinivas , Kyle Matoba , Himabindu Lakkaraju , Francois Fleuret

A computable expression for the rate-distortion (RD) function proposed by Heegard and Berger has eluded information theory for nearly three decades. Heegard and Berger's single-letter achievability bound is well known to be optimal for…

Information Theory · Computer Science 2012-12-12 Roy Timo , Tobias J. Oechtering , Michèle Wigger

Lossy gradient compression, with either unbiased or biased compressors, has become a key tool to avoid the communication bottleneck in centrally coordinated distributed training of machine learning models. We analyze the performance of two…

Machine Learning · Computer Science 2020-12-23 Sebastian U. Stich

The task of compression of data -- as stated by the source coding theorem -- is one of the cornerstones of information theory. Data compression usually exploits statistical redundancies in the data according to its prior distribution.…

Quantum Physics · Physics 2021-01-08 Matheus Capela , Fabio Costa