Related papers: Dp-rank and forbidden configurations
We derive that dpR(n) \leq dens(n) \leq dpR(n)+1, where dens(n) is the supremum of the VC density of all formulas in n parameters, and dpR(n) is the maximum depth of an ICT pattern in n variables. Consequently, strong dependence is…
This paper proposes a new algorithm for recovery of belief network structure from data handling hidden variables. It consists essentially in an extension of the CI algorithm of Spirtes et al. by restricting the number of conditional…
One of the earliest conjectures in computational learning theory-the Sample Compression conjecture-asserts that concept classes (equivalently set systems) admit compression schemes of size linear in their VC dimension. To-date this…
Gradient descent for matrix factorization exhibits an implicit bias toward approximately low-rank solutions. While existing theories often assume the boundedness of iterates, empirically the bias persists even with unbounded sequences. This…
Forbidden ordinal patterns are ordinal patterns (or `rank blocks') that cannot appear in the orbits generated by a map taking values on a linearly ordered space, in which case we say that the map has forbidden patterns. Once a map has a…
Gaussian MIMO channel under total transmit and interference power constraints (TPC and IPC) is considered. A closed-form solution for the optimal transmit covariance matrix in the general case is obtained using the KKT-based approach (up to…
Conventional treatment policies map patient covariates to a single recommended intervention in order to maximize expected clinical outcomes. Although a rich body of causal inference methods has been developed to estimate such policies,…
In theoretical cognitive science, there is a tension between highly structured models whose parameters have a direct psychological interpretation and highly complex, general-purpose models whose parameters and representations are difficult…
We show that the VC-density of any partitioned formula in a pair of ordered vector spaces is bounded above by twice the number of parameter variables. We also show that this bound is optimal and, as a by-product, we prove that no dense pair…
In this paper we relate t-designs to a forbidden configuration problem in extremal set theory. Let 1_t 0_l denote a column of t 1's on top of l 0's. We assume t>l. Let q. (1_t 0_l) denote the (t+l)xq matrix consisting of t rows of q 1's and…
Several authors have recently constructed characteristic classes for classes of infinite rank vector bundles appearing in topology and physics. These include the tangent bundle to the space of maps between closed manifolds, the infinite…
Despite the extreme popularity of deep learning in science and industry, its formal understanding is limited. This thesis puts forth notions of rank as key for developing a theory of deep learning, focusing on the fundamental aspects of…
This talk contains a summary of our work on dynamical CPT invariance and spontaneous CPT violation in string theories, including the possibility that stringy CPT violation could occur at levels detectable in the next generation of…
Iterative projection methods may become trapped at non-solutions when the constraint sets are nonconvex. Two kinds of parameters are available to help avoid this behavior and this study gives examples of both. The first kind of parameter,…
The architectures of deep neural networks (DNN) rely heavily on the underlying grid structure of variables, for instance, the lattice of pixels in an image. For general high dimensional data with variables not associated with a grid, the…
In a tie-breaker design (TBD), subjects with high values of a running variable are given some (usually desirable) treatment, subjects with low values are not, and subjects in the middle are randomized. TBDs are intermediate between…
The training dynamics of hidden layers in deep learning are poorly understood in theory. Recently, the Information Plane (IP) was proposed to analyze them, which is based on the information-theoretic concept of mutual information (MI). The…
Deep learning is renowned for its theory-practice gap, whereby principled theory typically fails to provide much beneficial guidance for implementation in practice. This has been highlighted recently by the benign overfitting phenomenon:…
A matrix is said to have factor width at most $k$ if it can be written as a sum of positive semidefinite matrices that are non-zero only in a single $k \times k$ principal submatrix. We explore the ``factor-width-$k$ rank'' of a matrix,…
Randomized Controlled Trials (RCTs), or A/B testing, have become the gold standard for optimizing various operational policies on online platforms. However, RCTs on these platforms typically cover a limited number of discrete treatment…