English
Related papers

Related papers: Unlabelled Sample Compression Schemes for Intersec…

200 papers

Learnable embedding vector is one of the most important applications in machine learning, and is widely used in various database-related domains. However, the high dimensionality of sparse data in recommendation tasks and the huge volume of…

Machine Learning · Computer Science 2024-02-14 Hailin Zhang , Penghao Zhao , Xupeng Miao , Yingxia Shao , Zirui Liu , Tong Yang , Bin Cui

Stochastic convex optimization is one of the most well-studied models for learning in modern machine learning. Nevertheless, a central fundamental question in this setup remained unresolved: "How many data points must be observed so that…

Machine Learning · Computer Science 2023-11-10 Daniel Carmon , Roi Livni , Amir Yehudayoff

Vapnik-Chervonenkis (VC) dimension is a fundamental measure of the generalization capacity of learning algorithms. However, apart from a few special cases, it is hard or impossible to calculate analytically. Vapnik et al. [10] proposed a…

Machine Learning · Statistics 2011-11-16 Daniel J. McDonald , Cosma Rohilla Shalizi , Mark Schervish

The field of compressed sensing has shown that a sparse but otherwise arbitrary vector can be recovered exactly from a small number of randomly constructed linear projections (or samples). The question addressed in this paper is whether an…

Information Theory · Computer Science 2010-01-26 Galen Reeves , Michael Gastpar

In the Minimum Description Length (MDL) principle, learning from the data is equivalent to an optimal coding problem. We show that the codes that achieve optimal compression in MDL are critical in a very precise sense. First, when they are…

Methodology · Statistics 2018-10-03 Ryan John Cubero , Matteo Marsili , Yasser Roudi

We study how neural networks compress uninformative input space in models where data lie in $d$ dimensions, but whose label only vary within a linear manifold of dimension $d_\parallel < d$. We show that for a one-hidden layer network…

Machine Learning · Computer Science 2021-05-07 Jonas Paccolat , Leonardo Petrini , Mario Geiger , Kevin Tyloo , Matthieu Wyart

Semi-supervised learning is attracting blooming attention, due to its success in combining unlabeled data. To mitigate potentially incorrect pseudo labels, recent frameworks mostly set a fixed confidence threshold to discard uncertain…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Lihe Yang , Zhen Zhao , Lei Qi , Yu Qiao , Yinghuan Shi , Hengshuang Zhao

Conformal prediction (CP) is a promising uncertainty quantification framework which works as a wrapper around a black-box classifier to construct prediction sets (i.e., subset of candidate classes) with provable guarantees. However,…

Machine Learning · Computer Science 2025-06-10 Yuanjie Shi , Hooman Shahrokhi , Xuesong Jia , Xiongzhi Chen , Janardhan Rao Doppa , Yan Yan

We initiate the rigorous study of classification in quasi-metric spaces. These are point sets endowed with a distance function that is non-negative and also satisfies the triangle inequality, but is asymmetric. We develop and refine a…

Machine Learning · Computer Science 2019-09-24 Lee-Ad Gottlieb , Shira Ozeri

One of the central issues in the hidden subgroup problem is to bound the sample complexity, i.e., the number of identical samples of coset states sufficient and necessary to solve the problem. In this paper, we present general bounds for…

Quantum Physics · Physics 2008-04-26 Masahito Hayashi , Akinori Kawachi , Hirotada Kobayashi

We study limits of the largest connected components (viewed as metric spaces) obtained by critical percolation on uniformly chosen graphs and configuration models with heavy-tailed degrees. For rank-one inhomogeneous random graphs, such…

Probability · Mathematics 2020-05-11 Shankar Bhamidi , Souvik Dhara , Remco van der Hofstad , Sanchayan Sen

A symmetric tensor category $\mathcal D$ over an algebraically closed field $k$ is incompressible if every tensor functor out of $\mathcal D$ is an embedding. E.g., the categories $Vec$ and $sVec$ of (super)vector spaces are incompressible.…

Category Theory · Mathematics 2023-06-19 Kevin Coulembier , Pavel Etingof , Victor Ostrik

Ensemble methods are among the state-of-the-art predictive modeling approaches. Applied to modern big data, these methods often require a large number of sub-learners, where the complexity of each learner typically grows with the size of…

Machine Learning · Computer Science 2018-10-29 Amichai Painsky , Saharon Rosset

A set of lower bounds on the continuum percolation threshold $\eta_c$ of overlapping convex hyperparticles of general nonspherical (anisotropic) shape with a specified orientational probability distribution in $d$-dimensional Euclidean…

Statistical Mechanics · Physics 2013-03-14 Salvatore Torquato , Yang Jiao

We present a one-shot method for compressing large labeled graphs called Random Edge Coding. When paired with a parameter-free model based on P\'olya's Urn, the worst-case computational and memory complexities scale quasi-linearly and…

Machine Learning · Computer Science 2023-05-18 Daniel Severo , James Townsend , Ashish Khisti , Alireza Makhzani

We solve exactly the general one-dimensional $O(N)$-invariant spin model taking values in the sphere $S^{N-1}$, with nearest-neighbor interactions, in finite volume with periodic boundary conditions, by an expansion in hyperspherical…

High Energy Physics - Lattice · Physics 2015-06-25 Attilio Cucchieri , Tereza Mendes , Andrea Pelissetto , Alan D. Sokal

We study computable probably approximately correct (CPAC) learning, where learners are required to be computable functions. It had been previously observed that the Fundamental Theorem of Statistical Learning, which characterizes PAC…

Machine Learning · Computer Science 2025-11-05 David Kattermann , Lothar Sebastian Krapp

Machine learning models with inputs in a Euclidean space $\mathbb{R}^d$, when implemented on digital computers, generalize, and their generalization gap converges to $0$ at a rate of $c/N^{1/2}$ concerning the sample size $N$. However, the…

Machine Learning · Computer Science 2026-05-14 Anastasis Kratsios , A. Martina Neuman , Gudmund Pammer

Data-free knowledge distillation~(DFKD) is an effective manner to solve model compression and transmission restrictions while retaining privacy protection, which has attracted extensive attention in recent years. Currently, the majority of…

Machine Learning · Computer Science 2025-10-07 Renrong Shao , Wei Zhang , Jun wang

We explore adhesive loose packings of dry small spherical particles of micrometer size using 3D discrete-element simulations with adhesive contact mechanics. A dimensionless adhesion parameter ($Ad$) successfully combines the effects of…

Soft Condensed Matter · Physics 2015-09-10 Wenwei Liu , Shuiqing Li , Adrian Baule , Hernán A. Makse