中文
相关论文

相关论文: Optimally compressing VC classes

200 篇论文

Given a set of vectors (the data) in a Hilbert space H, we prove the existence of an optimal collection of subspaces minimizing the sum of the square of the distances between each vector and its closest subspace in the collection. This…

经典分析与常微分方程 · 数学 2008-02-07 Akram Aldroubi , Carlos Cabrelli , Ursula Molter

List learning is a variant of supervised classification where the learner outputs multiple plausible labels for each instance rather than just one. We investigate classical principles related to generalization within the context of list…

机器学习 · 计算机科学 2026-03-24 Steve Hanneke , Shay Moran , Tom Waknine

In this short note, we show that the VC-dimension of the class of $k$-vertex polytopes in $\mathbb R^d$ is at most $8d^2k\log_2k$, answering an old question of Long and Warmuth.

计算几何 · 计算机科学 2020-04-15 Andrey Kupavskii

In-context learning has been extensively validated in large language models. However, the mechanism and selection strategy for in-context example selection, which is a crucial ingredient in this approach, lacks systematic and in-depth…

计算与语言 · 计算机科学 2024-05-21 Zhongxiang Sun , Kepu Zhang , Haoyu Wang , Xiao Zhang , Jun Xu

Dataset condensation aims to condense a large dataset with a lot of training samples into a small set. Previous methods usually condense the dataset into the pixels format. However, it suffers from slow optimization speed and large number…

计算机视觉与模式识别 · 计算机科学 2023-09-15 David Junhao Zhang , Heng Wang , Chuhui Xue , Rui Yan , Wenqing Zhang , Song Bai , Mike Zheng Shou

We study how to construct compressed datasets that suffice to recover optimal decisions in linear programs with an unknown cost vector $c$ lying in a prior set $\mathcal{C}$. Recent work by Bennouna et al. provides an exact geometric…

最优化与控制 · 数学 2026-05-25 Yuhan Ye , Saurabh Amin , Asuman Ozdaglar

We study a model of machine teaching where the teacher mapping is constructed from a size function on both concepts and examples. The main question in machine teaching is the minimum number of examples needed for any concept, the so-called…

组合数学 · 数学 2024-02-12 Brigt Håvardstun , Jan Kratochvíl , Joakim Sunde , Jan Arne Telle

This paper is dedicated to an efficient compression of weights and optimizer states (called checkpoints) obtained at different stages during a neural network training process. First, we propose a prediction-based compression approach, where…

机器学习 · 计算机科学 2025-06-16 Yuriy Kim , Evgeny Belyaev

We study the following basic machine learning task: Given a fixed set of $d$-dimensional input points for a linear regression problem, we wish to predict a hidden response value for each of the points. We can only afford to attain the…

机器学习 · 计算机科学 2018-06-07 Michał Dereziński , Manfred K. Warmuth

Let U be a monster model and let D be a subset of U. Let (U,D) denote theexpansion of U with a new predicate for D. Write e(D) for the collection of all subsets C of U such that (U,C) is elementary equivalent to (U,D). We prove that if e(D)…

逻辑 · 数学 2015-08-21 Domenico Zambella

The declustering problem is to allocate given data on parallel working storage devices in such a manner that typical requests find their data evenly distributed on the devices. Using deep results from discrepancy theory, we improve previous…

离散数学 · 计算机科学 2007-05-23 Benjamin Doerr , Nils Hebbinghaus , Sören Werth

We revisit the so-called sampling and discarding approach used to quantify the probability of constraint violation of a solution to convex scenario programs when some of the original samples are allowed to be discarded. Motivated by two…

最优化与控制 · 数学 2022-04-05 Licio Romao , Antonis Papachristodoulou , Kostas Margellos

The Vapnik-Chervonenkis (VC) dimension of a collection of subsets of a set is an important combinatorial concept in settings such as discrete geometry and machine learning. In this paper we prove that the VC dimension of the family of…

组合数学 · 数学 2017-11-28 Christian J. J. Despres

Many practical prediction algorithms represent inputs in Euclidean space and replace the discrete 0/1 classification loss with a real-valued surrogate loss, effectively reducing classification tasks to stochastic optimization. In this…

机器学习 · 计算机科学 2024-11-19 Bogdan Chornomaz , Shay Moran , Tom Waknine

We study how much a linear program (LP) can be compressed when solved repeatedly, given prior knowledge about its objective function. Existing data-driven projection methods learn low-dimensional surrogate LPs with approximate…

最优化与控制 · 数学 2026-05-26 Yuhan Ye , Omar Bennouna

Compressed sensing is the art of reconstructing structured $n$-dimensional vectors from substantially fewer measurements than naively anticipated. A plethora of analytic reconstruction guarantees support this credo. The strongest among them…

信息论 · 计算机科学 2018-12-20 Peter Jung , Richard Kueng , Dustin G. Mixon

We establish existence of global-in-time weak solutions to the one dimensional, compressible Navier-Stokes system for a viscous and heat conducting ideal polytropic gas (pressure $p=K\theta/\tau$, internal energy $e=c_v \theta$), when the…

偏微分方程分析 · 数学 2009-06-26 Helge Kristian Jenssen , Trygve Karper

In statistical setting of the pattern recognition problem the number of examples required to approximate an unknown labelling function is linear in the VC dimension of the target learning class. In this work we consider the question whether…

机器学习 · 计算机科学 2016-06-27 Daniil Ryabko

Traditionally, data compression deals with the problem of concisely representing a data source, e.g. a sequence of letters, for the purpose of eventual reproduction (either exact or approximate). In this work we are interested in the case…

信息论 · 计算机科学 2013-12-10 Amir Ingber , Tsachy Weissman

Recently, a series of works have started studying variations of concepts from learning theory for product spaces, which can be collected under the name high-arity learning theory. In this work, we consider a high-arity variant of sample…

机器学习 · 计算机科学 2026-05-15 Leonardo N. Coregliano , William Opich