English
Related papers

Related papers: Indexability, concentration, and VC theory

200 papers

We develop the technique of compactified correspondences and homotopies over one-dimensional base schemes, and illuminate the perfectness and the inverting of characteristic assumptions from the celebrating Voevodsky's strict homotopy…

Algebraic Geometry · Mathematics 2025-02-25 Andrei Druzhinin

We consider the distribution of a graph invariant of central similarity proximity catch digraphs (PCDs) based on one dimensional data. The central similarity PCDs are also a special type of parameterized random digraph family defined with…

Combinatorics · Mathematics 2015-03-17 Elvan Ceyhan

It has been demonstrated many times that the behavior of the human visual system is connected to the statistics of natural images. Since machine learning relies on the statistics of training data as well, the above connection has…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Alexander Hepburn , Valero Laparra , Raul Santos-Rodriguez , Johannes Ballé , Jesús Malo

We study the Vapnik-Chervonenkis (VC) density of definable families in certain stable first-order theories. In particular we obtain uniform bounds on VC density of definable families in finite U-rank theories without the finite cover…

Logic · Mathematics 2016-02-10 M. Aschenbrenner , A. Dolich , D. Haskell , D. Macpherson , S. Starchenko

In this paper, we study the distribution of parallelograms and rhombi in a given set in the plane over arbitrary finite fields $\mathbb{F}_q^2$. As an application, we improve a recent result due to Fitzpatrick, Iosevich, McDonald, and Wyman…

Combinatorics · Mathematics 2023-04-20 Thang Pham

In this work we study the quantitative relation between VC-dimension and two other basic parameters related to learning and teaching. Namely, the quality of sample compression schemes and of teaching sets for classes of low VC-dimension.…

Machine Learning · Computer Science 2016-11-28 Shay Moran , Amir Shpilka , Avi Wigderson , Amir Yehudayoff

As data volumes continue to grow, clustering and outlier detection algorithms are becoming increasingly time-consuming. Classical index structures for neighbor search are no longer sustainable due to the "curse of dimensionality". Instead,…

Databases · Computer Science 2021-05-12 Li Wang

Clustering analysis is of substantial significance for data mining. The properties of big data raise higher demand for more efficient and economical distributed clustering methods. However, existing distributed clustering methods mainly…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-07-03 Yifeng Xiao , Jiang Xue , Deyu Meng

We extend the definitions of complexity measures of functions to domains such as the symmetric group. The complexity measures we consider include degree, approximate degree, decision tree complexity, sensitivity, block sensitivity, and a…

Computational Complexity · Computer Science 2020-10-16 Neta Dafni , Yuval Filmus , Noam Lifshitz , Nathan Lindzey , Marc Vinyals

We propose In-Context Clustering (ICC), a flexible LLM-based procedure for clustering data from diverse distributions. Unlike traditional clustering algorithms constrained by predefined similarity measures, ICC flexibly captures complex…

Machine Learning · Computer Science 2025-10-10 Ying Wang , Mengye Ren , Andrew Gordon Wilson

High-Performance Computing (HPC) systems need to be constantly monitored to ensure their stability. The monitoring systems collect a tremendous amount of data about different parameters or Key Performance Indicators (KPIs), such as resource…

Artificial Intelligence · Computer Science 2023-12-12 Mohamed Soliman Halawa , Rebeca P. Díaz-Redondo , Ana Fernández-Vilas

Given two relations containing multiple measurements - possibly with uncertainties - our objective is to find which sets of attributes from the first have a corresponding set on the second, using exclusively a sample of the data. This…

Databases · Computer Science 2022-07-20 Alejandro Alvarez-Ayllon , Manuel Palomo-Duarte , Juan-Manuel Dodero

This paper deals with a clustering algorithm for histogram data based on a Self-Organizing Map (SOM) learning. It combines a dimension reduction by SOM and the clustering of the data in a reduced space. Related to the kind of data, a…

Machine Learning · Computer Science 2021-09-10 Guénaël Cabanes , Younès Bennani , Rosanna Verde , Antonio Irpino

We study the problem of finding the $k$ most similar trajectories to a given query trajectory. Our work is inspired by the work of Grossi et al. [6] that considers trajectories as walks in a graph. Each visited vertex is accompanied by a…

Data Structures and Algorithms · Computer Science 2020-10-20 Lutz Oettershagen , Anne Driemel , Petra Mutzel

Weakly supervised localization aims at finding target object regions using only image-level supervision. However, localization maps extracted from classification networks are often not accurate due to the lack of fine pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Xiaolin Zhang , Yunchao Wei , Yi Yang

We show how to place constraints on cluster physics by stacking the weak lensing signals from multiple clusters found through the Sunyaev-Zeldovich (SZ) effect. For a survey that covers about 200 sq. deg. both in SZ and weak lensing…

Astrophysics · Physics 2008-11-26 Carolyn Sealfon , Licia Verde , Raul Jimenez

Distribution shifts are problems where the distribution of data changes between training and testing, which can significantly degrade the performance of a model deployed in the real world. Recent studies suggest that one reason for the…

Machine Learning · Computer Science 2023-04-10 Takuro Kutsuna

In recent years, large-scale Bayesian learning draws a great deal of attention. However, in big-data era, the amount of data we face is growing much faster than our ability to deal with it. Fortunately, it is observed that large-scale…

Machine Learning · Computer Science 2022-02-15 Qianqian Song

Despite the prevalence of the attention sink phenomenon in Large Language Models (LLMs), where initial tokens disproportionately monopolize attention scores, its structural origins remain elusive. This work provides a \textit{mechanistic…

Machine Learning · Computer Science 2026-05-08 Siquan Li , Kaiqi Jiang , Jiacheng Sun , Tianyang Hu

Indexing is an effective way to support efficient query processing in large databases. Recently the concept of learned index, which replaces or complements traditional index structures with machine learning models, has been actively…

Databases · Computer Science 2022-08-01 Yao Tian , Tingyun Yan , Xi Zhao , Kai Huang , Xiaofang Zhou