English
Related papers

Related papers: Approximately Minwise Independence with Twisted Ta…

200 papers

Locality-sensitive hashing (LSH) is an effective randomized technique widely used in many machine learning tasks. The cost of hashing is proportional to data dimensions, and thus often the performance bottleneck when dimensionality is high…

Machine Learning · Computer Science 2023-09-28 Zongyuan Tan , Hongya Wang , Bo Xu , Minjie Luo , Ming Du

Over the last decade, the Dip-test of unimodality has gained increasing interest in the data mining community as it is a parameter-free statistical test that reliably rates the modality in one-dimensional samples. It returns a so called…

Machine Learning · Computer Science 2025-04-04 Lena G. M. Bauer , Collin Leiber , Christian Böhm , Claudia Plant

This paper defines the toroidal small world labeling problem that asks for a labeling of the vertices of a network such that the labels possess information that allows a compact routing scheme in the network. We consider the problem over a…

Data Structures and Algorithms · Computer Science 2019-11-13 Santiago Viertel , André Luís Vignatti

Given two nonincreasing $n$-tuples of real numbers $\lambda_n$, $\mu_n$, the Horn problem asks for a description of all nonincreasing $n$-tuples of real numbers $\nu_n$ such that there exist Hermitian matrices $X_n$, $Y_n$ and $Z_n$…

Probability · Mathematics 2026-03-24 Aalok Gangopadhyay , Hariharan Narayanan

Modern statistical analyses often involve testing large numbers of hypotheses. In many situations, these hypotheses may have an underlying tree structure that not only helps determine the order that tests should be conducted but also…

Methodology · Statistics 2019-03-19 Yunxiao Li , Yi-Juan Hu , Glen A. Satten

We present an efficient algorithm for the min-max correlation clustering problem. The input is a complete graph where edges are labeled as either positive $(+)$ or negative $(-)$, and the objective is to find a clustering that minimizes the…

Data Structures and Algorithms · Computer Science 2025-02-19 Nairen Cao , Steven Roche , Hsin-Hao Su

A perfect $H$-tiling in a graph $G$ is a collection of vertex-disjoint copies of a graph $H$ in $G$ that covers all vertices of $G$. Motivated by papers of Bush and Zhao and of Balogh, Treglown, and Wagner, we determine the threshold for…

Combinatorics · Mathematics 2024-11-20 Enrique Gomez-Leos , Ryan R. Martin

Unsupervised binary representation allows fast data retrieval without any annotations, enabling practical application like fast person re-identification and multimedia retrieval. It is argued that conflicts in binary space are one of the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-23 Fangrui Liu , Zheng Liu

We study smoothed analysis of distributed graph algorithms, focusing on the fundamental minimum spanning tree (MST) problem. With the goal of studying the time complexity of distributed MST as a function of the "perturbation" of the input…

Data Structures and Algorithms · Computer Science 2019-11-11 Soumyottam Chatterjee , Gopal Pandurangan , Nguyen Dinh Pham

We consider the weakly supervised binary classification problem where the labels are randomly flipped with probability $1- {\alpha}$. Although there exist numerous algorithms for this problem, it remains theoretically unexplored how the…

Machine Learning · Computer Science 2019-07-16 Xinyang Yi , Zhaoran Wang , Zhuoran Yang , Constantine Caramanis , Han Liu

In this work, we propose a new randomized algorithm for computing a low-rank approximation to a given matrix. Taking an approach different from existing literature, our method first involves a specific biased sampling, with an element being…

Data Structures and Algorithms · Computer Science 2014-10-16 Srinadh Bhojanapalli , Prateek Jain , Sujay Sanghavi

Tries are general purpose data structures for information retrieval. The most significant parameter of a trie is its height $H$ which equals the length of the longest common prefix of any two string in the set $A$ over which the trie is…

Data Structures and Algorithms · Computer Science 2020-03-10 Stefan Eckhardt , Sven Kosub , Johannes Nowak

Testing for association or dependence between pairs of random variables is a fundamental problem in statistics. In some applications, data are subject to selection bias that causes dependence between observations even when it is absent from…

Methodology · Statistics 2020-10-13 Yaniv Tenzer , Micha Mandel , Or Zuk

Thirty years ago, the Robin Hood collision resolution strategy was introduced for open addressing hash tables, and a recurrence equation was found for the distribution of its search cost. Although this recurrence could not be solved…

Data Structures and Algorithms · Computer Science 2016-05-16 Patricio V. Poblete , Alfredo Viola

In the paper we introduce the notion of twisted derivation of a bialgebra. Twisted derivations appear as infinitesimal symmetries of the category of representations. More precisely they are infinitesimal versions of twisted automorphisms of…

Quantum Algebra · Mathematics 2012-04-24 Alexei Davydov

Recent advances in random linear systems on finite fields have paved the way for the construction of constant-time data structures representing static functions and minimal perfect hash functions using less space with respect to existing…

Data Structures and Algorithms · Computer Science 2016-03-24 Marco Genuzio , Giuseppe Ottaviano , Sebastiano Vigna

Cluster randomized trials (CRTs) often enroll large numbers of participants, but due to logistical and fiscal challenges, only a subset of participants may be selected for measurement of certain outcomes, and those sampled may, purposely or…

We show how to assign labels of size $\tilde O(1)$ to the vertices of a directed planar graph $G$, such that from the labels of any three vertices $s,t,f$ we can deduce in $\tilde O(1)$ time whether $t$ is reachable from $s$ in the graph…

Data Structures and Algorithms · Computer Science 2023-07-17 Shiri Chechik , Shay Mozes , Oren Weimann

We propose a new and easily-realizable distributed hash table (DHT) peer-to-peer structure, incorporating a random caching strategy that allows for {\em polylogarithmic search time} while having only a {\em constant cache} size. We also…

Networking and Internet Architecture · Computer Science 2007-05-23 Nima Sarshar , Vwani Roychowdhury

We introduce a new quality measure to assess randomized low-discrepancy point sets of finite size $n$. This new quality measure, which we call "pairwise sampling dependence index", is based on the concept of negative dependence. A negative…

Statistics Theory · Mathematics 2021-09-08 C. Lemieux , J. Wiart