中文
相关论文

相关论文: Approximately Minwise Independence with Twisted Ta…

200 篇论文

Locality-sensitive hashing (LSH) is an effective randomized technique widely used in many machine learning tasks. The cost of hashing is proportional to data dimensions, and thus often the performance bottleneck when dimensionality is high…

机器学习 · 计算机科学 2023-09-28 Zongyuan Tan , Hongya Wang , Bo Xu , Minjie Luo , Ming Du

Over the last decade, the Dip-test of unimodality has gained increasing interest in the data mining community as it is a parameter-free statistical test that reliably rates the modality in one-dimensional samples. It returns a so called…

机器学习 · 计算机科学 2025-04-04 Lena G. M. Bauer , Collin Leiber , Christian Böhm , Claudia Plant

This paper defines the toroidal small world labeling problem that asks for a labeling of the vertices of a network such that the labels possess information that allows a compact routing scheme in the network. We consider the problem over a…

数据结构与算法 · 计算机科学 2019-11-13 Santiago Viertel , André Luís Vignatti

Given two nonincreasing $n$-tuples of real numbers $\lambda_n$, $\mu_n$, the Horn problem asks for a description of all nonincreasing $n$-tuples of real numbers $\nu_n$ such that there exist Hermitian matrices $X_n$, $Y_n$ and $Z_n$…

概率论 · 数学 2026-03-24 Aalok Gangopadhyay , Hariharan Narayanan

Modern statistical analyses often involve testing large numbers of hypotheses. In many situations, these hypotheses may have an underlying tree structure that not only helps determine the order that tests should be conducted but also…

统计方法学 · 统计学 2019-03-19 Yunxiao Li , Yi-Juan Hu , Glen A. Satten

We present an efficient algorithm for the min-max correlation clustering problem. The input is a complete graph where edges are labeled as either positive $(+)$ or negative $(-)$, and the objective is to find a clustering that minimizes the…

数据结构与算法 · 计算机科学 2025-02-19 Nairen Cao , Steven Roche , Hsin-Hao Su

A perfect $H$-tiling in a graph $G$ is a collection of vertex-disjoint copies of a graph $H$ in $G$ that covers all vertices of $G$. Motivated by papers of Bush and Zhao and of Balogh, Treglown, and Wagner, we determine the threshold for…

组合数学 · 数学 2024-11-20 Enrique Gomez-Leos , Ryan R. Martin

Unsupervised binary representation allows fast data retrieval without any annotations, enabling practical application like fast person re-identification and multimedia retrieval. It is argued that conflicts in binary space are one of the…

计算机视觉与模式识别 · 计算机科学 2020-11-23 Fangrui Liu , Zheng Liu

We study smoothed analysis of distributed graph algorithms, focusing on the fundamental minimum spanning tree (MST) problem. With the goal of studying the time complexity of distributed MST as a function of the "perturbation" of the input…

数据结构与算法 · 计算机科学 2019-11-11 Soumyottam Chatterjee , Gopal Pandurangan , Nguyen Dinh Pham

We consider the weakly supervised binary classification problem where the labels are randomly flipped with probability $1- {\alpha}$. Although there exist numerous algorithms for this problem, it remains theoretically unexplored how the…

机器学习 · 计算机科学 2019-07-16 Xinyang Yi , Zhaoran Wang , Zhuoran Yang , Constantine Caramanis , Han Liu

In this work, we propose a new randomized algorithm for computing a low-rank approximation to a given matrix. Taking an approach different from existing literature, our method first involves a specific biased sampling, with an element being…

数据结构与算法 · 计算机科学 2014-10-16 Srinadh Bhojanapalli , Prateek Jain , Sujay Sanghavi

Tries are general purpose data structures for information retrieval. The most significant parameter of a trie is its height $H$ which equals the length of the longest common prefix of any two string in the set $A$ over which the trie is…

数据结构与算法 · 计算机科学 2020-03-10 Stefan Eckhardt , Sven Kosub , Johannes Nowak

Testing for association or dependence between pairs of random variables is a fundamental problem in statistics. In some applications, data are subject to selection bias that causes dependence between observations even when it is absent from…

统计方法学 · 统计学 2020-10-13 Yaniv Tenzer , Micha Mandel , Or Zuk

Thirty years ago, the Robin Hood collision resolution strategy was introduced for open addressing hash tables, and a recurrence equation was found for the distribution of its search cost. Although this recurrence could not be solved…

数据结构与算法 · 计算机科学 2016-05-16 Patricio V. Poblete , Alfredo Viola

In the paper we introduce the notion of twisted derivation of a bialgebra. Twisted derivations appear as infinitesimal symmetries of the category of representations. More precisely they are infinitesimal versions of twisted automorphisms of…

量子代数 · 数学 2012-04-24 Alexei Davydov

Recent advances in random linear systems on finite fields have paved the way for the construction of constant-time data structures representing static functions and minimal perfect hash functions using less space with respect to existing…

数据结构与算法 · 计算机科学 2016-03-24 Marco Genuzio , Giuseppe Ottaviano , Sebastiano Vigna

Cluster randomized trials (CRTs) often enroll large numbers of participants, but due to logistical and fiscal challenges, only a subset of participants may be selected for measurement of certain outcomes, and those sampled may, purposely or…

统计方法学 · 统计学 2023-05-16 Joshua R. Nugent , Carina Marquez , Edwin D. Charlebois , Rachel Abbott , Laura B. Balzer

We show how to assign labels of size $\tilde O(1)$ to the vertices of a directed planar graph $G$, such that from the labels of any three vertices $s,t,f$ we can deduce in $\tilde O(1)$ time whether $t$ is reachable from $s$ in the graph…

数据结构与算法 · 计算机科学 2023-07-17 Shiri Chechik , Shay Mozes , Oren Weimann

We propose a new and easily-realizable distributed hash table (DHT) peer-to-peer structure, incorporating a random caching strategy that allows for {\em polylogarithmic search time} while having only a {\em constant cache} size. We also…

网络与互联网体系结构 · 计算机科学 2007-05-23 Nima Sarshar , Vwani Roychowdhury

We introduce a new quality measure to assess randomized low-discrepancy point sets of finite size $n$. This new quality measure, which we call "pairwise sampling dependence index", is based on the concept of negative dependence. A negative…

统计理论 · 数学 2021-09-08 C. Lemieux , J. Wiart