中文
相关论文

相关论文: Bottom-k and Priority Sampling, Set Similarity and…

200 篇论文

We show that for every $k\in\mathbb{N}$ and $\varepsilon>0$, for large enough alphabet $R$, given a $k$-CSP with alphabet size $R$, it is NP-hard to distinguish between the case that there is an assignment satisfying at least…

计算复杂性 · 计算机科学 2025-10-29 Dor Minzer , Kai Zhe Zheng

We study the Densest At-Least-$k$-Subgraph (DAL$k$S) problem, in which we are given an undirected graph $G$ and an integer $k$, and the goal is to find a subgraph of $G$ with at least $k$ vertices with maximum density. The best-known…

数据结构与算法 · 计算机科学 2026-05-26 Bundit Laekhanukit , Pasin Manurangsi , Ohad Trabelsi

We develop a general framework for estimating the $L_\infty(\mathbb{T}^d)$ error for the approximation of multivariate periodic functions belonging to specific reproducing kernel Hilbert spaces (RHKS) using approximants that are…

数值分析 · 数学 2019-09-06 Lutz Kämmerer

We consider the approximate recovery of multivariate periodic functions from a discrete set of function values taken on a rank-$s$ integration lattice. The main result is the fact that any (non-)linear reconstruction algorithm taking…

数值分析 · 数学 2016-08-02 Glenn Byrenheid , Lutz Kämmerer , Tino Ullrich , Toni Volkmer

In this paper, we investigate unconstrained and constrained sample-based federated optimization, respectively. For each problem, we propose a privacy preserving algorithm using stochastic successive convex approximation (SSCA) techniques,…

机器学习 · 计算机科学 2021-03-18 Chencheng Ye , Ying Cui

Low-rank approximation and column subset selection are two fundamental and related problems that are applied across a wealth of machine learning applications. In this paper, we study the question of socially fair low-rank approximation and…

机器学习 · 计算机科学 2024-12-10 Zhao Song , Ali Vakilian , David P. Woodruff , Samson Zhou

Kink model is developed to analyze the data where the regression function is twostage linear but intersects at an unknown threshold. In quantile regression with longitudinal data, previous work assumed that the unknown threshold parameters…

统计方法学 · 统计学 2020-09-07 Chuang Wan

Sample complexity of bias estimation is a lower bound on the runtime of any bias detection method. Many regulatory frameworks require the bias to be tested for all subgroups, whose number grows exponentially with the number of protected…

机器学习 · 计算机科学 2025-02-06 German Martinez Matilla , Jakub Marecek

Blasiok (SODA'18) recently introduced the notion of a subgaussian sampler, defined as an averaging sampler for approximating the mean of functions $f:\{0,1\}^m \to \mathbb{R}$ such that $f(U_m)$ has subgaussian tails, and asked for explicit…

计算复杂性 · 计算机科学 2019-09-19 Rohit Agrawal

In stochastic combinatorial optimization, algorithms differ in their adaptivity: whether or not they query realized randomness and adapt to it. Dean et al. (FOCS '04) formalize the adaptivity gap, which compares the performance of fully…

数据结构与算法 · 计算机科学 2026-03-03 Zohar Barak , Inbal Talgam-Cohen

Kernel $k$-means clustering can correctly identify and extract a far more varied collection of cluster structures than the linear $k$-means clustering algorithm. However, kernel $k$-means clustering is computationally expensive when the…

机器学习 · 计算机科学 2019-02-12 Shusen Wang , Alex Gittens , Michael W. Mahoney

Weighted Hamming distance, as a similarity measure between binary codes and binary queries, provides superior accuracy in search tasks than Hamming distance. However, how to efficiently and accurately find $K$ binary codes that have the…

计算机视觉与模式识别 · 计算机科学 2021-08-11 Zhenyu Weng , Yuesheng Zhu , Ruixin Liu

Finding an optimal matching in a weighted graph is a standard combinatorial problem. We consider its semi-bandit version where either a pair or a full matching is sampled sequentially. We prove that it is possible to leverage a rank-1…

机器学习 · 统计学 2021-08-03 Flore Sentenac , Jialin Yi , Clément Calauzènes , Vianney Perchet , Milan Vojnovic

We describe a new family of $k$-uniform hypergraphs with independent random edges. The hypergraphs have a high probability of being peelable, i.e. to admit no sub-hypergraph of minimum degree $2$, even when the edge density (number of edges…

数据结构与算法 · 计算机科学 2019-07-11 Martin Dietzfelbinger , Stefan Walzer

Hashing is a basic tool for dimensionality reduction employed in several aspects of machine learning. However, the perfomance analysis is often carried out under the abstract assumption that a truly random unit cost hash function is used,…

机器学习 · 统计学 2017-11-27 Søren Dahlgaard , Mathias Bæk Tejs Knudsen , Mikkel Thorup

Combining query answering and data science workloads has become prevalent. An important class of such workloads is top-k queries with a scoring function implemented as an opaque UDF - a black box whose internal structure and scores on the…

数据库 · 计算机科学 2025-03-27 Jiwon Chang , Fatemeh Nargesian

We present a 6-approximation algorithm for the minimum-cost $k$-node connected spanning subgraph problem, assuming that the number of nodes is at least $k^3(k-1)+k$. We apply a combinatorial preprocessing, based on the Frank-Tardos…

离散数学 · 计算机科学 2012-12-18 Joseph Cheriyan , Laszlo A. Vegh

In this paper, we apply an efficient top-$k$ shortest distance routing algorithm to the link prediction problem and test its efficacy. We compare the results with other base line and state-of-the-art methods as well as with the shortest…

社会与信息网络 · 计算机科学 2017-05-09 Andrei Lebedev , JooYoung Lee , Victor Rivera , Manuel Mazzara

Let $A$ be a set and $V$ a real Hilbert space. Let $H$ be a real Hilbert space of functions $f:A\to V$ and assume $H$ is continuously embedded in the Banach space of bounded functions. For $i=1,\cdots,n$, let $(x_i,y_i)\in A\times V$…

泛函分析 · 数学 2022-02-23 Karen Yeressian

The problem of finding a $k \times k$ submatrix of maximum volume of a matrix $A$ is of interest in a variety of applications. For example, it yields a quasi-best low-rank approximation constructed from the rows and columns of $A$. We show…

数值分析 · 数学 2019-02-07 Alice Cortinovis , Daniel Kressner , Stefano Massei