中文
相关论文

相关论文: Random Sampling for Group-By Queries

200 篇论文

In the current world, OLAP (Online Analytical Processing) is used intensively by modern organizations to perform ad hoc analysis of data, providing insight for better decision making. Thus, the performance for OLAP is crucial; however, it…

数据库 · 计算机科学 2022-04-15 Pritom Saha Akash , Wei-Cheng Lai , Po-Wen Lin

Subset sampling (also known as Poisson sampling), where the decision to include any specific element in the sample is made independently of all others, is a fundamental primitive in data analytics, enabling efficient approximation by…

数据库 · 计算机科学 2025-12-19 Aryan Esmailpour , Xiao Hu , Jinchao Huang , Stavros Sintos

Clustering is a fundamental task in machine learning and data analysis, but it frequently fails to provide fair representation for various marginalized communities defined by multiple protected attributes -- a shortcoming often caused by…

机器学习 · 计算机科学 2025-11-17 Diptarka Chakraborty , Kushagra Chatterjee , Debarati Das , Tien-Long Nguyen

Industry-grade database systems are expected to produce the same result if the same query is repeatedly run on the same input. However, the numerous sources of non-determinism in modern systems make reproducible results difficult to…

数据库 · 计算机科学 2020-04-07 Ingo Müller , Andrea Arteaga , Torsten Hoefler , Gustavo Alonso

We consider the problem of identifying the defectives from a population of items via a non-adaptive group testing framework with a random pooling-matrix design. We analyze the sufficient number of tests needed for approximate set…

信息论 · 计算机科学 2024-12-03 Sameera Bharadwaja H. , Chandra R. Murthy

The Central Limit Theorem provides a foundation for inferential statistics and hypothesis testing. It describes how standardized statistics behave under repeated sampling from large populations. However, if the size of the sample (n)…

统计方法学 · 统计学 2026-05-19 Mike Crowhurst

While large-scale pre-trained language models like BERT have advanced the state-of-the-art in IR, its application in query performance prediction (QPP) is so far based on pointwise modeling of individual queries. Meanwhile, recent studies…

信息检索 · 计算机科学 2022-04-26 Xiaoyang Chen , Ben He , Le Sun

Quantum computers are now on the brink of outperforming their classical counterparts. One way to demonstrate the advantage of quantum computation is through quantum random sampling performed on quantum computing devices. However, existing…

In computational science workflows, it is often the case that 1) objective functions for optimization involve multiple simulation outputs, and 2) those simulations can be performed (at least partially) in parallel. In this work, we…

最优化与控制 · 数学 2026-05-28 Matt Menickelly

We present a new algorithmic framework for grouped variable selection that is based on discrete mathematical optimization. While there exist several appealing approaches based on convex relaxations and nonconvex heuristics, we focus on…

统计方法学 · 统计学 2021-10-19 Hussein Hazimeh , Rahul Mazumder , Peter Radchenko

Crowdsourcing is becoming increasingly important in entity resolution tasks due to their inherent complexity such as clustering of images and natural language processing. Humans can provide more insightful information for these difficult…

数据库 · 计算机科学 2017-08-28 Vijaya Krishna Yalavarthi , Xiangyu Ke , Arijit Khan

Sampling from very large spatial populations is challenging. The solutions suggested in recent literature on this subject often require that the randomly selected units are well distributed across the study region by using complex…

统计方法学 · 统计学 2017-10-26 Roberto Benedetti , Federica Piersimoni

Several clustering frameworks with interactive (semi-supervised) queries have been studied in the past. Recently, clustering with same-cluster queries has become popular. An algorithm in this setting has access to an oracle with full…

数据结构与算法 · 计算机科学 2019-08-15 Barna Saha , Sanjay Subramanian

Clustered standard errors and approximate randomization tests are popular inference methods that allow for dependence within observations. However, they require researchers to know the cluster structure ex ante. We propose a procedure to…

计量经济学 · 经济学 2022-01-14 Yong Cai

The note studies the problem of selecting a good enough subset out of a finite number of alternatives under a fixed simulation budget. Our work aims to maximize the posterior probability of correctly selecting a good subset. We formulate…

最优化与控制 · 数学 2023-05-09 Gongbo Zhang , Bin Chen , Qing-shan Jia , Yijie Peng

Many data sources are naturally modeled by multiple weight assignments over a set of keys: snapshots of an evolving database at multiple points in time, measurements collected over multiple time periods, requests for resources served at…

数据库 · 计算机科学 2010-11-11 Edith Cohen , Haim Kaplan , Subhabrata Sen

We study the practical consequences of dataset sampling strategies on the performance of recommendation algorithms. Recommender systems are generally trained and evaluated on samples of larger datasets. Samples are often taken in a naive or…

信息检索 · 计算机科学 2021-07-13 Noveen Sachdeva , Carole-Jean Wu , Julian McAuley

Benchmarking methods that can be adapted to multi-qubit systems are essential for assessing the overall or "holistic" performance of nascent quantum processors. The current industry standard is Clifford randomized benchmarking (RB), which…

Quality assurance is one the most important challenges in crowdsourcing. Assigning tasks to several workers to increase quality through redundant answers can be expensive if asking homogeneous sources. This limitation has been overlooked by…

机器学习 · 计算机科学 2015-08-12 Besmira Nushi , Adish Singla , Anja Gruenheid , Erfan Zamanian , Andreas Krause , Donald Kossmann

Non-adaptive group testing refers to the problem of inferring a sparse set of defectives from a larger population using the minimum number of simultaneous pooled tests. Recent positive results for noiseless group testing have motivated the…

信息论 · 计算机科学 2021-07-16 Gabriel Arpino , Nicolò Grometto , Afonso S. Bandeira