中文
相关论文

相关论文: A Concentration of Measure Approach to Database De…

200 篇论文

Consider generalized adapted stochastic integrals with respect to independently scattered random measures with second moments. We use a decoupling technique, known as the "principle of conditioning", to study their stable convergence…

概率论 · 数学 2007-05-23 Giovanni Peccati , Murad S. Taqqu

Modern distributed systems often rely on so called weakly-consistent databases, which achieve scalability by sacrificing the consistency guarantee of distributed transaction processing. Such databases have been formalised in two different…

计算机科学中的逻辑 · 计算机科学 2017-08-02 Andrea Cerone , Alexey Gotsman , Hongseok Yang

The concept of matching dependencies (mds) is recently pro- posed for specifying matching rules for object identification. Similar to the functional dependencies (with conditions), mds can also be applied to various data quality…

数据库 · 计算机科学 2009-06-13 Shaoxu Song , Lei Chen

The concentration of empirical measures is studied for dependent data, whose joint distribution satisfies Poincar\'{e}-type or logarithmic Sobolev inequalities. The general concentration results are then applied to spectral empirical…

统计理论 · 数学 2010-11-30 S. G. Bobkov , F. Götze

We consider Markovian models on graphs with local dynamics. We show that, under suitable conditions, such Markov chains exhibit both rapid convergence to equilibrium and strong concentration of measure in the stationary distribution. We…

概率论 · 数学 2008-09-30 Malwina J. Luczak

We derive novel concentration inequalities that bound the statistical error for a large class of stochastic optimization problems, focusing on the case of unbounded objective functions. Our derivations utilize the following key tools: 1) A…

机器学习 · 统计学 2026-01-01 Jeremiah Birrell

Statistical tasks such as density estimation and approximate Bayesian inference often involve densities with unknown normalising constants. Score-based methods, including score matching, are popular techniques as they are free of…

机器学习 · 统计学 2021-12-22 Li K. Wenliang , Heishiro Kanagawa

Database alignment is a variant of the graph alignment problem: Given a pair of anonymized databases containing separate yet correlated features for a set of users, the problem is to identify the correspondence between the features and…

信息论 · 计算机科学 2023-07-06 Osman Emre Dai , Daniel Cullina , Negar Kiyavash

Bound-to-Bound Data Collaboration (B2BDC) provides a natural framework for addressing both forward and inverse uncertainty quantification problems. In this approach, QOI (quantity of interest) models are constrained by related experimental…

最优化与控制 · 数学 2019-04-02 Arun Hegde , Wenyu Li , James Oreluk , Andrew Packard , Michael Frenklach

Object counting models suffer when deployed across domains with differing density variety, since density shifts are inherently task-relevant and violate standard domain adaptation assumptions. To address this, we propose a theoretical…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Zhuonan Liang , Dongnan Liu , Jianan Fan , Yaxuan Song , Qiang Qu , Runnan Chen , Yu Yao , Peng Fu , Weidong Cai

We present an approach to computing consistent answers to queries possibly involving an aggregation operator in databases operating under a star schema and possibly containing missing values and inconsistent data. Our approach is based on…

数据库 · 计算机科学 2026-02-05 Dominique Laurent , Nicolas Spyratos

Measuring the concentration of random variables is a fundamental concept in probability and statistics. Here, we explore a type of concentration measure for continuous random variables with bounded support and use it to provide a notion of…

统计理论 · 数学 2024-06-06 S. Portnoy , N. Torrado , J. J. P. Veerman

Motivated by problems in high-dimensional statistics such as mixture modeling for classification and clustering, we consider the behavior of radial densities as the dimension increases. We establish a form of concentration of measure, and…

统计理论 · 数学 2016-09-13 Ery Arias-Castro , Xiao Pu

Increasing amounts of available data have led to a heightened need for representing large-scale probabilistic knowledge bases. One approach is to use a probabilistic database, a model with strong assumptions that allow for efficiently…

人工智能 · 计算机科学 2019-04-04 Tal Friedman , Guy Van den Broeck

We place ourselves in the setting of high-dimensional statistical inference, where the number of variables $p$ in a data set of interest is of the same order of magnitude as the number of observations $n$. More formally, we study the…

概率论 · 数学 2009-12-11 Noureddine El Karoui

This paper presents a study of the characteristics of transactional databases used in frequent itemset mining. Such characterizations have typically been used to benchmark and understand the data mining algorithms working on these…

数据库 · 计算机科学 2020-11-10 Christian Lezcano , Marta Arias

During the last two decades, concentration of measure has been a subject of various exciting developments in convex geometry, functional analysis, statistical physics, high-dimensional statistics, probability theory, information theory,…

信息论 · 计算机科学 2015-10-13 Maxim Raginsky , Igal Sason

The formal verification of large probabilistic models is important and challenging. Exploiting the concurrency that is often present is one way to address this problem. Here we study a restricted class of asynchronous distributed…

分布式、并行与集群计算 · 计算机科学 2014-08-06 Sumit Kumar Jha , Madhavan Mukund , Ratul Saha , P S Thiagarajan

Matching dependencies (MDs) were introduced to specify the identification or matching of certain attribute values in pairs of database tuples when some similarity conditions are satisfied. Their enforcement can be seen as a natural…

数据库 · 计算机科学 2010-08-30 Jaffer Gardezi , Leopoldo Bertossi , Iluju Kiringa

The field of property testing of probability distributions, or distribution testing, aims to provide fast and (most likely) correct answers to questions pertaining to specific aspects of very large datasets. In this work, we consider a…

数据结构与算法 · 计算机科学 2015-04-27 Clément L. Canonne