English
Related papers

Related papers: A Concentration of Measure Approach to Database De…

200 papers

We are concerned in clustering continuous data sets subject to non-ignorable missingness. We perform clustering with a specific semi-parametric mixture, under the assumption of conditional independence given the component. The mixture model…

Methodology · Statistics 2021-07-20 Marie Du Roy de Chaumaray , Matthieu Marbac

Schema matching constitutes a pivotal phase in the data ingestion process for contemporary database systems. Its objective is to discern pairwise similarities between two sets of attributes, each associated with a distinct data table. This…

We propose a norm of consistency for a mixed set of defeasible and strict sentences, based on a probabilistic semantics. This norm establishes a clear distinction between knowledge bases depicting exceptions and those containing outright…

Artificial Intelligence · Computer Science 2013-04-08 Moises Goldszmidt , Judea Pearl

We present an approach to model-based hierarchical clustering by formulating an objective function based on a Bayesian analysis. This model organizes the data into a cluster hierarchy while specifying a complex feature-set partitioning that…

Machine Learning · Computer Science 2013-01-18 Shivakumar Vaithyanathan , Byron E Dom

Consistent query answering is the problem of computing the answers from a database that are consistent with respect to certain integrity constraints that the database as a whole may fail to satisfy. Those answers are characterized as those…

Databases · Computer Science 2007-05-23 L. Bertossi , L. Bravo , E. Franconi , A. Lopatenko

A classical problem in causal inference is that of matching, where treatment units need to be matched to control units based on covariate information. In this work, we propose a method that computes high quality almost-exact matches for…

Machine Learning · Statistics 2021-02-16 Tianyu Wang , Marco Morucci , M. Usaid Awan , Yameng Liu , Sudeepa Roy , Cynthia Rudin , Alexander Volfovsky

A general stochastic approach to the description of coagulating aerosol system is developed. As the object of description one can consider arbitrary mesoscopic values (number of aerosol clusters, their size etc). The birth-and-death…

Chemical Physics · Physics 2011-10-12 V. V. Ryazanov

Recently, different works proposed a new way to mine patterns in databases with pathological size. For example, experiments in genome biology usually provide databases with thousands of attributes (genes) but only tens of objects…

Machine Learning · Computer Science 2009-02-10 Baptiste Jeudy , François Rioult

Process discovery algorithms automatically extract process models from event logs, but high variability often results in complex and hard-to-understand models. To mitigate this issue, trace clustering techniques group process executions…

Machine Learning · Computer Science 2025-12-11 Jari Peeperkorn , Johannes De Smedt , Jochen De Weerdt

This paper considers data-driven chance-constrained stochastic optimization problems in a Bayesian framework. Bayesian posteriors afford a principled mechanism to incorporate data and prior knowledge into stochastic optimization problems.…

Statistics Theory · Mathematics 2023-08-07 Prateek Jaiswal , Harsha Honnappa , Vinayak A. Rao

Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between…

Machine Learning · Statistics 2017-09-29 Sebastijan Dumancic , Hendrik Blockeel

Much is now known about the consistency of Bayesian updating on infinite-dimensional parameter spaces with independent or Markovian data. Necessary conditions for consistency include the prior putting enough weight on the correct…

Statistics Theory · Mathematics 2022-03-18 Cosma Rohilla Shalizi

Statistical heterogeneity is a measure of how skewed the samples of a dataset are. It is a common problem in the study of differential privacy that the usage of a statistically heterogeneous dataset results in a significant loss of…

Machine Learning · Computer Science 2024-12-02 Mary Scott , Graham Cormode , Carsten Maple

We consider a finite or countable collection of one-dimensional Brownian particles whose dynamics at any point in time is determined by their rank in the entire particle system. Using Transportation Cost Inequalities for stochastic…

Probability · Mathematics 2010-11-11 Soumik Pal , Mykhaylo Shkolnikov

The present work provides an original framework for random matrix analysis based on revisiting the concentration of measure theory from a probabilistic point of view. By providing various notions of vector concentration ($q$-exponential,…

Probability · Mathematics 2021-01-19 Cosme Louart , Romain Couillet

We use monads to relax the atomicity requirement for data in a database. Depending on the choice of monad, the database fields may contain generalized values such as lists or sets of values, or they may contain exceptions such as various…

Databases · Computer Science 2012-09-06 David I. Spivak

We apply random matrix theory to study the impact of measurement uncertainty on dynamic mode decomposition. Specifically, when the measurements follow a normal probability density function, we show how the moments of that density propagate…

Methodology · Statistics 2025-09-04 P. Algikar , P. Sharma , M. Netto , L. Mili

In order to increase the efficiency of the computer simulation of biological molecules, it is very common to impose holonomic constraints on the fastest degrees of freedom; normally bond lengths, but also possibly bond angles. However, as…

Chemical Physics · Physics 2011-12-19 Pablo Echenique , Claudio N. Cavasotto , Pablo García-Risueño

The growing expanse of e-commerce and the widespread availability of online databases raise many fears regarding loss of privacy and many statistical challenges. Even with encryption and other nominal forms of protection for individual…

Statistics Theory · Mathematics 2007-06-13 Stephen E. Fienberg

This paper develops new limit theory for data that are generated by networks or more generally display cross-sectional dependence structures that are governed by observable and unobservable characteristics. Strategic network formation…

Probability · Mathematics 2019-08-08 Guido M. Kuersteiner
‹ Prev 1 4 5 6 7 8 10 Next ›