中文
相关论文

相关论文: A Concentration of Measure Approach to Database De…

200 篇论文

Concentration of measure is studied, and obtained, for stable and related random vectors.

概率论 · 数学 2007-05-23 Christian Houdre , Philippe Marchal

This paper defines a constraint-based model dedicated to multidimensional databases. The model we define represents data through a constellation of facts (subjects of analyse) associated to dimensions (axis of analyse), which are possibly…

数据库 · 计算机科学 2010-05-20 Faiza Ghozzi , Franck Ravat , Olivier Teste , Gilles Zurfluh

We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…

统计理论 · 数学 2018-05-22 Takumi Saegusa

The community structure of complex networks reveals both their organization and hidden relationships among their constituents. Most community detection methods currently available are not deterministic, and their results typically depend on…

物理与社会 · 物理学 2012-03-29 Andrea Lancichinetti , Santo Fortunato

We define a stochastic model of a two-sided limit order book in terms of its key quantities \textit{best bid [ask] price} and the \textit{standing buy [sell] volume density}. For a simple scaling of the discreteness parameters, that keeps…

数理金融 · 定量金融 2015-01-06 Ulrich Horst , Michael Paulsen

Considering the high heterogeneity of the ontologies pub-lished on the web, ontology matching is a crucial issue whose aim is to establish links between an entity of a source ontology and one or several entities from a target ontology.…

人工智能 · 计算机科学 2015-01-26 Amira Essaid , Arnaud Martin , Grégory Smits , Boutheina Ben Yaghlane

In an earlier paper Rakonczai et al. (2014), we have emphasized the effective sample size for autocorrelated data. The simulations were based on the block bootstrap methodology. However, the discreteness of the usual block size did not…

统计理论 · 数学 2016-06-02 László Varga , András Zempléni

We consider the analysis of high dimensional data given in the form of a matrix with columns consisting of observations and rows consisting of features. Often the data is such that the observations do not reside on a regular grid, and the…

机器学习 · 统计学 2017-08-22 Gal Mishne , Ronen Talmon , Israel Cohen , Ronald R. Coifman , Yuval Kluger

Mathematical inequalities, combined with atomic-physics sum rules, enable one to derive lower and upper bounds for the Rosseland and/or Planck mean opacities. The resulting constraints must be satisfied, either for pure elements or…

原子物理 · 物理学 2023-11-17 Jean-Christophe Pain , Patricia Croset

We propose some axioms for hierarchical clustering of probability measures and investigate their ramifications. The basic idea is to let the user stipulate the clusters for some elementary measures. This is done without the need of any…

机器学习 · 统计学 2016-05-24 Philipp Thomann , Ingo Steinwart , Nico Schmid

The re-identification or de-anonymization of users from anonymized data through matching with publicly-available correlated user data has raised privacy concerns, leading to the complementary measure of obfuscation in addition to…

信息论 · 计算机科学 2022-09-16 Serhat Bakirtas , Elza Erkip

This thesis develops exact analytical tools to study strongly correlated stochastic systems, with a focus on extreme value statistics, gap statistics, and full counting statistics in multi-particle processes. A central contribution is the…

统计力学 · 物理学 2025-08-19 Marco Biroli

We consider the problem of sequential matching in a stochastic block model with several classes of nodes and generic compatibility constraints. When the probabilities of connections do not scale with the size of the graph, we show that…

概率论 · 数学 2026-01-14 Nahuel Soprano-Loto , Matthieu Jonckheere , Pascal Moyal

For noncorrelated random variables, we study a concentration property of the family of distributions of normalized sums formed by sequences of times of a given large length.

概率论 · 数学 2007-05-23 Sergey G. Bobkov

Over the past three decades, synthetic data methods for statistical disclosure control have continually evolved, but mainly within the domain of survey data sets. There are certain characteristics of administrative databases, such as their…

统计方法学 · 统计学 2022-05-13 James Edward Jackson , Robin Mitra , Brian Joseph Francis , Iain Dove

This paper provides conditions on the observation probability distribution in Bayesian localization and optimal filtering so that the conditional mean estimate satisfies convex stochastic dominance. Convex dominance allows us to compare the…

系统与控制 · 计算机科学 2019-10-29 Vikram Krishnamurthy

Mining association rules is a popular and well researched method for discovering interesting relations between variables in large databases. A practical problem is that at medium to low support values often a large number of frequent…

数据库 · 计算机科学 2008-12-18 Michael Hahsler , Christian Buchta , Kurt Hornik

System modeling is a classical approach to ensure their reliability since it is suitable both for a formal verification and for software testing techniques. In the context of model-based testing an approach combining random testing and…

软件工程 · 计算机科学 2018-06-14 Julien Bernard , Pierre-Cyrille Héam , Olga Kouchnarenko

Large, data centric applications are characterized by its different attributes. In modern day, a huge majority of the large data centric applications are based on relational model. The databases are collection of tables and every table…

信息检索 · 计算机科学 2012-06-28 Soumya Sen , Anjan Dutta , Agostino Cortesi , Nabendu Chaki

We formulate conditions for convergence of Laws of Large Numbers and show its links with of the parts of mathematical analysis such as summation theory, convergence of orthogonal series. We present also applications of the Law of Large…

概率论 · 数学 2018-09-07 Paweł J. Szabłowski