中文
相关论文

相关论文: A Concentration of Measure Approach to Database De…

200 篇论文

In applied probability, the normal approximation is often used for the distribution of data with assumed additive structure. This tradition is based on the central limit theorem for sums of (independent) random variables. However, it is…

概率论 · 数学 2020-10-27 Alexandra Dorofeeva , Victor Korolev , Alexander Zeifman

Concentration of measure is a phenomenon in which a random variable that depends in a smooth way on a large number of independent random variables is essentially constant. The random variable will "concentrate" around its median or…

概率论 · 数学 2015-08-25 Meg Walters

--- the companies populating a Stock market, along with their connections, can be effectively modeled through a directed network, where the nodes represent the companies, and the links indicate the ownership. This paper deals with this…

统计金融 · 定量金融 2018-07-26 Roy Cerqueti , Giulia Rotundo , Marcel Ausloos

Constrained sequential pattern mining aims at identifying frequent patterns on a sequential database of items while observing constraints defined over the item attributes. We introduce novel techniques for constraint-based sequential…

机器学习 · 计算机科学 2019-01-01 Amin Hosseininasab , Willem-Jan van Hoeve , Andre A. Cire

Metric clustering is fundamental in areas ranging from Combinatorial Optimization and Data Mining, to Machine Learning and Operations Research. However, in a variety of situations we may have additional requirements or knowledge, distinct…

While clustering is ubiquitously used across science and industry, uncertainty in cluster assignments is rarely quantified with rigorous guarantees. We propose a novel conformal inference framework for clustering that returns confidence…

统计方法学 · 统计学 2026-04-13 YoonHaeng Hur , Anirban Nath , Genevera Allen

The problem of appropriately matching items subject to compatibility constraints arises in a number of important applications. While most of the literature on matching theory focuses on a static setting with a fixed number of items, several…

概率论 · 数学 2022-01-04 Céline Comte

When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large…

统计理论 · 数学 2015-05-19 Alfred O. Hero , Bala Rajaratnam

Many analyses require linking records from two databases comprising overlapping sets of individuals. In the absence of unique identifiers, the linkage procedure often involves matching on a set of categorical variables, such as…

应用统计 · 统计学 2017-06-12 Nicole M. Dalzell , Jerome P. Reiter

We propose and illustrate a hierarchical Bayesian approach for matching statistical records observed on different occasions. We show how this model can be profitably adopted both in record linkage problems and in capture--recapture setups,…

应用统计 · 统计学 2011-07-29 Andrea Tancredi , Brunero Liseo

Self-consistency improves reasoning by aggregating diverse stochastic samples, yet the dynamics behind its efficacy remain underexplored. We reframe self-consistency as a dynamic distributional alignment problem, revealing that decoding…

计算与语言 · 计算机科学 2025-06-12 Yiwei Li , Ji Zhang , Shaoxiong Feng , Peiwen Yuan , Xinglin Wang , Jiayi Shi , Yueqi Zhang , Chuyi Tan , Boyuan Pan , Yao Hu , Kan Li

The objective of model reference control is to design a controller that regulates the system's behavior so as to match a specified reference model. This paper investigates necessary and sufficient conditions for model reference control from…

最优化与控制 · 数学 2024-11-01 Jiwei Wang , Simone Baldi , Henk J. van Waarde

This paper studies the sample complexity of searching over multiple populations. We consider a large number of populations, each corresponding to either distribution P0 or P1. The goal of the search problem studied here is to find one…

信息论 · 计算机科学 2016-11-17 Matthew L. Malloy , Gongguo Tang , Robert D. Nowak

Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a combinatorially large number of novel data? What signal in the…

Rapid growth of genetic databases means huge savings from improvements in their data compression, what requires better inexpensive statistical models. This article proposes automatized optimizations e.g. of Markov-like models, especially…

信息论 · 计算机科学 2022-05-04 Jarek Duda

The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…

统计方法学 · 统计学 2025-05-16 Luca Scrucca

Model selection in clustering requires (i) to specify a suitable clustering principle and (ii) to control the model order complexity by choosing an appropriate number of clusters depending on the noise level in the data. We advocate an…

信息论 · 计算机科学 2010-06-03 Joachim M. Buhmann

Threshold-type counts based on multivariate occupancy models with log concave marginals admit bounded size biased couplings under weak conditions, leading to new concentration of measure results for random graphs, germ-grain models in…

概率论 · 数学 2017-05-25 Jay Bartroff , Larry Goldstein , Ümit Işlak

In high-dimensional problems, choosing a prior distribution such that the corresponding posterior has desirable practical and theoretical properties can be challenging. This begs the question: can the data be used to help choose a good…

统计理论 · 数学 2019-09-25 Ryan Martin , Stephen G. Walker

Different ways of entering data into databases result in duplicate records that cause increasing of databases' size. This is a fact that we cannot ignore it easily. There are several methods that are used for this purpose. In this paper, we…

数据库 · 计算机科学 2011-12-15 Mohammad-Reza Feizi-Derakhshi , Azade Roohany