中文
相关论文

相关论文: Detecting duplicates in a homicide registry using …

200 篇论文

In this paper, we study the impact of combining profile and network data in a de-duplication setting. We also assess the influence of a range of prior distributions on the linkage structure. Furthermore, we explore stochastic gradient…

统计方法学 · 统计学 2021-11-17 Juan Sosa , Abel Rodriguez

This paper introduces the Partition Tree Weighting technique, an efficient meta-algorithm for piecewise stationary sources. The technique works by performing Bayesian model averaging over a large class of possible partitions of the data…

信息论 · 计算机科学 2012-11-22 Joel Veness , Martha White , Michael Bowling , András György

Markov Chain Monte Carlo (MCMC) is a well-established family of algorithms primarily used in Bayesian statistics to sample from a target distribution when direct sampling is challenging. Existing work on Bayesian decision trees uses MCMC.…

统计计算 · 统计学 2023-01-24 Efthyvoulos Drousiotis , Paul G. Spirakis , Simon Maskell

Bayesian change-point detection, together with latent variable models, allows to perform segmentation over high-dimensional time-series. We assume that change-points lie on a lower-dimensional manifold where we aim to infer subsets of…

机器学习 · 统计学 2020-11-04 Lorena Romero-Medrano , Pablo Moreno-Muñoz , Antonio Artés-Rodríguez

Generative retrieval represents a novel approach to information retrieval. It uses an encoder-decoder architecture to directly produce relevant document identifiers (docids) for queries. While this method offers benefits, current approaches…

信息检索 · 计算机科学 2024-09-30 Yubao Tang , Ruqing Zhang , Jiafeng Guo , Maarten de Rijke , Wei Chen , Xueqi Cheng

The uncertainty of classification outcomes is of crucial importance for many safety critical applications including, for example, medical diagnostics. In such applications the uncertainty of classification can be reliably estimated within a…

人工智能 · 计算机科学 2007-05-23 V. Schetinin , J. E. Fieldsend , D. Partridge , W. J. Krzanowski , R. M. Everson , T. C. Bailey , A. Hernandez

We present a procedure to diagnose model misspecification in situations where inference is performed using approximate Bayesian computation. We demonstrate theoretically, and empirically that this procedure can consistently detect the…

统计方法学 · 统计学 2022-10-25 Andrés Ramírez-Hassan , David T. Frazier

In ecology, the description of species composition and biodiversity calls for statistical methods that involve estimating features of interest in unobserved samples based on an observed one. In the last decade, the Bayesian nonparametrics…

统计方法学 · 统计学 2026-04-28 Alessandro Colombi , Raffaele Argiento , Federico Camerlenghi , Lucia Paci

Linear mixed models are widely used for analyzing hierarchically structured data involving missingness and unbalanced study designs. We consider a Bayesian clustering method that combines linear mixed models and predictive projections. For…

统计方法学 · 统计学 2021-07-07 Yinan Mao , David J. Nott

Hierarchical models are increasingly used in many applications. Along with this increased use comes a desire to investigate whether the model is compatible with the observed data. Bayesian methods are well suited to eliminate the many…

统计方法学 · 统计学 2008-02-08 M. J. Bayarri , M. E. Castellanos

The task of topical segmentation is well studied, but previous work has mostly addressed it in the context of structured, well-defined segments, such as segmentation into paragraphs, chapters, or segmenting text that originated from…

计算与语言 · 计算机科学 2022-12-06 Eitan Wagner , Renana Keydar , Amit Pinchevski , Omri Abend

Models with intractable likelihood functions arise in areas including network analysis and spatial statistics, especially those involving Gibbs random fields. Posterior parameter es timation in these settings is termed a doubly-intractable…

统计计算 · 统计学 2018-10-16 Lampros Bouranis , Nial Friel , Florian Maire

In today's data driven world, storing, processing, and gleaning insights from large-scale data are major challenges. Data compression is often required in order to store large amounts of high-dimensional data, and thus, efficient inference…

机器学习 · 统计学 2018-09-11 Denali Molitor , Deanna Needell

Data dispersed across multiple files are commonly integrated through probabilistic linkage methods, where even minimal error rates in record matching can significantly contaminate subsequent statistical analyses. In regression problems, we…

统计理论 · 数学 2024-09-18 Abhisek Chakraborty , Saptati Datta

Bipartite data is common in data engineering and brings unique challenges, particularly when it comes to clustering tasks that impose on strong structural assumptions. This work presents an unsupervised method for assessing similarity in…

机器学习 · 计算机科学 2017-02-17 Aaron Gerow , Mingyang Zhou , Stan Matwin , Feng Shi

A new Bayesian modelling framework is introduced for piece-wise homogeneous variable-memory Markov chains, along with a collection of effective algorithmic tools for change-point detection and segmentation of discrete time series. Building…

统计方法学 · 统计学 2025-01-14 Valentinian Lungu , Ioannis Papageorgiou , Ioannis Kontoyiannis

Complex data features, such as unmodelled censored event times and variables with time-dependent effects, are common in cancer recurrence studies and pose challenges for Bayesian survival modelling. Current methodologies for predictive…

统计方法学 · 统计学 2026-01-12 Saku Suorsa , Aki Vehtari

Binary classification problems can be naturally modeled as bipartite graphs, where we attempt to classify right nodes based on their left adjacencies. We consider the case of labeled bipartite graphs in which some labels and edges are not…

组合数学 · 数学 2018-11-13 R. W. R. Darling , Mark L. Velednitsky

We propose a distributed computing framework, based on a divide and conquer strategy and hierarchical modeling, to accelerate posterior inference for high-dimensional Bayesian factor models. Our approach distributes the task of…

统计方法学 · 统计学 2016-12-30 Gautam Sabnis , Debdeep Pati , Barbara Engelhardt , Natesh Pillai

We present a federated learning approach for Bayesian model-based clustering of large-scale binary and categorical datasets. We introduce a principled 'divide and conquer' inference procedure using variational inference with local merge and…

机器学习 · 统计学 2025-11-13 Jackie Rao , Francesca L. Crowe , Tom Marshall , Sylvia Richardson , Paul D. W. Kirk