English
Related papers

Related papers: Categorical Data Fusion Using Auxiliary Informatio…

200 papers

Analytical information needs, such as trend analysis and causal impact assessment, are prevalent across various domains including law, finance, science, and much more. However, existing information retrieval paradigms, whether based on…

Information Retrieval · Computer Science 2026-02-13 Yiteng Tu , Shuo Miao , Weihang Su , Yiqun Liu , Qingyao Ai

Suppose we have individual data from an internal study and various summary statistics from relevant external studies. External summary statistics have the potential to improve statistical inference for the internal population; however, it…

Methodology · Statistics 2026-02-06 Wenjie Hu , Ruoyu Wang , Wei Li , Wang Miao

The statistical matching problem is a data integration problem with structured missing data. The general form involves the analysis of multiple datasets that only have a strict subset of variables jointly observed across all datasets. The…

Methodology · Statistics 2019-04-01 Daniel Ahfock , Saumyadipta Pyne , Geoffrey J. McLachlan

Combining multiple predictors obtained from distributed data sources to an accurate meta-learner is promising to achieve enhanced performance in lots of prediction problems. As the accuracy of each predictor is usually unknown, integrating…

Machine Learning · Statistics 2024-08-16 Shiva Afshar , Yinghan Chen , Shizhong Han , Ying Lin

Missing data theory deals with the statistical methods in the occurrence of missing data. Missing data occurs when some values are not stored or observed for variables of interest. However, most of the statistical theory assumes that data…

In this work we introduce a mixture of GPs to address the data association problem, i.e. to label a group of observations according to the sources that generated them. Unlike several previously proposed GP mixtures, the novel mixture has…

Machine Learning · Statistics 2011-08-18 Miguel Lázaro-Gredilla , Steven Van Vaerenbergh , Neil Lawrence

When considering answering important questions with data, unsupervised data offers extensive insight opportunity and unique challenges. This study considers student survey data with a specific goal of clustering students into like groups…

Computers and Society · Computer Science 2018-12-14 Kathleen Campbell Garwood , Ph. D. , Arpit Arun Dhobale

We propose a Bayesian approach for model-based clustering of multivariate categorical data where variables are allowed to be associated within clusters and the number of clusters is unknown. The approach uses a two-layer mixture of finite…

Methodology · Statistics 2024-07-09 Gertraud Malsiner-Walli , Bettina Grün , Sylvia Frühwirth-Schnatter

The use of mutual information as a similarity measure in agglomerative hierarchical clustering (AHC) raises an important issue: some correction needs to be applied for the dimensionality of variables. In this work, we formulate the decision…

Machine Learning · Statistics 2016-08-07 Guillaume Marrelec , Arnaud Messé , Pierre Bellec

Most information dynamics and statistical causal analysis frameworks rely on the common intuition that causal interactions are intrinsically pairwise -- every 'cause' variable has an associated 'effect' variable, so that a 'causal arrow'…

Neurons and Cognition · Quantitative Biology 2019-09-06 Pedro A. M. Mediano , Fernando Rosas , Robin L. Carhart-Harris , Anil K. Seth , Adam B. Barrett

Simulation studies are commonly used to evaluate the performance of newly developed meta-analysis methods. For methodology that is developed for an aggregated data meta-analysis, researchers often resort to simulation of the aggregated data…

Applications · Statistics 2022-01-19 Edwin R. van den Heuvel , Osama Almalik , Zhuozhao Zhan

It is quite popular nowadays for researchers and data analysts holding different datasets to seek assistance from each other to enhance their modeling performance. We consider a scenario where different learners hold datasets with…

Machine Learning · Statistics 2024-05-15 Jiawei Zhang , Yuhong Yang , Jie Ding

In this paper we address a classification problem where two sources of labels with different levels of fidelity are available. Our approach is to combine data from both sources by applying a co-kriging schema on latent functions, which…

Machine Learning · Computer Science 2019-10-22 Nikita Klyuchnikov , Evgeny Burnaev

Inferring user characteristics such as demographic attributes is of the utmost importance in many user-centric applications. Demographic data is an enabler of personalization, identity security, and other applications. Despite that, this…

Social and Information Networks · Computer Science 2017-12-21 Yehezkel S. Resheff , Moni Shahar

We derive a method to enhance the evaluation for a text-based Emotion Aware Recommender that we have developed. However, we did not implement a suitable way to assess the top-N recommendations subjectively. In this study, we introduce an…

Information Retrieval · Computer Science 2021-02-12 John Kalung Leung , Igor Griva , William G. Kennedy

L1-norm regularized logistic regression models are widely used for analyzing data with binary response. In those analyses, fusing regression coefficients is useful for detecting groups of variables. This paper proposes a binomial logistic…

Methodology · Statistics 2023-12-15 Yuko Kakikawa , Shuichi Kawano

In the collaborative clustering framework, the hope is that by combining several clustering solutions, each one with its own bias and imperfections, one will get a better overall solution. The goal is that each local computation, quite…

Machine Learning · Computer Science 2021-03-25 Yohan Foucade , Younès Bennani

Multimodal data fusion is a key approach for enhancing diagnosis in medical applications. We propose an asymmetric fusion strategy starting from a primary modality and integrating secondary modalities by disentangling shared and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Jérémie Stym-Popper , Nathan Painchaud , Clément Rambour , Pierre-Yves Courand , Nicolas Thome , Olivier Bernard

The majority of existing methods for fake news detection universally focus on learning and fusing various features for detection. However, the learning of various features is independent, which leads to a lack of cross-interaction fusion…

Computation and Language · Computer Science 2020-04-22 Lianwei Wu , Yuan Rao

A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine…

Quantitative Methods · Quantitative Biology 2015-06-19 Ali Faisal , Jaakko Peltonen , Elisabeth Georgii , Johan Rung , Samuel Kaski
‹ Prev 1 4 5 6 7 8 10 Next ›