English
Related papers

Related papers: Categorical exploratory data analysis on goodness-…

200 papers

We propose NECA, a deep representation learning method for categorical data. Built upon the foundations of network embedding and deep unsupervised representation learning, NECA deeply embeds the intrinsic relationship among attribute values…

Machine Learning · Computer Science 2022-05-26 Xiaonan Gao , Sen Wu , Wenjun Zhou

Equation Discovery techniques have shown considerable success in regression tasks, where they are used to discover concise and interpretable models (\textit{Symbolic Regression}). In this paper, we propose a new ED-based binary…

Machine Learning · Computer Science 2025-10-29 Guus Toussaint , Arno Knobbe

Deciding whether a model provides a good description of data is often based on a goodness-of-fit criterion summarized by a p-value. Although there is considerable confusion concerning the meaning of p-values, leading to their misuse, they…

Data Analysis, Statistics and Probability · Physics 2013-05-29 Frederik Beaujean , Allen Caldwell , Daniel Kollar , Kevin Kroeninger

In a world increasingly awash with data, the need to extract meaningful insights from data has never been more crucial. Functional Data Analysis (FDA) goes beyond traditional data points, treating data as dynamic, continuous functions,…

Statistics Theory · Mathematics 2024-04-26 Sophie Dabo-Niang , Camille Frévent

Formal Concept Analysis (FCA) is a mathematical framework for knowledge representation and discovery. It performs a hierarchical clustering over a set of objects described by attributes, resulting in conceptual structures in which objects…

Artificial Intelligence · Computer Science 2025-08-12 Jessie Galasso

In this paper, we investigate the problem of mining numerical data in the framework of Formal Concept Analysis. The usual way is to use a scaling procedure --transforming numerical attributes into binary ones-- leading either to a loss of…

Artificial Intelligence · Computer Science 2011-11-28 Mehdi Kaytoue , Sergei O. Kuznetsov , Amedeo Napoli

`All models are wrong but some are useful' (George Box 1979). But, how to find those useful ones starting from an imperfect model? How to make informed data-driven decisions equipped with an imperfect model? These fundamental questions…

Econometrics · Economics 2023-03-09 Subhadeep , Mukhopadhyay

Significant progress has been made in developing identification and estimation techniques for missing data problems where modeling assumptions can be described via a directed acyclic graph. The validity of results using such techniques rely…

Methodology · Statistics 2023-06-13 Razieh Nabi , Rohit Bhattacharya

Functional data analysis (FDA) is a statistical framework that allows for the analysis of curves, images, or functions on higher dimensional domains. The goals of FDA, such as descriptive analyses, classification, and regression, are…

Methodology · Statistics 2023-12-12 Jan Gertheiss , David Rügamer , Bernard X. W. Liew , Sonja Greven

Despite being highly performant, deep neural networks might base their decisions on features that spuriously correlate with the provided labels, thus hurting generalization. To mitigate this, 'model guidance' has recently gained popularity,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Sukrut Rao , Moritz Böhle , Amin Parchami-Araghi , Bernt Schiele

Formal Concept Analysis (FCA) is a mathematical theory based on the formalization of the notions of concept and concept hierarchies. It has been successfully applied to several Computer Science fields such as data mining,software…

Artificial Intelligence · Computer Science 2009-05-29 Leonard Kwuida , Rokia Missaoui , Lahcen Boumedjout , Jean Vaillancourt

Subdata selection is a study of methods that select a small representative sample of the big data, the analysis of which is fast and statistically efficient. The existing subdata selection methods assume that the big data can be reasonably…

Methodology · Statistics 2024-05-01 Rakhi Singh

This paper explores Bayesian estimation for categorical data, focusing on simple yet effective models that provide a foundation for applying more advanced methods accurately and reliably in real-world applications. We begin by revisiting…

Methodology · Statistics 2025-09-03 Jan Kalina

Heterogeneous data pose serious challenges to data analysis tasks, including exploration and visualization. Current techniques often utilize dimensionality reductions, aggregation, or conversion to numerical values to analyze heterogeneous…

Graphics · Computer Science 2017-10-10 Mahsa Mirzargar , Ross T. Whitaker , Robert M. Kirby

Data imputation is an effective way to handle missing data, which is common in practical applications. In this study, we propose and test a novel data imputation process that achieve two important goals: (1) preserve the row-wise…

Machine Learning · Computer Science 2023-09-13 Katrina Chen , Xiuqin Liang , Zheng Ma , Zhibin Zhang

Nowadays data sets are available in very complex and heterogeneous ways. Mining of such data collections is essential to support many real-world applications ranging from healthcare to marketing. In this work, we focus on the analysis of…

Artificial Intelligence · Computer Science 2015-04-10 Aleksey Buzmakov , Elias Egho , Nicolas Jay , Sergei O. Kuznetsov , Amedeo Napoli , Chedy Raïssi

Dimensional analysis (DA) pays attention to fundamental physical dimensions such as length and mass when modelling scientific and engineering systems. It goes back at least a century to Buckingham's Pi theorem, which characterizes a…

Machine Learning · Computer Science 2023-12-19 G. Alexi Rodriguez-Arelis , William J. Welch

Causal decomposition analysis (CDA) is an approach for modeling the impact of hypothetical interventions to reduce disparities. It is useful for identifying foci that future interventions, including multilevel and multimodal interventions,…

Methodology · Statistics 2026-04-28 John W. Jackson , Ting-Hsuan Chang , Aster Meche , Trang Q. Nguyen

While modern day web applications aim to create impact at the civilization level, they have become vulnerable to adversarial activity, where the next cyber-attack can take any shape and can originate from anywhere. The increasing scale and…

Machine Learning · Statistics 2018-03-28 Tegjyot Singh Sethi , Mehmed Kantardzic

Symbolic data analysis (SDA) is an emerging area of statistics concerned with understanding and modelling data that takes distributional form (i.e. symbols), such as random lists, intervals and histograms. It was developed under the premise…

Computation · Statistics 2020-04-09 Boris Beranger , Huan Lin , Scott A. Sisson