English
Related papers

Related papers: Delta-Closure Structure for Studying Data Distribu…

200 papers

We investigate the behavior of extended urban traffic networks within the framework of percolation theory by using real and synthetic traffic data. Our main focus shifts from the statistical properties of the cluster size distribution…

Physics and Society · Physics 2021-07-21 Marco Cogoni , Giovanni Busonera

With rapidly increasing data, clustering algorithms are important tools for data analytics in modern research. They have been successfully applied to a wide range of domains; for instance, bioinformatics, speech recognition, and financial…

Data Structures and Algorithms · Computer Science 2015-12-01 Ka-Chun Wong

In this thesis, a detailed study shows that closed itemsets and minimal generators play a key role for concisely representing both frequent itemsets and association rules. These itemsets structure the search space into equivalence classes…

Databases · Computer Science 2019-11-05 Sadok Ben Yahia

Many methods in differentially private model training rely on computing the similarity between a query point (such as public or synthetic data) and private data. We abstract out this common subroutine and study the following fundamental…

Cryptography and Security · Computer Science 2024-03-15 Arturs Backurs , Zinan Lin , Sepideh Mahabadi , Sandeep Silwal , Jakub Tarnawski

Similarity-based clustering methods separate data into clusters according to the pairwise similarity between the data, and the pairwise similarity is crucial for their performance. In this paper, we propose {\em Clustering by Discriminative…

Machine Learning · Computer Science 2022-06-24 Yingzhen Yang , Ping Li

A discriminative structured analysis dictionary is proposed for the classification task. A structure of the union of subspaces (UoS) is integrated into the conventional analysis dictionary learning to enhance the capability of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Wen Tang , Ashkan Panahi , Hamid Krim , Liyi Dai

We develop a materials descriptor based on the electronic density of states and investigate the similarity of materials based on it. As an application example, we study the Computational 2D Materials Database that hosts thousands of…

Materials Science · Physics 2022-01-07 Martin Kuban , Santiago Rigamonti , Markus Scheidgen , Claudia Draxl

Spectral clustering requires the time-consuming decomposition of the Laplacian matrix of the similarity graph, thus limiting its applicability to large datasets. To improve the efficiency of spectral clustering, a top-down approach was…

Machine Learning · Computer Science 2024-12-19 Zhichang Xu , Zhiguo Long , Hua Meng

This paper develops new limit theory for data that are generated by networks or more generally display cross-sectional dependence structures that are governed by observable and unobservable characteristics. Strategic network formation…

Probability · Mathematics 2019-08-08 Guido M. Kuersteiner

We completely characterize $\Delta$- and local subexponentialities of positive-half compound Poisson distributions and extend the characterization on two-sided distributions. Moreover, $\Delta$-subexponentiality of infinitely divisible…

Probability · Mathematics 2023-02-21 Muneya Matsui , Toshiro Watanabe

Given the potential difficulties in obtaining large quantities of labelled data, many works have explored the use of deep semi-supervised learning, which uses both labelled and unlabelled data to train a neural network architecture. The…

Machine Learning · Computer Science 2021-09-02 Philip Sellars , Angelica Aviles-Rivero , Carola Bibiane Schönlieb

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability,…

Machine Learning · Statistics 2018-10-30 A. Adolfsson , M. Ackerman , N. C. Brownstein

This paper proposes a novel similarity measure for clustering sequential data. We first construct a common state-space by training a single probabilistic model with all the sequences in order to get a unified representation for the dataset.…

Machine Learning · Computer Science 2010-04-13 Darío García-García , Emilio Parrado-Hernández , Fernando Díaz-de-María

If the aphorism "All models are wrong"- George Box, continues to be true in data analysis, particularly when analyzing real-world data, then we should annotate this wisdom with visible and explainable data-driven patterns. Such annotations…

Machine Learning · Statistics 2020-12-07 Sabrina Enriquez , Fushing Hsieh

A main task in data analysis is to organize data points into coherent groups or clusters. The stochastic block model is a probabilistic model for the cluster structure. This model prescribes different probabilities for the presence of edges…

Machine Learning · Computer Science 2020-09-24 Alexander Jung

One basic requirement of many studies is the necessity of classifying data. Clustering is a proposed method for summarizing networks. Clustering methods can be divided into two categories named model-based approaches and algorithmic…

Machine Learning · Computer Science 2013-02-19 Raheleh Namayandeh , Farzad Didehvar , Zahra Shojaei

Clustering is a widely used unsupervised learning method for finding structure in the data. However, the resulting clusters are typically presented without any guarantees on their robustness; slightly changing the used data sample or…

Machine Learning · Statistics 2017-01-02 Andreas Henelius , Kai Puolamäki , Henrik Boström , Panagiotis Papapetrou

Lateral microsegregation in a monolayer of a binary mixture of particles or macromolecules is studied by MD simulations in a generic model with the interacting potentials inspired by effective interactions in biological or soft-matter…

Soft Condensed Matter · Physics 2025-05-26 M. Litniewski , W. T. Gozdz nd A. Ciach

Block encoding severs as an important data input model in quantum algorithms, enabling quantum computers to simulate non-unitary operators effectively. In this paper, we propose an efficient block-encoding protocol for sparse matrices based…

Quantum Physics · Physics 2025-07-30 Chunlin Yang , Zexian Li , Hongmei Yao , Zhaobing Fan , Guofeng Zhang , Jianshe Liu

We study the problem of deriving compressibility measures for Piecewise Linear Approximations (PLAs), i.e., error-bounded approximations of a set of two-dimensional increasing data points using a sequence of segments. Such approximations…

Data Structures and Algorithms · Computer Science 2025-09-12 Paolo Ferragina , Filippo Lari