English
Related papers

Related papers: An Efficient Rigorous Approach for Identifying Sta…

200 papers

Finding high-importance patterns in data is an emerging data mining task known as High-utility itemset mining (HUIM). Given a minimum utility threshold, a HUIM algorithm extracts all the high-utility itemsets (HUIs) whose utility values are…

Databases · Computer Science 2023-03-28 Shan Huang , Wensheng Gan , Jinbao Miao , Xuming Han , Philippe Fournier-Viger

The process of biomarker discovery is typically lengthy and costly, involving the phases of discovery, qualification, verification, and validation before clinical evaluation. Being able to efficiently identify the truly relevant markers in…

Applications · Statistics 2018-04-13 Lin-Yang Cheng , Bowei Xi

Synthetic datasets are important for evaluating and testing machine learning models. When evaluating real-life recommender systems, high-dimensional categorical (and sparse) datasets are often considered. Unfortunately, there are not many…

Information Retrieval · Computer Science 2024-12-11 Miha Malenšek , Blaž Škrlj , Blaž Mramor , Jure Demšar

The process of data mining produces various patterns from a given data source. The most recognized data mining tasks are the process of discovering frequent itemsets, frequent sequential patterns, frequent sequential rules and frequent…

Databases · Computer Science 2014-02-13 Thabet Slimani , Amor Lazzez

Knowledge of the association information between the attributes in a data set provides insight into the underlying structure of the data and explains the relationships (independence, synergy, redundancy) between the attributes and class (if…

Databases · Computer Science 2012-08-21 Pritam Chanda , Aidong Zhang , Murali Ramanathan

Image classifiers often use spurious patterns, such as "relying on the presence of a person to detect a tennis racket, which do not generalize. In this work, we present an end-to-end pipeline for identifying and mitigating spurious patterns…

Machine Learning · Computer Science 2022-08-19 Gregory Plumb , Marco Tulio Ribeiro , Ameet Talwalkar

This paper introduces a statistical method to decide whether two blocks in a pair of of images match reliably. The method ensures that the selected block matches are unlikely to have occurred "just by chance." The new approach is based on…

Computer Vision and Pattern Recognition · Computer Science 2017-12-08 Neus Sabater , Andrés Almansa , Jean-Michel Morel

Analyzing and finding anomalies in multi-dimensional datasets is a cumbersome but vital task across different domains. In the context of financial fraud detection, analysts must quickly identify suspicious activity among transactional data.…

Machine Learning · Computer Science 2024-10-29 Beatriz Feliciano , Rita Costa , Jean Alves , Javier Liébana , Diogo Duarte , Pedro Bizarro

Three variants of the statistical complexity function, which is used as a criterion in the problem of detection of a useful signal in the signal-noise mixture, are considered. The probability distributions maximizing the considered variants…

Statistics Theory · Mathematics 2023-11-30 Leonid Berlin , Andrey Galyaev , Pavel Lysenko

Level set estimation (LSE), the problem of identifying the set of input points where a function takes value above (or below) a given threshold, is important in practical applications. When the function is expensive-to-evaluate and…

Machine Learning · Statistics 2024-12-02 Yu Inatsu , Shion Takeno , Kentaro Kutsukake , Ichiro Takeuchi

We consider the "multi-frequency" periodogram, in which the putative signal is modelled as a sum of two or more sinusoidal harmonics with idependent frequencies. It is useful in the cases when the data may contain several periodic…

Instrumentation and Methods for Astrophysics · Physics 2013-10-30 Roman V. Baluev

In standardized educational testing, test items are reused in multiple test administrations. To ensure the validity of test scores, the psychometric properties of items should remain unchanged over time. In this paper, we consider the…

Applications · Statistics 2021-10-26 Yunxiao Chen , Yi-Hsuan Lee , Xiaoou Li

Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and…

Databases · Computer Science 2020-11-20 Ruoyu Wang , Xiaobo Hu , Daniel Sun , Guoqiang Li , Raymond Wong , Shiping Chen , Jianquan Liu

In this paper we describe a method to identify "relevant subsets" of variables, useful to understand the organization of a dynamical system. The variables belonging to a relevant subset should have a strong integration with the other…

Molecular Networks · Quantitative Biology 2015-02-09 Marco Villani , Andrea Roli , Alessandro Filisetti , Marco Fiorucci , Irene Poli , Roberto Serra

Given a labeled graph, the frequent-subgraph mining (FSM) problem asks to find all the $k$-vertex subgraphs that appear with frequency greater than a given threshold. FSM has numerous applications ranging from biology to network science, as…

Data Structures and Algorithms · Computer Science 2018-09-11 Cigdem Aslay , Muhammad Anis Uddin Nasir , Gianmarco De Francisci Morales , Aristides Gionis

Despite the widespread use of Transformer-based text embedding models in NLP tasks, surprising 'sticky tokens' can undermine the reliability of embeddings. These tokens, when repeatedly inserted into sentences, pull sentence similarity…

Computation and Language · Computer Science 2025-07-25 Kexin Chen , Dongxia Wang , Yi Liu , Haonan Zhang , Wenhai Wang

We introduce and study two new inferential challenges associated with the sequential detection of change in a high-dimensional mean vector. First, we seek a confidence interval for the changepoint, and second, we estimate the set of indices…

Methodology · Statistics 2023-03-03 Yudong Chen , Tengyao Wang , Richard J. Samworth

For applied intelligence, utility-driven pattern discovery algorithms can identify insightful and useful patterns in databases. However, in these techniques for pattern discovery, the number of patterns can be huge, and the user is often…

Databases · Computer Science 2022-06-14 Jinbao Miao , Wensheng Gan , Shicheng Wan , Yongdong Wu , Philippe Fournier-Viger

Detecting whether any anomalies exist within a dataset is crucial for effective anomaly detection, yet it remains surprisingly underexplored in anomaly detection literature. This paper presents a comprehensive study that addresses the…

Machine Learning · Computer Science 2025-08-14 Simon Klüttermann , Emmanuel Müller

An important problem in machine learning and statistics is to identify features that causally affect the outcome. This is often impossible to do from purely observational data, and a natural relaxation is to identify features that are…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , Amirata Ghorbani , James Zou
‹ Prev 1 8 9 10 Next ›