English
Related papers

Related papers: Indebted households profiling: a knowledge discove…

200 papers

One of the main challenges in data mining is choosing the optimal number of clusters without prior information. Notably, existing methods are usually in the philosophy of cluster validation and hence have underlying assumptions on data…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Ruilin Zhang , Haiyang Zheng , Hongpeng Wang

Statistical significance of network clustering has been an unresolved problem since it was observed that community detection algorithms produce false positives even in random graphs. After a phase transition between undetectable and…

Social and Information Networks · Computer Science 2016-05-03 Jeremi K. Ochab

We consider the problem of the statistical uncertainty of the correlation matrix in the optimization of a financial portfolio. We show that the use of clustering algorithms can improve the reliability of the portfolio in terms of the ratio…

Physics and Society · Physics 2008-12-02 Vincenzo Tola , Fabrizio Lillo , Mauro Gallegati , Rosario N. Mantegna

Macroeconomic factors have a critical impact on banking credit risk, which cannot be directly controlled by banks, and therefore, there is a need for an early credit risk warning system based on the macroeconomy. By comparing different…

Information Retrieval · Computer Science 2024-01-29 Hemlata Sharma , Aparna Andhalkar , Oluwaseun Ajao , Bayode Ogunleye

We aim to cluster financial assets in order to identify a small set of stocks to approximate the level of diversification of the whole universe of stocks. We develop a data-driven approach to clustering based on a correlation blockmodel in…

Portfolio Management · Quantitative Finance 2021-08-16 Wenpin Tang , Xiao Xu , Xun Yu Zhou

This paper proposes an uncertain data clustering approach to quantitatively analyze the complexity of prefabricated construction components through the integration of quality performance-based measures with associated engineering design…

Databases · Computer Science 2019-03-19 Wenying Ji , Simaan M. AbouRizk , Osmar R. Zaiane , Yitong Li

The paper presents findings from a comprehensive study examining the saving and credit behaviors of low-income households residing in unauthorized colonies within a metropolitan area. Utilizing a dual approach, the study engaged in…

General Economics · Economics 2024-12-05 Divya Sharma

During cluster analysis domain experts and visual analysis are frequently relied on to identify the optimal clustering structure. This process tends to be adhoc, subjective and difficult to reproduce. This work shows how competency…

Machine Learning · Computer Science 2020-06-02 Wiebke Toussaint , Deshendran Moodley

While clustering is ubiquitously used across science and industry, uncertainty in cluster assignments is rarely quantified with rigorous guarantees. We propose a novel conformal inference framework for clustering that returns confidence…

Methodology · Statistics 2026-04-13 YoonHaeng Hur , Anirban Nath , Genevera Allen

Causal relationships are commonly examined in manufacturing processes to support faults investigations, perform interventions, and make strategic decisions. Industry 4.0 has made available an increasing amount of data that enable…

Machine Learning · Computer Science 2022-08-03 Giovanni Menegozzo , Diego Dall'Alba , Paolo Fiorini

With the rapid development of online social media, online shopping sites and cyber-physical systems, heterogeneous information networks have become increasingly popular and content-rich over time. In many cases, such networks contain…

Databases · Computer Science 2012-02-01 Yizhou Sun , Charu C. Aggarwal , Jiawei Han

Nowadays small and medium-sized enterprises have become an essential part of the national economy. With the increasing number of such enterprises, how to evaluate their credit risk becomes a hot issue. Unlike big enterprises with massive…

Risk Management · Quantitative Finance 2022-05-03 Marui Du , Yue Ma , Zuoquan Zhang

Causal discovery procedures aim to deduce causal relationships among variables in a multivariate dataset. While various methods have been proposed for estimating a single causal model or a single equivalence class of models, less attention…

Methodology · Statistics 2024-10-08 Y. Samuel Wang , Mladen Kolar , Mathias Drton

Consider the following problem: given a database of records indexed by names (e.g., name of companies, restaurants, businesses, or universities) and a new name, determine whether the new name is in the database, and if so, which record it…

Databases · Computer Science 2018-06-29 Bahare Fatemi , Seyed Mehran Kazemi , David Poole

Vulnerable individuals have a limited ability to make reasonable financial decisions and choices and, thus, the level of care that is appropriate to be provided to them by financial institutions may be different from that required for other…

Computers and Society · Computer Science 2021-06-14 Tasos Spiliotopoulos , Dave Horsfall , Magdalene Ng , Kovila Coopamootoo , Aad van Moorsel , Karen Elliott

The rapid development of artificial intelligence methods contributes to their wide applications for forecasting various financial risks in recent years. This study introduces a novel explainable case-based reasoning (CBR) approach without a…

Computational Finance · Quantitative Finance 2021-07-20 Wei Li , Florentina Paraschiv , Georgios Sermpinis

Spectral clustering methods which are frequently used in clustering and community detection applications are sensitive to the specific graph constructions particularly when imbalanced clusters are present. We show that ratio cut (RCut) or…

Machine Learning · Statistics 2016-11-18 Cem Aksoylar , Jing Qian , Venkatesh Saligrama

Credit risk modeling has permeated our everyday life. Most banks and financial companies use this technique to model their clients' trustworthiness. While machine learning is increasingly used in this field, the resulting large-scale…

Cryptography and Security · Computer Science 2020-10-07 Yuli Zheng , Zhenyu Wu , Ye Yuan , Tianlong Chen , Zhangyang Wang

We develop an approach to incorporate additional knowledge, in the form of general purpose integrity constraints (ICs), to reduce uncertainty in probabilistic databases. While incorporating ICs improves data quality (and hence quality of…

Databases · Computer Science 2009-07-10 Naveen Ashish , Sharad Mehrotra , Pouria Pirzadeh

Novel Class Discovery (NCD) is the problem of trying to discover novel classes in an unlabeled set, given a labeled set of different but related classes. The majority of NCD methods proposed so far only deal with image data, despite tabular…