English
Related papers

Related papers: Open Data Analytical Model for Human Development I…

200 papers

With the large amount of data generated every day, public sentiment is a key factor for various fields, including marketing, politics, and social research. Understanding the public sentiment about different topics can provide valuable…

Computation and Language · Computer Science 2024-10-18 Mayimunah Nagayi , Clement Nyirenda

The popular K-means clustering algorithm potentially suffers from a major weakness for further analysis or interpretation. Some cluster may have disproportionately more (or fewer) points from one of the subpopulations in terms of some…

Machine Learning · Computer Science 2026-02-10 Guancheng Zhou , Haiping Xu , Hongkang Xu , Chenyu Li , Donghui Yan

High-resolution estimates of population health indicators are critical for precision public health. We propose a method for high-resolution estimation that fuses distinct data sources: an unbiased, low-resolution data source (e.g.…

Methodology · Statistics 2025-08-21 Amy Guan , Marissa Reitsma , Roshni Sahoo , Joshua Salomon , Stefan Wager

Clustering is a commonly used method for exploring and analysing data where the primary objective is to categorise observations into similar clusters. In recent decades, several algorithms and methods have been developed for analysing…

Machine Learning · Computer Science 2021-02-17 Bryar A. Hassan , Tarik A. Rashid

Health-related data analysis plays an important role in self-knowledge, disease prevention, diagnosis, and quality of life assessment. With the advent of data-driven solutions, a myriad of apps and Internet of Things (IoT) devices…

Computers and Society · Computer Science 2018-09-07 Vero Estrada-Galinanes , Katarzyna Wac

In the following paper, we describe results from mining citations, mentions, and links to open government data (OGD) in peer-reviewed literature. We inductively develop a method for categorizing how OGD are used by different research…

Computers and Society · Computer Science 2018-03-28 An Yan , Nicholas Weber

Panel data analysis is an important topic in statistics and econometrics. Traditionally, in panel data analysis, all individuals are assumed to share the same unknown parameters, e.g. the same coefficients of covariates when the linear…

Statistics Theory · Mathematics 2017-06-09 Heng Lian , Xinghao Qiao , Wenyang Zhang

Efficient extraction of useful knowledge from these data is still a challenge, mainly when the data is distributed, heterogeneous and of different quality depending on its corresponding local infrastructure. To reduce the overhead cost,…

Databases · Computer Science 2017-04-17 Nhien-An Le-Khac , M-Tahar Kechadi

This paper presents thirteen datasets for binary, multiclass and multilabel classification based on the European Court of Human Rights judgments since its creation. The interest of such datasets is explained through the prism of the…

Machine Learning · Computer Science 2019-02-05 Alexandre Quemy

Privacy policy documents are often lengthy, complex, and difficult for non-expert users to interpret, leading to a lack of transparency regarding the collection, processing, and sharing of personal data. As concerns over online privacy…

Cryptography and Security · Computer Science 2025-07-08 Vijayalakshmi Ramasamy , Seth Barrett , Gokila Dorai , Jessica Zumbach

National Health and Nutritional Status Survey (NHANSS) is conducted annually by the Ministry of Health in Negara Brunei Darussalam to assess the population health and nutritional patterns and characteristics. The main aim of this study was…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Usman Khalil , Owais Ahmed Malik , Daphne Teck Ching Lai , Ong Sok King

Understanding how people move in the urban area is important for solving urbanization issues, such as traffic management, urban planning, epidemic control, and communication network improvement. Leveraging recent availability of large…

Social and Information Networks · Computer Science 2019-05-27 Yuren Zhou , Billy Pik Lik Lau , Chau Yuen , Bige Tunçer , Erik Wilhelm

Many data mining tasks cannot be completely addressed by auto- mated processes, such as sentiment analysis and image classification. Crowdsourcing is an effective way to harness the human cognitive ability to process these machine-hard…

Databases · Computer Science 2018-10-22 Chengliang Chai , Ju Fan , Guoliang Li , Jiannan Wang , Yudian Zheng

In recent years, the use of databases that analyze trends, sentiments or news to make economic projections or create indicators has gained significant popularity, particularly with the Google Trends platform. This article explores the…

Econometrics · Economics 2025-03-31 Juan Tenorio , Heidi Alpiste , Jakelin Remón , Arian Segil

Climate change is a critical issue that will be in the political agenda for the next decades. While it is important for this topic to be discussed at higher levels, it is also of paramount importance that the populations became aware of the…

Applications · Statistics 2026-05-20 Gianpaolo Zammarchi , Paolo Maranzano

We propose a clustering procedure to group K populations into subgroups with the same dependence structure. The method is adapted to paired population and can be used with panel data. It relies on the differences between orthogonal…

Methodology · Statistics 2022-11-14 Yves Ismaël Ngounou Bakam , Denys Pommeret

The open data movement constitutes an approach to achieving accountability for government organizations, and is aligned with one of the sustainable development goals outlined by the United Nations. In the area of health care, government…

Computers and Society · Computer Science 2017-10-31 A. Ravishankar Rao , Daniel Clarke

Publicly available data from open sources (e.g., United States Census Bureau (Census), World Health Organization (WHO), Intergovernmental Panel on Climate Change (IPCC)) are vital resources for policy makers, students and researchers across…

Recently introduced privacy legislation has aimed to restrict and control the amount of personal data published by companies and shared to third parties. Much of this real data is not only sensitive requiring anonymization, but also…

Databases · Computer Science 2020-07-20 Mostafa Milani , Yu Huang , Fei Chiang

High-throughput microarray and sequencing technology have been used to identify disease subtypes that could not be observed otherwise by using clinical variables alone. The classical unsupervised clustering strategy concerns primarily the…

Methodology · Statistics 2020-07-23 Peng Liu , Yusi Fang , Zhao Ren , Lu Tang , George C. Tseng