Related papers: Regionalization of China's PM2.5 through Robust Sp…
Air pollution remains a leading global health threat, with fine particulate matter (PM2.5) contributing to millions of premature deaths annually. Chemical transport models (CTMs) are essential tools for evaluating how emission controls…
Analyzing air pollution data is challenging as there are various analysis focuses from different aspects: feature (what), space (where), and time (when). As in most geospatial analysis problems, besides high-dimensional features, the…
We propose the DPSM method, a density-based node clustering approach that automatically determines the number of clusters and can be applied in both data space and graph space. Unlike traditional density-based clustering methods, which…
Air pollution, a pressing global problem, threatens public health, environmental sustainability, and climate stability. Achieving accurate and scalable forecasting across spatially distributed monitoring stations is challenging due to…
Clustering is indispensable for data analysis in many scientific disciplines. Detecting clusters from heavy noise remains challenging, particularly for high-dimensional sparse data. Based on graph-theoretic framework, the present paper…
The development of public transportation networks and associated transit oriented development policies are efficient tools to mitigate urban sprawl and its negative environmental impacts, especially in terms of commuting emissions. We study…
Understanding the dynamics of traffic clusters is crucial for enhancing urban transportation systems, particularly in managing congestion and free-flow states. This study applies computational percolation theory to analyze the formation and…
Regionalization aims to partition a spatial domain into contiguous regions that share similar characteristics, enabling more effective spatial analysis, policy making, and resource management. Existing approaches for spatial regionalization…
In this paper, we consider feature screening for ultrahigh dimensional clustering analyses. Based on the observation that the marginal distribution of any given feature is a mixture of its conditional distributions in different clusters, we…
Air pollution is known to be a major threat for human and ecosystem health. A proper understanding of the factors generating pollution and of the behavior of air pollution in time is crucial to support the development of effective policies…
Generalized Chinese Remainder Theorem (CRT) is a well-known approach to solve ambiguity resolution related problems. In this paper, we study the robust CRT reconstruction for multiple numbers from a view of statistics. To the best of our…
In recent years, the world has become increasingly concerned with air pollution. Particularly in the global north, countries are implementing systems to monitor air pollution on a large scale to aid decision-making. Such efforts are…
Accurate and reliable data stream plays an important role in air quality assessment. Air pollution data collected from monitoring networks, however, could be biased due to instrumental error or other interventions, which covers up the real…
Understanding the global organization of complicated and high dimensional data is of primary interest for many branches of applied sciences. It is typically achieved by applying dimensionality reduction techniques mapping the considered…
Classical approaches in cluster analysis are typically based on a feature space analysis. However, many applications lead to datasets with additional spatial information and a ground truth with spatially coherent classes, which will not…
With the development of urbanization, the scale of urban road network continues to expand, especially in some Asian countries. Short-term traffic state prediction is one of the bases of traffic management and control. Constrained by the…
Clustering large, mixed data is a central problem in data mining. Many approaches adopt the idea of k-means, and hence are sensitive to initialisation, detect only spherical clusters, and require a priori the unknown number of clusters. We…
We analyze the patterns of traffic jams in urban networks of five large cities and an urban agglomeration region in China using real data based on a recently developed jam tree model. This model focuses on the way traffic jams spread…
The description of complex configuration is a difficult issue. We present a powerful technique for cluster identification and characterization. The scheme is designed to treat with and analyze the experimental and/or simulation data from…
This paper builds the clustering model of measures of market microstructure features which are popular in predicting stock returns. In a 10-second time-frequency, we study the clustering structure of different measures to find out the best…