English

Using Machine Learning in Analyzing Air Quality Discrepancies of Environmental Impact

Computers and Society 2025-06-24 v1 Machine Learning

Abstract

In this study, we apply machine learning and software engineering in analyzing air pollution levels in City of Baltimore. The data model was fed with three primary data sources: 1) a biased method of estimating insurance risk used by homeowners loan corporation, 2) demographics of Baltimore residents, and 3) census data estimate of NO2 and PM2.5 concentrations. The dataset covers 650,643 Baltimore residents in 44.7 million residents in 202 major cities in US. The results show that air pollution levels have a clear association with the biased insurance estimating method. Great disparities present in NO2 level between more desirable and low income blocks. Similar disparities exist in air pollution level between residents' ethnicity. As Baltimore population consists of a greater proportion of people of color, the finding reveals how decades old policies has continued to discriminate and affect quality of life of Baltimore citizens today.

Keywords

Cite

@article{arxiv.2506.17319,
  title  = {Using Machine Learning in Analyzing Air Quality Discrepancies of Environmental Impact},
  author = {Shuangbao Paul Wang and Lucas Yang and Rahouane Chouchane and Jin Guo and Michael Bailey},
  journal= {arXiv preprint arXiv:2506.17319},
  year   = {2025}
}

Comments

IEEE 2024 International Conference on AI x Data & Knowledge Engineering (AIxDKE)