0% found this document useful (0 votes)
4 views12 pages

Mega

This study investigates the impact of land use characteristics on air pollutant concentrations in South Korea, utilizing data from 443 air quality monitoring stations. It establishes a model to analyze the relationship between six pollutants and various factors, revealing that land cover changes significantly affect air quality, with spatial ranges of influence varying for each pollutant. The findings provide critical insights for urban planning and policymaking aimed at improving urban air quality.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views12 pages

Mega

This study investigates the impact of land use characteristics on air pollutant concentrations in South Korea, utilizing data from 443 air quality monitoring stations. It establishes a model to analyze the relationship between six pollutants and various factors, revealing that land cover changes significantly affect air quality, with spatial ranges of influence varying for each pollutant. The findings provide critical insights for urban planning and policymaking aimed at improving urban air quality.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Atmospheric Pollution Research 16 (2025) 102498

Contents lists available at ScienceDirect

Atmospheric Pollution Research


journal homepage: [Link]/locate/apr

Impact of land use characteristics on air pollutant concentrations


considering the spatial range of influence
Lee Gunwon a , Han Yuhan b , Geunhan Kim c,*
a
Department of Architecture, Korea University, Seoul, 02841, Republic of Korea
b
Department of Geoinformatics, University of Seoul, Seoul, 02504, Republic of Korea
c
Department of Environmental Planning, Korea Environment Institute, Sejong, 30147, Republic of Korea

A R T I C L E I N F O A B S T R A C T

Keywords: Prediction models ranging from statistical probability to machine learning techniques have been employed to
Urban air pollution improve and manage urban air quality. However, the number of air quality monitoring stations (AQMS) for the
Land cover collection of air quality information is limited. This study established a model that explains the relationship
Multiple regression analysis
between six air pollutants–SO2, CO, O3, NO2, PM10, and PM2.5–measured by approximately 443 AQMS in South
XGBoost
Korea and factors, such as the vegetation index, topography, and land cover elements. The model analyzed the
impact of land cover changes on air pollutant concentrations and derived scenarios predicting changes in the air
quality due to land use changes. Despite the relatively small sample size of approximately 360 AQMS, multiple
regression analysis demonstrated higher explanatory power compared with Xtreme Gradient Boosting, a
representative machine learning technique. The optimal spatial range for explaining air pollutant concentrations
varied for each air pollutant. The highest R2 in the multiple regression analysis was 0.34 at a distance of 12,000
m for SO2; 0.27 at 11,000 m for CO; 0.50 at 6000 m for O3; 0.70 at 18,000 m for NO2; 0.49 at 18,000 m for PM10;
and 0.48 at 11,000 m for PM2.5. Certain land cover characteristics were found to significantly affect air quality,
whereas small-scale restoration had a minimal impact on air quality improvement, and large-scale development
substantially increased pollutant concentrations. This study provides essential information for urban planning
and policymaking aimed at improving urban air quality.

1. Introduction Urban air pollution is driven by several factors, including the use of
fossil fuels (Johnston et al., 2011), the scale of the city and its associated
Continuous urbanization has led to the development of industries production activities (Capello and Camagni, 2000), and the conversion
within cities and a proportional increase in population, which increases of mountainous areas for urban development (Lee et al., 2016). These
automobile traffic. Urbanization has caused rapid changes in land use activities are primarily driven by socioeconomic activities aimed at
within cities, leading to environmental issues, such as air quality dete­ addressing basic production and consumption needs, as well as the
rioration (Balew and Korme, 2020). Urban air pollution significantly essential requirements for food, clothing, and shelter of urban residents
negatively affects the quality of life and health of urban residents (Faiz, (Irga et al., 2015; Yang et al., 2017). These activities vary depending on
1993; Akimoto, 2003). Large cities with high population densities, such the urban spatial structure and land-use patterns. Therefore, to improve
as Seoul, face crucial air pollution problems (Wang et al., 2004). The and manage urban air quality, the land use, land cover, and spatial
World Health Organization (WHO) estimates that exposure to air structure of the city must be considered (Chen et al., 2022).
pollution results in seven million premature deaths annually and the loss Several studies have attempted to identify urban elements that
of millions of healthy life years (World Health Organization, 2021). directly or indirectly cause air pollution in cities. Among the various

Abbreviations: (AQMS), air quality monitoring stations; (WHO), World Health Organization; (ESA), European Space Agency; (NDBI), Normalized Difference Built-
up Index; (NDVI), Normalized Difference Vegetation Index; (DEM), digital elevation model; (NSDI), National Spatial Data Infrastructure Portal; (EGIS), Environ­
mental Geographic Information Service; (MLR), Multiple Linear Regression; (SHAP), Shapley Additive Explanations; (OLS), Ordinary Least Squares; (XGBoost),
Xtreme Gradient Boosting; (XAI), explainable artificial intelligence; (RMSE), Root Mean Square Error; (MAE), Mean Absolute Error.
* Corresponding author.
E-mail addresses: rhyme2997@[Link] (L. Gunwon), dbgks25@[Link] (H. Yuhan), ghkim@[Link] (G. Kim).

[Link]
Received 20 October 2024; Received in revised form 5 March 2025; Accepted 5 March 2025
Available online 6 March 2025
1309-1042/© 2025 Turkish National Committee for Air Pollution Research and Control. Published by Elsevier B.V. This is an open access article under the CC
BY-NC-ND license ([Link]
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

urban elements, urban transportation systems and spatial structures is inherently limited. Therefore, studies should examine whether the
(Baldauf et al., 2013), population distribution and density (Van Der unconditional application of machine learning or deep learning tech­
Waals, 2000), and terrain shape (Hanna et al., 1982; Roberts et al., niques is appropriate in cases where the number of available samples is
1994) are associated with urban air pollution. Other studies have limited (Smith et al., 2013; Luan et al., 2020; Rajput et al., 2023) or
focused on the relationship between land use, which is one of the most whether regression analysis is more a suitable alternative
representative urban components, and air pollution. Patterns and (Bonilla-Bedoya et al., 2021).
changes in land use affect air pollutant dispersion and air quality (Huang This study aimed to differentiate itself from previous studies by
et al., 2013; Wei and Ye, 2014; Zahari et al., 2016; Huang and Du, 2018). focusing on the following aspects: South Korea has 443 AQMS nation­
Rapid changes in land use can lead to sudden increases in air pollution wide, which were used for the analysis in this study. Based on the limited
(Du et al., 2010; Tao et al., 2015; Hien et al., 2020). Furthermore, the AQMS data, we examined the suitability of the model and reviewed the
composition and concentration of air pollutants vary depending on land spatial extent required to adequately explain air quality. By examining
use and the spatial ranges over which air pollutants disperse (Nagar spatial ranges to the AQMSs from near (1000 m) to distant (20,000 m),
et al., 2017; Hsu et al., 2018; Harrison, 2020; Yu and Park, 2021). we aimed to identify the land use and meteorological characteristics that
Additionally, as wind (affected by urban patterns and building heights) affect air quality within specific ranges and examine their explanatory
influences these air pollutants, their concentrations vary depending on power differences. Additionally, we verified how land use and vegeta­
the layout of the city and the height of buildings (Blocken et al., 2016; tion indices within these ranges impacted specific air quality measures.
Tuckett-Jones and Reade, 2017). Finally, we aimed to determine whether the established model could
To address these issues, previous studies have employed prediction predict changes in air pollutant concentrations resulting from land-cover
models ranging from statistical probability techniques to machine changes. This study aimed to provide crucial foundational information
learning. Commonly used statistical probability techniques include for urban planning and policymaking aimed at improving urban air
correlation analysis and multiple linear regression models, whereas quality by analyzing the impact level of independent variables related to
machine learning algorithms include ANN (Wang et al., 2019; Park land cover and land use on air pollutant concentrations.
et al., 2020; Chen et al., 2021), KNN (Bozdağ et al., 2020; Tella and
Balogun, 2021), SVM (Yang et al., 2018; Su et al., 2020; Mogollón-Sotelo 2. Materials and methods
et al., 2021; Zhang et al., 2021), Random Forest (RF; Yuchi et al., 2019;
Shao et al., 2020; Ma et al., 2021a,b), Ensemble (Lim et al., 2019; Van 2.1. Study area
Roode et al., 2019; Adams et al., 2020; Huang et al., 2022), and Xtreme
Gradient Boosting (XGBoost; Hu et al., 2017; AlThuwaynee et al., 2021; This study focused on South Korea (Fig. 1), covering a total area of
Zhao et al., 2021). Machine learning algorithms are powerful tools for approximately 118,118.94 km2 (Kim and Kim, 2022). The temporal
modeling complex relationships and interactions within data. They have scope of this study was based on 2019 data. This study was limited to the
been proven effective in predicting various scenarios more accurately year 2019 to ensure that the analysis reflects typical air quality condi­
and uncovering patterns that traditional research methods may overlook tions, excluding the unusual disruptions caused by the COVID-19
(Ma et al., 2024). pandemic, which significantly reduced air pollution levels at the
However, a vast amount of data is required for performing machine beginning of 2020. Numerous studies have reported a substantial
learning or deep learning analysis and the installation of air quality reduction in air pollutant concentrations during the early stages of the
monitoring stations (AQMS) for the collection of air quality information pandemic due to decreased anthropogenic activity (Chossière et al.,

Fig. 1. Study area Korea and Areas Affected by Land-Use Change Scenarios: Dongjak-daero and Surrounding Greenbelt.

2
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

2021; Dutheil et al., 2020). Therefore, using data from 2020 onward used as dependent variables. After preprocessing, the annual average
could introduce biases in the analysis, as the changes in air quality values were calculated. AirKorea provided AQMS data for South Korea
during this period were driven by temporary behavioral and policy shifts on an hourly basis, which was collected as point data for analysis. A total
rather than long-term land-use changes. of 443 AQMSs (Air Quality Monitoring Stations) distributed nationwide
Dongjak-daero serves as a representative case of urban trans­ were initially collected for analysis. After preprocessing, 351 to 365
formation in South Korea, where rapid urbanization has led to increased AQMSs were used for each type of air quality measure. Some monitoring
traffic congestion, reduced green spaces, and significant air quality stations were excluded to ensure data reliability and consistency, as
challenges. To address these issues, an underground road and drainage stations with significant NoData values were removed during the pre­
tunnel have been planned to alleviate congestion and flooding while processing stage.
restoring surface-level green spaces, effectively converting the area into
a park (Chosun Biz, 2023). Additionally, the potential impact of the 2.2.2. Independent variables
removal of greenbelt development restrictions in Seoul on neighboring The independent variables included vegetation indices (Fig. 2(a) and
areas must be considered. This study analyzed the implications of these 2(b)), topography (Fig. 2(c) and (d)), and level-3 land cover areas (Fig. 2
urban development initiatives, particularly their effects on air quality (e)). The vegetation indices included NDBI and NDVI derived from
and the balance between urban expansion and environmental sustain­ Sentinel-2 satellite images.
ability, offering insights applicable to cities facing similar challenges. In this study, Sentinel-2A/MSI L1C imagery, captured by the Euro­
Fig. 1 illustrates the study area, highlighting the regions affected by pean Space Agency (ESA) on May 23, 2019, was used as raw data to
major land-use change scenarios, including Dongjak-daero (Scenario 1) calculate the vegetation indices. The cloud-free imagery was pre­
and the potential greenbelt removal (Scenario 2). The left panel provides processed using the Semi-Automatic Classification tool. The indices
a national-scale view of South Korea, where red dots indicate air quality were calculated using Sentinel-2 Band 4 (RED), Band 8 (NIR), and Band
monitoring stations distributed across the country. The right panel 11 (SWIR) as follows:
presents a zoomed-in map of Seoul, detailing the specific areas of in­ ( )
SWIR − NIR
terest for this study. The blue line represents the planned Dongjak-daero NDBI = and (1)
SWIR + NIR
underground road and drainage tunnel project (Scenario 1). The red-
shaded areas represent regions where greenbelt restrictions are ( )
NIR − Red
assumed to be lifted, allowing for potential urban expansion (Scenario NDVI = . (2)
NIR + Red
2). The green-shaded areas indicate existing greenbelt zones that remain
protected. And the yellow triangles mark the locations of air quality Where SWIR: Short-Wave Infrared, NIR: Near-Infrared, Red: Sentinel
monitoring stations that will be used in Scenario 2. Red Band.
Topographical characteristics, including elevation and slope, were
2.2. Data also used in this study. Elevation data was derived from digital elevation
model (DEM) datasets obtained from the National Spatial Data Infra­
Table 1 summarizes the air quality monitoring data used as the structure Portal [NSDI, available online: NSDI ([Link]) (accessed on
dependent variables for analyzing air quality concentrations. The data December 6, 2023)], and the slope was calculated using DEM data.
on vegetation indices [Normalized Difference Built-up Index (NDBI) and Additionally, a Level-3 land cover map (10 m resolution), provided
Normalized Difference Vegetation Index (NDVI)], topography (Eleva­ by the Environmental Geographic Information Service [available online:
tion, Slope), and Level-3 land cover were used as independent variables. EGIS ([Link]) (accessed on December 6, 2023)] was used. The Level-3
land cover map of South Korea classifies land features, such as resi­
2.2.1. Dependent variables dential, commercial, industrial areas, and green spaces, into 41 cate­
Air pollutant concentration data (CO, NO2, O3, SO2, PM10, and gories following standardized criteria. This map was created at a 1:5000
PM2.5) from January to December 2019, obtained from AirKorea, were scale and has been updated annually by the Ministry of Environment
since 2019 (Mun and Kil, 2024). The AQMSs used in this study are
Table 1 mainly located in urban areas, leading to an uneven distribution of land
Air quality monitoring data used in air quality concentration analysis. cover types within the buffer zones. Fig. 2(e) shows the Level-3 land
Data Spatial Source cover classification used in this study.
Resolution (Year)

Dependent Air Quality CO (ppm) Point AirKorea


Variable Concentration NO2 (ppm) (2019)
2.3. Methods
O3 (ppm)
SO2 (ppm) 2.3.1. Research procedure
PM10 (㎍/㎥) This study analyzed the relationship between air pollutant concen­
PM2.5 (㎍/㎥)
trations (CO, NO2, O3, SO2, PM10, and PM2.5) and variables, such as
Independent Vegetation NDBI 10 × 10 m Sentinel-2
Variable Index NDVI (2019) vegetation indices, topography, and land use. Zonal statistics were
Topography Elevation Digital performed within 20 buffer zones (1000–20,000-m at 1000-m intervals)
Slope Terrain around the AQMSs. Independent variables with a VIF greater than 10
Model were excluded due to multicollinearity (Fig. 3).
(2019)
Level-3 land Urbanized Level-3 land
Numerous studies have examined the relationship between air
cover area area cover map pollutant concentrations and surrounding land use/cover by analyzing
Agricultural (2019) spatial ranges extending up to 1–20 km to capture broader spatial im­
area pacts (Yang and Jiang, 2021; Ma et al., 2024). In this study, we used 1
Forest area
km increments for our analyses, creating a balance between fine-scale
Grassland
area sensitivity and computational feasibility. Although it is theoretically
Wetland area possible to employ finer increments (e.g., 100 m), such high-resolution
Bare land analyses would require substantially more computing resources and
area longer processing times. Consequently, a 1 km scale was deemed
Water area
appropriate for balancing fine-scale sensitivity with computational

3
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Fig. 2. Variables used in the study: (a) NDBI, (b) NDVI, (c) Elevation, (d) Slope, and (e) Level-3 land cover.

efficiency, a choice that aligns with several existing studies in the field where β0 : intercept, βi : coefficients for each independent variable xi , xi :
(Liu et al., 2021; Phillips et al., 2021). independent variable ϵ: error term.
Multiple Linear Regression (MLR) and XGBoost models were devel­ Separate models were built for SO2, CO, O3, NO2, PM10, and PM2.5
oped and the best model for each pollutant was selected. The coefficients using vegetation indices, topography, and land-cover features as inde­
from the MLR model were examined to identify significant correlations pendent variables.
between land cover characteristics and pollutant concentrations. Addi­
tionally, Shapley Additive Explanations (SHAP) values were used to 2.3.3. XGBoost
interpret the contribution and importance of each feature in the models. XGBoost is a scalable machine learning system for tree boosting that
The optimal buffer radius and key influencing factors were identified improves on Gradient Boosting by addressing challenges, such as over­
for each pollutant. Two scenarios were then simulated: converting fitting and learning speed (Chen and Guestrin, 2016). XGBoost was used
Dongjak-daero into a green space and urbanizing the surrounding to handle tabular data due to its efficiency and ability to prevent over­
greenbelt, with the predicted air quality changes listed in Tables 2 and 3. fitting through several methods. First, it uses a regularized loss function
that reduces model complexity and prevents overfitting, as follows
2.3.2. Multiple linear regression (Dong et al., 2022):
MLR is a statistical method used to analyze the relationship between
a dependent variable and multiple independent variables. MLR fits the
∑ ∑ 1
L (ϕ) = l(̂
y i , yi ) + Ω(fk ), Ω(f) = γT + λ‖w‖2 , (4)
data into a linear equation to determine the contribution of each inde­ i k
2
pendent variable to the dependent variable and, ultimately, identifies ( )
the best-fit line that minimizes the difference between the observed and where L (ϕ): total objective function, l ̂y i , yi : loss function, Ω(f): reg­
predicted values (ordinary least squares; OLS) (Gulati et al., 2023), ularization term to control model complexity, γ and λ: the regularization
defined as follows: parameters(L1,L2), T: number of leaves, w: leaf weights.
Second, shrinkage reduces the influence of earlier trees, functioning
y = β0 + β1 x1 + β2 x2 + … + βn xn + ϵ, (3)
similarly to a learning rate adjustment. Finally, column subsampling

4
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Fig. 3. Flowchart explaining the methods used in this study.

selects a subset of features for training, further reducing overfitting risk Scenario 2: Greenbelt Development (Urbanization of Adjacent Green
(Dong et al., 2022). Areas)
Another key reason for selecting XGBoost was its ability to handle The second scenario assumed that greenbelt areas surrounding
multicollinearity effectively, which is a common issue in air quality Dongjak-daero would be converted into urban developments, such as
datasets where several input variables can be highly related. The tree- residential, commercial, and industrial facilities. This assumption was
based structure of the model inherently reduces the impact of multi­ based on ongoing policy discussions in South Korea regarding the
collinearity by prioritizing the most informative features during the relaxation of greenbelt regulations to address housing shortages in
splitting process. Additionally, regularization methods (L1 and L2) urban areas. The types of developments assumed in this scenario were
further help mitigate the effects of multicollinearity by shrinking the chosen to reflect typical patterns observed in urban expansion projects.
coefficients of less relevant features to zero, thereby improving model It was further assumed that such developments would produce increased
stability and performance. emissions from vehicle traffic and industrial activities, contributing to
To address potential overfitting in the XGBoost model, hyper­ higher air pollutant concentrations. To ensure consistency, it was
parameter tuning was conducted using GridSearchCV with 5-fold cross- assumed that existing industrial emission controls would remain con­
validation (Fig. 4). The parameter grid included n_estimators, learnin­ stant during the analysis period.
g_rate, reg_alpha, reg_lambda, max_depth, and subsample. Additionally,
the dataset was split into 80% training and 20% testing data to ensure 3. Results
unbiased model validation. StandardScaler was applied for scaling and
variables with high multicollinearity (VIF >10) were removed before 3.1. MLR vs. XGBoost
training. The hyperparameter tuning results are shown in Fig. 4.
We calculated the mean values of the landscape and topography
2.3.4. Scenario assumptions characteristics, along with the total land cover area, within buffer zones
In this study, two land-use change scenarios were designed to eval­ ranging from 1000- to 20,000-m at 1000-m intervals for the following
uate the impact of land-use alterations on air quality. The assumptions air pollutants: SO2, CO, O3, NO2, PM10, and PM2.5. Independent vari­
behind each scenario were established based on real-world urban ables with a VIF >10 were excluded from model training. During this
planning discussions and policies to ensure practical relevance (Fig. 5). process, NDVI and certain land cover variables were removed due to
Scenario 1: Road Greening (Dongjak-daero Conversion to Green high multicollinearity, as they did not significantly improve model
Space) performance.
The first scenario assumed that the surface of Dongjak-daero would MLR consistently showed higher explanatory power for predicting
be converted into green space after constructing the underground road air pollutant concentrations compared with XGBoost. Conversely,
and drainage tunnel. This assumption was based on recent urban plan­ XGBoost demonstrated overfitting, as seen in certain cases. For instance,
ning initiatives aimed at reducing urban heat islands, improving air in the 1000 m model for CO, the R2 value for the training model was
quality, and restoring green spaces in metropolitan areas. 0.1652, whereas that of the test model decreased to 0.0107 despite L1
The types of vegetation selected for conversion (deciduous forests, and L2 regularization.
coniferous forests, and mixed forests) were chosen to reflect typical Fig. 6 illustrates the R2 values of MLR (blue line) and XGBoost (green
urban greening projects due to their air pollution mitigation potential line) across buffer zones from 1000 to 20,000 m for each dependent
(Nowak et al., 2006). The selection also accounted for vegetation variable. The red box highlights the buffer zone where MLR achieved its
maturity, as mature vegetation can more effectively capture air pollut­ highest R2 value for prediction, indicating optimal model performance
ants and provide cooling effects. It was assumed that traffic previously at each distance.
using Dongjak-daero would not significantly increase on neighboring Tables S.1–6 present the explanatory power (R2) of multiple linear
roads, allowing the analysis to focus solely on the direct environmental regression and XGBoost for each dependent variable across buffer dis­
benefits of converting the road surface into green space. tances. The optimal buffer distance for SO2 was determined to be

5
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Table 2
Predicted air quality changes due to land use changes (Greening) in Dongjak-daero.
Hangang- Dosan- Seocho- Yongsan- Gangnam- Gwanak- Dongjak- Gangnam- Gwacheon- Byeoryang- Dongjak-
daero daero gu gu gu gu daero daero dong dong gu

SO2 Original 0.0043 0.0038 0.0038 0.0033 0.0047 0.0043 0.0056 0.0039 0.0033 0.0036 0.0033
Pred No 0.0042 0.0040 0.0040 0.0041 0.0039 0.0040 0.0039 0.0037 0.0036 0.0034 0.0039
Change
Deciduous 0.0042 0.0040 0.0040 0.0041 0.0039 0.0040 0.0039 0.0037 0.0036 0.0034 0.0039
Forest
Coniferous 0.0042 0.0040 0.0040 0.0041 0.0039 0.0040 0.0039 0.0037 0.0036 0.0034 0.0039
Forest
Mixed – – – – – – – – – – –
Forest
CO Original 0.5793 0.7905 0.3749 0.4941 0.4678 0.4509 0.5995 0.6461 0.5942 0.6176 0.4718
Pred No 0.5589 0.5282 0.5357 0.5489 0.5411 0.5362 0.5282 0.5285 0.5331 0.5523 0.5280
Change
Deciduous 0.5592 0.5420 0.5360 0.5491 0.5414 0.5364 0.5285 0.5288 0.5334 0.5525 0.5282
Forest
Coniferous 0.5589 0.5417 0.5357 0.5489 0.5411 0.5362 0.5282 0.5286 0.5331 0.5523 0.5280
Forest
Mixed – – – – – – – – – – –
Forest
O3 Original 0.0170 0.0195 0.0270 0.0229 0.0220 0.0258 0.0170 0.0161 0.0244 0.0232 0.0241
Pred No 0.0234 0.0216 0.0190 0.0225 0.0197 0.0229 0.0216 0.0213 0.0252 0.0273 0.0222
Change
Deciduous 0.0234 0.0203 0.0190 0.0225 0.0197 0.0228 0.0216 0.0213 0.0251 0.0273 0.0221
Forest
Coniferous 0.0234 0.0203 0.0190 0.0225 0.0197 0.0228 0.0216 0.0213 0.0251 0.0273 0.0222
Forest
Mixed 0.0234 0.0203 0.0190 0.0225 0.0197 0.0228 0.0216 0.0213 0.0251 0.0273 0.0221
Forest
NO2 Original 0.0397 0.0303 0.0300 0.0318 0.0275 0.0305 0.0470 0.0469 0.0256 0.0294 0.0302
Pred No 0.0342 0.0342 0.0347 0.0341 0.0327 0.0335 0.0342 0.0329 0.0328 0.0320 0.0342
Change
Deciduous – – – – – – – – – – –
Forest
Coniferous 0.0342 0.0339 0.0347 0.0341 0.0327 0.0335 0.0342 0.0329 0.0328 0.0320 0.0342
Forest
Mixed – – – – – – – – – – –
Forest
PM10 Original 48.5521 45.9760 43.3098 33.8829 39.8950 48.6406 46.6591 46.2964 42.6036 46.7572 43.5449
Pred No 45.0509 44.4488 44.6999 44.7248 44.2351 44.6850 44.4488 43.8219 43.9249 43.5393 44.4312
Change
Deciduous – – – – – – – – – – –
Forest
Coniferous 45.0436 44.4967 44.6924 44.7174 44.2273 44.6791 44.4436 43.8170 43.9177 43.5320 44.4237
Forest
Mixed – – – – – – – – – – –
Forest
PM2.5 Original 27.5614 25.8506 25.5377 23.7683 24.7205 27.5339 24.8568 25.5783 22.2822 22.1938 26.4567
Pred No 26.1165 24.2000 24.7671 25.7238 25.2860 24.8759 24.2000 24.2616 23.9556 23.8665 24.0840
Change
Deciduous 26.1240 25.1916 24.7759 25.7313 25.2939 24.8850 24.2079 24.2733 23.9631 23.8741 24.0915
Forest
Coniferous 26.1161 25.1837 24.7680 25.7233 25.2860 24.8771 24.2000 24.2654 23.9552 23.8661 24.0836
Forest
Mixed – – – – – – – – – – –
Forest

12,000 m, yielding the highest explanatory power (R2 = 0.3402). The cover characteristics. The results revealed that certain land cover types
differences in R2 values among the 11,000-m (R2 = 0.34), 12,000-m (R2 significantly affect pollutant levels, consistent with the expected spatial
= 0.3402), and 13,000-m (R2 = 0.34) distances are minimal. For other distribution of emission sources and dispersion processes. The results for
pollutants, the optimal buffer distances varied: CO peaked at 11,000 m each pollutant, as summarized in Tables S.7–12, indicate distinct re­
(R2 = 0.2733), O3 showed the highest R2 at 6000 m (R2 = 0.5024), NO2 lationships between specific land cover types and air pollutant levels.
reached its maximum explanatory power at 18,000 m (R2 = 0.7037), For SO2, industrial facilities and tidal flats exhibited significant
PM10 had its highest R2 at 18,000 m (R2 = 0.4898), and PM2.5 peaked at positive correlations with SO2 concentrations. Conversely, orchards and
11,000 m (R2 = 0.4822). These optimal buffer distances were chosen to deciduous forests were significantly negatively correlated with SO2.
develop models that effectively capture air pollutant variations across However, variables, such as elevation and single housing did not reach
different spatial scales. statistical significance, indicating that the current data cannot reliably
support their effects on SO2 levels.
Similarly, the CO model showed positive relationships with single
3.2. Correlations between land cover characteristics and air pollutants housing, deciduous forests, and other artificial barren areas, whereas the
river variable was significantly negatively correlated. For O3, significant
In this study, multiple linear regression was employed to examine the negative predictors included apartment housing, deciduous forests, and
relationships between air pollutant concentrations and various land

6
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Table 3
Predicted air quality changes due to land use changes (urbanization) in the greenbelt.
Hangang- Dosan- Seocho- Yongsan- Gangnam- Gwanak- Dongjak- Gangnam- Gwacheon- Byeoryang- Dongjak-
daero daero gu gu gu gu daero daero dong dong gu

SO2 Original 0.0043 0.0038 0.0038 0.0033 0.0047 0.0043 0.0056 0.0039 0.0033 0.0036 0.0033
Pred No 0.0042 0.0040 0.0040 0.0041 0.0039 0.0040 0.0039 0.0037 0.0036 0.0034 0.0039
Change
Single 0.0043 0.0044 0.0045 0.0045 0.0043 0.0042 0.0044 0.0042 0.0041 0.0039 0.0044
housing
Apartment – – – – – – – – – – –
Housing
Industrial 0.0050 0.0061 0.0062 0.0058 0.0059 0.0051 0.0062 0.0060 0.0059 0.0057 0.0061
facilities
CO Original 0.5792 0.7905 0.3749 0.4941 0.4678 0.4509 0.5995 0.6461 0.5942 0.6176 0.4718
Pred No 0.5589 0.5282 0.5357 0.5489 0.5411 0.5362 0.5282 0.5285 0.5331 0.5523 0.5280
Change
Single 0.5810 0.6247 0.6233 0.5940 0.6118 0.5746 0.6148 0.6133 0.6215 0.6376 0.6099
housing
Apartment – – – – – – – – – – –
Housing
Industrial 0.5581 0.5371 0.5309 0.5465 0.5360 0.5351 0.5235 0.5239 0.5282 0.5476 0.5234
facilities
O3 Original 0.0170 0.0195 0.0270 0.0229 0.0220 0.0258 0.0170 0.0161 0.0244 0.0232 0.0241
Pred No 0.0234 0.0216 0.0190 0.0225 0.0197 0.0229 0.0216 0.0213 0.0252 0.0273 0.0222
Change
Single 0.0234 0.0207 0.0198 0.0225 0.0204 0.0238 0.0229 0.0243 0.0275 0.0290 0.0235
housing
Apartment 0.0234 0.0151 0.0083 0.0225 0.0113 0.0088 0.0047 − 0.0155 − 0.0052 0.0042 0.0039
Housing
Industrial 0.0234 0.0201 0.0185 0.0225 0.0194 0.0221 0.0208 0.0198 0.0238 0.0262 0.0213
facilities
NO2 Original 0.0397 0.0303 0.0300 0.0318 0.0275 0.0305 0.0470 0.0469 0.0256 0.0294 0.0302
Pred No 0.0342 0.0342 0.0347 0.0341 0.0327 0.0335 0.0342 0.0329 0.0328 0.0320 0.0342
Change
Single – – – – – – – – – – –
housing
Apartment 0.0493 0.0491 0.0498 0.0493 0.0479 0.0486 0.0494 0.0480 0.0479 0.0471 0.0493
Housing
Industrial 0.0406 0.0403 0.0410 0.0405 0.0391 0.0398 0.0406 0.0393 0.0392 0.0384 0.0405
facilities
PM10 Original 48.5521 45.9760 43.3098 33.8829 39.8950 48.6406 46.6591 46.2964 42.6036 46.7572 43.5449
Pred No 45.0509 44.4488 44.6999 44.7248 44.2351 44.6850 44.4488 43.8219 43.9249 43.5393 44.4312
Change
Single – – – – – – – – – – –
housing
Apartment 50.2786 49.7317 49.9274 49.9524 49.4623 49.9140 49.6786 49.0520 49.1526 48.7670 49.6587
Housing
Industrial 48.2692 47.7223 47.9180 47.9430 47.4529 47.9046 47.6692 47.0426 47.1432 46.7576 47.6493
facilities
PM2.5 Original 27.5614 25.8505 25.5377 23.7683 24.7205 27.5338 24.8568 25.5783 22.2822 22.1937 26.4567
Pred No 26.1165 24.2000 24.7671 25.7237 25.2859 24.8759 24.2000 24.2616 23.9556 23.8665 24.0840
Change
Single 27.4480 30.1654 30.0212 28.4307 29.5354 27.1849 29.3990 29.3539 29.2625 28.9907 29.0015
housing
Apartment – – – – – – – – – – –
Housing
Industrial 26.0577 24.8614 24.4303 25.5554 24.9484 24.7930 23.8679 23.9427 23.6112 23.5382 23.7649
facilities

rivers, with other variables failing to achieve significance. In the case of selected AQMSs. The stations were chosen based on their proximity to
NO2, strong positive associations were found for apartment housing and Dongjak-daero and their location within the minimum buffer distance to
industrial facilities, whereas tidal flats also showed a significant nega­ reflect the impact of land cover changes (Fig. 1).
tive effect. Moreover, the PM10 and PM2.5 models revealed that housing In the first scenario, the road surface of Dongjak-daero was trans­
and agricultural areas are generally associated with higher particulate formed into green space by converting it into land cover types, such as
matter concentrations, whereas forested areas contribute to their deciduous, coniferous, and mixed forests, ensuring no multicollinearity
reduction, as indicated by significant negative coefficients. issues. The goal was to assess the potential reduction in air pollutant
concentrations by replacing an urban area with green space. Predicted
pollutant concentrations, such as SO2, CO, NO2, O3, PM10, and PM2.5,
3.3. Air quality restoration cases were compared before and after the change (Table 2).
The analysis showed relatively modest improvements in air quality,
In this study, two distinct scenarios were analyzed to observe how suggesting that greening smaller urban roads, such as Dongjak-daero
land-use changes impact air pollutant concentrations. We first identified alone may have limited effectiveness in reducing pollutants. The
the effective range of influence for land cover and meteorological largest observed reduction was in CO, with a notable decrease of 0.0123
characteristics on air pollutants. Based on this, we predicted changes in ppm after conversion to coniferous forest, whereas changes in other
air pollutant concentrations when these factors were altered around

7
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

studies that explored the correlation between urban air quality and
urban environments (Johnson et al., 2010; Lai et al., 2021). This sug­
gests that rather than unconditionally applying machine learning or
deep learning models for prediction and relationship analysis, re­
searchers should assess experimental conditions, such as the sample size,
and determine the appropriate model through empirical testing. This
also implies that simple linear regression analysis is a viable candidate
model for such analyses.
Furthermore, the analysis results indicate that the explanatory power
of air quality concentration varies with the spatial range and urban
environment, including land use and land cover, depending on the
characteristics of each air quality parameter. For example, the explan­
atory power for O3 was significant at 6000 m, whereas that for NO2 and
PM10 was significant at 18,000 m. This suggests that urban and envi­
ronmental planning should consider spatial influences on air quality.
Previous studies (Weng and Yang, 2006; Li et al., 2015) have often
considered narrow buffer zones of only a few hundred meters.
Conversely, this study analyzed the relationship between the air quality
and urban environment over a broader buffer range of 6–18 km. This
broader spatial analysis highlights the fact that extensive urban envi­
ronments can significantly affect air quality and should be considered in
future urban and environmental planning. Additionally, this study found
Fig. 4. Air quality monitoring data used in air quality concentration analysis. that certain air quality parameters exhibited a strong explanatory rela­
tionship with the surrounding urban environment, including land use,
whereas others exhibited a weaker association. Therefore, future
pollutants were minimal.
land-use changes or land-use planning should consider specific air
The second scenario assessed the impact of converting adjacent
quality parameters that can be effectively analyzed for environmental
greenbelt areas into urban developments, such as single housing,
impact assessments.
apartment buildings, and industrial facilities. This scenario had more
In this study, the analysis was performed based on the direct distance
significant negative effects on air quality compared to road greening
that demonstrated the highest explanatory power. Similar analyses often
(Table 3). For instance, when the greenbelt was converted into industrial
consume several computing resources to calculate the urban environ­
areas, there was a substantial increase in SO2 (0.001936 ppm) and PM10
ment within large areas that span a radius of several kilometers.
(5.2331 μg/m3), demonstrating that urbanization has a considerable
Therefore, when there is a notable difference in explanatory power, it is
adverse impact on pollutant concentrations. Furthermore, converting
necessary to select a suitable consensus range for analysis that considers
greenbelts into single housing increased PM2.5 concentrations by
various factors, such as computing resources and analysis time. For
4.314206 μg/m3, underscoring the crucial role that green spaces play in
example, for the case of SO2 conducted in this study, the optimal buffer
maintaining better air quality.
distance for SO2 was determined to be 12,000 m, yielding the highest
explanatory power (R2 = 0.3402). However, the differences in R2 values
4. Discussion
among the 11,000-m (R2 = 0.34), 12,000-m (R2 = 0.3402), and 13,000-
m (R2 = 0.34) distances were minimal. If there is no significant differ­
Linear regression analysis provided better explanatory power than
ence in explanatory power, 1100 m can be considered the optimal
XGBoost, a representative machine learning model, in explaining urban
analysis range, as it balances efficiency in analysis time and the utili­
environmental factors, including land use within a defined spatial range,
zation of computing resources.
based on data from 351 to 365 AQMSs. This result is probably due to the
The analysis of the urban environment within a certain range around
small sample size used in the analysis, which is consistent with previous

Fig. 5. Land-use change Scenarios(Scenario 1: Road greening, scenario 2: Greenbelt development).

8
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Fig. 6. Explanatory Power (R2) of Multiple Linear Regression and XGBoost by Buffer Distance for each air pollutant.

each AQMS and the characteristics of the air pollutant concentrations other barren lands. This result is consistent with those of previous
showed that the impact on air quality varied according to the charac­ studies, indicating that barren lands contribute to fine dust generation
teristics of the urban environment. This implies that urban planning (Pinho et al., 2008).
should consider the urban environment and spatial range, including However, for PM2.5, the same type of forest showed different re­
surrounding land use and land cover. lations: coniferous forests exhibited a negative correlation, whereas
Consistent with previous research, NO2 showed a strong positive deciduous forests showed a positive correlation. When constructing
correlation with surrounding apartment housing, suggesting that NOx, land-cover maps in South Korea, security facilities, such as military in­
although naturally occurring, significantly increased owing to human stallations, airports, and power plants, are often reclassified as agricul­
industrial activities, especially from combustion facilities, such as tural or barren land. This classification may have influenced the
boilers in commercial and residential buildings (Yue et al., 2018). For analysis, necessitating additional reviews to validate the impact of these
PM10, agricultural areas, such as paddy fields (Li and Huang, 2020), and factors. Therefore, future research should focus on examining the spe­
apartment housing showed a strong positive correlation. Conversely, the cific effects of these facilities on air quality.
forest area exhibited a negative relation with PM10. Further, PM2.5 Many of our findings align with existing knowledge, such as the
concentrations were higher in areas with a high density of single-family positive association between industrial facilities with SO2, NO2, and
housing and, similar to PM10, lower in areas with extensive forest cover. particulate matter, and the mitigating effect of forested areas on PM10
Additionally, paddy fields, unconsolidated upland fields, and consoli­ and PM2.5. However, certain results appear less intuitive. In particular,
dated paddy fields were positively correlated with PM2.5, similar to air pollutants, such as CO and SO2 exhibited low explanatory power in

9
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

the linear regression analysis, lacking sufficient explanatory power in number of AQMS (360) is limited, it is difficult to generalize the results.
terms of their relationship with land cover. For example, deciduous Moreover, while this study mainly compared regression analysis and
forests showed a positive correlation with CO, however, a negative machine learning approaches, there is a potential for more in-depth
correlation with SO2. Examining the results of models with low findings by applying complex methodologies such as spatiotemporal
explanatory power reveals that several predictor variables had low p- data analysis, network analysis, or multilevel modeling. These meth­
values, indicating weak explanatory relationships between specific land odologies could offer a more nuanced explanation of the complex in­
cover types and air pollutant concentrations. Similarly, factors expected teractions between urban structure and air pollution, warranting
to reduce pollution levels, such as high elevation or proximity to rivers, consideration in future research.
may correspond to high-traffic transport corridors or small residential The findings of this study also provide important insights for urban
combustion areas, contributing to increased levels of CO, NO2, and planning and policymaking. First, the fact that air pollutant concentra­
PM10, as observed in Jain et al. (2021). These cases suggest that land tions vary depending on spatial scale indicates that specific spatial
cover variables may act as proxies for unmeasured factors (e.g., heating ranges must be considered in urban design. For example, NO2 and PM10
fuel, traffic intensity, or tourism activities), complicating purely linear exhibited effects over a wider area of about 18 km, whereas O3 showed
interpretations. This underscores the need to consider nonlinearity and measurable impacts even within a narrower range of around 6 km. This
regional drivers of air quality, as well as the necessity of incorporating highlights the need to develop tailored policies that consider the diverse
additional emissions and meteorological data to completely explain the land-use patterns within a city.
localized surges of specific pollutants, such as CO. Therefore, when Moreover, small-scale urban development efforts, such as green
predicting the concentrations of air pollutants that have weak explan­ space restoration, had a limited impact on improving air quality,
atory power in relation to land cover, additional predictor variables whereas the urbanization of green belts had a significantly negative
must be incorporated and more sophisticated modeling approaches must effect. These findings emphasize the need to preserve green belts and
be considered to better capture the complex interactions between land enforce strict regulations to minimize environmental harm during large-
cover and air pollution. scale development projects. Such insights can be used to inform various
Using the model developed in this study, we predicted changes in air environmental impact assessments and formulate climate change
pollutant concentrations caused by changes in land cover. However, adaptation policies.
restoration performed in small areas provided minimal improvement in As certain air pollutants showed low explanatory power in the model
air quality. Conversely, the significant development in the area was presented in this study, future research should utilize a broader range of
predicted to increase air pollutant concentrations. Notably, when large- predictor variables to better explain these air pollutant concentrations.
scale development projects were undertaken in previously green spaces, Additionally, more sophisticated modeling approaches should be
air pollutant concentrations were predicted to increase significantly. considered to better capture the complex interactions between land
These results are consistent with those of most previous studies that have cover and air pollution. Given the limitations of currently available
examined changes in air quality due to land use and land cover changes. official datasets, including time-series data, future studies should focus
However, this study demonstrated the ability to apply Level-3 land use on collecting more diverse and comprehensive datasets to enhance
changes, distinguishing it from previous studies that primarily examined model accuracy and interpretability.
land use and land cover changes over broader areas.
Finally, vegetation indices, such as NDVI, which indirectly assess the CRediT authorship contribution statement
quality of green spaces, were excluded from the regression analysis due
to multicollinearity issues identified during the pre-analysis. This is Lee Gunwon: Writing – review & editing, Writing – original draft,
because most forests in South Korea exhibit high vegetation vitality (Lee Methodology, Conceptualization. Han Yuhan: Writing – review &
and Park, 2020). Additionally, as the forests of South Korea are pri­ editing, Writing – original draft, Visualization, Validation, Software,
marily located in high-altitude mountainous areas, we included Resources, Formal analysis, Data curation. Geunhan Kim: Writing –
high-altitude regions in the analysis. Consequently, when predicting review & editing, Writing – original draft, Supervision, Project admin­
future land use changes, this could lead to the exclusion of green space istration, Methodology, Investigation, Funding acquisition,
quality. Therefore, future research should develop methodologies that Conceptualization.
incorporate the quality of green spaces into their analysis.
Funding sources
5. Conclusion
Declaration of generative AI in scientific writing.
This study aimed to provide crucial baseline information for urban
planning and policymaking by analyzing the impact of land cover- and Declaration of competing interest
land use-related independent variables on air pollutant concentrations.
The findings offer valuable insights into urban air quality, including The authors declare the following financial interests/personal re­
more effective urban planning and environmental management strate­ lationships which may be considered as potential competing interests:
gies that can be developed. However, several limitations exist in Geunhan Kim reports financial support was provided by Korea Agency
deriving and interpreting the results. for Infrastructure Technology Advancement (KAIA). If there are other
First, the limited number of AQMSs led to a higher explanatory authors, they declare that they have no known competing financial in­
power for linear regression analysis compared to machine learning terests or personal relationships that could have appeared to influence
models. Therefore, increasing the number of AQMSs and data samples the work reported in this paper.
will improve the performance of machine-learning models. Future
research should focus on collecting additional data from various envi­ Acknowledgements
ronments to construct more reliable models. Second, vegetation indices,
such as NDVI, were excluded from the linear regression model due to This paper is based on the findings of the research project (2025-014
high multicollinearity, which prevented the consideration of green (R)) which was conducted by the Korea Environment Institute (KEI) and
space quality. Therefore, future research should develop methods to supported by a Korea Agency for Infrastructure Technology Advance­
incorporate greenspace quality into the analysis. ment (KAIA) grant funded by the Ministry of Land, Infrastructure and
Similar to other studies, this research also has some methodological Transport (RS-2023-00242291). This paper is based on the findings of
constraints. First, because the dataset is confined to South Korea and the the research project (2025-014(R)) which was conducted by the Korea

10
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Environment Institute (KEI) and supported by a Korea Agency for Hsu, C.Y., Wu, C.D., Hsiao, Y.P., Chen, Y.C., Chen, M.J., Lung, S.C.C., 2018. Developing
land-use regression models to estimate PM2.5-bound compound concentrations.
Infrastructure Technology Advancement (KAIA) grant funded by the
Remote Sens. 10. [Link]
Ministry of Land, Infrastructure and Transport (RS-2023-00242291). Hu, K., Rahman, A., Bhrugubanda, H., Sivaraman, V., 2017. HazeEst: machine learning
based metropolitan air pollution estimation from fixed and mobile sensors. IEEE
Appendix A. Supplementary data Sens. J. 17, 3517–3525. [Link]
Huang, C., Sun, K., Hu, J., Xue, T., Xu, H., Wang, M., 2022. Estimating 2013–2019 NO2
exposure with high spatiotemporal resolution in China using an ensemble model.
Supplementary data to this article can be found online at [Link] Environ. Pollut. 292, 118285. [Link]
org/10.1016/[Link].2025.102498. Huang, Y.K., Luvsan, M.E., Gombojav, E., Ochir, C., Bulgan, J., Chan, C.C., 2013. Land
use patterns and SO2 and NO2 pollution in Ulaanbaatar, Mongolia. Environ. Res.
124, 1–6. [Link]
References Huang, Z., Du, X., 2018. Urban land expansion and air pollution: evidence from China. J.
Urban Plan. Dev 144. [Link]
Adams, M.D., Massey, F., Chastko, K., Cupini, C., 2020. Spatial modelling of particulate Irga, P.J., Burchett, M.D., Torpy, F.R., 2015. Does urban forestry have a quantitative
matter air pollution sensor measurements collected by community scientists while effect on ambient air quality in an urban environment? Atmos. Environ. 120,
cycling, land use regression with spatial cross-validation, and applications of 173–181. [Link]
machine learning for data correction. Atmos. Environ. 230, 117479. [Link] Jain, S., Presto, A.A., Zimmerman, N., 2021. Spatial modeling of daily PM2. 5, NO2, and
10.1016/[Link].2020.117479. CO concentrations measured by a low-cost sensor network: comparison of linear,
Akimoto, H., 2003. Global air quality and pollution. Science 302, 1716–1719. https:// machine learning, and hybrid land use models. Environ. Sci. Technol. 55 (13),
[Link]/10.1126/science.1092666. 8631–8641.
AlThuwaynee, O.F., Kim, S.W., Najemaden, M.A., Aydda, A., Balogun, A.L., Fayyadh, M. Johnson, M., Isakov, V., Touma, J.S., Mukerjee, S., Özkaynak, H., 2010. Evaluation of
M., Park, H.J., 2021. Demystifying uncertainty in PM10 susceptibility mapping using land-use regression models used to predict air quality concentrations in an urban
variable drop-off in extreme-gradient boosting (XGB) and random forest (RF) area. Atmos. Environ. 44, 3660–3668. [Link]
algorithms. Environ. Sci. Pollut. Res. Int. 28, 43544–43566. [Link] atmosenv.2010.06.041.
10.1007/s11356-021-13255-4. Johnston, F., Hanigan, I., Henderson, S., Morgan, G., Bowman, D., 2011. Extreme air
Baldauf, R.W., Heist, D., Isakov, V., Perry, S., Hagler, G.S.W., Kimbrough, S., Shores, R., pollution events from bushfires and dust storms and their association with mortality
Black, K., Brixey, L., 2013. Air quality variability near a highway in a complex urban in Sydney, Australia 1994–2007. Environ. Res. 111, 811–816. [Link]
environment. Atmos. Environ. 64, 169–178. [Link] 10.1016/[Link].2011.05.007.
atmosenv.2012.09.054. Kim, M., Kim, G., 2022. Modeling and predicting urban expansion in South Korea using
Balew, A., Korme, T., 2020. Monitoring land surface temperature in Bahir Dar city and its explainable artificial intelligence (XAI) model. Appl. Sci. 12, 9169. [Link]
surrounding using Landsat images. Egypt. J. Remote Sens. Space Sci. 23, 371–386. 10.3390/app12189169.
[Link] Lai, S.B.S., Binti Md Shahri, N.H.N.B.M., Mohamad, M.B., Rahman, H.A.B.A., Rambli, A.
Blocken, B., Vervoort, R., van Hooff, T., 2016. Reduction of outdoor particulate matter B., 2021. Comparing the performance of AdaBoost, XGBoost, and logistic regression
concentrations by local removal in semi-enclosed parking garages: a preliminary for imbalanced data. Math. Stat. 9, 379–385. [Link]
case study for Eindhoven city center. J. Wind Eng. Ind. Aerodyn. 159, 80–98. ms.2021.090320.
[Link] Lee, H.J., Chatfield, R.B., Strawa, A.W., 2016. Enhancing the applicability of satellite
Bonilla-Bedoya, S., Zalakeviciute, R., Coronel, D.M., Durango-Cordero, J., Molina, J.R., remote sensing for PM2.5 estimation using MODIS Deep Blue AOD and land use
Macedo-Pezzopane, J.E., Herrera, M.Á., 2021. Spatiotemporal variation of forest regression in California, United States. Environ. Sci. Technol. 50, 6546–6555.
cover and its relation to air quality in urban Andean socio-ecological systems. Urban [Link]
For. Urban Green. 59, 127008. [Link] Lee, P.S.H., Park, J., 2020. An effect of urban forest on urban thermal environment in
Bozdağ, A., Dokuz, Y., Gökçek, Ö.B., 2020. Spatial prediction of PM10 concentration Seoul, South Korea, based on landsat imagery analysis. Forests 11, 630. [Link]
using machine learning algorithms in Ankara, Turkey. Environ. Pollut. 263, 114635. org/10.3390/f11060630.
[Link] Li, J., Huang, X., 2020. Impact of land-cover layout on particulate matter 2.5 in urban
Capello, R., Camagni, R., 2000. Beyond optimal city size: an evaluation of alternative areas of China. Int. J. Digit. Earth. 13, 474–486. [Link]
urban growth patterns. Urban Stud. 37, 1479–1496. [Link] 17538947.2018.1530310.
00420980020080221. Li, X., Liu, W., Chen, Z., Zeng, G., Hu, C., León, T., Liang, J., Huang, G., Gao, Z., Li, Z.,
Chen, B., You, S., Ye, Y., Fu, Y., Ye, Z., Deng, J., Wang, K., Hong, Y., 2021. An Yan, W., 2015. The application of semicircular-buffer-based land use regression
interpretable self-adaptive deep neural network for estimating daily spatially models incorporating wind direction in predicting quarterly NO2 and PM10
continuous PM2.5 concentrations across China. Sci. Total Environ. 768, 144724. concentrations. Atmos. Environ. 103, 18–24.
[Link] Lim, C.C., Kim, H., Vilcassim, M.J.R., Thurston, G.D., Gordon, T., Chen, L.C., Lee, K.,
Chen, T., Guestrin, C., 2016. Xgboost: a scalable tree boosting system in Proc. In: 22nd Heimbinder, M., Kim, S.Y., 2019. Mapping urban air quality using mobile sampling
ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 785–794 with low-cost sensors and machine learning in Seoul, South Korea. Environ. Int. 131,
Chen, Y., Xu, Y., Wang, F., Shi, F., 2022. Mapping the emission of air pollution sources 105022. [Link]
based on land-use classification: a case study of Shengzhou, China. Land Use Policy Liu, Z., Guan, Q., Lin, J., Yang, L., Luo, H., Wang, N., 2021. A new buffer selection
117, 106083. [Link] strategy for land use regression model of PM 2.5 in Xi’an, China. Environ. Sci. Pollut.
Chossière, G.P., Xu, H., Dixit, Y., Isaacs, S., Eastham, S.D., Allroggen, F., et al., 2021. Air Control Ser. 28, 21245–21255.
pollution impacts of COVID-19–related containment measures. Sci. Adv. 7 (21), Luan, J., Zhang, C., Xu, B., Xue, Y., Ren, Y., 2020. The predictive performances of
eabe1178. random forest models with limited sample size and different species traits. Fish. Res.
Chosun, Biz, 2023. A tunnel combining road and rainwater drainage to be built in South 227, 105534. [Link]
Korea to ease traffic congestion and prevent flooding. Chosun Biz. Retrieved from. Ma, M., Yao, G., Guo, J., Bai, K., 2021a. Distinct spatiotemporal variation patterns of
[Link] surface ozone in China due to diverse influential factors. J. Environ. Manag. 288,
5REGJOEAA4I2GFUXNQ/. 112368. [Link]
Dong, J., Chen, Y., Yao, B., Zhang, X., Zeng, N., 2022. A neural network boosting Ma, R., Ban, J., Wang, Q., Zhang, Y., Yang, Y., He, M.Z., Li, S., Shi, W., Li, T., 2021b.
regression model based on XGBoost. Appl. Soft Comput. 125, 109067. [Link] Random forest model based fine scale spatiotemporal O3 trends in the Beijing-
org/10.1016/[Link].2022.109067. Tianjin-Hebei region in China, 2010 to 2017. Environ. Pollut. 276, 116635. https://
Du, N., Ottens, H., Sliuzas, R., 2010. Spatial impact of urban expansion on surface water [Link]/10.1016/[Link].2021.116635.
bodies—a case study of Wuhan, China. Landsc. Urban Plann. 94, 175–185. https:// Ma, X., Zou, B., Deng, J., Gao, J., Longley, I., Xiao, S., Guo, B., Wu, Y., Xu, T., Xu, X.,
[Link]/10.1016/[Link].2009.10.002. Yang, X., Wang, X., Tan, Z., Wang, Y., Morawska, L., Salmond, J., 2024.
Dutheil, F., Baker, J.S., Navel, V., 2020. COVID-19 and air pollution: the worst is yet to A comprehensive review of the development of land use regression approaches for
come. Environ. Sci. Pollut. Control Ser. 27 (35), 44647–44649. modeling spatiotemporal variations of ambient air pollution: a perspective from
Faiz, A., 1993. Automotive emissions in developing countries: relative implications for 2011 to 2023. Environ. Int. 183, 108430. [Link]
global warming, acidification, and urban air quality. Transp. Res. A. 27, 167–186. envint.2024.108430.
Gulati, S., Bansal, A., Pal, A., Mittal, N., Sharma, A., Gared, F., 2023. Estimating PM2.5 Mogollón-Sotelo, C., Casallas, A., Vidal, S., Celis, N., Ferro, C., Belalcazar, L., 2021.
utilizing multiple linear regression and ANN techniques. Sci. Rep. 13, 22578. A support vector machine model to forecast ground-level PM2.5 in a highly
[Link] populated city with a complex terrain. Air Qual. Atmos. Health 14, 399–409.
Hanna, S.R., Briggs, G.A., Hosker Jr, R.P., 1982. Handbook on atmospheric diffusion. [Link]
National Oceanic and Atmospheric Administration, Atmospheric Turbulence and Mun, D.C., Kil, S.H., 2024. Research on valuation of ecosystem services for water quality
Diffusion Laboratory. Oak Ridge, Tennessee (DOE/TIC-11223). improvement using unmanned aerial vehicles-Focusing on Purchased land in
Harrison, R.M., 2020. Airborne particulate matter. Philos. Trans. A Math. Phys. Eng. Sci. Gwangdong-ri area, Gwangju city (Gyeonggi). J. Korean Soc. Environ. Restor.
378, 20190319. [Link] Technol. 27, 1–16.
Hien, P.D., Men, N.T., Tan, P.M., Hangartner, M., 2020. Impact of urban expansion on Nagar, P.K., Singh, D., Sharma, M., Kumar, A., Aneja, V.P., George, M.P., Agarwal, N.,
the air pollution landscape: a case study of Hanoi. Vietnam. Sci. Total Environ. 702, Shukla, S.P., 2017. Characterization of PM2.5 in Delhi: role and impact of secondary
134635. [Link] aerosol, burning of biomass, and municipal solid waste and crustal matter. Environ.
Sci. Pollut. Res. Int. 24, 25179–25189. [Link]
3.

11
L. Gunwon et al. Atmospheric Pollution Research 16 (2025) 102498

Nowak, D.J., Crane, D.E., Stevens, J.C., 2006. Air pollution removal by urban trees and China. Energy Build. 36, 1299–1308. [Link]
shrubs in the United States. Urban For. Urban Green. 4 (3–4), 115–123. enbuild.2003.09.013.
Park, Y., Kwon, B., Heo, J., Hu, X., Liu, Y., Moon, T., 2020. Estimating PM2.5 Wei, Y.D., Ye, X., 2014. Urbanization, urban land expansion and environmental change
concentration of the conterminous United States via interpretable convolutional in China Stoch. Stoch. Environ. Res. Risk Assess. 28, 757–765. [Link]
neural networks. Environ. Pollut. 256, 113395. [Link] 10.1007/s00477-013-0840-9.
envpol.2019.113395. Weng, Q., Yang, S., 2006. Urban air pollution patterns, land use, and thermal landscape:
Phillips, B.B., Bullock, J.M., Osborne, J.L., Gaston, K.J., 2021. Spatial extent of road an examination of the linkage using GIS. Environ. Monit. Assess. 117, 463–489.
pollution: a national analysis. Sci. Total Environ. 773, 145589. [Link]
Pinho, P., Augusto, S., Martins-Loução, M.A., Pereira, M.J., Soares, A., Máguas, C., World Health Organization, 2021. New WHO global air quality guidelines aim to save
Branquinho, C., 2008. Causes of change in nitrophytic and oligotrophic lichen millions of lives from air pollution. [Link]
species in a Mediterranean climate: impact of land cover and atmospheric pollutants. w-who-global-air-quality-guidelines-aim-to-save-millions-of-lives-from-air-pollution
Environ. Pollut. 154, 380–389. [Link] .
Rajput, D., Wang, W.J., Chen, C.C., 2023. Evaluation of a decided sample size in machine Yang, S., Kim, H., Kim, S.N., Ahn, K., 2018. What is achieved and lost in living in a
learning applications. BMC Bioinf. 24, 48. [Link] mixed-income neighborhood? Findings from South Korea. J. Hous. Built Environ. 33,
05156-9. 807–828. [Link]
Roberts, P.T., Fryer-Taylor, R.E.J., Hall, D.J., 1994. Wind-tunnel studies of roughness Yang, W., Jiang, X., 2021. Evaluating the influence of land use and land cover change on
effects in gas dispersion. Atmos. Environ. 28, 1861–1870. [Link] fine particulate matter. Sci. Rep. 11 (1), 17612.
1352-2310(94)90325-5. Yang, X.F., Zheng, Y.X., Geng, G.N., Liu, H., Man, H.Y., Lv, Z.F., He, K.B., de Hoogh, K.,
Shao, Y., Ma, Z., Wang, J., Bi, J., 2020. Estimating daily ground-level PM2.5 in China 2017. Development of PM2.5 and NO2 models in a LUR framework incorporating
with random-forest-based spatiotemporal kriging. Sci. Total Environ. 740, 139761. satellite remote sensing and air quality model data in Pearl River Delta region,
[Link] China. Environ. Pollut. 226, 143–153. [Link]
Smith, P.F., Ganesh, S., Liu, P., 2013. A comparison of random forest regression and envpol.2017.03.079.
multiple linear regression for prediction in neuroscience. J. Neurosci. Methods 220, Yu, G.H., Park, S., 2021. Chemical characterization and source apportionment of PM2.5
85–91. [Link] at an urban site in Gwangju, Korea. Atmos. Pollut. Res. 12, 101092. [Link]
Su, X., An, J., Zhang, Y., Zhu, P., Zhu, B., 2020. Prediction of ozone hourly 10.1016/[Link].2021.101092.
concentrations by support vector machine and kernel extreme learning machine Yuchi, W., Gombojav, E., Boldbaatar, B., Galsuren, J., Enkhmaa, S., Beejin, B.,
using wavelet transformation and partial least squares methods. Atmos. Pollut. Res. Naidan, G., Ochir, C., Legtseg, B., Byambaa, T., Barn, P., Henderson, S.B., Janes, C.
11, 51–60. [Link] R., Lanphear, B.P., McCandless, L.C., Takaro, T.K., Venners, S.A., Webster, G.M.,
Tao, W., Liu, J., Ban-Weiss, G.A., Hauglustaine, D.A., Zhang, L., Zhang, Q., Cheng, Y., Allen, R.W., 2019. Evaluation of random forest regression and multiple linear
Yu, Y., Tao, S., 2015. Effects of urban land expansion on the regional meteorology regression for predicting indoor fine particulate matter concentrations in a highly
and air quality of eastern China. Atmos. Chem. Phys. 15, 8597–8614. [Link] polluted city. Environ. Pollut. 245, 746–753. [Link]
org/10.5194/acp-15-8597-2015. envpol.2018.11.034.
Tella, A., Balogun, A.L., 2021. GIS-based air quality modelling: spatial prediction of Yue, T., Gao, X., Gao, J., Tong, Y., Wang, K., Zuo, P., Zhang, X., Tong, L., Wang, C.,
PM10 for Selangor State, Malaysia using machine learning algorithms. Environ. Sci. Xue, Y., 2018. Emission characteristics of NOx, CO, NH3 and VOCs from gas-fired
Pollut. Res. 1-17. industrial boilers based on field measurements in Beijing city, China. Atmos.
Tuckett-Jones, B., Reade, T., 2017. City Air Quality at Height – Lessons for Developers & Environ. 184, 1–8. [Link]
Planners WSP | Parsons Brinckerhoff. Zahari, M.A.Z., Majid, M.R., Ho, C.S., Kurata, G., Nadhirah, N., Irina, S.Z., 2016.
Van Der Waals, J., 2000. The compact city and the environment: a review. Tijdschr. Relationship between land use composition and PM10 concentrations in Iskandar
Econ. Soc. Geogr. 91, 111–121. [Link] Malaysia. Clean Technol. Environ. Policy 18, 2429–2439. [Link]
Van Roode, S., Ruiz-Aguilar, J.J., González-Enrique, J., Turias, I.J., 2019. An artificial s10098-016-1263-3.
neural network ensemble approach to generate air pollution maps. Environ. Monit. Zhang, P., Ma, W., Wen, F., Liu, L., Yang, L., Song, J., Wang, N., Liu, Q., 2021. Estimating
Assess. 191, 727. [Link] PM2.5 concentration using the machine learning GA-SVM method to improve the
Wang, W., Zhao, S., Jiao, L., Taylor, M., Zhang, B., Xu, G., Hou, H., 2019. Estimation of land use regression model in Shaanxi, China. Ecotoxicol. Environ. Saf. 225, 112772.
PM2.5 concentrations in China using a spatial back propagation neural network. Sci. [Link]
Rep. 9, 13788. [Link] Zhao, B., Yu, L., Wang, C., Shuai, C., Zhu, J., Qu, S., Taiebat, M., Xu, M., 2021. Urban air
Wang, Z., Bai, Z., Yu, H., Zhang, J., Zhu, T., 2004. Regulatory standards related to pollution mapping using fleet vehicles as mobile monitors and machine learning.
building energy conservation and indoor-air-quality during rapid urbanization in Environ. Sci. Technol. 55, 5579–5588. [Link]

12

You might also like