Spatial
Spatial
Urban air pollution changes quickly over time and can also spread across nearby
cities. This study analyzes the spatio-temporal dynamics of urban air pollution and
builds an Early Warning System (EWS) to answer a practical question: Will tomorrow
be an extreme pollution day? Hourly data are collected from the AQICN API (August
2025–January 2026) for 28 cities and aggregated to daily series. Extreme events are de-
fined in a data-driven way using a city-specific 80th percentile threshold of daily AQI.
The analysis combines statistical profiling of AQI distributions, temporal methods (lag
dependence, autocorrelation, and wavelet analysis), and spatial/spatio-temporal meth-
ods (Global Moran’s I, LISA hotspots, and variograms). Weather–pollution regimes are
also identified using DBSCAN on standardized meteorology and gaseous-pollutant vari-
ables. Finally, logistic regression models are evaluated to predict extreme tomorrow.
Results show that adding continuous meteorology/chemistry drivers improves prediction
compared with using temporal features alone, while regime labels mainly support in-
terpretation. Overall, the framework provides a clear and interpretable approach for
understanding and forecasting extreme pollution risk.
2
Contents
1 Introduction 6
1.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.2 Problem Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.3 Project Objectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.4 Report Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2 Methodology 8
2.1 Data and Study Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.1.1 Data source and time range . . . . . . . . . . . . . . . . . . . . . 8
2.1.2 Study cities (core vs extended) and variables . . . . . . . . . . . . 8
2.1.3 Train/test split and standardization . . . . . . . . . . . . . . . . . 9
2.2 Preprocessing and Feature Engineering . . . . . . . . . . . . . . . . . . . 9
2.2.1 Cleaning and aggregation (hourly to daily) . . . . . . . . . . . . . 9
2.2.2 Extreme definition and target construction . . . . . . . . . . . . . 9
2.2.3 Temporal features . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.2.4 Multivariate features . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.3 Analytical Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.3.1 Statistical profiling . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.3.2 Temporal methods . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.3.3 Spatial Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.3.4 Spatio-Temporal Analysis . . . . . . . . . . . . . . . . . . . . . . 12
2.3.5 Regime discovery . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.3.6 Predictive models . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
3 Results 14
3.1 Statistical Profiling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
3.2 Temporal Dynamics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
3.2.1 Temporal AQI Dynamics . . . . . . . . . . . . . . . . . . . . . . . 17
3.2.2 Lag Dependency - Temporal Persistence . . . . . . . . . . . . . . 18
3.2.3 Multi-Scale Temporal Dynamics . . . . . . . . . . . . . . . . . . . 20
3.3 Spatial Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
3.3.1 Step 1: Testing for Spatial Autocorrelation . . . . . . . . . . . . . 22
3.3.2 Step 2: Local Hotspot Identification . . . . . . . . . . . . . . . . . 23
3.3.3 Step 3: Estimating Spatial Range . . . . . . . . . . . . . . . . . . 24
3.3.4 Synthesis of Spatial Structure . . . . . . . . . . . . . . . . . . . . 25
3.4 Spatio-Temporal Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.4.1 Step 1: Temporal Evolution of Spatial Autocorrelation . . . . . . 26
3.4.2 Step 2: Spatial Stability Analysis . . . . . . . . . . . . . . . . . . 27
3.4.3 Step 3: Seasonal Variogram Analysis (The Heterogeneity Paradox) 28
3
CONTENTS
4
List of Figures
5
Chapter 1
Introduction
1.1 Background
Urban air pollution is one of the most serious environmental challenges in modern
cities. Rapid urbanization, industrial development, and increasing transportation activ-
ities have contributed to the deterioration of air quality in many metropolitan regions.
High concentrations of pollutants such as particulate matter (PM) and nitrogen oxides
(NO2 ) pose serious risks to public health, environmental sustainability, and urban living
conditions.
Air pollution is also highly dynamic across both space and time. Pollution levels can
fluctuate from day to day due to changing meteorological conditions, human activities,
and atmospheric transport processes. Pollutants may also travel across long distances and
affect neighboring cities, creating regional pollution patterns that are difficult to capture
with simple descriptive analyses. Therefore, beyond describing pollution patterns, an
operational goal is to forecast short-term extreme-risk days to support early warnings.
Answering this question requires an integrated framework that links statistical profiling,
temporal memory, spatial structure, and multivariate atmospheric conditions.
6
CHAPTER 1. INTRODUCTION
7
Chapter 2
Methodology
8
CHAPTER 2. METHODOLOGY
9
CHAPTER 2. METHODOLOGY
label forward by one day within each city. The resulting task is a binary classification
problem: predicting whether tomorrow will be an extreme day.
10
CHAPTER 2. METHODOLOGY
where µ is the mean of the AQI time series. The ACF (1) provides a statistical measure
of how strongly pollution levels persist from one day to the next, with higher values
indicating stronger temporal dependence.
Multi-Scale Temporal Dynamics While lag analysis and autocorrelation assess im-
mediate, day-to-day serial dependence, air pollution dynamics are often driven by complex
processes that operate across broader and varying time scales, from short-term weather
cycles to seasonal trends. To capture these multi-scale variations, wavelet analysis is
applied to the AQI time series. Unlike traditional spectral methods that assume station-
arity, wavelet analysis transforms the one-dimensional time series into a two-dimensional
time-frequency representation, allowing for the detection of localized oscillatory behav-
iors.
In this study, the wavelet power spectrum is generated to visualize the intensity
of these temporal oscillations across different frequencies. To summarize this complex
information into a single comparable feature across cities, the mean wavelet energy is
calculated. This metric quantifies the overall strength of multi-scale temporal fluctuations
in the AQI series.
11
CHAPTER 2. METHODOLOGY
12
CHAPTER 2. METHODOLOGY
13
Chapter 3
Results
Figure 3.1: Evidence of right-skewed AQI distributions. The mean–median gap across
cities indicates that extreme pollution events pull the mean toward the right tail. The
Beijing example illustrates how the right tail increases the mean above the median.
The distribution fitting results indicate that the Normal distribution generally pro-
vides the poorest fit to the AQI data. In most cities, the Normal model produces sig-
nificantly higher AIC and BIC values compared with the other candidate distributions.
14
CHAPTER 3. RESULTS
This confirms that the assumption of symmetric Gaussian behavior does not accurately
describe the statistical structure of daily AQI.
In contrast, the Gamma and Lognormal distributions consistently provide better fits.
These models capture the asymmetric structure of AQI data, where most observations
occur at moderate pollution levels while a smaller number of observations extend toward
very high values.
Figure 3.2 summarizes the distribution fitting comparison across the eight cities. The
horizontal axis shows the ∆AIC between the best and the second-best models. Larger
values indicate stronger statistical evidence supporting the winning distribution.
Figure 3.2: Distribution fitting results across the eight cities. The plot shows the ∆AIC
between the best and the second-best model. Larger values indicate stronger statistical
evidence for the winning distribution.
The analysis demonstrates that AQI does not follow a Gaussian distribution but
instead shows right-skewed and heavy-tailed characteristics. Because of this behavior,
extreme pollution events are defined using a percentile-based threshold rather than a
fixed AQI value.
Specifically, the 80th percentile of the AQI distribution is used as the threshold for
extreme pollution events in each city. Figure 3.3 illustrates these city-specific thresholds.
The red vertical line in each histogram indicates the AQI level that separates typical days
from extreme pollution days.
15
CHAPTER 3. RESULTS
Figure 3.3: City-specific 80th percentile thresholds for extreme pollution events. The red
vertical line indicates the AQI level used to define extreme pollution days in each city.
The resulting extreme pollution indicator is later used as the binary target variable
for the Early Warning System.
16
CHAPTER 3. RESULTS
Figure 3.4: Temporal AQI dynamics for (Beijing, Hamburg, and Seoul)
From Figure 3.4, clear differences in temporal pollution patterns can be observed
between the cities. Beijing exhibits the highest variability with frequent and intense
spikes, indicating unstable pollution levels during the observed period. Seoul also shows
noticeable fluctuations with several peaks during the later months, suggesting periods of
increased pollution. In contrast, Hamburg presents relatively lower AQI levels and more
moderate variations, reflecting generally cleaner air conditions compared to the other
cities.
To complement the visual analysis, Table 1 summarizes the average AQI level for each
city during the study period.
The summary statistics in Table 3.1 reveal clear differences in air pollution levels
among the selected cities. Beijing records the highest median AQI value (72.0), indicating
17
CHAPTER 3. RESULTS
more severe pollution conditions compared to the other locations. In contrast, Hamburg
shows the lowest median AQI value (27.5), suggesting relatively cleaner air conditions.
Several cities fall within a moderate pollution range. Seoul (54.5) and Warsaw (51.0)
display moderately elevated AQI levels, while Nagoya (46.0) remains slightly lower but
still within the mid-range. Meanwhile, Fukuoka (42.0) and Zhongshan (42.0) present sim-
ilar median AQI values, indicating relatively moderate air quality conditions. Kawasaki
(38.0) shows somewhat lower pollution levels compared with most Asian cities in the
sample.
Overall, the results highlight substantial variation in air pollution levels across the
studied cities. Beijing clearly stands out as the most polluted city in terms of median
AQI, while Hamburg consistently exhibits the cleanest air conditions. These differences
provide an important baseline for further analysis of the temporal dynamics of AQI across
locations.
While the time-series visualization provides an overall view of how AQI levels fluctuate
over time, it does not explicitly quantify the relationship between consecutive observa-
tions. To better understand whether current pollution levels are influenced by recent
conditions, the next analysis examines the lag dependency in the AQI series.
Figure 3.5: Lag relationship between the current AQI and the previous day’s AQI for
selected cities
The scatter plot illustrates a positive relationship between the AQI of a given day and
the AQI of the previous day across the selected cities. In most cases, higher AQI values
tend to follow periods with already elevated pollution levels, indicating that air pollution
conditions often persist across consecutive days.
18
CHAPTER 3. RESULTS
Table 3.2 summarizes the temporal persistence indicators of AQI across the selected
cities. While the mean Lag1 AQI reflects the general level of pollution carried over from
the previous day, the ACF1 statistic directly measures the strength of serial dependence
in the time series.
Cities such as Warsaw (0.7209), Zhongshan (0.6220), and Seoul (0.6108) exhibit rel-
atively high autocorrelation values, indicating strong persistence in daily air pollution
levels. Hamburg (0.5711) also shows a noticeable degree of temporal dependence. In
contrast, Kawasaki (0.2864) and Nagoya (0.3179) display weaker autocorrelation values,
suggesting more variable day-to-day pollution dynamics. Beijing (0.3997) and Fukuoka
(0.3898) fall within a moderate range of temporal dependence.
Overall, the combined evidence from Lag1 AQI and ACF1 indicators suggests that
air pollution levels tend to exhibit temporal persistence across most cities, although the
strength of this dependence varies across urban environments.
While lag dependency and autocorrelation analysis reveal the degree of short-term
persistence in AQI values, they mainly capture relationships between consecutive obser-
vations. Such measures provide limited insight into broader temporal structures that may
emerge across longer time horizons. To further investigate these multi-scale temporal dy-
namics and identify oscillatory patterns occurring at different time frequencies, wavelet
analysis is applied to the AQI time series.
19
CHAPTER 3. RESULTS
The wavelet power spectrum for Beijing shows several high-energy regions, particu-
larly at shorter periods between roughly 4 and 12 days. This indicates strong short-term
oscillations in AQI levels. Overall, the pattern suggests pronounced multi-scale variability
in Beijing’s air pollution dynamics.
Table 3.3 summarizes the average wavelet energy for each city, reflecting the strength
of multi-scale temporal oscillations in AQI. Beijing shows the highest mean wavelet energy
(0.1400), indicating stronger temporal variability across different time scales. Nagoya
(0.1260), Fukuoka (0.1251), and Kawasaki (0.1242) also exhibit relatively high energy
levels, suggesting noticeable oscillatory patterns. In contrast, Seoul (0.1010) and Warsaw
(0.0992) present lower values, indicating weaker multi-scale fluctuations.
20
CHAPTER 3. RESULTS
Overall, the results suggest that the strength of temporal oscillations in AQI varies
across cities, with some urban environments exhibiting more pronounced multi-scale dy-
namics than others.
Conclusion
This section investigated the temporal dynamics of air quality across multiple cities
using several complementary time-series analysis techniques. The results show that AQI
levels exhibit clear temporal structure, including persistence across consecutive days,
short-term fluctuations in pollution levels, and multi-scale oscillatory behavior.
Lag analysis and autocorrelation results indicate that air pollution often persists from
one day to the next, although the strength of this dependence varies between cities. The
analysis of daily AQI differences further reveals that short-term variability differs across
locations, with some cities showing more volatile pollution dynamics. In addition, wavelet
analysis highlights the presence of multi-scale temporal oscillations, suggesting that air
pollution patterns are influenced by processes operating at different time scales.
Overall, these findings provide a comprehensive characterization of the temporal be-
havior of AQI and allow the extraction of several informative temporal indicators. These
indicators serve as useful quantitative features for the subsequent stages of the analyt-
ical framework, where they can support deeper comparative analysis and data-driven
modeling of urban air pollution patterns.
21
CHAPTER 3. RESULTS
Figure 3.7: Time series of Global Moran’s I showing consistently positive values across
all months.
As shown in Figure 3.7, the Moran’s I value remains consistently positive (averaging
around 0.38) and never drops to zero. Consequently, air pollution is absolutely not
a random event. It operates as a persistently connected network. A dirty city will
inevitably pull its neighbors into a pollution cluster.
22
CHAPTER 3. RESULTS
Figure 3.8: LISA Cluster Map: Red dots represent High-High Hotspots, Blue dots rep-
resent Low-Low Coldspots, and Grey dots represent not significant areas.
Figure 3.8 reveals a stark geographic reality. The “High-High” hotspots (red dots)
are densely concentrated in South Asia (especially Delhi, Mumbai, etc.). These cities are
not only highly polluted themselves but are surrounded by other heavily polluted areas.
Environmental responsibility is not shared equally. South Asia acts as the
“pollution engine” of the region. Any predictive model or policy must treat these anchored
hotspots as the primary risk sources.
23
CHAPTER 3. RESULTS
Figure 3.9: Spatial Range: The Variogram curve plateaus at 5250 km, showing the
maximum distance of spatial correlation.
As shown in Figure 3.9 and our computational output, the model returned highly
specific parameters quantifying the spatial structure of pollution:
• Nugget (0.00): Zero micro-scale variation (no major sensor noise or hyperlocal
interruptions).
• Total Sill (2619.43): The maximum spatial variance (the plateau where correla-
tion ends).
• Range (5250 km): The exact distance where correlation drops to zero based on
the Spherical model cutoff.
The “blast radius” of air pollution is massive. Spatial correlation extends to exactly
5250 km—beyond this distance, cities become statistically independent. This Range pro-
vides a quantitative boundary for setting up regional environmental treaties and defining
the radar scope for our Early Warning System.
Note: With 28 globally distributed cities, this 5250 km range reflects large-scale re-
gional spatial structure rather than local atmospheric dispersion.
24
CHAPTER 3. RESULTS
2. Anchored epicenters: LISA hotspots are heavily concentrated in South and East
Asia.
We have now mapped the spatial architecture of pollution. The analytical question
shifts: Does this rigid spatial architecture stay still, or does it flex and bend
with the seasons?
25
CHAPTER 3. RESULTS
Figure 3.10: Seasonal Pattern: Moran’s I spikes dramatically during the Winter months,
proving seasonal elasticity.
The data reveals a massive Winter Amplification effect. During summer, the
spatial clustering is moderate. However, as winter approaches, the Moran’s I score jumps
by nearly 39.2%.
This indicates that winter does not merely make the air colder; it forces cities to bind
together. Temperature inversions (a common winter phenomenon) act like a giant lid,
trapping pollution close to the ground and enhancing regional coordination.
26
CHAPTER 3. RESULTS
*Why Bangkok? Bangkok acts as a sensitive “buffer zone” on the edge of the Asian
pollution core. During summer, clean ocean monsoons push it out of the network. In
winter, shifting continental winds and regional transport pull it directly into the massive
High-High cluster.
Ultimately, air pollution has an extremely rigid core structure. The core hotspots—
Delhi, Beijing, Mumbai—remain
27
CHAPTER 3. RESULTS
Figure 3.11: Seasonal Range Elasticity: The Winter curve reaches its Sill much faster
due to intense spatial heterogeneity.
The output revealed a fascinating paradox: The Winter Range actually contracts
by 13.6% (an absolute reduction of 905 km) compared to the Summer Range.
The mathematical parameters behind these curves provide the definitive proof of this
phenomenon:
Table 3.5: Seasonal Variogram Parameters: Quantifying the Spatial Heterogeneity Para-
dox
28
CHAPTER 3. RESULTS
*Note on Winter Nugget (0): The micro-scale variance drops to zero primarily due to
the preprocessing pipeline and the nature of the dataset. Seasonal aggregation effectively
smooths out random, short-term sensor noise. Additionally, given the sparse intercity
distance between the 28 globally distributed cities, hyperlocal fluctuations are structurally
filtered out.
Quantitative Breakdown:
• Range Contraction: The spatial range shrinks by 13.6% (an absolute reduction
of 905 km).
• Sill Explosion: The spatial variance (Sill) in winter is nearly 3.7 times higher
than in summer (2526.66 vs 689.41).
Urban air pollution operates as a geographically anchored network with seasonal elas-
ticity. The structural core is locked in place (96.4% stability), but winter tightens the
screws, amplifying the intensity and hyper-concentrating the danger into extreme regional
hotspots.
We now understand the spatial and temporal physics of pollution. However, the
atmosphere is driven by dozens of variables (Temperature, Humidity, Wind, PM2.5,
NO2). To feed this into an Early Warning System, how can we compress this chaotic,
multi-dimensional weather data into simple, predictable “Weather Regimes”?
29
CHAPTER 3. RESULTS
Table 3.6: VIF results for the 8 standardized drivers (core cities).
Variable VIF
co mean 1.84
no2 mean 1.72
so2 mean 1.46
wind mean 1.39
humid mean 1.38
press mean 1.38
o3 mean 1.26
Figure 3.12: kNN distance plot (k = 5) used to select the DBSCAN eps.
30
CHAPTER 3. RESULTS
The DBSCAN visualization (Figure 3.13) is consistent with this structure: one large
cluster dominates, with only a few isolated points. For interpretation, DBSCAN outputs
are mapped to two labels, Normal Baseline (main cluster) and Rare or Anomaly (noise
points). The figure is a 2D projection, while clustering is performed in the full 8D
standardized feature space.
Figure 3.13: DBSCAN clustering result for core cities (eps=1.5, minPts=5).
Table 3.7: Next-day extreme rate by DBSCAN regime (core cities, no PCA).
31
CHAPTER 3. RESULTS
Model ROC-AUC
A (Temporal baseline) 0.687
B (+ Weather Regime) 0.680
C (+ Raw 8 vars) 0.744
Model PR-AUC
A (Temporal baseline) 0.496
B (+ Weather Regime) 0.493
C (+ Raw 8 vars) 0.557
ROC curves (visual comparison). Figure 3.14 confirms the table results. The curve
for Model C lies above Models A and B over most of the range, meaning it achieves higher
sensitivity for the same false-positive rate. Models A and B nearly overlap, consistent
with their similar ROC-AUC values.
32
CHAPTER 3. RESULTS
Figure 3.14: ROC curves on the test set for Models A–C. Model C dominates most of
the curve, consistent with its higher ROC-AUC.
Although Model C is the best performer, overall operational metrics remain moder-
ate (precision and recall around 0.5–0.57). This is expected because next-day extreme
events are relatively rare and are influenced by unobserved factors (e.g., abrupt emission
changes and local micro-meteorology), which limits perfect separation. In practice, the
threshold τ = 0.45 represents a trade-off: some false alarms (FP) are tolerated to reduce
missed extreme days (FN), which is often preferred in early warning settings. Over-
all, the strongest EWS performance is achieved by combining temporal memory features
with continuous meteorology/chemistry drivers (Model C). The regime label is useful for
interpretation but does not improve predictive accuracy in this evaluation setting.
33
Chapter 4
4.1 Conclusion
This report developed an integrated framework to characterize spatio-temporal dy-
namics of urban air pollution and to support an Early Warning System (EWS) for pre-
dicting whether extreme tomorrow occurs. Statistical profiling indicates that daily AQI
is typically right-skewed and heavy-tailed, motivating a city-specific, data-driven extreme
definition based on the 80th percentile of daily AQI. Temporal analysis shows short-term
persistence, supporting operational memory features (Lag1, Delta AQI, extreme lag1).
Spatial and spatio-temporal modules highlight non-random geographic structure and sea-
sonal elasticity in clustering.
Multivariate regimes are discovered using DBSCAN on eight standardized meteorol-
ogy drivers. Rare/anomalous regime days exhibit higher next-day extreme risk than
baseline days, providing interpretable atmospheric context. Predictive validation shows
that the best performance is achieved by combining temporal memory with continu-
ous drivers (Model C), improving ROC-AUC and PR-AUC over the temporal baseline.
Adding the binary regime label alone does not improve discrimination, suggesting it is
more useful for interpretation than for prediction once temporal memory is included.
4.2 Limitations
The study period is limited (August 2025–January 2026), and evaluation uses a
single time-aware split rather than rolling backtesting. The percentile-based extreme
threshold may differ from health-based regulatory cutoffs. Predictive performance is also
constrained by unobserved drivers. DBSCAN regimes are summarized into baseline vs
rare/noise, which may oversimplify atmospheric diversity.
34
References
[1] World Air Quality Index Project, “Aqicn: World air quality index project,” https:
//[Link]/api/, 2025.
[2] R. S. Bivand, E. Pebesma, and V. Gomez-Rubio, Applied Spatial Data Analysis with
R, 2nd ed. Springer, 2013.
[4] N. Cressie, Statistics for Spatial Data. John Wiley & Sons, 1993.
35