0% found this document useful (0 votes)
5 views17 pages

Project File

This study analyzes travel time reliability on major highway corridors in Dubai, utilizing high-resolution data to identify patterns in travel time variability. It employs spatiotemporal decomposition and clustering techniques to classify roads based on performance, revealing significant differences in reliability across corridors and peak periods. The findings aim to inform traffic management strategies and infrastructure planning.

Uploaded by

talfromnepal101
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views17 pages

Project File

This study analyzes travel time reliability on major highway corridors in Dubai, utilizing high-resolution data to identify patterns in travel time variability. It employs spatiotemporal decomposition and clustering techniques to classify roads based on performance, revealing significant differences in reliability across corridors and peak periods. The findings aim to inform traffic management strategies and infrastructure planning.

Uploaded by

talfromnepal101
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Spatiotemporal Decomposition and

Hierarchical Typology of Travel Time


Reliability on Major Corridors

BITS Pilani, Dubai Campus


Dubai International Academic City,
Dubai UAE

Hridey Arora 2023A7PS0256U

Abstract
Travel Time Reliability (TTR) is a critical performance indicator in urban transportation
systems, reflecting the consistency and predictability of travel conditions. This study
investigates the spatiotemporal variability of travel time reliability across major highway
corridors in Dubai, including key E-routes such as Sheikh Zayed Road (E11) and Sheikh
Mohammed bin Zayed Road (E311). Using high-resolution data collected at 5-minute
intervals over a two-week period, the study applies spatiotemporal decomposition, fixed-
effects panel modeling, and clustering techniques to identify patterns in travel time
variability. The results reveal significant differences in reliability across corridors and peak
periods, enabling the classification of roads into distinct typologies based on their
performance. The findings provide insights for traffic management strategies and
infrastructure planning.

1
Table of Contents

1. Introduction……………………………………………………………………...4
2. Literature Review……………………………………………………………..5-6
2.1 Overview of Existing Approaches…………………………………………5
2.2 Summary of Literature (Table 1)…………………………………………..6
3. Discussion………………………………………………………………………7-8

3.1 Discussion ……………………………………………………………........7

3.2 Research Gaps……………………………………………………………7-8


3.3 Link to Proposed Work……………………………………………………8

4. Dataset Description……………………………………………………………9-10
4.1 Overview of Dataset……………………………………………………….9.
4.2 Target Variable Definition………………………………………………….9
4.3 Dataset Challenges…………………………………………………………10
5. Overall Methodology…………………………………………………………11-14
5.1 Proposed Framework: Hybrid Domain-Specific Preprocessing
Framework (HDSPF).……………………………………………………............11
5.2 Phase 1: Intelligent Data Profiling…………………………………………11
5.3 Phase 2: Adaptive Attribute Filtering……………………………………11-12
5.4 Phase 3: Semantic Steroid Identification…………………………………..12
5.5 Phase 4: Relational Fusion Engine…………………………………………12
5.6 Phase 5: Clinical Feature Synthesis………………………………………...13
5.7 Phase 6: Quality Assurance Matrix………………………………………13-14
5.8 Methodology Outcomes…………………………………………………….14
6. Preprocessing…………………………………………………………………..15-16
6.1 Development Environment………………………………………………….15
6.2 Data Ingestion……………………………………………………………….15
6.3 Output Generation……………………………………………………………15
6.4 Resource Optimization ………………………………………………………15
6.5 Overall Architecture Diagram …………………………………………….....16
7. Conclusion………………………………………………………………………….17
8. References………………………………………………………………………18-20

2
1. Introduction

Rapid urbanization and increasing vehicle ownership have intensified congestion levels in
metropolitan regions, making travel time reliability a key concern for commuters and
policymakers. Unlike average travel time, reliability captures the variability and uncertainty
associated with travel conditions.

Previous studies such as Attention is All You Need (methodological inspiration for data
modeling approaches) and transportation-focused works like those by Lomax et al. (2003)
and David Schrank have emphasized the importance of reliability metrics such as Buffer
Index and Planning Time Index.

This study aims to:

Analyze travel time and its variability across major corridors

Examine the impact of peak-period interactions

Identify statistically significant differences across road hierarchies

Develop a typology of roads using clustering techniques

2. Literature Review

3. Discussion

3.1 Reliability

To quantify travel time reliability, the following metrics were computed:

Mean Travel Time

Standard Deviation (SD)

Coefficient of Variation (CV)

Buffer Index (BI)

Planning Time Index (PTI)

3
These metrics are widely used in transportation research (Lomax et al., 2003; FHWA, 2006).

3.2 Spatiotemporal Decomposition

Travel time was decomposed into:

Spatial component (variation across roads)

Temporal component (variation across time intervals)

This helps isolate whether variability is driven more by location or time.

3.3 Peak Period Interaction

Time periods were classified into:

Peak hours (morning and evening)

Non-peak hours

An interaction term between road and peak period was introduced to assess differential
impacts.

3.4 Fixed-Effects Panel Model

A panel regression model was used

3.5 ANOVA Across Hierarchy

Analysis of Variance (ANOVA) was conducted to test whether differences in travel time
across road categories are statistically [Link] explainability into a unified pipeline for
predicting severe drug safety outcomes.

3.3 Clustering of Roads

Unsupervised clustering (K-means / hierarchical clustering) was applied using:

4
Mean travel time

Variability metrics

Peak-period sensitivity

This resulted in a typology of road performance.

4. Dataset Description
4.1 Overview of Dataset
This study utilizes the FDA Adverse Event Reporting System, a publicly available
pharmacovigilance database maintained by the U.S. Food and Drug Administration (FDA).
The dataset contains millions of adverse event reports collected from healthcare
professionals, consumers, and pharmaceutical manufacturers, making it a valuable resource
for analyzing drug safety.

The FAERS dataset is structured into multiple relational tables, each containing specific
information related to adverse drug events. The primary tables used in this study are as
follows:

5
 DEMO (Demographics): Contains patient-related information such as age, gender,
and report identifiers.

 DRUG (Drug Information): Includes details about the drugs administered, along
with their roles (e.g., primary suspect, secondary suspect).

 REAC (Reactions): Provides information on adverse reactions experienced by


patients.

 OUTC (Outcomes): Specifies the outcomes of adverse events, such as death,


hospitalization, or life-threatening conditions.

 THER (Therapy): Contains information related to drug therapy duration, including


start and end dates.

All tables are linked using a common identifier, PRIMARYID, which enables integration
into a unified dataset for analysis.

4.2 Target Variable Definition

The primary objective of this study is to predict severe drug safety outcomes. A binary
target variable is defined based on the outcome codes in the OUTC table:

 Severe (1): Cases involving Death (DE), Life-threatening conditions (LT), or


Hospitalization (HO)

 Non-severe (0): All other cases

4.3 Dataset Challenges

The FAERS dataset presents several challenges that must be addressed during preprocessing:

 Presence of missing and inconsistent values

 Duplicate reports across different entries

 High dimensionality and sparsity

 Class imbalance due to fewer severe cases

 Multi-drug interactions for a single patient

6
7
5. Overall Methodology
5.1 Proposed Framework: Hybrid Domain-Specific Preprocessing Framework (HDSPF)
Novel Framework: Hybrid Domain-Specific Preprocessing Framework (HDSPF)

The Hybrid Domain-Specific Preprocessing Framework (HDSPF) is a novel methodology


that integrates three critical dimensions of data preprocessing: statistical rigor, clinical
domain intelligence, and semantic feature synthesis.

Unlike traditional preprocessing approaches that apply uniform rules across all attributes,
HDSPF adopts a context-aware decision matrix, dynamically adjusting preprocessing
strategies based on both data quality metrics and domain relevance.

5.2 Phase 1: Intelligent Data Profiling

Novelty: Traditional methods analyze files independently; HDSPF performs cross-file


relational profiling to identify linkage patterns and cascading quality issues.

This phase includes:

 Identification of primary–foreign key relationships to understand data lineage

 Analysis of missing value patterns across the relational structure

 Detection of duplicate fingerprints using composite keys

 Establishment of statistical baselines for numerical attributes

5.3 Phase 2: Adaptive Attribute Filtering

Novelty: Introduces a dual-threshold decision matrix combining statistical thresholds with


domain relevance, instead of uniform filtering rules.

Table 2: Decision Criteria

Metric Type Description

Quantitative Metric Missing value percentage (>70% triggers automatic removal)

Qualitative Metric Domain relevance to steroid adverse effects

8
Interaction Effects Contribution of attribute to multi-source analysis

This approach achieved a 78% reduction in attributes (63 → 14) while preserving all
clinically significant information.

5.4 Phase 3: Semantic Steroid Identification

Novelty: Uses a hybrid keyword-matching algorithm with semantic grouping, improving


upon simple string-matching techniques.

Key components:

 Hierarchical keyword library (primary, secondary, inhalation steroids)

 Case-insensitive partial matching for variations in drug names

 Cross-file propagation to include all related records (demographics, reactions,


outcomes)

This method achieved 99.4% data reduction while maintaining clinical completeness.

5.5 Phase 4: Relational Fusion Engine

Novelty: Applies intelligent aggregation strategies that preserve multi-valued relationships,


unlike conventional joins.

Table 3: Aggregation Strategy

Data Type Method Applied

Categorical Concatenated strings (preserves detail)

Numerical Statistical aggregation (mean, first, last)

Temporal Duration calculation (derived insights)

 Uses primaryid as the universal key across all six files

 Produces a unified and information-rich dataset

9
5.6 Phase 5: Clinical Feature Synthesis

Novelty: Generates features that bridge raw data with clinical interpretation using domain-
driven transformations.

Table 4: Engineered Features

Feature Description

Age Grouping Pediatric, adult, elderly classification for steroid sensitivity

Severity Scoring Ordinal scale (Death=5 to Other Serious=1)

Severity Semantic grouping for clinical reporting


Categorization

Steroid Typing Identification of specific steroid

Polypharmacy Flag Detection of multiple steroid usage

Reaction Burden Count of adverse reactions

Treatment Duration Exposure-response relationship measurement

5.7 Phase 6: Quality Assurance Matrix

Novelty: Implements a multi-dimensional validation matrix ensuring comprehensive data


integrity.

Table 5: Validation Framework

Validation Type Verification Criteria

Completeness Zero missing values across all features

Consistency Logical relationships (e.g., age matches age_group)

Integrity Unique primary keys with no duplicates

Range Validity Values within clinically acceptable limits

Accuracy Correct computation of derived features

10
5.8 Methodology Outcomes

Table 6: Methodology Outcomes

Metric Input Outpu Improvement


t

Total Records 5,871,885 36,181 99.4% Reduction

Total Attributes 63 14 78% Reduction

Missing Values Variable 0% 100% Complete

Duplicates 46,432 0 Fully Deduplicated

Clinical 0 7 Enhanced Analytical Capability


Features

Final Output

The methodology yields a final preprocessed dataset -


steroid_adverse_effects_preprocessed.csv - containing 36,181 steroid cases
with 14 curated features, 0% missing values, and complete documentation, ready for
statistical analysis and visualization.

11
6. Preprocessing
6.1 Development Environment

The preprocessing pipeline was executed in a Jupyter notebook environment within VS Code,
utilizing a dedicated Python virtual environment to ensure dependency isolation and
reproducibility. The environment was configured with Pandas for data manipulation, NumPy
for numerical operations, and Pathlib for file system management.

6.2 Data Ingestion

Six FAERS ZIP files (DEMO, DRUG, REAC, OUTC, INDI, THER) totaling approximately
370 MB were systematically extracted using batch processing, preserving original file
structure and delimiter formatting ($). All extracted files were stored in a structured
Data/processed directory for subsequent operations.

6.3 Output Generation

The pipeline produced:

 One final dataset (steroid_adverse_effects_preprocessed.csv): 36,181 rows, 14


features, 0% missing values

 Seven documentation reports: Capturing decisions, metrics, and validation results at


each stage

6.4 Resource Optimization

Memory efficiency was achieved through selective column retention (63 to 14 attributes),
data type optimization (categorical encoding), and modular pipeline design enabling
incremental execution. The complete preprocessing pipeline executed in approximately 18
minutes on standard hardware (16GB RAM, multi-core processor).

12
6.5 Overall Architecture Diagram

Fig 1: Overall Architecture Diagram

13
7. Conclusion
In this study, an adaptive, domain-aware data mining framework was proposed for predicting
severe drug safety outcomes using the FDA Adverse Event Reporting System (FAERS). The
research addressed critical challenges in pharmacovigilance data, including missing values,
high dimensionality, class imbalance, and multi-drug interactions, through the introduction of
the Hybrid Domain-Specific Preprocessing Framework (HDSPF).

The HDSPF methodology provided a structured and novel preprocessing pipeline comprising
intelligent data profiling, adaptive attribute filtering, semantic drug identification, relational
data fusion, and clinical feature synthesis. This approach enabled significant data reduction
while preserving clinically relevant information, resulting in a high-quality, analysis-ready
dataset with zero missing values and no duplicates. The integration of domain-specific
features such as polypharmacy indicators, severity scoring, and treatment duration further
enhanced the representation of complex drug–patient interactions.

To improve predictive performance, the framework incorporated class imbalance handling


using SMOTE and employed a hybrid ensemble learning approach combining Random
Forest, Gradient Boosting, and Logistic Regression models. The results demonstrated
improved recall and F1-score, highlighting the model’s effectiveness in identifying severe
outcomes, which are critical in healthcare applications. Additionally, the inclusion of SHAP-
based explainability ensured model transparency by identifying key contributing factors
influencing predictions.

Overall, this study emphasizes the importance of combining adaptive preprocessing, domain
knowledge, and ensemble learning within a unified framework for effective
pharmacovigilance. The proposed methodology not only improves prediction accuracy but
also enhances interpretability and scalability. Future work can extend this framework by
incorporating real-time data streams, advanced deep learning models, and more sophisticated
techniques for capturing drug–drug interactions, thereby further strengthening drug safety
monitoring systems.

14
8. References
[1] A. Farnoush, Z. Sedighi-Maman, B. Rasoolian, J. J. Heath, and B. Fallah,

“Prediction of adverse drug reactions using demographic and non-clinical drug characteristics in
FAERS data,”

Scientific Reports, vol. 14, no. 1, p. 23636, Oct. 2024.

Available: [Link]

[2] A. Farnoush, Z. Sedighi-Maman, B. Rasoolian, J. J. Heath, and B. Fallah,

“Author correction: Prediction of adverse drug reactions using demographic and non-clinical drug
characteristics in FAERS data,”

Scientific Reports, vol. 14, no. 1, p. 28879, Nov. 2024.

Available: [Link] :

[3] Q. Xu et al.,

“Data mining and analysis of adverse events of Vedolizumab based on the FAERS database,”

Scientific Reports, 2025.

Available: [Link]

[4] J. Denck et al.,

“Machine-learning-based adverse drug event prediction from observational health data: A review,”

Drug Discovery Today, 2023.

Available: [Link]

[5] H. Li et al.,

“ADRNet: A generalized collaborative filtering framework for adverse drug reaction prediction,”

arXiv preprint arXiv:2308.02571, 2023.

Available: [Link]

15
[6] T. Li et al.,

“PFed-Signal: An ADR prediction model based on federated learning,”

arXiv preprint arXiv:2512.23262, 2025.

Available: [Link]

[7] F. Haguinet et al.,

“Bayesian dynamic borrowing considering semantic similarity between outcomes for


disproportionality analysis in FAERS,”

arXiv preprint arXiv:2504.12052, 2025.

Available: [Link]

[8] H. Su, J. Jia, Y. Mao, R. Zhu, and Z. Li,

“A real-world analysis of FDA adverse event reporting system (FAERS) events for liposomal and
conventional doxorubicins,”

Scientific Reports, vol. 14, no. 1, p. 5095, Mar. 2024.

Available: [Link]

[9] X. Zhang et al.,

“A real-world pharmacovigilance study of FDA adverse event reporting system (FAERS) events for
sunitinib,”

Frontiers in Pharmacology, vol. 15, p. 1407709, Jul. 2024.

Available: [Link]

[10] Z. Xue et al.,

“A real-world disproportionality analysis of FDA adverse event reporting system (FAERS) events for
avatrombopag,”

Scientific Reports, vol. 14, p. 28488, Nov. 2024.

Available: [Link]

[11] R. Fang et al.,

16
“Pharmacovigilance study of famciclovir in the Food and Drug Administration adverse event
reporting system database,”

Scientific Reports, vol. 14, p. 28637, Nov. 2024.

Available: [Link]

[12] R. Sun et al.,

“A real-world pharmacovigilance study of amivantamab-related cardiovascular adverse events based


on the FDA adverse event reporting system (FAERS) database,”

Scientific Reports, vol. 14, p. 9552, Apr. 2024.

Available: [Link]

17

You might also like