MIHIR PATEL
Halifax, NS | (902) 916-0913 | mihir20011patel@[Link] | LinkedIn | GitHub
SUMMARY
Data Engineer with 2+ years of hands-on experience building Databricks, PySpark, and SQL ETL pipelines on
Azure, plus Angular dashboard delivery and Java/SQL back-end work. Combines dimensional data modeling, data
quality testing, and lineage discipline with full-stack fluency across banking-grade and healthcare data platforms.
SKILLS
Languages: Python (PySpark), SQL, Java (17), JavaScript, Bash
Frameworks & ETL: Databricks, dbt Cloud, Apache Spark/PySpark, Spring Boot, Angular (consumer), MLflow
Data Management: Dimensional modeling, ERD design, data profiling, data quality testing (uniqueness, not-null,
schema-consistency), metadata enrichment, data provenance & lineage, semantic layers
Databases: MySQL, Azure Blob Storage, Amazon DynamoDB, Google BigQuery, Amazon S3, OpenSearch
Cloud: Azure (Blob Storage, Databricks), AWS (Lambda, S3, API Gateway, SageMaker, DynamoDB, EventBridge)
Pipeline Orchestration: Databricks Jobs, AWS EventBridge, dbt Cloud scheduling (Azure Data Factory)
DevOps: Git, SVN, GitLab CI, Terraform, CloudFormation, Docker, Agile/Scrum
BI: Tableau, Power BI, Looker Studio
PROFESSIONAL EXPERIENCE
DATA ENGINEER — CO-OP Nova Scotia Health
Halifax, NS, Canada January 2026 – Present
• Built end-to-end Databricks + PySpark ETL pipelines ingesting operational metrics from Azure Blob Storage
into forecasting models across provincial facilities.
• Co-developing a scheduled hourly orchestration pipeline monitoring Azure Blob Storage for new uploads —
pattern directly analogous to ADF trigger-based ingestion.
• Tracked experiments and artifacts with MLflow for reproducibility, capturing model lineage across facility- and
zone-specific tuning cycles.
• Delivered forecasts to an Angular dashboard with zone/facility filters used by hospital administrators for data-
driven capacity planning.
ASSOCIATE DATA ENGINEER Aptologics Private Limited
Ahmedabad, Gujarat, India August 2023 – August 2024
• Built and owned end-to-end Databricks + dbt ETL pipelines across 6–7 client engagements, from source ingestion
through dimensionally modeled BI-ready outputs.
• Authored 20+ production dbt models with documented data lineage and data-quality tests (uniqueness, not-null,
schema-consistency) enforcing enterprise data standards.
• Integrated Salesforce Cloud and multiple source systems into unified datasets; diagnosed and remediated a silent
duplication defect from an upstream schema change.
• Delivered 8–12 Tableau and Power BI dashboards on dbt semantic layers, enriching metadata to formalize 3
variants of “qualified pipeline” for stakeholder governance.
ASSOCIATE SYSTEM ENGINEER Tata Consultancy Services
Ahmedabad, Gujarat, India June 2022 – August 2023
• Improved page response times 25–30% via SQL optimization on production MySQL serving 2.9M citizens,
eliminating redundant DB calls and adding HikariCP pooling.
• Rewrote an N+1 query pattern (400+ queries) into a single joined query with composite index, cutting admin page
load from 4–5s to under 1s.
• Owned Java/Spring Boot integration of YOTI identity verification with audit logging for regulatory com-
pliance, plus retry, timeout, and fallback resilience.
• Progressed from reviewee to reviewing most backend commits within 5 months; coordinated bi-weekly Agile releases
with the onsite Scotland team.
PROJECTS
DALScooter: Serverless Platform with Warehouse Analytics | AWS Lambda, DynamoDB, BigQuery, Terraform,
Python June – August 2025
• Modeled data layer across 5 DynamoDB tables with indexes tuned to access patterns; applied least-privilege IAM
scoping per Lambda for enterprise-grade access control.
• Designed scheduled EventBridge → Lambda → S3 → BigQuery export pipeline (ADF-analogous orchestra-
tion) isolating warehouse analytics from OLTP workload.
• Provisioned full AWS footprint with Terraform across 10 services, enabling reproducible infrastructure and au-
ditable change management via version control.
Recipe Suggestion AI: RAG Pipeline on AWS | SageMaker, OpenSearch, S3, Flask, Python June 2025
• Built end-to-end RAG pipeline — embedded recipe corpus via SageMaker endpoint and indexed vectors in
OpenSearch k-NN for similarity retrieval.
• Developed Flask retrieval tier performing query embedding, top-5 k-NN retrieval, and prompt stitching to a second
SageMaker generation endpoint.
• Scoped OpenSearch domain to backend EC2 role only, enforcing least-privilege access controls consistent with
enterprise data governance principles.
Serverless Web Archiver | AWS Lambda, S3, DynamoDB, EventBridge, CloudFormation July 2025
• Built Python archiving Lambda with Requests + BeautifulSoup4 to fetch URLs, walk CSS/JS dependencies, and
preserve versioned snapshots with timestamped S3 keys for lineage.
• Designed hybrid orchestration — on-submit archive plus EventBridge scheduled re-archiving — with SNS
notifications capturing completion and failure provenance.
• Backed metadata across 2 DynamoDB tables (composite-key schema) and 6 Lambdas, each with tightly scoped
IAM, supporting auditable lifecycle tracking.
EDUCATION
Dalhousie University Master of Applied Computer Science (Sep 2024 – Present)
Halifax, NS, Canada
Gujarat Technological University B.S. Computer Science and Engineering (Jun 2018 – May 2022)
Gujarat, India