Amir Sohail
910803270
mdamirsohail890@[Link]
Profile:
Azure Data Engineer having good Experience and understanding of distribution file systems of Big
Data and its Ecosystems (Azure Data bricks, Azure Data factory, Azure Data Lake).
Experience Summary:
Around 4+ years of strong experience in software development using Azure Cloud, SQL,
PySpark, Python technologies.
Hands on experience on Azure Data Factory V2.
Hands on experience on Creating ADF Pipeline for Full Load and Incremental Load
Hands on experience on Creating the linked services for the different services
Hands on experience on creating the datasets for the different sources.
Hands on experience on installing the self-hosted Integration runtime to migrate data from on
premise to cloud.
Experience on Spark, Sql, Python.
Processed flat files in various file formats CSV, Parquet and Json etc.
Created Azure Data Factory V2 Linked Services to connect to different sources like on premise
SQL server, blob storage and SQL Data warehouse.
Created the dynamic pipelines to copy multiple tables using dynamic parameters
Attending the spring retrospect meetings to review the sprint flaws and improvements.
Having good working experience of Storages like Azure Data Lake Storage and Blob Storage.
Developed various spark applications to perform ETL workloads on terabytes of data.
Familiar with Data Ingestion pipeline design and advanced data processing tools.
Experienced in version control and source code management tools like GIT
Experience with relational databases such as MySQL and SQL Server.
Technical Skills:
Cloud Technologies Azure Data Factory V2, Azure Data bricks, Azure Data Lake Storage
Languages & Scripting SQL
Database/ RDBMS MS-SQL server, MySQL
Operating Systems Windows
Version Control Azure GIT
Database Query Tools SSMS, Storage Explorer
Work Experience:
Currently working with DXC Technologies as Software Engineer from Dec 2018 to till Date.
Project Details
Project:
Project Title : Retail Store.
Client : John Lewis, US.
Role : Azure Data Engineer
Environment : Azure Data Factory V2, Azure Databricks, Azure Data Lake Storage
Responsibilities:
Attending Planning meeting to understand the tasks for the sprint.
Attending the DSM (Daily Standup Meetings) to update the task status.
Implemented ADF pipelines using ETL activities to load data from on premise source system to ADLS Gen2.
Created Linked Services for multiple source system (i.e.: SQL Server, ADLS Gen2 and Blob Storage)
Configured and implemented the Azure Data Factory Triggers and Scheduled the Pipelines.
Converted the csv files to parquet files using pyspark in azure databricks.
Mounted Azure data lake storage gen2 on azure databricks.
Performed data quality checks using pyspark in databricks notebooks.
Configured the logic apps to handle email notification to the end users and key shareholders with the help of
web services activity.
Created Azure Data Factory V2 Linked Services to connect to Azure Databricks.
Created the azure data factory V2 pipelines to invoke the Azure databricks notebooks.
Created the azure data factory V2 pipelines for daily, weekly and monthly full load.
Created triggers in azure data factory V2 to run pipelines on scheduled time.
Created the stored procedure to updated the audit logs of the ADF V2 copy activity.
Created the dynamic pipelines to copy multiple tables using dynamic parameters.
Attending the spring retrospect meetings to review the sprint flaws and improvements.