0% found this document useful (0 votes)
23 views3 pages

Aggrify: ETL Management in Databricks

The candidate has extensive experience developing ETL pipelines to ingest data from various sources and process it using big data tools like Hadoop, Spark, Hive and Python. They have worked on multiple projects involving building data pipelines, migrating existing code to optimize performance, and automating workflows. The candidate is proficient in developing data processing logic, scheduling jobs, and working with tools for code collaboration and project management.

Uploaded by

aniruddha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
23 views3 pages

Aggrify: ETL Management in Databricks

The candidate has extensive experience developing ETL pipelines to ingest data from various sources and process it using big data tools like Hadoop, Spark, Hive and Python. They have worked on multiple projects involving building data pipelines, migrating existing code to optimize performance, and automating workflows. The candidate is proficient in developing data processing logic, scheduling jobs, and working with tools for code collaboration and project management.

Uploaded by

aniruddha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Technical Skills:

Java/J2EE Java, J2EE, Struts, Liferay, Apache SOLR


Big Data tools Horton Works, Cloudera, Hadoop, HDFS, Map Reduce, Scala, Pig, Sqoop, Flume, Pig,
Hive, Impala, PySpark, Apache Nifi, Oozie, H2o
Cloud Technologies AWS, Microsoft AZURE, Data Bricks, Open-Source Delta Lake
Databases Oracle, PostgreSQL, DB2, SQL Server & MongoDB
Operating Systems Windows, UNIX
Scripting & HTML Python, JavaScript, Shell Scripting & HTML
IDE IBM RAD 7.0, IntelliJ, Eclipse 3.1, and Net Beans
Code Collaboration tools Bit Bucket, GITHUB, SVN, JIRA, Screwdriver, AZURE DevOps
Academics:
Bachelor of Technology from J.N.T University, Hyderabad.
Certifications:
 Microsoft Certified Azure Data Engineer Associate.
Description:
The Purpose of Aggrify is to track and manage ETL jobs that are run daily throughout Johnson and Johnson. The Aggrify application
has several reports that slice and present the data in different ways to help analyze the transformation of data across the company.

Business Deliverable - Name of the application as defined in the AGGRIFY database. A single Business Deliverable can have
multiple SLT IDs.  This is to allow the application team to group their tasks and batches under separate but unique SLT IDs. 
However, to allow grouping on the report, the teams can use the same business deliverable name for all their SLT IDS.

Environment: AWS, EMR, EC2, Hadoop 2.7.3, Spark3.3, Python 3.8, Open-Source Delta Lake, Kafka, Oracle, VS code.

Description: It is a digital platform that provides health plans to have a better understanding of their population, best care, and
reduced cost. It is a simple-to-use product used to analyze, monitor, intervene, and improve quality of care management.

Roles & Responsibilities:


 Design and implement ETL pipelines to consume data from data sources.
 Extensively worked on Migration of existing application from Cloudera to Azure Platform.
 Working with Databricks notebooks using Databricks utilities, magic commands etc.
 Working with Databricks Tables, Databricks File System (DBFS) etc.
 Developed code in Azure Data Bricks for the curation of the source data.
 Automation of the pipelines in Azure.

Project Title: Rx Surveillance Tool

Description: The Rx Surveillance project monitors invoice claims to ensure they are correctly processed based on the set of criteria or
rules, thousands of claims everyday are identified as outliers. Businesses need to verify that each claim has been adjudicated properly.
This tool works as a single repository of all the outliers claims and allows users to filter and writeback comments. Writeback allows
users to add comments to a claim or mass update to multiple claims. Comment updates are then saved in real time to a database.

Role and Responsibilities:


 Design and implement ETL pipelines to consume data from data sources.
 Developed spark SQL code using Scala to migrate existing Impala code to enable faster and in memory computing.
 Developed scripts to load and process the data using Hive QL
 Worked on using Jenkins, an open- source automated build platform designed for CI/CD i.e., tests pull requests, builds the
merged commits, and deploying the code to respective non-prod and prod environments.
 Customized scheduling using AutoSys to run complete end to end flow of self-contained Model, which have a functionality
of integrated Spark functionality.
 Production Support
Environment: Microsoft AZURE, Cloudera Hadoop 6.3.3, Impala 3.2.0, Hive 2.1.1, Oozie, Spark 2.4, Scala 2.11.2, GIT HUB, JIRA,
Jenkins, IntelliJ, HUE, AutoSys, Data Bricks.

Description: NSP (Network Service Personalization) is one of the critical system applications in the Network Personalization program
that creates an interface that is specific to the needs of the Network Repair Bureau (NRB) technicians and ties into the functions they
perform day-to-day.
NSP leverages scoring data produced by Network Data Lake (NDL a Hadoop Platform) based on specific Businesses identified Key
Performance Indicators (KPIs) and presents all the customer experience in form of Interactive Grafana dashboard. NSP also combines
NRB ticket information, Provisioning Data, and record details in a single pane of glass used for network outage troubleshooting. It
includes automation of process streaming and removing unnecessary troubleshooting steps and downstream communication with other
applications to prevent redundant tickets from reaching the NRB.
The main objective of this project is to achieve better control over automation of Tickets/Customer Experience.

Role and Responsibilities:


 Design and implement ETL pipelines to consume data from data sources.
 Developed spark SQL code using Scala to migrate existing pig scripts to enable faster and in memory computing.
 Developed spark code to load the data from Hive tables and csv files and publish the data to Pulsar broker using batch and
stream processing.
 Developed scripts to load and process the data using Hive QL, Pig.
 Worked with Verizon Data science team to provide and integrate the KPIs scoring data into Spark code for scoring.
 Worked on using Screwdriver pipeline, an open-source automated build platform designed for CI/CD i.e., tests pull requests,
builds the merged commits, and deploying the code to respective non-prod and prod environments.
 Customized scheduling in Oozie to run complete end to end flow of self-contained Model, which have a functionality of
integrated Hive, Spark, and Pig jobs.
 Developed UDF in spark Scala for custom functionality.

Environment: Hadoop 2.8, HDFS, Pig 0.14.0, Hive 1.2, Oozie, Spark 2.4, Scala 2.11, GIT HUB, JIRA, Screwdriver, IntelliJ, HUE,
Jenkins

Description: Aetna is a US based health care company, which sells traditional, and consumer directed health care insurance plans and
related services, such as medical, pharmaceutical, dental, behavioral health, long-term care, and disability plans. On average Aetna
receives 1 million claims each day. The sheer number of providers, members and plan types makes the pricing of these claims
incredibly complex. Through misinterpretation of provider contracts and human errors a small number of claims are paid improperly.
As part of the Data Science team all the data from the critical domains like Aetna Medicare, Traditional Group membership, Member,
Plan, Claim, Provider are migrated to Hadoop Environment. All the demographic information is moved from MySQL to Hadoop and
the analysis is done on the data. Also, the claims data will be moved from MQ to Hadoop and after processing the claims sending the
response back to MQ and build history to track all the changes corresponding to the processed claims.

Role and Responsibilities:


 Design and implement data pipelines to consume data from heterogeneous data sources and build an integrated health
Insurance Claims view of data. Use Hortonworks data platform, which consists of Red Hat Linux edge nodes with
availability of Hadoop Distributed File System (HDFS). Data processing and storage is done across a 1000 node cluster.
 Developed scripts to load and process the data in Hive, Pig
 Worked on performance tuning of Hive and Pig queries to improve data processing and retrieving
 Worked closely with data scientists to provide the required data for building the model features.
 Created multiple Hive tables, implemented partitioning, dynamic partitioning in Hive for efficient data access.
 Developed scripts in scheduling the jobs using Zeke Framework.
 Worked on migration of pig scripts to Pyspark code to enable faster and in memory computing. Perform ad hoc analytics on
large/diverse data using PySpark.
 Customized scheduling in Oozie to run complete end to end flow of integrated Hive and Pig jobs.
 Extensively used GIT as a code repository for managing day agile project development process and to keep track of the
issues and blockers.
 Sharing the Production Outcome Analysis report from different jobs to the end users on a weekly basis.
 Developed UDF in Pyspark for the custom functionalities.

Environment: Hadoop 2.7, HDFS, Pig 0.14.0, Hive 0.13.0, Sqoop, Flume, Apache NIFI, Oozie, GIT, JIRA, Pyspark, H2o, Netezza,
MY SQL, Aginity.

Description: “iTunes OPS Reporting” is a near real time data warehouse and reporting solution for iTunes Online Store and acts as a
reporting system for external and operational reporting needs. It also publishes data to downstream systems like Piano (for Label
Reporting) and ICA (for Business Objects Reporting, Campaign List Pull and Analytics). De-normalized data is used for publishing
various reports to the users. In addition, this project caters to the need of the ITS (iTunes Store) Business user groups. A lot of
complex analytical expertise is required which involves a lot of domain knowledge, detailed understanding of the features of iTunes,
its data flow and measuring the accuracy of the system in place.

Role and Responsibilities


 Worked on Analyzing the Tera Data Procedures.
 Worked on Developing the Graffle Design Documents for the Teradata Procedures in Hive
 Creation of design document for implementing HQL for the same.
 Developed HQL scripts for creating the tables and populating the data.
 Worked on testing the Map Side Joins for Performance.
 Developed Oozie scripts for automation.
 Production Support.

Environment: Hadoop 2.x7, HDFS, Hive 0.13.0, Oozie, GIT, JIRA, Java, Teradata

Common questions

Powered by AI

Apache NiFi is utilized for data flow management within ETL workflows, enabling the automation of data ingestion, transformation, and routing across various systems. In the context of the Hadoop environment described, NiFi's role is to facilitate seamless data integration between heterogeneous systems by managing complex data flow operations and ensuring that the data pipeline is both dynamic and scalable. It helps streamline the movement of large volumes of data, ensuring reliability and efficiency within the organization's data infrastructure .

Data lakes play a crucial role in integrating healthcare data across heterogeneous data sources by providing a centralized repository that can store structured, semi-structured, and unstructured data at any scale. In the Aetna project, the data lake facilitates the aggregation of diverse data types—from claims to member information—enabling comprehensive analysis and closer alignment with business needs. This approach supports real-time analytics and data-driven decision-making, improving efficiency and accuracy in healthcare operations. Moreover, data lakes' capability to handle large and diverse datasets makes them ideal for integrating vast amounts of healthcare data, enhancing interoperability and insight generation .

Jenkins aids in the deployment of ETL pipelines by automating the continuous integration and continuous deployment (CI/CD) processes, ensuring there are consistent and reliable code builds and deployments. It facilitates automated testing of code changes, reducing the chances of integration issues and enabling faster feedback loops for developers. Jenkins manages the deployment lifecycle, from building and testing to deploying the ETL code into various non-production and production environments. This continuous automation reduces manual intervention, accelerates release cycles, and enhances the overall reliability and quality of ETL deployments .

Customizing scheduling with Oozie contributes to the efficiency of ETL processes by automating and orchestrating complex workflows that integrate multiple Hadoop ecosystem components such as Hive, Spark, and Pig. Oozie's ability to define job dependencies ensures that tasks are executed in the correct sequence, optimizing resource utilization and minimizing idle times between jobs. It also supports time-based and data-triggered workflows, allowing for flexible scheduling that aligns with data arrival patterns and processing SLAs. This flexibility in workflow management enhances operational efficiency in distributed data environments, reducing latency and improving throughput .

The diversification of programming languages such as Python, Scala, and Java within ETL processes provides multiple benefits to big data projects. Each language offers unique strengths: Python is known for its simplicity and extensive libraries that support data science and machine learning tasks, Scala integrates seamlessly with Apache Spark for efficient in-memory data processing, and Java brings robustness, performance, and extensive ecosystem support. This diversification allows project teams to select the best tool for specific tasks, improving flexibility and capability to optimize performance, reduce development time, and leverage existing skill sets within the team .

The integration of Azure Data Bricks and Apache Spark significantly enhances data processing capabilities by leveraging Spark's fast, in-memory processing capabilities alongside Azure's scalable cloud infrastructure. This setup allows for efficient data transformations and analytics on large datasets. Azure Data Bricks automates cluster management, reducing the operational complexity, and integrating with Azure's ecosystem provides seamless connectivity with other Azure services. This integration supports advanced ETL tasks by allowing developers to use Spark’s rich set of libraries and APIs, thereby optimizing the performance of ETL pipelines in cloud environments .

Migrating existing applications from Cloudera to Azure Platform presents several challenges, including data compatibility issues, differences in platform services, and the need to refactor applications to leverage Azure-specific features. Addressing these challenges involves thorough planning and assessment to ensure seamless data transformation and compatibility. Organizations should adopt a staged migration approach, starting with less critical components, to mitigate risks. Additionally, leveraging Azure's service equivalent to Cloudera's tools can ease integration and minimize downtime. The use of Azure Data Bricks and other supportive tools must be optimized to handle data processing tasks traditionally managed by Cloudera .

Deploying a unified data platform helps manage large-scale data operations in a network service environment by centralizing data processing and storage, thereby ensuring consistency and reliability. In the NSP project, the unified platform integrates disparate data sources to provide a single view of operational metrics and customer experiences. This facilitates efficient troubleshooting and performance monitoring, while the centralized approach enhances data security and management control. Additionally, the platform supports automation processes and reduces redundancies, resulting in more streamlined operations and better strategic decision-making based on comprehensive data analytics .

Using HiveQL and UDFs in Hive significantly enhances the efficiency of data processing in a Hadoop environment. HiveQL provides a SQL-like interface for querying and manipulating data, which is user-friendly and allows for complex data operations with relatively simple syntax, enhancing developer productivity. UDFs empower developers to implement custom code for operations not natively supported by HiveQL, providing greater flexibility in handling data transformations. This combination allows for more expressive data pipelines tailored to specific business needs, leading to better performance optimization especially in scenarios involving complex analytical workflows within large datasets .

Utilizing Spark SQL code with Scala improves ETL operation performance by taking advantage of Spark's in-memory computation capabilities, which are faster compared to disk-based operations in Impala. Spark SQL, running on Spark engines, allows for parallel processing of datasets, thus enabling high-speed analytical queries and transformations. Scala's tight integration with Spark further enhances performance through more efficient memory use and execution. This results in reduced processing time for large datasets when compared to Impala scripts, which are generally less efficient in handling complex transformations and large-scale data processing tasks .

You might also like