0% found this document useful (0 votes)
6 views30 pages

Optimizing Databricks Cluster Performance

This document provides an overview of performance tuning for Databricks and Apache Spark, focusing on optimizing hardware configurations. Key topics include the importance of compute resources, differences between standard and premium Databricks workspaces, and best practices for cluster configuration. Additionally, it discusses the impact of hardware choices on performance, including CPU, memory, and storage considerations.

Uploaded by

LearnMSBI First
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views30 pages

Optimizing Databricks Cluster Performance

This document provides an overview of performance tuning for Databricks and Apache Spark, focusing on optimizing hardware configurations. Key topics include the importance of compute resources, differences between standard and premium Databricks workspaces, and best practices for cluster configuration. Additionally, it discusses the impact of hardware choices on performance, including CPU, memory, and storage considerations.

Uploaded by

LearnMSBI First
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Understanding Databricks/Spark

Performance Tuning
Lesson 02: Easiest Fix – Optimizing the Hardware
by Bryan Cafferky from my YouTube channel
Where We're Going?
➢ Why Compute Resources Matter
➢ Databricks Workspace – Standard or Premium
➢ Databricks/Apache Spark Cluster Architecture
➢ Hardware Under the Cluster Architecture
➢ Optimizing the Cluster Configuration Step By Step
➢ Shuffles & Spills
➢ Wrap Up
Why Compute Resources Matter
➢ Never Talked About in Depth
➢ Cost vs. Performance Trade Off
➢ Single Easiest Thing You Can Change to Improve Performance
➢ There are Interdependencies Between Performance Optimization &
the Compute Selection/Options
Databricks Workspace
Databricks Workspace

Premium vs. Standard


➢ Unity Catalog Requires Premium
➢ Delta Live Tables Requires Premium
➢ Photon Does NOT Require Premium
Unity Catalog

Premium Required for Unity Catalog

[Link]
Delta Live Tables (DLT)

Premium Required for DLT

[Link]
Cluster Architecture

[Link]
Cluster Architecture
External Storage
Worker Nodes

RAM RAM RAM RAM RAM RAM RAM RAM RAM

Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core
Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk
Do Count Do Count Do Count Do Count Do Count Do Count Do Count Do Count Do Count

Send Back Result Partition Data by City


and Copy One City to Each Node Executor

Constraints Distributing the Data Over the Cluster


Driver Node
Core Core
➢ Hardware/Resources Disk Disk
➢ Software (Spark/DBR) External
Storage
➢ Environment Configurations SELECT City, Count(*) FROM PhoneBook
Group By City
➢ Your Code/Application Order By City
➢ Data Source & Format
➢ Data Distribution

Phone Book
Hardware Focus
5 External Storage External Storage External Storage

1 Worker Worker
Network Traffic
6
CORE CORE CORE CORE

Memory Memory Memory Memory


2 Storage | Working Storage | Working
Storage | Working Storage | Working
Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk
3

5
1 Worker Size/CPU Type Driver
2 Memory Usage/Allocation External
Storage
3 Local Disk Speed
CORE CORE
4 Driver Configuration
Memory Memory
5 External Disk 4
Storage | Working Storage | Working
6 Number of Nodes Disk Disk Disk Disk Disk Disk
Hardware – CPU & Memory
➢ CPU/GPU Type
➢ Workload Type

➢ Number of Cores
➢ Core = Executor = Task = Partition
➢ # of Cores = Degree of Parallelism

➢ Memory
➢ Enough to Support Your Workload
➢ Avoid Spills
➢ Split Between Working & Storage Memory
Hardware – Input/Output
➢ Network
➢ Latency (Local vs. Cross Region vs. On Prem)
➢ Bandwidth
➢ External Resource Latency (example: Cosmos DB vs Azure SQL)

➢ Storage (Local, Remote (External))


➢ Type/Speed (HDD, SSD)
➢ Locality

➢ I/O
➢ Spills
➢ Waiting for Disk Writes

➢ Databricks Runtime Version (Effects Features)


Optimizing the Cluster Configuration
Cluster – Configuration Settings

We'll Walk Through


Filling Out This Screen.

[Link]
[Link]
Cluster – Configuration Best Practices
Use Case Notes Recommendation
Analysis Interactive Development ➢ Single Node with High Memory and Cores
➢ Likely require reading the same data
repeatedly, so recommended node types
are storage optimized with disk cache
enabled.
Basic Batch ETL No Wide Transformations Compute Optimized
Complex ETL Has Wide Transformations Compute Optimized with Less Nodes
ML Training Experimentation/Dev Single Node Type with High Memory and
Cores
ML Training Production Minimal Worker Nodes Storage Optimized with Disk Caching Enabled
or
GPU (lacks disk caching)

[Link]
Cluster Options
Use Case Notes Recommendation
Spot Instances Saves money by uses available
capacity.
[Link]
us/azure/virtual-machines/spot-vms
Serverless Compute Near instant Cluster. Like having a set of VMs on standby. When
Unity Catalog must be enabled. you need them, they are allocated to your
work almost instantly.
[Link]
us/azure/databricks/compute/serverless

[Link]
us/azure/databricks/release-
notes/serverless#limitations

[Link]
features/serverless-security
Photon Vastly faster processing.
Cluster – Access Mode

Single User for Credential Pass Through.

Credential Passthrough
Automatically Passes Your
Credentials Through to Backend
Resources like ADLS.
Cluster – Runtime and Node Types

Replaces most of the Scala Node Execution code.

LTS = Long Term Support


Photon – a New Execution Engine on Databricks

Photon is Only Available on Databricks

[Link]
Cluster – Delta cache accelerated

Improves Reading of Parquet and


Delta Files.

Renamed to disk cache.

[Link]
Cluster – Databricks Runtime
Standard Machine Learning

Databricks Runtime
Cluster – Worker Type
Worker Type

Worker Type
Larger
Workers
means less
nodes
required.
Cluster – Driver Type
Driver Type

Driver Type Make this


larger if you
want to do a
lot of work on
the driver, i.e.,
collecting
data, local
processing.
Cluster – Advanced Settings

Modify Spark
Configuration Settings
to Improve
performance
Cluster – Shuffles & Spills
Getting to the Spark UI from your Notebook

Click for the Spark UI

Limiting Memory to Force a Spill


Getting to the Spark UI from your Notebook

Generated Spark Job & Stages.


Click on View to go to the Spark UI.
Cluster – Shuffle & Spills

Shuffle Read/Writes
Cluster – Shuffle & Spills

Spills Degrade Performance.

Spill (Memory) and Spill (Disk)


only show when > 0
Wrapping Up
➢ Why Compute Resources Matter
➢ Databricks
Thank You! – Standard or Premium
Workspace
➢ Databricks/Apache Spark Cluster Architecture
➢ Hardware Under the Cluster Architecture
➢ Optimizing the Cluster Configuration Step By Step
➢ Shuffles & Spills

You might also like