Understanding Databricks/Spark
Performance Tuning
Lesson 02: Easiest Fix – Optimizing the Hardware
by Bryan Cafferky from my YouTube channel
Where We're Going?
➢ Why Compute Resources Matter
➢ Databricks Workspace – Standard or Premium
➢ Databricks/Apache Spark Cluster Architecture
➢ Hardware Under the Cluster Architecture
➢ Optimizing the Cluster Configuration Step By Step
➢ Shuffles & Spills
➢ Wrap Up
Why Compute Resources Matter
➢ Never Talked About in Depth
➢ Cost vs. Performance Trade Off
➢ Single Easiest Thing You Can Change to Improve Performance
➢ There are Interdependencies Between Performance Optimization &
the Compute Selection/Options
Databricks Workspace
Databricks Workspace
Premium vs. Standard
➢ Unity Catalog Requires Premium
➢ Delta Live Tables Requires Premium
➢ Photon Does NOT Require Premium
Unity Catalog
Premium Required for Unity Catalog
[Link]
Delta Live Tables (DLT)
Premium Required for DLT
[Link]
Cluster Architecture
[Link]
Cluster Architecture
External Storage
Worker Nodes
RAM RAM RAM RAM RAM RAM RAM RAM RAM
Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core Core
Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk
Do Count Do Count Do Count Do Count Do Count Do Count Do Count Do Count Do Count
Send Back Result Partition Data by City
and Copy One City to Each Node Executor
Constraints Distributing the Data Over the Cluster
Driver Node
Core Core
➢ Hardware/Resources Disk Disk
➢ Software (Spark/DBR) External
Storage
➢ Environment Configurations SELECT City, Count(*) FROM PhoneBook
Group By City
➢ Your Code/Application Order By City
➢ Data Source & Format
➢ Data Distribution
Phone Book
Hardware Focus
5 External Storage External Storage External Storage
1 Worker Worker
Network Traffic
6
CORE CORE CORE CORE
Memory Memory Memory Memory
2 Storage | Working Storage | Working
Storage | Working Storage | Working
Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk Disk
3
5
1 Worker Size/CPU Type Driver
2 Memory Usage/Allocation External
Storage
3 Local Disk Speed
CORE CORE
4 Driver Configuration
Memory Memory
5 External Disk 4
Storage | Working Storage | Working
6 Number of Nodes Disk Disk Disk Disk Disk Disk
Hardware – CPU & Memory
➢ CPU/GPU Type
➢ Workload Type
➢ Number of Cores
➢ Core = Executor = Task = Partition
➢ # of Cores = Degree of Parallelism
➢ Memory
➢ Enough to Support Your Workload
➢ Avoid Spills
➢ Split Between Working & Storage Memory
Hardware – Input/Output
➢ Network
➢ Latency (Local vs. Cross Region vs. On Prem)
➢ Bandwidth
➢ External Resource Latency (example: Cosmos DB vs Azure SQL)
➢ Storage (Local, Remote (External))
➢ Type/Speed (HDD, SSD)
➢ Locality
➢ I/O
➢ Spills
➢ Waiting for Disk Writes
➢ Databricks Runtime Version (Effects Features)
Optimizing the Cluster Configuration
Cluster – Configuration Settings
We'll Walk Through
Filling Out This Screen.
[Link]
[Link]
Cluster – Configuration Best Practices
Use Case Notes Recommendation
Analysis Interactive Development ➢ Single Node with High Memory and Cores
➢ Likely require reading the same data
repeatedly, so recommended node types
are storage optimized with disk cache
enabled.
Basic Batch ETL No Wide Transformations Compute Optimized
Complex ETL Has Wide Transformations Compute Optimized with Less Nodes
ML Training Experimentation/Dev Single Node Type with High Memory and
Cores
ML Training Production Minimal Worker Nodes Storage Optimized with Disk Caching Enabled
or
GPU (lacks disk caching)
[Link]
Cluster Options
Use Case Notes Recommendation
Spot Instances Saves money by uses available
capacity.
[Link]
us/azure/virtual-machines/spot-vms
Serverless Compute Near instant Cluster. Like having a set of VMs on standby. When
Unity Catalog must be enabled. you need them, they are allocated to your
work almost instantly.
[Link]
us/azure/databricks/compute/serverless
[Link]
us/azure/databricks/release-
notes/serverless#limitations
[Link]
features/serverless-security
Photon Vastly faster processing.
Cluster – Access Mode
Single User for Credential Pass Through.
Credential Passthrough
Automatically Passes Your
Credentials Through to Backend
Resources like ADLS.
Cluster – Runtime and Node Types
Replaces most of the Scala Node Execution code.
LTS = Long Term Support
Photon – a New Execution Engine on Databricks
Photon is Only Available on Databricks
[Link]
Cluster – Delta cache accelerated
Improves Reading of Parquet and
Delta Files.
Renamed to disk cache.
[Link]
Cluster – Databricks Runtime
Standard Machine Learning
Databricks Runtime
Cluster – Worker Type
Worker Type
Worker Type
Larger
Workers
means less
nodes
required.
Cluster – Driver Type
Driver Type
Driver Type Make this
larger if you
want to do a
lot of work on
the driver, i.e.,
collecting
data, local
processing.
Cluster – Advanced Settings
Modify Spark
Configuration Settings
to Improve
performance
Cluster – Shuffles & Spills
Getting to the Spark UI from your Notebook
Click for the Spark UI
Limiting Memory to Force a Spill
Getting to the Spark UI from your Notebook
Generated Spark Job & Stages.
Click on View to go to the Spark UI.
Cluster – Shuffle & Spills
Shuffle Read/Writes
Cluster – Shuffle & Spills
Spills Degrade Performance.
Spill (Memory) and Spill (Disk)
only show when > 0
Wrapping Up
➢ Why Compute Resources Matter
➢ Databricks
Thank You! – Standard or Premium
Workspace
➢ Databricks/Apache Spark Cluster Architecture
➢ Hardware Under the Cluster Architecture
➢ Optimizing the Cluster Configuration Step By Step
➢ Shuffles & Spills