DPR401
Look under the hood of AWS DeepRacer
with Amazon SageMaker
Don Barber
Sr Solutions Architect Manager
AWS
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Agenda
Real-world examples of reinforcement learning (RL)
Under the hood recap
Train and evaluate AWS DeepRacer using Amazon SageMaker notebooks
How to customize your training
Dive into making custom tracks!
Workshop cleanup
More resources
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Set up the environment
To get started, log into your AWS account console, then go to
[Link]
Skip to “Using your own AWS Account.”
Follow step-by-step instructions on the AWS DeepRacer 400-level workshop
materials to train your models directly in Amazon SageMaker and Amazon
RoboMaker.
Follow the directions through “Run Notebook” and pause there.
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Real-world use cases
for applying
reinforcement learning
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Recap – AWS DeepRacer learns to drive fast
AWS
Action = left RoboMaker Reward = lap
completion
RL agent
AWS DeepRacer Evo
Agent takes actions and gets rewards for
completing a lap without crashing
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Capcom builds fun games faster with managed
services on AWS
Action = Unity Reward = level
move forward completion
RL agent (gamer)
[Link]
examples/tree/master/reinforcement_learning/rl_unity_ray
Agent plays various levels of computer games and
evaluates new games prior to production
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
MathWorks simulation environments enable
industrial workflows
MathWorks Simulink
Reward =
State
Action =
Deploy trained model Action
Simulation/real
application
energy
SageMaker
pitch
State
Amazon EC2
SageMaker Action
Reward simulation hosts
training job
Temp
endpoint
RL agent (SCADA)
Client simulation
application
drives the
training
Customer VPC
Agent learns to control the wind turbines to maximize the
energy generation
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
More ideas
• Datacenter cooling
• AI-powered stock buying/selling
• Self-learning bots
• Personalized recommendations
• Path-finding for cleaning robots
• Teaching cars to drive safer
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
AWS DeepRacer
under the hood
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Under the hood
• 4WD 1:18-scale car
• Intel Atom processor
• Intel distribution of OpenVINO toolkit
• Stereo camera (4MP)
• 360-degree 12-meters scanning radius LiDAR Sensor
• System memory – 4 GB RAM
• 802.11ac Wi-Fi
• Ubuntu 20.04 LTS
• ROS 2 Foxy on device
ROS Melodic in simulation
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Iteration and convergence
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
How does learning happen?
State
Reward
Agent Environment Model
Action
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Rewards are maximized by a neural network
Goal – Maximum
cumulative rewards
0.4 ± 𝛿 0.3 ± 𝛿
New
weights
New
weights
J(q)
Gather episode Calculate the gradient of Update the policy for
using current policy estimated cumulative reward gradient ascent
Neural network learns by maximizing discounted
cumulative rewards
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Reward Function Editor
Code editor: Python 3 syntax
Three example reward functions
Code validation via AWS Lambda
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
AWS DeepRacer with simulator architecture
AWS Cloud
AWS DeepRacer
VPC
Amazon
SageMaker
Amazon
CloudWatch
Amazon
AWS S3
DeepRacer
VPC
Amazon
AWS
Kinesis Video
RoboMaker
Streams
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Model evaluation architecture
AWS Cloud
AWS DeepRacer
VPC
AWS RoboMaker
Simulation
Amazon
Kinesis Video
Streams
AWS DeepRacer
Amazon
Amazon S3 CloudWatch
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
AWS DeepRacer vehicle software architecture
ROS nodes
Model ROS messages
optimizer
Optimized
Model model Stored file
1D
LiDAR LiDAR
node data
Sensor Combined Inference Inference Navigation Autonomous
fusion sensor data engine results node drive
node
Control
Media Video node
engine M-JPEG
Web Web Manual
server server drive
Sensors video publisher Servo & motor
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Train and evaluate your model in
Amazon SageMaker
Continue with the workshop through “Evaluate your models”
Pause at Workshop Checkpoint #4
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Customize your training
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
What does SageMaker and RoboMaker facilitate
beyond the AWS DeepRacer console?
• Extend your reward function even further – other modules or languages
• Modify neural network architectures beyond shallow neural networks
• Multiple simulation applications to collect experience in shorter time
• Train and evaluate models in custom race tracks
• Add new sensors beyond image-based cameras and LiDAR
• Change your training instance
• Run DeepRacer in other regions
• Customize racing videos
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Modify your reward function
Browse to src/artifacts/rewards
Ideas:
• Add your own Python modules
• Call out to other languages
• Store and retrieve external state
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Modify the Training Algorithm
The default training algorithm is Proximal Policy Optimization (PPO)
Ideas:
• Try the built-in Soft Actor Critic (SAC) by modifying the model_metadata.json
file
• Modify the src/markov files to add new reinforcement learning algorithms, such
as TRPO, DDPG, A3C, NAF, ACER, or ACKTR
• Change training to stop on progress rather than time
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Modifying neural
network architecture
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Quick ML primer – One node network
m
3 x
x=3m
Input Network Output parameter
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
More nodes and connections = more math
And more sophisticated decision making!
3 a e
m x
f
g
5 b
h
x=m(3ae+5bf+7cg+11dh)
i
7 c
j
k
y=q(3ai+5bj+7ck+11dl)
q y
11 d
l
Input Network Output
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
How does the neural network architecture look?
Image
2-D CNN image embedders
Slow, left
Camera
Slow, right
Sensor input Input embedders
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Middleware Action output
Adding new sensors
2-D CNN image embedders
Stereo image Left
camera
Slow, left
Slow, right
Right
camera
…
1-D dense layer
0, 0, 0, 0, 1, 1, 1, 0
Sectorized LiDAR
data
Sensor input Input embedders
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Middleware Action output
Now with a deeper network More Nodes and Connections = more
parameters = more math = more time!
2-D CNN image embedders
Stereo image Left
camera
Slow, left
Slow, right
Right
camera
…
1-D dense layer
0, 0, 0, 0, 1, 1, 1, 0
Sectorized LiDAR
data
Sensor input Input embedders
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Middleware Action output
Deeper networks
Pros Cons
Better approximation of the reward More math, hence, slower
function leads to higher rewards convergence over time (Hint: Consider
and faster convergence in number a GPU instance)
of episodes
Approximates reward function with
discontinuities with lower error May become too specialized and not
(on-track/off-track) vs. gradually generalize to real world as well as the
decaying penalties shallower neural networks
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Shallow versus deep learning
Shallow Deep 80%
60%
60 % 60 %
Still 60% progress after 80% progress after
400 episodes 400 episodes – but takes
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
longer!
How to gather data faster?
State
Action
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Multiple Rollouts!
State State
Action Action
num_simulation_workers = 4 State
State
Action Action
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Customize the world
Explore
src/deepracer_simulation_environment /
share/deepracer_simulation_environment
• Waypoints in routes/
• 3d models in meshes/
File format is COLLADA
Customize for additional textures, lighting,
etc!
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Custom Tracks
Run ./make_csv.py
Then ./[Link]
Then ./new_track.py [Link]
Use an image editor to create a Custom_track.png
Zip up the png, npy, and dae files, along with the textures folder.
Upload to the notebook.
Extract and copy/modify xml files.
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Logs and videos
Explore your S3 bucket to find log files and output videos
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Clean up
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Clean up the environment
Deprovision the resources in the AWS account so those resources do not continue
to be charged.
Follow the “Workshop clean-up” instructions in the workshop
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Thank you!
Don Barber
donbarb@[Link]
Please complete survey at
[Link]
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.