0% found this document useful (0 votes)
5 views38 pages

DPR401

The document provides an overview of AWS DeepRacer, detailing its architecture, training processes, and customization options using Amazon SageMaker and RoboMaker. It includes real-world applications of reinforcement learning, instructions for setting up the environment, and methods for modifying training algorithms and reward functions. Additionally, it covers the creation of custom tracks and emphasizes the importance of resource cleanup after the workshop.

Uploaded by

jcoronelcortes
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views38 pages

DPR401

The document provides an overview of AWS DeepRacer, detailing its architecture, training processes, and customization options using Amazon SageMaker and RoboMaker. It includes real-world applications of reinforcement learning, instructions for setting up the environment, and methods for modifying training algorithms and reward functions. Additionally, it covers the creation of custom tracks and emphasizes the importance of resource cleanup after the workshop.

Uploaded by

jcoronelcortes
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DPR401

Look under the hood of AWS DeepRacer


with Amazon SageMaker
Don Barber
Sr Solutions Architect Manager
AWS

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Agenda

Real-world examples of reinforcement learning (RL)

Under the hood recap

Train and evaluate AWS DeepRacer using Amazon SageMaker notebooks

How to customize your training

Dive into making custom tracks!

Workshop cleanup

More resources
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Set up the environment

To get started, log into your AWS account console, then go to


[Link]
Skip to “Using your own AWS Account.”

Follow step-by-step instructions on the AWS DeepRacer 400-level workshop


materials to train your models directly in Amazon SageMaker and Amazon
RoboMaker.

Follow the directions through “Run Notebook” and pause there.

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Real-world use cases
for applying
reinforcement learning

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Recap – AWS DeepRacer learns to drive fast
AWS
Action = left RoboMaker Reward = lap
completion

RL agent

AWS DeepRacer Evo

Agent takes actions and gets rewards for


completing a lap without crashing
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Capcom builds fun games faster with managed
services on AWS
Action = Unity Reward = level
move forward completion

RL agent (gamer)

[Link]
examples/tree/master/reinforcement_learning/rl_unity_ray

Agent plays various levels of computer games and


evaluates new games prior to production
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
MathWorks simulation environments enable
industrial workflows
MathWorks Simulink
Reward =
State

Action =
Deploy trained model Action
Simulation/real
application

energy
SageMaker

pitch
State
Amazon EC2
SageMaker Action
Reward simulation hosts
training job
Temp
endpoint

RL agent (SCADA)
Client simulation
application
drives the
training

Customer VPC

Agent learns to control the wind turbines to maximize the


energy generation
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
More ideas

• Datacenter cooling
• AI-powered stock buying/selling
• Self-learning bots
• Personalized recommendations
• Path-finding for cleaning robots
• Teaching cars to drive safer

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
AWS DeepRacer
under the hood

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Under the hood

• 4WD 1:18-scale car

• Intel Atom processor

• Intel distribution of OpenVINO toolkit

• Stereo camera (4MP)

• 360-degree 12-meters scanning radius LiDAR Sensor

• System memory – 4 GB RAM

• 802.11ac Wi-Fi

• Ubuntu 20.04 LTS

• ROS 2 Foxy on device


ROS Melodic in simulation

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Iteration and convergence

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
How does learning happen?

State

Reward

Agent Environment Model

Action

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Rewards are maximized by a neural network

Goal – Maximum
cumulative rewards
0.4 ± 𝛿 0.3 ± 𝛿

New
weights
New
weights
J(q)

Gather episode Calculate the gradient of Update the policy for


using current policy estimated cumulative reward gradient ascent

Neural network learns by maximizing discounted


cumulative rewards
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Reward Function Editor

Code editor: Python 3 syntax

Three example reward functions

Code validation via AWS Lambda

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
AWS DeepRacer with simulator architecture
AWS Cloud

AWS DeepRacer

VPC

Amazon
SageMaker

Amazon
CloudWatch
Amazon
AWS S3
DeepRacer
VPC

Amazon
AWS
Kinesis Video
RoboMaker
Streams

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Model evaluation architecture
AWS Cloud

AWS DeepRacer

VPC

AWS RoboMaker

Simulation

Amazon
Kinesis Video
Streams

AWS DeepRacer

Amazon
Amazon S3 CloudWatch

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
AWS DeepRacer vehicle software architecture
ROS nodes

Model ROS messages


optimizer

Optimized
Model model Stored file
1D
LiDAR LiDAR
node data
Sensor Combined Inference Inference Navigation Autonomous
fusion sensor data engine results node drive
node
Control
Media Video node
engine M-JPEG
Web Web Manual
server server drive
Sensors video publisher Servo & motor

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Train and evaluate your model in
Amazon SageMaker

Continue with the workshop through “Evaluate your models”

Pause at Workshop Checkpoint #4

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Customize your training

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
What does SageMaker and RoboMaker facilitate
beyond the AWS DeepRacer console?
• Extend your reward function even further – other modules or languages
• Modify neural network architectures beyond shallow neural networks
• Multiple simulation applications to collect experience in shorter time
• Train and evaluate models in custom race tracks
• Add new sensors beyond image-based cameras and LiDAR
• Change your training instance
• Run DeepRacer in other regions
• Customize racing videos

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Modify your reward function

Browse to src/artifacts/rewards

Ideas:
• Add your own Python modules
• Call out to other languages
• Store and retrieve external state

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Modify the Training Algorithm

The default training algorithm is Proximal Policy Optimization (PPO)


Ideas:
• Try the built-in Soft Actor Critic (SAC) by modifying the model_metadata.json
file
• Modify the src/markov files to add new reinforcement learning algorithms, such
as TRPO, DDPG, A3C, NAF, ACER, or ACKTR
• Change training to stop on progress rather than time

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Modifying neural
network architecture

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Quick ML primer – One node network

m
3 x
x=3m
Input Network Output parameter

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
More nodes and connections = more math
And more sophisticated decision making!
3 a e
m x
f
g
5 b
h
x=m(3ae+5bf+7cg+11dh)
i
7 c
j

k
y=q(3ai+5bj+7ck+11dl)
q y
11 d
l

Input Network Output

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
How does the neural network architecture look?

Image

2-D CNN image embedders


Slow, left
Camera
Slow, right

Sensor input Input embedders


© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Middleware Action output
Adding new sensors

2-D CNN image embedders

Stereo image Left


camera

Slow, left
Slow, right
Right
camera

1-D dense layer

0, 0, 0, 0, 1, 1, 1, 0
Sectorized LiDAR
data

Sensor input Input embedders


© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Middleware Action output
Now with a deeper network More Nodes and Connections = more
parameters = more math = more time!

2-D CNN image embedders

Stereo image Left


camera

Slow, left
Slow, right
Right
camera

1-D dense layer

0, 0, 0, 0, 1, 1, 1, 0
Sectorized LiDAR
data

Sensor input Input embedders


© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Middleware Action output
Deeper networks

Pros Cons
Better approximation of the reward More math, hence, slower
function leads to higher rewards convergence over time (Hint: Consider
and faster convergence in number a GPU instance)
of episodes

Approximates reward function with


discontinuities with lower error May become too specialized and not
(on-track/off-track) vs. gradually generalize to real world as well as the
decaying penalties shallower neural networks

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Shallow versus deep learning

Shallow Deep 80%


60%

60 % 60 %

Still 60% progress after 80% progress after


400 episodes 400 episodes – but takes
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
longer!
How to gather data faster?

State

Action

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Multiple Rollouts!
State State

Action Action
num_simulation_workers = 4 State
State

Action Action
© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Customize the world

Explore
src/deepracer_simulation_environment /
share/deepracer_simulation_environment

• Waypoints in routes/
• 3d models in meshes/

File format is COLLADA

Customize for additional textures, lighting,


etc!

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Custom Tracks

Run ./make_csv.py
Then ./[Link]
Then ./new_track.py [Link]
Use an image editor to create a Custom_track.png
Zip up the png, npy, and dae files, along with the textures folder.
Upload to the notebook.
Extract and copy/modify xml files.

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Logs and videos

Explore your S3 bucket to find log files and output videos

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Clean up

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Clean up the environment

Deprovision the resources in the AWS account so those resources do not continue
to be charged.

Follow the “Workshop clean-up” instructions in the workshop

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.
Thank you!
Don Barber
donbarb@[Link]

Please complete survey at


[Link]

© 2023, Amazon Web Services, Inc. or its affiliates. All rights reserved.

You might also like