0% found this document useful (0 votes)
8 views803 pages

Edge AI Engineering

The document is a comprehensive guide on Edge AI Engineering using Raspberry Pi, detailing its significance, advantages, and applications. It covers setup, hardware specifications, software installation, and image classification fundamentals, along with practical projects. The book is designed for individuals interested in learning about Edge AI and Raspberry Pi technology.

Uploaded by

Marcelo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views803 pages

Edge AI Engineering

The document is a comprehensive guide on Edge AI Engineering using Raspberry Pi, detailing its significance, advantages, and applications. It covers setup, hardware specifications, software installation, and image classification fundamentals, along with practical projects. The book is designed for individuals interested in learning about Edge AI and Raspberry Pi technology.

Uploaded by

Marcelo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2026

Edge AI Engineering
Hands-on with the Raspberry Pi

Marcelo Rovai

2026-05-24
Table of contents

Preface 3

Acknowledgments 5

Introduction 6
Edge AI Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
Why Edge AI Matters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
The Raspberry Pi Advantage . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
What You’ll Learn . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Who This Book Is For . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

About this Book 8


Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
Structure and Organization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9

Classification of AI Applications 11
Fixed Function AI vs. Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Fixed Function AI (Reactive) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Generative AI (Proactive) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
Summary Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
The Edge AI Advantage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13

Setup 15
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
Raspberry Pi Models (covered in this book) . . . . . . . . . . . . . . . . . . . . 17
Engineering Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
Hardware Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Raspberry Pi Zero 2W . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
Installing the Operating System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
The Operating System (OS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
Installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
Initial Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23

2
Remote Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
SSH Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
To shut down the Raspi via terminal: . . . . . . . . . . . . . . . . . . . . . . . . 26
Transfer Files between the Raspberry Pi and a computer . . . . . . . . . . . . . 26
Increasing Memory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
Increasing zram for the Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . 32
Increasing memory for the Rasp-Zero . . . . . . . . . . . . . . . . . . . . . . . . 33
Notes specific to Zero 2 W Trixie . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Installing a Camera . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Installing a Camera Module on the CSI port . . . . . . . . . . . . . . . . . . . 35
Installing a USB WebCam . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
Running the Raspi Desktop remotely . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
Updating and Installing Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
Model-Specific Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Raspberry Pi Zero (Raspi-Zero) . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Raspberry Pi 4 or 5 (Raspi-4 or Raspi-5) . . . . . . . . . . . . . . . . . . . . . . 52
Measuring Temperature and Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Check CPU Temperature . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
Understand Power Measurement on Pi 5 . . . . . . . . . . . . . . . . . . . . . . 53
Script: Average Temperature and Power . . . . . . . . . . . . . . . . . . . . . . 54
Run a measurement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
Interpreting the Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57

Image Classification Fundamentals 60


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
Applications in Real-World Scenarios . . . . . . . . . . . . . . . . . . . . . . . . 61
Advantages of Running Classification on Edge Devices like Raspberry Pi . . . . 61
Setting Up the Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
Updating the Raspberry Pi . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
Installing Required System Level Libraries . . . . . . . . . . . . . . . . . . . . . 62
Setting up a Virtual Environment . . . . . . . . . . . . . . . . . . . . . . . . . . 63
Install Python Packages (inside Virtual Environment) . . . . . . . . . . . . . . 63
Setting up Jupyter Notebook . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
Installing LiteRT . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
Creating a working directory: . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
Verifying the Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
Making inferences with Mobilenet V2 . . . . . . . . . . . . . . . . . . . . . . . . . . 73
Define a general Image Classification function . . . . . . . . . . . . . . . . . . . 79
Testing the classification with the Camera . . . . . . . . . . . . . . . . . . . . . 81
Exploring a Model Trained from Zero . . . . . . . . . . . . . . . . . . . . . . . 83
Conclusion: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86

3
Custom Image Classification Project 88
Image Classification Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
The Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
Data Collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
Training the model with Edge Impulse Studio . . . . . . . . . . . . . . . . . . . . . . 98
Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
The Impulse Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
Image Pre-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
Model Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
Model Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
Trading off: Accuracy versus speed . . . . . . . . . . . . . . . . . . . . . . . . . 104
Model Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Live Image Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
Summary: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123

Object Detection: Fundamentals 125


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
Object Detection Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
Pre-Trained Object Detection Models Overview . . . . . . . . . . . . . . . . . . . . . 131
Setting Up the TFLite Environment . . . . . . . . . . . . . . . . . . . . . . . . 132
Creating a Working Directory: . . . . . . . . . . . . . . . . . . . . . . . . . . . 133
Inference and Post-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133
EfficientDet . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140
Object Detection on a live stream . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 146
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 146

Custom Object Detection Project 148


Object Detection Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
The Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
Raw Data Collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150
Labeling Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153
Training an SSD MobileNet Model on Edge Impulse Studio . . . . . . . . . . . . . . 163
Uploading the annotated data . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163
The Impulse Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164
Preprocessing all dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166
Model Design, Training, and Test . . . . . . . . . . . . . . . . . . . . . . . . . . 168
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170
Inference and Post-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
Training a FOMO Model at Edge Impulse Studio . . . . . . . . . . . . . . . . . . . . 182
How FOMO works? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183

4
Impulse Design, new Training and Testing . . . . . . . . . . . . . . . . . . . . . 185
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 188
Inference and Post-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 190
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196

Computer Vision Applications with YOLO 197


Exploring a YOLO Models using Ultralitics . . . . . . . . . . . . . . . . . . . . . . . 197
Talking about the YOLO Model . . . . . . . . . . . . . . . . . . . . . . . . . . 198
Available Model Sizes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
Installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
Installing Ultralytics/Yolo . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 203
Testing the YOLO . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204
Export Models to NCNN format . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206
Exploring YOLO with Python . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
Inference Arguments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
Exploring other Computer Vision Applications . . . . . . . . . . . . . . . . . . . . . 214
Environment Setup and Dependencies . . . . . . . . . . . . . . . . . . . . . . . 214
Model Configuration and Loading . . . . . . . . . . . . . . . . . . . . . . . . . . 215
Performance Characteristics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215
Results Object Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 216
Bounding Box Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 216
8. Visualization and Customization . . . . . . . . . . . . . . . . . . . . . . 216
Exploring Other Computer Vision Tasks . . . . . . . . . . . . . . . . . . . . . . . . . 219
Image Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219
Instance Segmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220
Training YOLO on a Customized Dataset . . . . . . . . . . . . . . . . . . . . . . . . 224
Object Detection Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224
The Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 225
Inference with the trained model, using the Raspi . . . . . . . . . . . . . . . . . 230
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 233
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 234

Counting objects with YOLO 235


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 236
Estimating the number of Bees . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 237
Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 238
Pre-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 242
Training YOLOv8 on a Customized Dataset . . . . . . . . . . . . . . . . . . . . 245
Inference with the trained model, using the Rasp-Zero . . . . . . . . . . . . . . . . . 249
Considerations about the Post-Processing . . . . . . . . . . . . . . . . . . . . . . . . 253
Script For Reading the SQLite Database . . . . . . . . . . . . . . . . . . . . . . 257
Adding Environment data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 258

5
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 259
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 260

Image Classification with EXECUTORCH 261


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261
What is EXECUTORCH? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262
Why EXECUTORCH for Edge AI? . . . . . . . . . . . . . . . . . . . . . . . . . 262
Framework Comparison: EXECUTORCH vs TensorFlow Lite . . . . . . . . . . 263
Setting Up the Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 264
Updating the Raspberry Pi . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 264
Installing Required System-Level Libraries . . . . . . . . . . . . . . . . . . . . . 264
Setting up a Virtual Environment . . . . . . . . . . . . . . . . . . . . . . . . . . 266
Install Python Packages . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 268
PyTorch and EXECUTORCH Installation . . . . . . . . . . . . . . . . . . . . . . . . 269
Installing PyTorch for Raspberry Pi . . . . . . . . . . . . . . . . . . . . . . . . 269
Installing EXECUTORCH Runtime . . . . . . . . . . . . . . . . . . . . . . . . 269
Verifying the Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 270
Image Classification using MobileNet V2 . . . . . . . . . . . . . . . . . . . . . . . . . 271
Working directory: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271
Making inference with Torch . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271
Exporting Models to EXECUTORCH Format . . . . . . . . . . . . . . . . . . . . . . 274
Understanding the Export Process . . . . . . . . . . . . . . . . . . . . . . . . . 275
Exporting MobileNet V2 to ExecuTorch . . . . . . . . . . . . . . . . . . . . . . 275
Model Quantization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
Model Size/Performance Comparison . . . . . . . . . . . . . . . . . . . . . . . . 288
Making Inferences with EXECUTORCH . . . . . . . . . . . . . . . . . . . . . . . . . 288
Setting up Jupyter Notebook . . . . . . . . . . . . . . . . . . . . . . . . . . . . 289
The Project folder . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 289
Loading and Running a Model . . . . . . . . . . . . . . . . . . . . . . . . . . . 290
Setup and Verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 290
Download Test Image . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 291
Load EXECUTORCH Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 292
Download ImageNet Labels . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
Image Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
Define preprocessing pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
Apply preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
Add batch dimension: [1, 3, 224, 224] . . . . . . . . . . . . . . . . . . . . . . . . 295
Run Inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 295
Process and Display Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 296
Create Reusable Classification Function . . . . . . . . . . . . . . . . . . . . . . . . . 297
Classification Function Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
Using the XNNPACK accelerated backend . . . . . . . . . . . . . . . . . . . . . . . . 300
Quantized model - XNNPACK accelerated backend . . . . . . . . . . . . . . . . . . . 301

6
Camera Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303
Image Capture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303
Performance Benchmarking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 306
Basic (Float32): mobilenet_v2.pte . . . . . . . . . . . . . . . . . . . . . . . . . 308
XNNPACK Backend (Flot32): mobilenet_v2_xnnpack.pte . . . . . . . . . . . 309
Quantization (INT8): mobilenet_v2_quantized_xnnpack.pte . . . . . . . . . . 310
Performance Comparison Table . . . . . . . . . . . . . . . . . . . . . . . . . . . 311
Exploring Custom Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 312
Exporting a Custom Trained Model . . . . . . . . . . . . . . . . . . . . . . . . 313
Running Custom Models on Raspberry Pi . . . . . . . . . . . . . . . . . . . . . 314
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317
Key Takeaways . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318
Performance Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Code Repository . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Official Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Books . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319

Beyond CPU - Hardware Acceleration for Edge AI 320


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 321
Why Hardware Acceleration? . . . . . . . . . . . . . . . . . . . . . . . . . . . . 321
Our goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 322
Hardware Installation and Verification . . . . . . . . . . . . . . . . . . . . . . . . . . 322
Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 322
Installation and Cooling Considerations . . . . . . . . . . . . . . . . . . . . . . 323
Verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
Software Installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
Create Project Directory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
Python Version Management with Pyenv . . . . . . . . . . . . . . . . . . . . . 324
MemryX Drivers and SDK Installation . . . . . . . . . . . . . . . . . . . . . . . 326
Install Tools (Inside Virtual Environment) . . . . . . . . . . . . . . . . . . . . . 328
Verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 329
Our First Accelerated Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 330
Understanding the MX3 Workflow . . . . . . . . . . . . . . . . . . . . . . . . . 330
Download and Compile MobileNetV2 . . . . . . . . . . . . . . . . . . . . . . . . 335
Benchmarking Performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 336
Building an Inference Application . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
Prepare Directory Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
Download Test Image . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
Understanding Input Requirements . . . . . . . . . . . . . . . . . . . . . . . . . 339
Getting the Labels . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 339
Image Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 340
Prepare Input Tensor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 341

7
Run Inference on MemryX Accelerator (MXA) . . . . . . . . . . . . . . . . . . 341
Decode the MXA Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 342
Comparing CPU vs. MXA Performance . . . . . . . . . . . . . . . . . . . . . . 343
Measuring Latency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 343
Testing with Larger Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 345
Clean Shutdown . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Folders Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Performance Comparison Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Key Observations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
When to Use the MX3? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
YOLOv8 Object Detection with MX3 Hardware Acceleration . . . . . . . . . . . . . 347
Model Export and Compilation . . . . . . . . . . . . . . . . . . . . . . . . . . . 349
Understanding YOLOv8 Output Format . . . . . . . . . . . . . . . . . . . . . . 352
Complete Inference Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 352
Going deeper in the functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 355
Making Inferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 358
Inference with a custom model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 360
Adjusting Confidence Threshold . . . . . . . . . . . . . . . . . . . . . . . . . . 364
Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Batch Processing (Optimization) . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Thermal Management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Confidence Threshold Tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Exploring MemryX eXamples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Clone the MemryX eXamples Repository . . . . . . . . . . . . . . . . . . . . . 366
Troubleshooting Common Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 366
Device Not Detected . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 366
Compilation Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 368
Thermal Throttling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369
Python Version Conflicts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369
Low FPS / Poor Performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . 370
Import Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 370
Model Accuracy Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Next Steps and Extensions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Project Ideas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Advanced Topics to Explore . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
References and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Official Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Code and Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Background Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Community and Support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373

8
Text Generation with RNNs 374
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
What Are We Actually Building? . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Neural Network Architectures Background . . . . . . . . . . . . . . . . . . . . . . . . 375
The Human Brain Analogy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Recurrent Neural Networks (RNN) . . . . . . . . . . . . . . . . . . . . . . . . . 375
The Memory Problem and GRU Solution . . . . . . . . . . . . . . . . . . . . . 377
Dataset Preparation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 378
Data Preprocessing Steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 379
Tokenization and Vocabulary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
Character-Level Tokenization . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
Vocabulary Building Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
Creating the Character Dictionary . . . . . . . . . . . . . . . . . . . . . . . . . 381
Training Sequences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 382
The Sliding Window Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . 382
Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Creating Training Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Character Embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 384
From Sparse to Dense Representation . . . . . . . . . . . . . . . . . . . . . . . 384
Learning Character Relationships . . . . . . . . . . . . . . . . . . . . . . . . . . 384
Visualization and Understanding . . . . . . . . . . . . . . . . . . . . . . . . . . 385
Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
RNN Architecture Components . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
Model Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
Memory and Processing Flow . . . . . . . . . . . . . . . . . . . . . . . . . . . . 387
Why GRU over Basic RNN? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Training Process: Teaching the Model to Write . . . . . . . . . . . . . . . . . . . . . 388
The Learning Objective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Hardware and Time Requirements . . . . . . . . . . . . . . . . . . . . . . . . . 388
Monitoring Progress . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Preventing Overfitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 389
Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 389
Training Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
Text Generation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
The Generation Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
Temperature Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
Generation Example (Temperature = 0.5) . . . . . . . . . . . . . . . . . . . . . 392
Example Output Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 392
Challenges and Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
Context Window Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
Character vs. Word Level Trade-offs . . . . . . . . . . . . . . . . . . . . . . . . 393
Coherence Challenges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393

9
Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Potential Improvements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Connecting to Modern Language Models . . . . . . . . . . . . . . . . . . . . . . . . . 394
Scale Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Architectural Evolution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
Training Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
Summary: Why Modern Models Perform Better? . . . . . . . . . . . . . . . . . 395
Conclusion: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 396
Resourses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397

Knowledge Distillation in Practice 398


Introduction to Knowledge Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . 399
Why Knowledge Distillation Matters . . . . . . . . . . . . . . . . . . . . . . . . 400
The Teacher-Student Paradigm . . . . . . . . . . . . . . . . . . . . . . . . . . . 400
Theoretical Foundations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 401
Soft Targets vs. Hard Targets . . . . . . . . . . . . . . . . . . . . . . . . . . . . 401
The Temperature Parameter . . . . . . . . . . . . . . . . . . . . . . . . . . . . 401
Distillation Loss Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 402
Implementation with TensorFlow and MNIST . . . . . . . . . . . . . . . . . . . . . . 403
Dataset Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 403
Teacher Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 403
Student Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 406
Knowledge Distillation Implementation . . . . . . . . . . . . . . . . . . . . . . . 407
Training Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 407
Results and Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 408
Visualizing Distillation Benefits . . . . . . . . . . . . . . . . . . . . . . . . . . . 408
Advanced Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 409
Feature-Based Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 409
Attention-Based Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 410
Practical Tips for Advanced Distillation . . . . . . . . . . . . . . . . . . . . . . 410
Scaling to Large Language Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 410
Challenges in LLM Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . . 410
LLM Distillation Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 411
From MNIST to Real-World LLMs: Llama 3.2 Case Study . . . . . . . . . . . . 411
Applications in Modern LLM Development . . . . . . . . . . . . . . . . . . . . 413
Practical Guidelines and Best Practices . . . . . . . . . . . . . . . . . . . . . . . . . 414
Choosing the Right Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . 414
Hyperparameter Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 414
Common Pitfalls and Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . 415
Performance Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 415
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 415
Key Takeaways . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 415
From MNIST to LLMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416

10
Final Thoughts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416

Small Language Models (SLM) 418


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419
Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419
Raspberry Pi Active Cooler . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 420
Generative AI (GenAI) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 421
Large Language Models (LLMs) . . . . . . . . . . . . . . . . . . . . . . . . . . 422
Closed vs Open Models: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 423
Small Language Models (SLMs) . . . . . . . . . . . . . . . . . . . . . . . . . . . 424
Ollama . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 426
Installing Ollama . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 427
Meta Llama 3.2 1B/3B . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 429
Google Gemma . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 433
Microsoft Phi3.5 3.8B . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 439
MoonDream . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 441
Llava-Phi-3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 443
Ollama Commands . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 446
Inspecting Model parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . 447
Inspecting local resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 450
Ollama Python Library . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 451
Ollama Generate . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 456
[Link]() . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 463
Image Description: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 465
Calling Direct API . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 470
Error Handling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471
Connection Management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471
Features & Functionality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 472
Dependencies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 472
Advanced Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 472
When to Use Each? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 473
Performance Difference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 474
Going Further . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 474
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 475
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 477

SLM: Basic Optimization Techniques 478


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 478
Function Calling Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 480
But what exactly is “function calling”? . . . . . . . . . . . . . . . . . . . . . . . 480
Using the SLM for calculations . . . . . . . . . . . . . . . . . . . . . . . . . . . 480
The Solution: Function Calling / Tool Use . . . . . . . . . . . . . . . . . . . . . 482

11
Best Practices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 482
Function Calling Solution for Calculations . . . . . . . . . . . . . . . . . . . . . . . . 483
Define the Tool (Function Schema) . . . . . . . . . . . . . . . . . . . . . . . . . 483
Implement the Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 483
Project: Calculating Distances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 484
Running with other models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491
Adding images . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491
Retrievel Augmentation Generation (RAG) . . . . . . . . . . . . . . . . . . . . . . . 496
A simple RAG project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 497
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 504
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 504

Vision-Language Models at the Edge 506


Why Florence-2 at the Edge? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 507
Florence-2 Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 507
Technical Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 509
Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 509
Training Dataset (FLD-5B) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 510
Key Capabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 511
Practical Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 511
Comparing Florence-2 with other VLMs . . . . . . . . . . . . . . . . . . . . . . 512
Setup and Installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 512
Environment configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 513
Testing the installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 516
4. Defining the Prompt . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 520
7. Generating the Output . . . . . . . . . . . . . . . . . . . . . . . . . . . . 521
Florence-2 Tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 525
Exploring computer vision and vision-language tasks . . . . . . . . . . . . . . . . . . 527
Caption . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 529
DETAILED_CAPTION . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 529
MORE_DETAILED_CAPTION . . . . . . . . . . . . . . . . . . . . . . . . . . 530
OD - Object Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531
DENSE_REGION_CAPTION . . . . . . . . . . . . . . . . . . . . . . . . . . . 533
CAPTION_TO_PHRASE_GROUNDING . . . . . . . . . . . . . . . . . . . . 534
Cascade Tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 535
OPEN_VOCABULARY_DETECTION . . . . . . . . . . . . . . . . . . . . . . 536
Referring expression segmentation . . . . . . . . . . . . . . . . . . . . . . . . . 537
Region to Segmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 540
Region to Texts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 541
OCR . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 542
Latency Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 546
Fine-Tunning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 547

12
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 548
Key Advantages of Florence-2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . 548
Trade-offs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 549
Best Use Cases . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 549
Future Implications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 550
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 550

Audio and Vision AI Pipeline 551


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 551
The Audio to Audio AI Pipeline Architecture . . . . . . . . . . . . . . . . . . . . . . 552
Understanding Multimodal AI Systems . . . . . . . . . . . . . . . . . . . . . . . 552
System Architecture Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . 552
Edge AI Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 553
Hardware Setup and Audio Capture . . . . . . . . . . . . . . . . . . . . . . . . . . . 553
Audio Hardware Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 553
Testing Basic Audio Functionality . . . . . . . . . . . . . . . . . . . . . . . . . 554
Python Audio Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 555
Test Audio Python Script . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 556
Speech Recognition (SST) with Moonshine . . . . . . . . . . . . . . . . . . . . . . . 558
Why Moonshine for Edge Deployment . . . . . . . . . . . . . . . . . . . . . . . 558
Model Selection Strategy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 558
Implementation and Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . 559
SLM Integration and Response Generation . . . . . . . . . . . . . . . . . . . . . . . 559
Connecting STT Output to SLMs . . . . . . . . . . . . . . . . . . . . . . . . . 559
Optimizing Prompts for Voice Interaction . . . . . . . . . . . . . . . . . . . . . 560
Text-to-Speech (TTS) with PIPER . . . . . . . . . . . . . . . . . . . . . . . . . . . . 561
Voice Model Selection and Installation . . . . . . . . . . . . . . . . . . . . . . . 561
Handling Long Text and Special Cases . . . . . . . . . . . . . . . . . . . . . . . 564
Pipeline Integration and Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . 564
Building the Complete System . . . . . . . . . . . . . . . . . . . . . . . . . . . 564
Performance Optimization Strategies . . . . . . . . . . . . . . . . . . . . . . . . 565
Memory Management Considerations . . . . . . . . . . . . . . . . . . . . . . . . 565
Designing Resilient Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 565
Debugging Complex Pipelines . . . . . . . . . . . . . . . . . . . . . . . . . . . . 565
The full Python Script . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 566
The Image to Audio AI Pipeline Architecture . . . . . . . . . . . . . . . . . . . . . . 568
Troubleshooting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 573
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 573
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 574

Physical Computing with Raspberry Pi 576


From Sensors to Smart Analysis with Small Language Models . . . . . . . . . . 576
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 577

13
Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 578
Accessing the GPIOs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 578
Pin Numbering Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
Safety First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
GPIO Zero Library . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 580
“Hello World”: Blinking an LED . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 581
Understanding the LED Circuit . . . . . . . . . . . . . . . . . . . . . . . . . . . 581
Testing with GPIO Zero . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 582
Installing all LEDs (the “actuators”) . . . . . . . . . . . . . . . . . . . . . . . . 584
Sensors Installation and setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 586
Button . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 586
Installing Adafruit CircuitPython . . . . . . . . . . . . . . . . . . . . . . . . . . 590
DHT22 - Temperature & Humidity Sensor . . . . . . . . . . . . . . . . . . . . . 592
Addendum: DHT22 on Raspberry Pi 5 with Debian Trixie . . . . . . . . . . . . 597
Installing the BMP280: Barometric Pressure & Altitude Sensor . . . . . . . . . 600
Measuring Weather and Altitude With BMP280 . . . . . . . . . . . . . . . . . 605
Playing with Sensors and Actuators . . . . . . . . . . . . . . . . . . . . . . . . . . . 608
Testing the Notebook setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
Initialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
GPIO Input and Output . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 611
Getting and displaying Sensor Data . . . . . . . . . . . . . . . . . . . . . . . . 615
Widgets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 617
Advanced GPIO Zero Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 618
PWM LED Brightness Control . . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Composite Devices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Other Useful GPIO Zero Devices . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Interacting an SLM with the Physical world . . . . . . . . . . . . . . . . . . . . . . . 620
Other Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 627
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Key Achievements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Technical Insights . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Practical Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Challenges and Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Future Enhancements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Final Thoughts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 630

Experimenting with SLMs for IoT Control 631


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 631
Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 633
Hardware Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 633
Software Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 635
Basic Sensor Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 635

14
SLM Basic Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 638
Act on Output (Actuators) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 642
Creating new functions for the Actuation: . . . . . . . . . . . . . . . . . . . . . 643
Prompting Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 646
Key Changes in the code: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 646
System Overview: Enhanced IoT Environmental Monitoring with SLM Control 647
The new code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 651
Code Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 654
TEST1: Temp above the threshold . . . . . . . . . . . . . . . . . . . . . . . . . 654
TEST2: Temp below the threshold . . . . . . . . . . . . . . . . . . . . . . . . . 656
TEST3: Alarm Button pressed . . . . . . . . . . . . . . . . . . . . . . . . . . . 657
Interacting with IoT Systems, using Natural Language Commands . . . . . . . . . . 657
What will change? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 658
How It will Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 659
Key Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 661
Running the System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Special Commands . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Languages . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Advantages of This Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Error Handling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Tips for Best Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665
Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665
Adding Data Logging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
Running the System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 667
Example Queries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 667
Data Files . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
sensor_readings.csv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
command_history.csv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
Tips . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 669
Examples: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 670
Prompt Optimization and Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . 671
Quick Solution: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 672
New Code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 674
Key Changes Under the Hood: . . . . . . . . . . . . . . . . . . . . . . . . . . . 674
Using Pydantic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 675
2. Implementing Pydantic in the Code . . . . . . . . . . . . . . . . . . . . . . . 676
Next Steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 679

Conclusion 681
Resourses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 682

15
Advancing EdgeAI: Beyond Basic SLMs 683
Understanding SLM Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 684
1. Knowledge Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 684
2. Reasoning Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
3. Inconsistent Outputs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
4. Domain Specialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
Techniques for Enhancing SLM at the Edge . . . . . . . . . . . . . . . . . . . . . . . 686
Optimizing Prompting Strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Chain-of-Thought Prompting . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Few-Shot Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Task Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 688
Building Agents with SLMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 689
General Knowledge Router . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 700
Improving Agent Reliability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703
1. Function Calling with Pydantic . . . . . . . . . . . . . . . . . . . . . . . . . 703
2. Response Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 705
Retrieval-Augmented Generation (RAG) . . . . . . . . . . . . . . . . . . . . . . . . . 707
Understanding RAG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 707
Implementing a Naive RAG System . . . . . . . . . . . . . . . . . . . . . . . . 708
Instalation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 709
Key Components of the Naive RAG System . . . . . . . . . . . . . . . . . . . . 709
Advantages of RAG for Edge AI . . . . . . . . . . . . . . . . . . . . . . . . . . 711
Optimizing RAG for Edge Devices . . . . . . . . . . . . . . . . . . . . . . . . . 712
Application: Enhanced Weather Station with RAG . . . . . . . . . . . . . . . . 712
Using the RAG System for Edge AI Engineering . . . . . . . . . . . . . . . . . 714
Testing Different Models and Chunk Sizes . . . . . . . . . . . . . . . . . . . . . 718
Advanced Agentic RAG System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 721
System Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 722
Key Workflow . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 723
Important Code Sections . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 724
Detailed Workflow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 726
Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 728
Fine-Tuning SLMs for Edge Deployment . . . . . . . . . . . . . . . . . . . . . . . . . 730
Preparing for Fine-Tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 731
Setting Up a Fine-Tuning Process . . . . . . . . . . . . . . . . . . . . . . . . . 731
Real implementation: Supervised Fine-Tuning (SFT) . . . . . . . . . . . . . . . 732
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 734
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 734

LiteRT-LM: Production-Ready LLM Inference at the Edge 736


Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 737
What is LiteRT-LM? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 738
From TensorFlow Lite to LiteRT . . . . . . . . . . . . . . . . . . . . . . . . . . 738

16
Architecture Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 739
The .litertlm Format . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
Supported Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
LiteRT-LM vs Ollama: When to Use Each . . . . . . . . . . . . . . . . . . . . . . . . 741
Installation on Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Creating a Virtual Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Installing the Package . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 743
Installing the Hugging Face CLI (for model downloads) . . . . . . . . . . . . . 743
Downloading a Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 744
Quick Sanity Check with the CLI . . . . . . . . . . . . . . . . . . . . . . . . . . 744
Python API . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
The Engine and Conversation Objects . . . . . . . . . . . . . . . . . . . . . . . 745
1. Simple Single-Turn Query . . . . . . . . . . . . . . . . . . . . . . . . . . . . 746
2. Interactive Chat with Persistent History . . . . . . . . . . . . . . . . . . . . 747
3. Tool Calling (Function Calling) . . . . . . . . . . . . . . . . . . . . . . . . . 749
4. Streaming with Performance Metrics . . . . . . . . . . . . . . . . . . . . . . 752
Complete Example: Putting it All Together . . . . . . . . . . . . . . . . . . . . . . . 755
Going Further . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759

Edge AI Engineering - Weekly Labs 761


Week 1: Introduction and Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
Lab 1: Raspberry Pi Configuration . . . . . . . . . . . . . . . . . . . . . . . . . 761
Lab 2: Development Environment Setup . . . . . . . . . . . . . . . . . . . . . . 761
Week 2: Image Classification Fundamentals . . . . . . . . . . . . . . . . . . . . . . . 762
Lab 3: Working with Pre-trained Models . . . . . . . . . . . . . . . . . . . . . . 762
Lab 4: Custom Dataset Creation . . . . . . . . . . . . . . . . . . . . . . . . . . 763
Week 3: Custom Image Classification . . . . . . . . . . . . . . . . . . . . . . . . . . 763
Lab 5: Edge Impulse Model Training . . . . . . . . . . . . . . . . . . . . . . . . 763
Lab 6: Model Deployment to Raspberry Pi . . . . . . . . . . . . . . . . . . . . 764
Week 4: Object Detection Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . 764
Lab 7: Pre-trained Object Detection . . . . . . . . . . . . . . . . . . . . . . . . 764
Lab 8: EfficientDet and FOMO Models . . . . . . . . . . . . . . . . . . . . . . 765
Week 5: Custom Object Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . . 766
Lab 9: Dataset Creation and Annotation . . . . . . . . . . . . . . . . . . . . . . 766
Lab 10: Training Models in Edge Impulse . . . . . . . . . . . . . . . . . . . . . 766
Week 6: Advanced Object Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . 767
Lab 11: FOMO Model Training . . . . . . . . . . . . . . . . . . . . . . . . . . . 767
Lab 12: YOLO Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . 767
Week 7: Object Counting Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 768
Lab 13: Custom YOLO Training . . . . . . . . . . . . . . . . . . . . . . . . . . 768

17
Lab 14: Fixed-Function AI Integration (Optional) . . . . . . . . . . . . . . . . 768
Week 8: Introduction to Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . 769
Lab 15: Raspberry Pi Configuration for SLMs . . . . . . . . . . . . . . . . . . . 769
Lab 16: Ollama Installation and Testing . . . . . . . . . . . . . . . . . . . . . . 769
Week 9: SLM Python Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 770
Lab 17: Ollama Python Library . . . . . . . . . . . . . . . . . . . . . . . . . . . 770
Lab 18: Function Calling and Structured Outputs . . . . . . . . . . . . . . . . 771
Week 10: Retrieval-Augmented Generation . . . . . . . . . . . . . . . . . . . . . . . 772
Lab 19: RAG Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 772
Lab 20: Advanced RAG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 772
Week 11: Vision-Language Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 773
Lab 21: Florence-2 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 773
Lab 22: Vision Tasks with Florence-2 . . . . . . . . . . . . . . . . . . . . . . . 773
Week 12: Physical Computing Basics . . . . . . . . . . . . . . . . . . . . . . . . . . . 774
Lab 23: Sensor and Actuator Integration . . . . . . . . . . . . . . . . . . . . . . 774
Lab 24: Jupyter Notebook Integration . . . . . . . . . . . . . . . . . . . . . . . 775
Week 13: SLM-Physical Computing Integration . . . . . . . . . . . . . . . . . . . . . 775
Lab 25: Basic SLM Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 775
Lab 26: SLM-IoT Control System . . . . . . . . . . . . . . . . . . . . . . . . . . 776
Week 14: Advanced Edge AI Techniques . . . . . . . . . . . . . . . . . . . . . . . . . 777
Lab 27: Building Agents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 777
Lab 28: Advanced Prompting and Validation . . . . . . . . . . . . . . . . . . . 777
Week 15: Final Project Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778
Lab 29: Agentic RAG System . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778
Lab 30: Final Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778
Hardware Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 779
Basic Setup (Weeks 1-7) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 779
Generative AI (Weeks 8-15) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 779
Physical Computing (Weeks 12-15) . . . . . . . . . . . . . . . . . . . . . . . . . 779
Software Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Development Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Computer Vision and DL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Physical Computing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Assessment Criteria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 781
Tips for Success . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 781

References 782
To learn more: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
Online Courses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
Books . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
Projects Repository . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782

18
AiEng4D 783

About the author 784

19
Preface

In the rapidly evolving technology landscape, the convergence of artificial intelligence and edge
computing is one of the most exciting frontiers. This intersection promises to revolutionize how
we interact with the world around us, bringing intelligence and decision-making capabilities
directly to the devices we use every day. At the heart of this revolution lies the Raspberry Pi,
a powerful yet accessible single-board computer (SBC) that has democratized computing and
now stands poised to do the same for edge AI.
This book, which serves as the official textbook for IESTI05 Edge AI Engineering at the
Federal University of Itajubá (UNIFEI) in Brazil, embodies both a passion for technology
and a conviction in its capacity to address real-world problems. While developed to support
UNIFEI’s engineering curriculum, the content is valuable for all learners, whether in academic
settings or pursuing independent study.
“Edge AI Engineering: Hands-on with the Raspberry Pi” is not just about theory or abstract
concepts. It’s about getting your hands dirty, writing code, training models, and seeing your
creations come to life. Each chapter blends foundational knowledge with practical application,
focusing on what’s possible with the Raspberry Pi platform.
From the compact Raspberry Pi Zero to the more powerful Pi 5, we explore how these incredible
devices can become the brains of intelligent systems—recognizing images, understanding speech,
detecting objects, and even running small language models. Each project serves as a stepping
stone, building your skills and confidence as you progress.
Beyond the technical skills, this book aims to instill something more valuable – a sense of
curiosity and possibility. The field of edge AI is still in its infancy, with new applications
and techniques emerging daily. By mastering the fundamentals presented here, you’ll be
well-equipped to explore these frontiers, perhaps even pushing the boundaries of what’s possible
on edge devices.
Whether you’re a student seeking to understand AI’s practical applications, a professional
looking to expand your skill set, or an enthusiast eager to add intelligence to your projects, we
hope this book serves as both a guide and an inspiration.
As you embark on this journey, remember that every expert was once a beginner. The learning
path is filled with challenges and moments of joy and discovery. Embrace both, and let your
creativity guide you.

20
Thank you for joining us on this exciting adventure into edge machine learning. Let’s begin
exploring what’s possible when we bring AI to the edge, one Raspberry Pi at a time.
Happy coding, and may your models always converge!
Prof. Marcelo Rovai
May, 2026

21
Acknowledgments

I extend my deepest gratitude to the entire AiEng4D Academic Network, comprised of distin-
guished professors, researchers, and professionals. Notable contributions from Marco Zennaro,
Ermanno Petrosemoli, Brian Plancher, José Alberto Ferreira, Jesus Lopez, Diego Mendez,
Shawn Hymel, Dan Situnayake, Pete Warden, and Laurence Moroney have been instrumental
in advancing our understanding of Embedded Machine Learning (TinyML) and Edge AI.
Special commendation is reserved for Professor Vijay Janapa Reddi of Harvard University. His
steadfast belief in the transformative potential of open-source communities, coupled with his
invaluable guidance and teachings, has been a beacon and a cornerstone of our efforts from the
beginning.
Acknowledging these individuals, we pay tribute to the collective wisdom and dedication that
have enriched this field and our work.

Google Nano Banana and OpenAI’s GPT were used to generate some of the images
in the book. Claude Sonnet and Perplexity helped with code and text reviews.

22
Introduction

Edge AI Engineering

In today’s rapidly evolving technological landscape, the convergence of artificial intelligence


and edge computing represents one of the most promising frontiers of innovation. Edge AI—the
practice of running AI algorithms locally on hardware devices rather than in the cloud—
transforms how we interact with technology daily, enabling more responsive, private, and
efficient intelligent systems.
This book, “Edge AI Engineering: Hands-on with the Raspberry Pi,” is your practical guide to
this exciting field. We’ll explore fixed-function AI (reactive systems that process specific inputs)
and generative AI (proactive systems that create new content) through hands-on projects using
the versatile and accessible Raspberry Pi platform.

Why Edge AI Matters

Traditional AI deployment often relies on cloud infrastructure, which requires constant con-
nectivity and introduces latency. Edge AI addresses these limitations by bringing intelligence
directly to where data is generated and actions occur. This approach offers several compelling
advantages:

• Reduced latency: Process data locally for near-instantaneous responses


• Enhanced privacy: Keep sensitive information on your device rather than sending it to
remote servers
• Network independence: Maintain functionality even without internet connectivity
• Lower bandwidth usage: Process data locally, sending only relevant results when
needed
• Energy efficiency: Optimize processing for resource-constrained environments

The Raspberry Pi Advantage

The Raspberry Pi, with its combination of affordability, processing capability, and extensive
GPIO options, provides an ideal platform for exploring Edge AI concepts. From the compact
Raspberry Pi Zero 2W to the more powerful Pi 5, these devices offer:

23
• Sufficient computational power for running optimized AI models
• A complete Linux-based operating system for straightforward development
• Extensive connectivity options for integrating with sensors and actuators
• A vibrant community and ecosystem of libraries and tools
• An accessible entry point for students, hobbyists, and professionals alike

What You’ll Learn

This book takes a progressive approach to Edge AI engineering, starting with foundational
concepts and building toward more advanced applications:

1. Essential setup and configuration: Prepare your Raspberry Pi for Edge AI develop-
ment
2. Computer vision applications: Implement image classification and object detection
systems
3. Small Language Models (SLMs): Run and optimize language models directly on
your Raspberry Pi
4. Vision-Language Models: Explore multimodal AI with Florence-2
5. Physical computing integration: Connect AI systems with sensors and actuators
6. Advanced optimization techniques: Enhance model performance through methods
like RAG, agents, and function calling

Each chapter includes detailed explanations, step-by-step instructions, and practical projects
demonstrating real-world applications of Edge AI concepts.

Who This Book Is For

Whether you’re a student exploring AI for the first time, an educator developing a curriculum,
a maker building innovative projects, or a professional seeking to expand your skills, this book
provides the knowledge and hands-on experience needed to successfully implement Edge AI
solutions on the Raspberry Pi platform.
Join us on this journey to the edge of AI innovation, where we’ll bridge theory and practice
through engaging, accessible projects that demonstrate the transformative potential of intelligent
edge computing.

24
About this Book

Several chapters (Labs) in this book also accompany the open-source book Machine Learning
Systems by Professor Vijay Janapa Reddi from Harvard, which we invite you to read.

“Edge AI Engineering: Hands-on with the Raspberry Pi” is designed as a practical, project-
based learning resource that bridges theoretical AI concepts with tangible implementations.
This book is part of the open-source Machine Learning Systems initiative, democratizing access
to AI education and applications.

Key Features

1. Progressive Learning Path: The book structure follows a natural progression from
basic to advanced concepts, beginning with foundational computer vision applications
and advancing to generative AI techniques.

25
2. Model-Specific Optimizations: Each chapter provides targeted guidance for specific
Raspberry Pi models, helping you maximize performance whether you’re using a Pi Zero
2W or a Pi 5.
3. Open-Source Foundation: We emphasize accessible tools and frameworks, including
Edge Impulse Studio, TensorFlow Lite, PyTorch, Transformers, and Ollama, ensuring
you can continue your learning journey with widely available resources.
4. Practical Problem-Solving: Rather than abstract exercises, each project addresses
real-world challenges that demonstrate Edge AI’s practical value.
5. Resource Optimization Techniques: Learn essential strategies for deploying AI on
resource-constrained devices, balancing performance needs with hardware limitations.
6. Cross-Domain Applications: Explore implementations spanning computer vision,
natural language processing, and physical computing, showcasing the versatility of Edge
AI.

Structure and Organization

The book is organized into two main sections:

1. Fixed Function AI (Computer Vision): Chapters covering image classification,


object detection, and specialized applications like object counting. We also explore the
section “Hardware Acceleration for Fixed-Function AI.”
2. Generative AI (Language and Vision Models): Chapters exploring Small Lan-
guage Models, Vision-Language Models, physical computing integration, and advanced
optimization techniques.

Each chapter follows a consistent format that includes:

• Conceptual background and theory


• Step-by-step implementation guides
• Practical projects with complete code
• Performance optimization strategies
• Ideas for further exploration

Prerequisites

While designed to be accessible, readers will benefit from:

• Basic Python programming knowledge


• Familiarity with Linux command-line basics
• Elementary understanding of machine learning concepts
• Previous experience with Raspberry Pi (helpful but not required)

26
By completing this book, you’ll possess the skills to design, implement, and optimize Edge
AI applications across a wide range of use cases, leveraging the unique capabilities of the
Raspberry Pi platform to bring intelligence to the edge.

27
Classification of AI Applications

As we embark on our journey through Edge AI Engineering with the Raspberry Pi, it’s essential
to understand the fundamental classification of AI applications that form the structure of this
book. Our exploration is divided into two parts, each representing a different paradigm in
artificial intelligence implementation.

Fixed Function AI vs. Generative AI

AI applications can be broadly categorized into two approaches that represent different capa-
bilities, interaction models, and implementation strategies:

Fixed Function AI (Reactive)

Fixed Function AI, or Reactive AI, operates by analyzing specific inputs according to predeter-
mined patterns and rules and then producing consistent outputs for given scenarios. These
systems:

• Respond to specific triggers: They activate only when presented with particular
inputs.
• Follow defined patterns: Their behavior is predictable and consistent.
• Excel at structured tasks: They perform exceptionally well at classification, detection,
and pattern recognition
• Operate within boundaries: Their capabilities are limited to their specific program-
ming.

In the first part of this book (Chapters 2-4), we explore fixed-function AI through computer
vision applications:

• Image classification for identifying objects in images


• Object detection for locating and labeling multiple objects
• Specialized detection applications like counting objects

These applications demonstrate how edge devices can deliver reliable, efficient AI in constrained
environments, focusing on specific, well-defined tasks.

28
Generative AI (Proactive)

Generative AI, also known as Proactive AI, represents a fundamental shift in capability. These
systems can:

• Create new content: They generate novel text, images, or solutions.


• Understand context: They interpret and respond to nuanced situations.
• Engage in dialogue: They maintain contextual awareness across interactions.
• Adapt to novel scenarios: They apply knowledge to situations beyond their explicit
training.

The second part of this book (Chapters 5-9) explores Generative AI at the edge:

• Small Language Models that bring conversational AI to edge devices


• Vision-Language Models that combine visual and textual understanding
• Physical computing integration that connects AI to the real world
• Advanced techniques to enhance edge AI capabilities

This progression from Fixed Function to Generative AI mirrors the evolution of artificial
intelligence itself—from specialized systems designed for specific tasks to more flexible, creative
systems capable of addressing a broader range of challenges.

Summary Table

AI Type Core Focus Example Applications Typical Use Cases


Fixed Data analysis, Fraud detection, spam Banking, healthcare
Function assessment, filters, image recognition diagnostics, security
(Reactive) automation systems
Generative Content creation, ChatGPT, DALL·E, Content creation,
(Proactive) anticipation, predictive maintenance, customer support, design,
dialogue smart assistants automation

Conclusion

• Fixed Function (Reactive) AI is ideal for applications requiring efficiency, predictabil-


ity, and low resource use, where tasks are well-defined and do not require creative output
or adaptation.
• Generative (Proactive) AI is suited for scenarios demanding creativity, personalization,
and anticipatory actions. It enables richer, more human-like interactions and innovative
solutions across industries.

29
The Edge AI Advantage

Both Fixed Function and Generative AI gain unique benefits when deployed at the edge:

• Reduced latency: Processing happens locally, eliminating network delays


• Enhanced privacy: Sensitive data remains on-device
• Operational reliability: Systems function regardless of network connectivity
• Resource efficiency: Optimized models utilize limited hardware effectively

By understanding these fundamental classifications, you’ll realize how different AI approaches


serve distinct purposes and how each can be effectively implemented on edge devices like the
Raspberry Pi.
As we progress through this book, this classification framework will help contextualize each
project and technique, connecting individual implementations to broader AI concepts and
applications.
#
Fixed Function AI (Reactive)

30
31
Setup

Figure 1: DALL·E prompt - An electronics laboratory environment inspired by the 1950s,


with a cartoon style. The lab should have vintage equipment, large oscilloscopes,
old-fashioned tube radios, and large, boxy computers. The Raspberry Pi 5 board is
prominently displayed, accurately shown in its real size, similar to a credit card, on
a workbench. The Pi board is surrounded by classic lab tools like a soldering iron,
resistors, and wires. The overall scene should be vibrant, with exaggerated colors and
playful details characteristic of a cartoon. No logos or text should be included.

32
This chapter will guide you through setting up the Raspberry Pi Zero 2 W (Raspi-Zero) and the
Raspberry Pi 5 (Raspi-5) models. We’ll cover hardware setup, operating system installation,
initial configuration, and tests.

The general instructions for the Raspi-5 also apply to the older Raspberry Pi
versions, such as the Raspi-3 and Raspi-4.

Introduction

The Raspberry Pi is a powerful and versatile single-board computer that has become an essential
tool for engineers across various disciplines. Developed by the Raspberry Pi Foundation, these
compact devices offer a unique combination of affordability, computational power, and extensive
GPIO (General Purpose Input/Output) capabilities, making them ideal for prototyping,
embedded systems development, and advanced engineering projects.

Key Features

1. Computational Power: Despite their small size, Raspberry Pis offer significant pro-
cessing capabilities, with the latest models featuring multi-core ARM processors and up
to 8GB of RAM.
2. GPIO Interface: The 40-pin GPIO header enables direct interaction with sensors,
actuators, and other electronic components, facilitating hardware-software integration.
3. Extensive Connectivity: Built-in Wi-Fi, Bluetooth, Ethernet, and multiple USB ports
enable a wide range of communication and networking projects.
4. Low-Level Hardware Access: Raspberry Pis provide access to interfaces such as I2C,
SPI, and UART, enabling detailed control and communication with external devices.
5. Real-Time Capabilities: With proper configuration, Raspberry Pis can be used for soft
real-time applications, making them suitable for control systems and signal processing
tasks.
6. Power Efficiency: Low power consumption enables battery-powered and energy-efficient
designs, especially in models like the Pi Zero.

33
Raspberry Pi Models (covered in this book)

1. Raspberry Pi Zero 2 W (Raspi-Zero):


• Ideal for: Compact embedded systems
• Key specs: 1GHz single-core CPU (ARM Cortex-A53), 512MB RAM, minimal
power consumption (usually, around 600mW in Idle. Also, the power-off (“zombie
current”) is very low, at around 45 mA (225 mW), significantly lower than on full-size
Raspberry Pi models.
2. Raspberry Pi 5 (Raspi-5):
• Ideal for more demanding applications, such as edge computing, computer vision,
and edgeAI applications, including LLMs.
• Key specs: 2.4GHz quad-core CPU (ARM Cortex A-76), up to 8GB RAM, PCIe
interface for expansions, higher power consumption (usually, around 3-3.5W in Idle)

Engineering Applications

1. Embedded Systems Design: Develop and prototype embedded systems for real-world
applications.
2. IoT and Networked Devices: Create interconnected devices and explore protocols
like MQTT, CoAP, and HTTP/HTTPS.
3. Control Systems: Implement feedback control loops, PID controllers, and interface
with actuators.
4. Computer Vision and AI: Utilize libraries like OpenCV and TensorFlow Lite for
image processing and machine learning at the edge.
5. Data Acquisition and Analysis: Collect sensor data, perform real-time analysis, and
create data logging systems.
6. Robotics: Build robot controllers, implement motion planning algorithms, and interface
with motor drivers.
7. Signal Processing: Perform real-time signal analysis, filtering, and DSP applications.
8. Network Security: Set up VPNs, firewalls, and explore network penetration testing.

This lab will guide you through setting up the most common Raspberry Pi models, so you can
get started on your machine learning project quickly. We’ll cover hardware setup, operating
system installation, and initial configuration, focusing on preparing your Pi for Machine
Learning applications.

34
Hardware Overview

Raspberry Pi Zero 2W

• Processor: 1GHz quad-core 64-bit Arm Cortex-A53 CPU


• RAM: 512MB SDRAM
• Wireless: 2.4GHz 802.11 b/g/n wireless LAN, Bluetooth 4.2, BLE
• Ports: Mini HDMI, micro USB OTG, CSI-2 camera connector
• Power: 5V/2.5A via micro USB port 12.5W Micro USB power supply

35
Raspberry Pi 5

• Processor:
– Pi 5: Quad-core 64-bit Arm Cortex-A76 CPU @ 2.4GHz
– Pi 4: Quad-core Cortex-A72 (ARM v8) 64-bit SoC @ 1.5GHz
• RAM: 2GB, 4GB, or 8GB options (8GB recommended for AI tasks)
• Wireless: Dual-band 802.11ac wireless (2.4 GHz and 5 GHz), Bluetooth 5.0
• Ports: 2 × micro HDMI ports, 2 × USB 3.0 ports, 2 × USB 2.0 ports, CSI camera port,
DSI display port
• Power: 5V/5A, 5V/3A limits peripherals to 600mA, via USB-C connector 27W USB-C
power supply

In the labs, we will use different names to address the Raspberry Pi: Raspi, Raspi-5,
Raspi-Zero, etc. Usually, “Raspi” or “Raspberry Pi” is used when the instructions
or comments apply to all models.

36
Installing the Operating System

The Operating System (OS)

An operating system (OS) is essential software that manages computer hardware and software
resources, providing standard services for computer programs. It is the core software that runs
on a computer, serving as an intermediary between hardware and application software. The
OS oversees the computer’s memory, processes, device drivers, files, and security protocols.

1. Key functions:
• Process management: Allocating CPU time to different programs
• Memory management: Allocating and freeing up memory as needed
• File system management: Organizing and keeping track of files and directories
• Device management: Communicating with connected hardware devices
• User interface: Providing a way for users to interact with the computer
2. Components:
• Kernel: The core of the OS that manages hardware resources
• Shell: The user interface for interacting with the OS
• File system: Organizes and manages data storage
• Device drivers: Software that allows the OS to communicate with hardware

The Raspberry Pi runs a specialized version of Linux designed for embedded systems. This
operating system, typically a variant of Debian called Raspberry Pi OS (formerly Raspbian), is
optimized for the Pi’s ARM-based architecture and limited resources.

The latest version of Raspberry Pi OS (Dec 25) is based on Debian Trixie.

Key features:

1. Lightweight: Tailored to run efficiently on the Pi’s hardware.


2. Versatile: Supports a wide range of applications and programming languages.
3. Open-source: Allows for customization and community-driven improvements.
4. GPIO support: Enables interaction with sensors and other hardware through the Pi’s
pins.
5. Regular updates: Continuously improved for performance and security.

Embedded Linux on the Raspberry Pi provides a full-featured operating system in a compact


package, making it ideal for projects ranging from simple IoT devices to more complex edge
machine-learning applications. Its compatibility with standard Linux tools and libraries makes
it a powerful platform for development and experimentation.

37
Installation

To use the Raspberry Pi, we will need an operating system. By default, Raspberry Pis check
for an operating system on any SD card inserted in the slot, so we should install an operating
system using Raspberry Pi Imager.

In November 2025, the Raspberry Pi Imager 2.0 was launched. It brings a new
wizard interface, the opportunity to pre-configure Raspberry Pi Connect, and
improved accessibility for screen readers and other assistive technologies.

Raspberry Pi Imager is a tool for downloading and writing images on macOS, Windows, and
Linux. It includes many popular operating system images for Raspberry Pi. We will also use
the Imager to preconfigure credentials and remote access settings.
Follow the steps to install the OS on your Raspberry Pi.

1. Download and install the Raspberry Pi Imager on your computer.


2. Insert a microSD card into your computer (a 32GB SD card is recommended) .
3. Open Raspberry Pi Imager and select your Raspberry Pi model.

38
4. Choose the appropriate operating system:

• For Raspi-Zero: For example, you can select under Raspberry Pi OS (Other),
Raspberry Pi OS Lite (64-bit).

Due to the Raspberry Pi Zero’s limited SDRAM (512 MB), the recommended
OS is the 32-bit version. However, to run some machine learning models, such
as the YOLO from Ultralitics, we should use the 64-bit version. Although the
Raspi-Zero can run a desktop, we will choose the LITE version (no Desktop) to
reduce the RAM needed for regular operation.

• For Raspi-5: We can select the full 64-bit version, which includes a desktop:
Raspberry Pi OS (64-bit)

39
5. Select your microSD card as the storage device.
6. Click Next, then go to the Customization tab. The imager will guide you through setting
the hostname, the Raspberry Pi username and password, configuring WiFi, and enabling
SSH (Very important!).
7. Write the image to the microSD card.

In the examples here, we will use different hostnames depending on the device used:
raspi, raspi-5, raspi-Zero, etc. Please replace it with the one you’re currently using.

Initial Configuration

1. Insert the microSD card into your Raspberry Pi.


2. Connect the power to boot up the Raspberry Pi. For the Rasp-5, use a 5V/3A (or
5A) external Power supply. For the Rasp-Zero, use a 5V/2.5A external Power supply
(Optionally, for light use, you can use the computer’s USB port to power the Rasp-Zero).

40
3. Please wait for the initial boot process to complete (it may take a few minutes).

You can find the most common Linux commands for the Raspberry Pi here or here.

Remote Access

SSH Access

The easiest way to interact with the Raspi-Zero is via SSH (“Headless”). You can use a
Terminal (MAC/Linux), PuTTy (Windows), or any other.

The Raspberry Pi and the notebook should be on the same WiFi network. Note
that the Raspberry Pi 5 supports dual-band 802.11ac Wi-Fi (2.4 GHz and 5 GHz),
whereas the Raspberry Pi Zero 2W supports only 2.4 GHz.

1. Find your Raspberry Pi’s IP address (for example, check your router).
2. On your computer, open a terminal and connect via SSH:
ssh username@[raspberry_pi_ip_address]

Alternatively, if you do not have the IP address, you can try, for example, ssh
mjrovai@[Link] , ssh mjrovai@[Link] , etc.: bash ssh username@[Link]
When you see the prompt:
mjrovai@rpi-5:~ $

It means you are interacting with your Raspberry Pi remotely.


It is a good practice to regularly update/upgrade the system. For that, you should run:
sudo apt-get update
sudo apt upgrade
sudo reboot # Reboot to ensure all updates take effect

41
You should confirm the Raspberry Pi IP address. On the terminal, you can use:
hostname -I

Also, check the current Python version:

python3 --version

Once we use the latest Raspberry Pi OS (based on Debian Trixie), it should be: 3.13:

As of today (January 2026), some packages, such as ExecuTorch, officially support only Python
versions 3.10-3.12. Python 3.13.5 is too new and will likely cause compatibility issues. Since
Debian Trixie ships with Python 3.13 by default, we’ll need to install a compatible Python
version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python. We will do it in the appropriate Lab.
For now, we will keep the system Python.

If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.

42
To shut down the Raspi via terminal:

When you want to turn off your Raspberry Pi, there are better ideas than just pulling the
power cord. This is because the Raspi may still be writing data to the SD card, in which case
merely powering down may result in data loss or, even worse, a corrupted SD card.
For a safety shutdown, use the command line:

sudo shutdown -h now

To avoid potential data loss and SD card corruption, wait a few seconds after
shutdown for the Raspberry Pi’s LED to stop blinking and go dark before removing
power. Once the LED goes out, it’s safe to power down.

Transfer Files between the Raspberry Pi and a computer

Transferring files between the Raspberry Pi and our main computer can be done using a USB
drive, via the terminal (scp), or an FTP program over the network.

Using Secure Copy Protocol (scp):

Copy files to your Raspberry Pi


Let’s create a text file on our computer, for example, [Link].

43
You can use any text editor. In the same terminal, an option is nano.

To copy the file named [Link] from your personal computer to a user’s home folder on your
Raspberry Pi, run the following command from the directory containing [Link], replacing
the <username> placeholder with the username you use to log in to your Raspberry Pi and the
<pi_ip_address> placeholder with your Raspberry Pi’s IP address:

$ scp [Link] <username>@<pi_ip_address>:~/

Note that ~/ means we will move the file to the ROOT of our Raspberry Pi. You
can choose any folder in your Raspberry Pi. But you should create the folder before
you run scp, since scp won’t create folders automatically.

For example, let’s transfer the file [Link] to the ROOT of my Raspberry Pi Zero, which
has an IP of [Link]:

44
scp [Link] mjrovai@[Link]:~/

I use a different profile to differentiate the terminals. The action above occurs on your
computer. Now, let’s go to our Raspi (using SSH) and check if the file is there:

Copy files from your Raspberry Pi


To copy a file named [Link] from a user’s home directory on a Raspberry Pi to the current
directory on another computer, run the following command on your Host Computer:

$ scp <username>@<pi_ip_address>:[Link] .

For example:
On the Raspi, let’s create a copy of the file with another name:

cp [Link] test_2.txt

And on the Host Computer (in my case, a Mac)

45
scp mjrovai@[Link]:test_2.txt .

Transferring files using FTP

Transferring files via FTP, such as FileZilla FTP Client, is also possible and much easier to use.
Follow the instructions to install the program on your Desktop, then use the Raspberry Pi’s IP
address as the Host. For example:

s[Link]

Enter your Raspberry Pi username and password. Pressing Quickconnect opens two windows,
one for your host computer desktop (right) and another for the Raspberry Pi (left).

46
Increasing Memory

Using htop, a cross-platform interactive process viewer, we can easily monitor resources on our
Raspberry Pi in real time, including the list of processes, the running CPUs, and the memory
usage. To lunch hop, enter with the command on the terminal:

htop

47
Regarding memory, among the Raspberry Pi family, the Raspberry Pi Zero has the least SRAM
(500 MB), compared to 2GB to 16GB on the Raspberry Pi 4 or 5.
On any Raspberry Pi, it is possible to increase/modify the system’s available memory using
“Swap.” Swap memory, also known as swap space, is a technique used in computer operating
systems to temporarily store data from RAM (Random Access Memory) on the SD card when
the physical RAM is fully utilized. This allows the operating system (OS) to continue running
even when RAM is full, which can prevent system crashes or slowdowns.
Swap memory benefits devices with limited RAM, such as the Raspberry Pi Zero. Increasing
swap can help run more demanding applications or processes, but it’s essential to balance this
with the potential performance impact of frequent disk access.
We can check the swap memory using htopor by the command:

sudo swapon --show

48
On the Debian Trixie, 2 MB of swap memory is configured by default on the Rasp-5
and 512MB on the Raspi-Zero

Increasing zram for the Raspberry Pi 5

1. Create the drop‑in folder (if it does not exist):


sudo mkdir -p /etc/rpi/[Link].d

2. Criar um arquivo de config, por exemplo, [Link] :


sudo nano /etc/rpi/[Link].d/[Link]

3. Set the zram swap (in this case 4Gi):


[Main]
Mechanism=zram+file

[Zram]
RamMultiplier=2
MaxSizeMiB=4096

• RamMultiplier=2: until 2x the physical RAM (limited by MaxSizeMiB ). •


MaxSizeMiB=4096: Set to 4 GB, even if the multiplier permits higher memory.
4. Reboot
sudo reboot

5. Apply and check:


swapon --show
free -h

We should see /dev/zram0 with the new size, 4.0Gi.

49
Increasing memory for the Rasp-Zero

By default, the Rapi-Zero’s SWAP (zram) memory is 512MB, which may be insufficient for
running more complex and demanding Machine Learning applications (for example, YOLO).
So, let’s increase it to 2MB:
On Zero 2 W with Raspberry Pi OS Trixie, you configure rpi-swap by dropping a small config
file into /etc/rpi/[Link].d/. The same mechanism works on Pi 5 and Zero‑2.
Below are two common patterns: fixed swap file (on SD/SSD) and fixed zram size (in RAM).
#### Fixed‑size swap file (e.g., 2 GB on SD)

1. Create the override directory (if not already there):


sudo mkdir -p /etc/rpi/[Link].d

2. Create an override file:


sudo nano /etc/rpi/[Link].d/[Link]

3. Put this in it (example: 2 GB = 2048 MiB):


[Main]
Mechanism=swapfile

[File]
FixedSizeMiB=2048

• Mechanism=swapfile tells rpi-swap to use a classic swap file.

• FixedSizeMiB is the exact size we want, in MiB.

4. Apply:
sudo systemctl restart rpi-swap
swapon --show

We should now see /var/swap with the size we set.

Fixed‑size zram (better for Zero‑2’s SD card)

If we want to keep everything in compressed RAM (no SD wear), we can set a fixed zram size
instead:

50
1. Create the same drop‑in file:
sudo nano /etc/rpi/[Link].d/[Link]

2. Set the zram swap:


[Main]
Mechanism=zram

[Zram]
FixedSizeMiB=2048

3. Apply and check:


sudo systemctl restart rpi-swap
swapon --show

We should see /dev/zram0 with the new size.

Notes specific to Zero 2 W Trixie

• Default Trixie on Zero‑2 uses rpi-swap with zram sized roughly to RAM×1, capped by
MaxSizeMiB (usually 2048), and no swap file unless you add it.
• If we create multiple drop‑ins in /etc/rpi/[Link].d/, they are merged; keep it
simple with one clear override file per board/use‑case.

To keep htop running, open another terminal window to continuously interact with
the Raspberry Pi.

Installing a Camera

The Raspberry Pi is an excellent device for computer vision applications that require a camera.
We can install a camera module connected to the Raspberry Pi CSI (Camera Serial Interface)
port, or a standard USB webcam on the micro-USB port using a USB OTG adapter (Raspi-Zero)
or directly on the USB port on the Raspi-5.

USB Webcams generally have inferior image quality compared to camera modules
that connect to the CSI port. They can not be controlled using the raspistill and
rasivid commands in the terminal, or the picamera recording package in Python.
Nevertheless, there may be reasons you want to connect a USB camera to your
Raspberry Pi, such as setting up multiple cameras with a single Raspberry Pi,
avoiding long cables, or simply because you have such a camera on hand.

51
Installing a Camera Module on the CSI port

There are now several Raspberry Pi camera modules. The original 5-megapixel model was
releasedin 2013, followed by the 8-megapixel Camera Module 2, released in 2016. The latest
camera model is the 12-megapixel Camera Module 3, released in 2023.
The original 5MP camera (Arducam OV5647) is no longer available from Raspberry Pi but
can be found from several alternative suppliers. Below is an example of such a camera on a
Raspberry Pi Zero.

Here is another example of a v2 Camera Module, which has a Sony IMX219 8-megapixel
sensor:

52
First, try to list the installed cameras:

rpicam-hello --list-cameras

If the command is not recognized, install

sudo apt install libcamera-apps

Try to list the installed camera again. If Ok, you should see something like:

53
Any camera module will work on the Raspberry Pi, but for that, it is possible that the
[Link] file must be updated:

sudo nano /boot/firmware/[Link]

At the bottom of the file, for example, to use the 5MP Arducam OV5647 camera, add the
line:

dtoverlay=ov5647,cam0

Or for the v2 module, wich has the 8MP Sony IMX219 camera:

dtoverlay=imx219,cam0

Save the file (CTRL+O [ENTER] CRTL+X) and reboot the Raspi:

Sudo reboot

After the boot, you can see if the camera is listed:

rpicam-hello --list-cameras

54
libcamerais an open-source software library that supports camera systems directly
on Linux for Arm processors. It minimizes the amount of proprietary code running
on the Broadcom GPU.

Let’s capture a JPEG image with a resolution of 640 x 480 for testing and save it to a file
named test_cli_camera.jpg

rpicam-jpeg --output test_cli_camera.jpg --width 640 --height 480

To view the saved file, we should use ls, which lists the contents of the current directory:

As before, we can use scp to view the image:

55
Alternatively, you can transfer it to your desktop using FileZilla.

Installing a USB WebCam

1. Power off the Raspi:

sudo shutdown -h no

2. Connect the USB Webcam (USB Camera Module 30 fps, 1280x720) to your Raspberry
Pi (in this example, I am using a Raspberry Pi Zero, but the instructions work for all
Raspberry Pis).

56
3. Power on again and run the SSH
4. To check if your USB camera is recognized, run:

lsusb

You should see your camera listed in the output.

57
5. To take a test picture with your USB camera, use:

fswebcam test_image.jpg

This will save an image named “test_image.jpg” in your current directory.

6. Since we are using SSH to connect to our Rapsi, we must transfer the image to our main
computer so we can view it. We can use FileZilla or SCP for this:

Open a terminal on your host computer and run:

scp mjrovai@[Link]:~/test_image.jpg .

Replace “mjrovai” with your username and “raspi-zero” with Pi’s hostname.

7. If the image quality isn’t satisfactory, you can adjust various settings; for example, define
a resolution that is suitable for YOLO (640x640):

58
fswebcam -r 640x640 --no-banner test_image_yolo.jpg

This captures a higher-resolution image without the default banner.

An ordinary USB Webcam can also be used:

59
And verified using lsusb

Video Streaming

For stream video (which is more resource-intensive), we can install and use mjpg-streamer:
First, install Git:

60
sudo apt install git

Now, we should install the necessary dependencies for mjpg-streamer, clone the repository, and
proceed with the installation:

sudo apt install cmake libjpeg62-turbo-dev


git clone [Link]
cd mjpg-streamer/mjpg-streamer-experimental
make
sudo make install

Then start the stream with:

mjpg_streamer -i "input_uvc.so" -o "output_http.so -w ./www"

We can then access the stream by opening a web browser and navigating to:
[Link] In my case: [Link]
We should see a webpage with options to view the stream. Click on the link that says “Stream”
or try accessing:

[Link]

61
Running the Raspi Desktop remotely

While we’ve primarily interacted with the Raspberry Pi via SSH terminal commands, we
can access the full graphical desktop environment remotely if we have installed the complete
Raspberry Pi OS (for example, Raspberry Pi OS (64-bit). This can be particularly useful
for tasks that benefit from a visual interface. To enable this functionality, we must set up a
VNC (Virtual Network Computing) server on the Raspberry Pi. Here’s how to do it:

1. Enable the VNC Server:

62
• Connect to your Raspberry Pi via SSH.
• Run the Raspberry Pi configuration tool by entering:
sudo raspi-config

• Navigate to Interface Options using the arrow keys.

• Select VNC and Yes to enable the VNC server.

• Exit the configuration tool (use [Tab]), saving changes when prompted.

63
2. Install a VNC Viewer on Your Computer:

• Download and install a VNC viewer application on your main computer. Popular
options include RealVNC Viewer, TightVNC, or VNC Viewer by RealVNC. We will
install VNC Viewer by RealVNC.

3. Once installed, confirm the Raspberry Pi’s IP address. For example, on the terminal,
you can use:
hostname -I

4. Connect to Your Raspberry Pi:

• Open your VNC viewer application.

64
• Enter your Raspberry Pi’s IP address and hostname.
• When prompted, enter your Raspberry Pi’s username and password.

65
5. The Raspberry Pi 5 Desktop should appear on your computer monitor.

66
6. Adjust Display Settings (if needed):

• Once connected, adjust the display resolution for optimal viewing. This can be done
through the Raspberry Pi’s desktop settings or by modifying the [Link] file.
• Let’s do it using the desktop settings. Reach the menu (the Raspberry Icon at the
left upper corner) and select the best screen definition for your monitor:

67
Updating and Installing Software

1. Update your system:


sudo apt update && sudo apt upgrade -y

2. Install essential software:


sudo apt install python3-pip -y

3. Enable pip for Python projects:


sudo rm /usr/lib/python3.11/EXTERNALLY-MANAGED

AVOID using PIP for System-Level dependencies libraries


Use sudo apt install (outside of a virtual environment (venv) for:

• System-level dependencies and libraries


• Hardware interface libraries (as a camera)
• Development headers and build tools
• Anything that needs to interface directly with hardware

68
Rule of thumb: Use sudo apt install only for system dependencies and hardware interfaces.
Use pip install (without sudo) inside an activated virtual environment for everything else.
Inside the vent, PIP or PIP3 are the same.

Model-Specific Considerations

Raspberry Pi Zero (Raspi-Zero)

• Limited processing power, best for lightweight projects


• It is better to use a headless setup (SSH) to conserve resources.
• Consider increasing swap space for memory-intensive tasks.
• It can be used for Image Classification and Object Detection Labs, but not for the LLM
(SLM).

Raspberry Pi 4 or 5 (Raspi-4 or Raspi-5)

• Suitable for more demanding projects, including AI and machine learning.


• It can run the whole desktop environment smoothly.
• The Raspberry Pi 4 can be used for Image Classification and Object Detection Labs, but
it will not work well with LLMs (SLM) due to high latency.
• For the Raspberry Pi 5, consider using an active cooler to manage temperature during
intensive tasks, as in the LLMs (SLMs) lab.

Remember to adjust your project requirements based on the specific Raspberry Pi model you’re
using. The Raspi-Zero is great for low-power, space-constrained projects, while the Raspi-4 or
5 models are better suited for more computationally intensive tasks.
Here’s a concise, copy‑pasteable tutorial you can share or publish.

Measuring Temperature and Power

In this section, we will explore how to measure (and monitor) CPU temperature and power
consumption on a Raspberry Pi 5 using only built‑in tools and a small shell script.

69
Check CPU Temperature

Quick one‑shot reading

vcgencmd measure_temp

Example:

temp=47.2'C

This is the SoC (CPU/GPU) temperature in degrees Celsius.

Continuous temperature monitor

watch -n 2 vcgencmd measure_temp

• Updates every 2 seconds.

• Press Ctrl+C to stop.

Reading from sysfs (for scripts)

cat /sys/class/thermal/thermal_zone0/temp

The result is in millidegrees Celsius (e.g., 47200 = 47.2 °C).


We should convert to human‑readable:

echo $(($(cat /sys/class/thermal/thermal_zone0/temp) / 1000))'°C'

Understand Power Measurement on Pi 5

Raspberry Pi 5 exposes internal PMIC rails via vcgencmd pmic_read_adc.

70
vcgencmd pmic_read_adc

Example (truncated):

3V3_SYS_A current(1)=0.06245952A
1V8_SYS_A current(2)=0.16102850A
...
3V3_SYS_V volt(13)=3.29687500V
1V8_SYS_V volt(14)=1.79687500V
...

To get power, we need to sum


𝑃𝑖 = 𝑉 𝑖 × 𝐼 𝑖
over all rails. This gives board core power, not exact wall power (USB devices and PSU
losses are not fully included).

Script: Average Temperature and Power

Let’s create a script that:

• Samples once per second for N seconds.

• Computes the average CPU temperature.

• Sums all PMIC rails to get power.

• Applies a linear correction to estimate real board power, based on the RPi5‑power
calibration.

Create the script

nano avg_temp_power.sh

Paste:

71
#!/bin/bash
# avg_temp_power.sh
# Measure average CPU temperature and power on Raspberry Pi 5

SECONDS=50 # number of samples (1 per second)

TEMP_FILE=$(mktemp)
PMIC_FILE=$(mktemp)

echo "Sampling for ${SECONDS} seconds..."

for i in $(seq 1 ${SECONDS}); do


# --- Temperature (°C) ---
temp_c=$(vcgencmd measure_temp | cut -d= -f2 | cut -d"'" -f1)
echo "$temp_c" >> "${TEMP_FILE}"

# --- PMIC power sum (W) ---


pmic_line=$(vcgencmd pmic_read_adc)
volts=($(echo "$pmic_line" | grep -o "volt([^)]*)=[0-9.]*V" | sed 's/.*=//;s/V//'))
currents=($(echo "$pmic_line" | grep -o "current([^)]*)=[0-9.]*A" | sed 's/.*=//;s/A//'))

sum_pmic=0
for idx in "${!volts[@]}"; do
v=${volts[$idx]}
i_amp=${currents[$idx]}
sum_pmic=$(awk -v a="$sum_pmic" -v b="$v" -v c="$i_amp" \
'BEGIN{printf "%.8f", a + b*c}')
done

echo "$sum_pmic" >> "${PMIC_FILE}"


sleep 1
done

# --- Compute statistics ---


awk -v corr_a=1.1451 -v corr_b=0.5879 '
function stats(file, x, n, sum, sum2, mean, std) {
while ((getline x < file) > 0) {
n++
sum += x
sum2 += x*x
}
close(file)

72
if (n > 1) {
mean = sum / n
std = sqrt( (sum2 - (sum*sum)/n) / (n-1) )
} else {
mean = sum
std = 0
}
return mean "|" std
}
BEGIN{
split(stats("'"${TEMP_FILE}"'"), t, "|")
t_mean = t[1]
t_std = t[2]

split(stats("'"${PMIC_FILE}"'"), p, "|")
p_pmic_mean = p[1]
p_pmic_std = p[2]

# Apply linear correction for real power estimate (RPi5-power calibration)


p_real_mean = corr_a * p_pmic_mean + corr_b
p_real_std = corr_a * p_pmic_std

printf "Average_temperature = %.2f +/- %.2f °C\n", t_mean, t_std


printf "Average_pmic_power = %.3f +/- %.3f W\n", p_pmic_mean, p_pmic_std
printf "Estimated_real_power = %.3f +/- %.3f W\n", p_real_mean, p_real_std
}
' /dev/null

rm -f "${TEMP_FILE}" "${PMIC_FILE}"

Save and exit.

Make it executable

chmod +x avg_temp_power.sh

Run a measurement

Idle baseline:

73
./avg_temp_power.sh

Example output (30–50 s window, active fan):

Sampling for 50 seconds...


Average_temperature = 47.25 +/- 0.40 °C
Average_pmic_power = 2.312 +/- 0.298 W
Estimated_real_power = 3.235 +/- 0.341 W

Then start a heavy workload (e.g., Llama 3.2:3B) in another terminal and run the script again
during inference. You should see higher temperature and power, often around 70–75 °C and
10–11 W with active cooling.

Interpreting the Results

Typical Raspberry Pi 5 ranges with active cooling:

74
State Temp (°C) Estimated real power (W)
Idle CLI 40–50 3.0–3.5
Light desktop 50–60 4–5
Moderate CPU load 55–70 5–7
Heavy CPU/GPU 70–80 7–10

• Throttling usually starts around 80 °C, with a hard limit at 85 °C.[5]


• If we are well below that under full load, your cooling solution is working properly.

75
76
Image Classification Fundamentals

Figure 2: DALL·E prompt - “Create a Cartoon with style from the 50’s doing Image Classifi-
cation on a Raspberrry Pi - based on the image uploaded.”

77
Introduction

Image classification is a fundamental task in computer vision that involves categorizing an


image into one of several predefined classes. It’s a cornerstone of artificial intelligence, en-
abling machines to interpret and understand visual information in ways that mimic human
perception.
Image classification is the assignment of a label or category to an entire image based on its visual
content. This task is crucial in computer vision and has numerous applications across various
industries. Image classification’s importance lies in its ability to automate visual understanding
tasks that would otherwise require human intervention.

Applications in Real-World Scenarios

Image classification has found its way into numerous real-world applications, revolutionizing
various sectors:

• Healthcare: Assisting in medical image analysis, such as identifying abnormalities in


X-rays or MRIs.
• Agriculture: Monitoring crop health and detecting plant diseases through aerial imagery.
• Automotive: Enabling advanced driver assistance systems and autonomous vehicles to
recognize road signs, pedestrians, and other vehicles.
• Retail: Powering visual search capabilities and automated inventory management systems.
• Security and Surveillance: Enhancing threat detection and facial recognition systems.
• Environmental Monitoring: Analyzing satellite imagery for deforestation, urban planning,
and climate change studies.

Advantages of Running Classification on Edge Devices like Raspberry Pi

Implementing image classification on edge devices such as the Raspberry Pi offers several
compelling advantages:

1. Low Latency: Processing images locally eliminates the need to send data to cloud servers,
significantly reducing response times.
2. Offline Functionality: Classification can be performed without an internet connection,
making it suitable for remote or connectivity-challenged environments.
3. Privacy and Security: Sensitive image data remains on the local device, addressing data
privacy concerns and compliance requirements.
4. Cost-Effectiveness: Eliminates the need for expensive cloud computing resources, espe-
cially for continuous or high-volume classification tasks.

78
5. Scalability: Enables distributed computing architectures in which multiple devices can
operate independently or in a network.
6. Energy Efficiency: Optimized models on dedicated hardware can be more energy-efficient
than cloud-based solutions, which is crucial for battery-powered or remote applications.
7. Customization: Deploying specialized or frequently updated models tailored to specific
use cases is more manageable.

We can create more responsive, secure, and efficient computer vision solutions by leveraging
the power of edge devices such as the Raspberry Pi for image classification. This approach
opens new possibilities for integrating intelligent visual processing across diverse applications
and environments.
In the following sections, we’ll explore how to implement and optimize image classification
on the Raspberry Pi, leveraging these advantages to build powerful, efficient computer vision
systems.

Setting Up the Environment

Updating the Raspberry Pi

First, ensure your Raspberry Pi is up to date:

sudo apt update


sudo apt upgrade -y
sudo reboot # Reboot to ensure all updates take effect

Installing Required System Level Libraries

Install Python tools and camera libraries

sudo apt install -y python3-pip python3-venv python3-picamera2


sudo apt install -y libcamera-dev libcamera-tools libcamera-apps

Picamera2]([Link] a Python library for interacting with


Raspberry Pi’s camera, is based on the libcamera camera stack, and the Raspberry Pi Foundation
maintains it. The Picamera2 library is supported on all Raspberry Pi models, from the Pi Zero
to the Pi 5.
Testing picamera2 Instalation

79
Setting up a Virtual Environment

Create a virtual environment with access to system packages to manage dependencies:

python3 -m venv ~/tflite_env --system-site-packages

Activate the environment:

source ~/tflite_env/bin/activate

To exit the virtual environment, use:

deactivate

Install Python Packages (inside Virtual Environment)

Ensure you’re in the virtual environment (venv)

pip install numpy pillow # Image processing


pip install matplotlib # For displaying images
pip install opencv-python # Computer vision

Verify installation

pip list | grep -E "(numpy|pillow|opencv|picamera)"

80
System vs pip Package Installation Rule

Use sudo apt install (outside of venv) for:

• System-level dependencies and libraries


• Hardware interface libraries (as a camera)
• Development headers and build tools
• Anything that needs to interface directly with hardware

Use pip install (inside venv) for:

• Pure Python packages


• Application-specific libraries
• Packages that don’t need system-level access

Rule of thumb: Use sudo apt install only for system dependencies and hard-
ware interfaces. Use pip install (without sudo) inside an activated virtual
environment for everything else. Inside the vent, PIP or PIP3 are the same.
The virtual environment will automatically include both system packages and
pip-installed packages thanks to the --system-site-packages flag.

Setting up Jupyter Notebook

Let’s set up Jupyter Notebook optimized for headless Raspberry Pi camera work and develop-
ment:

pip install jupyter jupyterlab notebook


jupyter notebook --generate-config

To run Jupyter Notebook, run the command (change the IP address for yours):

81
jupyter notebook --ip=[Link] --no-browser

On the terminal, you can see the local URL address to open the notebook:

You can access it from another device by entering the Raspberry Pi’s IP address and the
provided token in a web browser (you can copy the token from the terminal).

Define the working directory in the Raspi and create a new Python 3 notebook. For example:

82
cd Documents
mkdir Python

It is possible to create folders directly in the Jupyter Notebook

Create a new Notebook and test the code below:


Import Libraries

import time
import numpy as np
from PIL import Image
import [Link] as plt
from picamera2 import Picamera2

Load an image from the internet, for example (note that it is possible to run a command line
from the Notebook, using ! before the command:

!wget [Link]

An image ([Link] will be downloaded to the current directory.


Load and show the image:

img_path = "[Link]"
img = [Link](img_path)

# Display the image


[Link](figsize=(6, 6))
[Link](img)
[Link]("Original Image")
[Link]()

We can see the image displayed on the Notebook:

83
Now, let’s use the camera to capture a local image:

from picamera2 import Picamera2


import time

# Initialize camera
picam2 = Picamera2()
[Link]()

# Wait for camera to warm up

84
[Link](2)

# Capture image
picam2.capture_file("class3_test.jpg")
print("Image captured: class3_test.jpg")

# Stop camera
[Link]()
[Link]()

And use a similar code as before to show it (adapting the img_path and title):

85
Installing LiteRT

We are interested in inference, which involves running trained models on a device to make
predictions from input data. To perform an inference with a model, we must run it through an
interpreter. For that, we will use LiteRT, Google’s on-device framework for high-performance
ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimiza-
tion.

LiteRT continues TensorFlow Lite’s legacy as the trusted, high-performance runtime


for on-device AI.

LiteRT features advanced GPU/NPU acceleration, delivers superior ML & GenAI performance,
making on-device ML inference easier than ever.
For installation on the Raspi, let’s use the command:

pip install ai-edge-litert

Verifying the instalation:

pip show ai-edge-litert

Creating a working directory:

If you are working on the Raspi-Zero with the minimum OS (No Desktop), you do not have a
user-pre-defined directory tree (you can check it with ls. So, let’s create one:

86
mkdir Documents
cd Documents/
mkdir TFLITE
cd TFLITE/
mkdir IMG_CLASS
cd IMG_CLASS
mkdir models
cd models

On the Raspi-5, the /Documents should be there.

Get a pre-trained Image Classification model:


An appropriate pre-trained model is crucial for successful image classification on resource-
constrained devices like the Raspberry Pi. MobileNet is designed for mobile and embedded
vision applications with a good balance between accuracy and speed. Versions: MobileNetV1,
MobileNetV2, MobileNetV3. Let’s download the V2:

wget [Link]

tar xzf mobilenet_v2_1.0_224_quant.tgz

Get its labels:

wget [Link]

In the end, you should have the models in its directory:

We will only need the mobilenet_v2_1.0_224_quant.tflite model and the


[Link]. We can delete the other files.

87
Verifying the Setup

Let’s test our setup by running a simple Python script on TFLITE/IMG_CLASS folder:

from ai_edge_litert.interpreter import Interpreter


import numpy as np
from PIL import Image

print("NumPy:", np.__version__)
print("Pillow:", Image.__version__)

# Try to create a LiteRT Interpreter


model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
print("LiteRT Interpreter created successfully!")

We can create the Python script using nano on the terminal, saving it with CTRL+0 + ENTER +
CTRL+X

And run it with the command:

88
python setup_test.py

Or you can run it directly on the Notebook:

89
Making inferences with Mobilenet V2

In the last section, we set up the environment, including downloading a popular pre-trained
model, Mobilenet V2, trained on ImageNet’s 224x224 images (1.2 million) for 1,001 classes
(1,000 object categories plus 1 background). The model was converted to a compact 3.5MB
tflite format, making it suitable for the limited storage and memory of a Raspberry Pi.

In the IMG_CLASS working directory, let’s start a new notebook to follow all the steps to classify
one image:
Import the needed libraries:

import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter

Load the TFLite model and allocate tensors:

model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()

INFO: Created TensorFlow Lite XNNPACK delegate for CPU.

The message means LiteRTsuccessfully enabled an optimized CPU backend (XNNPACK) for
our model, which is good and expected.
What XNNPACK is

• XNNPACK is a library of highly optimized operators (conv, FC, etc.) for running neural
networks on CPUs, especially ARM and x86.

90
• LiteRT can “delegate” supported ops to XNNPACK so they run using these faster kernels
instead of the default reference CPU implementation.

So, it means the interpreter has attached the XNNPACK delegate and will run all compatible
parts of the graph on the CPU using it.

• On devices like the Raspberry Pi, this usually results in lower inference latency at the
cost of slightly longer delivery time and a bit more RAM for packed weights.
• We are currently using CPU acceleration, not GPU, which is the standard/optimal path
for many TFLite/LiteRT models on Pi-class hardware.

Get input and output tensors.

input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

Input details provide information on how the model should be fed an image. The shape of (1,
224, 224, 3) informs us that an image with dimensions (224x224x3) should be input one by one
(Batch Dimension: 1).

The output details indicate that the inference will produce an array of 1,001 integer values.
Those values result from image classification, where each value is the probability that the
corresponding label is associated with the image.

91
Let’s also inspect the dtype of the input details of the model

input_dtype = input_details[0]['dtype']
input_dtype

dtype('uint8')

This shows that the input image should be represented as raw pixels (0-255).
Let’s get a test image. We can either transfer it from our computer or download one for testing,
as we did before. Let’s first create a folder under our working directory:

mkdir images
cd images
wget [Link]

Let’s load and display the image:

# Load the image


img_path = "./images/[Link]"
img = [Link](img_path)

# Display the image


[Link](figsize=(8, 8))
[Link](img)
[Link]("Original Image")
[Link]()

92
We can see the image size by running the command:

width, height = [Link]

That shows that the image is an RGB image with a width and height of 1600 pixels each. To
use our model, we should reshape it to (224, 224, 3) and add a batch dimension of 1, as defined
in the input details: (1, 224, 224, 3). The inference result, as shown in the output details, will
be an array of size 1001, as shown below:

93
So, let’s reshape the image, add the batch dimension, and see the result:

img = [Link]((input_details[0]['shape'][1], input_details[0]['shape'][2]))


input_data = np.expand_dims(img, axis=0)
input_data.shape

The input_data shape is as expected: (1, 224, 224, 3)


Let’s confirm the dtype of the input data:

input_data.dtype

dtype('uint8')
The input data dtype is ‘uint8’, which is compatible with the dtype expected for the model.
Using the input_data, let’s run the interpreter and get the predictions (output):

interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
predictions = interpreter.get_tensor(output_details[0]['index'])[0]

The prediction is an array with 1001 elements. Let’s get the Top-5 indices where their elements
have high values:

top_k_results = 5
top_k_indices = [Link](predictions)[::-1][:top_k_results]
top_k_indices

94
The top_k_indices is an array with 5 elements: array([283, 286, 282])
So, 283, 286, 282, 288, and 479 are the image’s most probable classes. Having the index, we
must find to which class it belongs (such as car, cat, or dog). The text file downloaded with
the model includes a label for each index from 0 to 1,000. Let’s use a function to load the .txt
file as a list:

def load_labels(filename):
with open(filename, 'r') as f:
return [[Link]() for line in [Link]()]

And get the list, printing the labels associated with the indexes:

labels_path = "./models/[Link]"
labels = load_labels(labels_path)

print(labels[286])
print(labels[283])
print(labels[282])
print(labels[288])
print(labels[479])

As a result, we have:

Egyptian cat
tiger cat
tabby
lynx
carton

At least four of the top indices are related to felines. The prediction content is the probability
associated with each one of the labels. As we saw in the output details, those values are
quantized and should be dequantized:

scale, zero_point = output_details[0]['quantization']


dequantized_output = ([Link](np.float32) - zero_point) * scale
dequantized_output

array([1.8662813e-06, 3.0599874e-06, 1.8146475e-05, ..., 6.9421253e-07,


2.0032754e-05, 4.2967865e-04], shape=(1001,), dtype=float32)

95
The output (positive and negative numbers) shows that the output probably does not have a
Softmax. Checking the model documentation ([Link] MobileNet
V2 typically doesn’t include a softmax layer at the output. It usually ends with a 1x1 convolution
followed by average pooling and a fully connected layer. So, for getting the probabilities (0 to
1), we should apply Softmax:

exp_output = [Link](dequantized_output - [Link](dequantized_output))


probabilities = exp_output / [Link](exp_output)

Let’s print the top-5 probabilities:

print (probabilities[286])
print (probabilities[283])
print (probabilities[282])
print (probabilities[288])
print (probabilities[479])

0.265947
0.39499295
0.17906114
0.08961108
0.022443123

For clarity, let’s create a function to relate the labels to the probabilities:

for i in range(top_k_results):
print("\t{:20}: {}%".format(
labels[top_k_indices[i]],
(int(probabilities[top_k_indices[i]]*100))))

tiger cat : 39%


Egyptian cat : 26%
tabby : 17%
lynx : 8%
carton : 2%

Define a general Image Classification function

Let’s create a general function to give an image as input, and we get the Top-5 possible
classes:

96
def image_classification(img_path, model_path, labels, top_k_results=5):
# load the image
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')

# Load the LiteRT model


interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()

# Get input and output tensors


input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

# Preprocess
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims(img, axis=0)

# Inference on Raspi
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()

# Obtain results and map them to the classes


predictions = interpreter.get_tensor(output_details[0]['index'])[0]

# Get indices of the top k results


top_k_indices = [Link](predictions)[::-1][:top_k_results]

# Get quantization parameters


scale, zero_point = output_details[0]['quantization']

# Dequantize the output and apply softmax


dequantized_output = ([Link](np.float32) - zero_point) * scale
exp_output = [Link](dequantized_output - [Link](dequantized_output))
probabilities = exp_output / [Link](exp_output)

print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {}%".format(
labels[top_k_indices[i]],

97
(int(probabilities[top_k_indices[i]]*100))))

And loading some images for testing, we have:

Testing the classification with the Camera

Let’s modify the Python script used before to capture an image from the camera (size: 224x224),
saving it in the images folder:

from picamera2 import Picamera2


import time

def capture_image(image_path):

# Initialize camera
picam2 = Picamera2() # default is index 0

# Configure the camera


config = picam2.create_still_configuration(main={"size": (224, 224)})
[Link](config)
[Link]()

98
# Wait for camera to warm up
[Link](2)

# Capture image
picam2.capture_file(image_path)
print("Image captured: "+"image_path")

# Stop camera
[Link]()
[Link]()

Now, let’s capture an image and sent it for inference:

img_path = './images/cam_img_test.jpg'
model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
labels = load_labels("./models/[Link]")
capture_image(img_path)
image_classification(img_path, model_path, labels, top_k_results=5)

99
Exploring a Model Trained from Zero

Let’s get a TFLite model trained from scratch. For that, we can follow the Notebook:

100
CNN to classify Cifar-10 dataset
In the notebook, we trained a model using the CIFAR10 dataset, which contains 60,000 images
from 10 classes of CIFAR (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and
truck). CIFAR has 32x32 color images (3 color channels) where the objects are not centered
and can have the object with a background, such as airplanes that might have a cloudy sky
behind them! In short, small but real images.
The CNN trained model (cifar10_model.keras) had a size of 2.0MB. Using the TFLite Converter,
the model [Link] became with 674MB (around 1/3 of the original size).

Runing the notebook 20_Cifar_10_Image_Classification.ipynb with the trained model:


[Link], and following the same steps as we did with the MobileNet model, we
can see that the inference result is inferior in terms of probability when compared with the
MobileNetV2.
Below are examples of images using the General Function for Image Classification on a
Raspberry Pi Zero, as shown in the last section.

Conclusion:

This chapter has established a solid foundation for understanding and implementing image
classification on Raspberry Pi devices using Python and LiteRT. Throughout this journey, we

101
have explored the essential components that make edge-based computer vision both practical
and powerful.
We began by understanding the theoretical foundations of image classification and its real-
world applications across diverse sectors, from healthcare to environmental monitoring. The
advantages of running classification on edge devices like the Raspberry Pi—including low
latency, offline functionality, enhanced privacy, and cost-effectiveness—make it an attractive
solution for many practical applications.
The hands-on experience of setting up the development environment provided crucial insights
into the requirements and constraints of embedded systems. We successfully configured LiteRT,
installed essential Python libraries, and established a working directory structure that serves
as the foundation for computer vision projects.
Working with the pre-trained MobileNet V2 model demonstrated several key concepts:

• Model Architecture Understanding: We explored how pre-trained models like


MobileNet V2 are optimized for mobile and embedded applications, achieving an excellent
balance between accuracy and computational efficiency.
• Quantization Benefits: The 3.5MB quantized model showed how compression tech-
niques make sophisticated neural networks feasible on resource-constrained devices without
significant accuracy loss.
• Inference Pipeline: We implemented the complete inference workflow, from image pre-
processing (resizing to 224×224, handling data types) to post-processing (dequantization,
softmax application, and top-k prediction extraction).
• Performance Considerations: The chapter highlighted the importance of understand-
ing model input/output specifications, memory management, and the trade-offs between
accuracy and speed on edge devices.

The practical implementation using cameras connected to a Raspberry Pi demonstrated the


seamless integration between hardware and software components. The ability to capture
images directly from the device and perform real-time classification showcases the potential for
autonomous and IoT applications.
Key technical achievements include:

• Successfully setting up LiteRT and dependencies


• Implementing proper image preprocessing pipelines
• Understanding quantized model operations and dequantization processes
• Creating reusable functions for image classification tasks
• Integrating camera capture with inference workflows

This foundational knowledge prepares us for more advanced topics, including custom model
training and deployment. The skills developed here—understanding model architectures,
implementing inference pipelines, and working with embedded Python environments—are
transferable to a wide range of computer vision applications.

102
The chapter serves as a stepping stone toward building more sophisticated AI systems on edge
devices, demonstrating that powerful computer vision capabilities are accessible even on modest
hardware platforms when properly optimized and implemented.

Resources

• Dataset Example
• Setup Test Notebook on a Raspi
• Image Classification Notebook on a Raspi
• CNN to classify Cifar-10 dataset at CoLab
• Cifar 10 - Image Classification on a Raspi
• Python Scripts

103
104
Custom Image Classification Project

Figure 3: DALL·E prompt - A cover image for an ‘Image Classification’ chapter in a Raspberry
Pi tutorial, designed in the same vintage 1950s electronics lab style as previous covers.
The scene should feature a Raspberry Pi connected to a camera module, with the
camera capturing a photo of the small blue robot provided by the user. The robot should
be placed on a workbench, surrounded by classic lab tools like soldering irons, resistors,
and wires. The lab background should include vintage equipment like oscilloscopes
and tube radios, maintaining the detailed and nostalgic feel of the era. No text or
logos should be included.
105
Image Classification Project

In this chapter, we will develop a complete Image Classification project using the Edge Impulse
Studio. As we did with the MobiliNet V2, the trained and converted TFLite model will be
used for inference using a Python script.
Here is a typical ML workflow that we will use in our project:

The Goal

The first step in any ML project is to define its goal. In this case, it is to detect and classify
two specific objects present in one image. For this project, we will use two small toys: a robot
and a small Brazilian parrot (named Periquito). We will also collect images of a background
where those two objects are absent.

Data Collection

Once we have defined our Machine Learning project goal, the next and most crucial step is
collecting the dataset. We can use a phone for the image capture, but we will use the Raspi
here. Let’s set up a simple web server on our Raspberry Pi to view the QVGA (320 x 240)
captured images in a browser.

106
1. First, let’s install Flask, a lightweight web framework for Python:
pip install flask

2. Go to the working folder (IMG_CLASS) and create a new Python script combining image
capture with a web server. We’ll call it get_img_data.py:

from flask import Flask, Response, render_template_string,


request, redirect, url_for
from picamera2 import Picamera2
import io
import threading
import time
import os
import signal

app = Flask(__name__)

# Global variables
base_dir = "dataset"
picam2 = None
frame = None
frame_lock = [Link]()
capture_counts = {}
current_label = None
shutdown_event = [Link]()

def initialize_camera():
global picam2
picam2 = Picamera2()
config = picam2.create_preview_configuration(
main={"size": (320, 240)}
)
[Link](config)
[Link]()
[Link](2) # Wait for camera to warm up

def get_frame():
global frame
while not shutdown_event.is_set():
stream = [Link]()
picam2.capture_file(stream, format='jpeg')
with frame_lock:

107
frame = [Link]()
[Link](0.1) # Adjust as needed for smooth preview

def generate_frames():
while not shutdown_event.is_set():
with frame_lock:
if frame is not None:
yield (b'--frame\r\n'
b'Content-Type: image/jpeg\r\n\r\n' +
frame + b'\r\n')
[Link](0.1) # Adjust as needed for smooth streaming

def shutdown_server():
shutdown_event.set()
if picam2:
[Link]()
# Give some time for other threads to finish
[Link](2)
# Send SIGINT to the main process
[Link]([Link](), [Link])

@[Link]('/', methods=['GET', 'POST'])


def index():
global current_label
if [Link] == 'POST':
current_label = [Link]['label']
if current_label not in capture_counts:
capture_counts[current_label] = 0
[Link]([Link](base_dir, current_label),
exist_ok=True)
return redirect(url_for('capture_page'))
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture - Label Entry</title>
</head>
<body>
<h1>Enter Label for Dataset</h1>
<form method="post">
<input type="text" name="label" required>
<input type="submit" value="Start Capture">

108
</form>
</body>
</html>
''')

@[Link]('/capture')
def capture_page():
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture</title>
<script>
var shutdownInitiated = false;
function checkShutdown() {
if (!shutdownInitiated) {
fetch('/check_shutdown')
.then(response => [Link]())
.then(data => {
if ([Link]) {
shutdownInitiated = true;
[Link](
'video-feed').src = '';
[Link](
'shutdown-message')
.[Link] = 'block';
}
});
}
}
setInterval(checkShutdown, 1000); // Check
every second
</script>
</head>
<body>
<h1>Dataset Capture</h1>
<p>Current Label: {{ label }}</p>
<p>Images captured for this label: {{ capture_count
}}</p>
<img id="video-feed" src="{{ url_for('video_feed')
}}" width="640"
height="480" />

109
<div id="shutdown-message" style="display: none;
color: red;">
Capture process has been stopped.
You can close this window.
</div>
<form action="/capture_image" method="post">
<input type="submit" value="Capture Image">
</form>
<form action="/stop" method="post">
<input type="submit" value="Stop Capture"
style="background-color: #ff6666;">
</form>
<form action="/" method="get">
<input type="submit" value="Change Label"
style="background-color: #ffff66;">
</form>
</body>
</html>
''', label=current_label, capture_count=capture_counts.get(
current_label, 0))

@[Link]('/video_feed')
def video_feed():
return Response(generate_frames(),
mimetype='multipart/x-mixed-replace;
boundary=frame')

@[Link]('/capture_image', methods=['POST'])
def capture_image():
global capture_counts
if current_label and not shutdown_event.is_set():
capture_counts[current_label] += 1
timestamp = [Link]("%Y%m%d-%H%M%S")
filename = f"image_{timestamp}.jpg"
full_path = [Link](base_dir, current_label,
filename)

picam2.capture_file(full_path)

return redirect(url_for('capture_page'))

@[Link]('/stop', methods=['POST'])

110
def stop():
summary = render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture - Stopped</title>
</head>
<body>
<h1>Dataset Capture Stopped</h1>
<p>The capture process has been stopped.
You can close this window.</p>
<p>Summary of captures:</p>
<ul>
{% for label, count in capture_counts.items() %}
<li>{{ label }}: {{ count }} images</li>
{% endfor %}
</ul>
</body>
</html>
''', capture_counts=capture_counts)

# Start a new thread to shutdown the server


[Link](target=shutdown_server).start()

return summary

@[Link]('/check_shutdown')
def check_shutdown():
return {'shutdown': shutdown_event.is_set()}

if __name__ == '__main__':
initialize_camera()
[Link](target=get_frame, daemon=True).start()
[Link](host='[Link]', port=5000, threaded=True)

3. Run this script:

python get_img_data.py

4. Access the web interface:

• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]

111
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your
Raspberry Pi’s IP address). For example: [Link]

This Python script creates a web-based interface for capturing and organizing image datasets
using a Raspberry Pi and its camera. It’s handy for machine learning projects that require
labeled image data.

Key Features:

1. Web Interface: Accessible from any device on the same network as the Raspberry Pi.
2. Live Camera Preview: This shows a real-time feed from the camera.
3. Labeling System: Allows users to input labels for different categories of images.
4. Organized Storage: Automatically saves images in label-specific subdirectories.
5. Per-Label Counters: Keeps track of how many images are captured for each label.
6. Summary Statistics: Provides a summary of captured images when stopping the
capture process.

Main Components:

1. Flask Web Application: Handles routing and serves the web interface.
2. Picamera2 Integration: Controls the Raspberry Pi camera.
3. Threaded Frame Capture: Ensures smooth live preview.
4. File Management: Organizes captured images into labeled directories.

Key Functions:

• initialize_camera(): Sets up the Picamera2 instance.


• get_frame(): Continuously captures frames for the live preview.
• generate_frames(): Yields frames for the live video feed.
• shutdown_server(): Sets the shutdown event, stops the camera, and shuts down the
Flask server
• index(): Handles the label input page.
• capture_page(): Displays the main capture interface.
• video_feed(): Shows a live preview to position the camera
• capture_image(): Saves an image with the current label.
• stop(): Stops the capture process and displays a summary.

112
Usage Flow:

1. Start the script on your Raspberry Pi.


2. Access the web interface from a browser.
3. Enter a label for the images you want to capture and press Start Capture.

4. Use the live preview to position the camera.


5. Click Capture Image to save images under the current label.

6. Change labels as needed for different categories, selecting Change Label.

113
7. Click Stop Capture when finished to see a summary.

Technical Notes:

• The script uses threading to handle concurrent frame capture and web serving.
• Images are saved with timestamps in their filenames for uniqueness.
• The web interface is responsive and can be accessed from mobile devices.

Customization Possibilities:

• Adjust image resolution in the initialize_camera() function. Here we used QVGA


(320 × 240).
• Modify the HTML templates for a different look and feel.
• Add additional image processing or analysis steps in the capture_image() function.

Number of samples on Dataset:

Get around 60 images from each category (periquito, robot and background). Try to capture
different angles, backgrounds, and light conditions.
On the Raspi, we will end with a folder named dataset, which contains three sub-folders:
periquito, robot, and background, one for each class of images.
You can use Filezilla to transfer the created dataset to your main computer.

114
Training the model with Edge Impulse Studio

We will use the Edge Impulse Studio to train our model. Go to the Edge Impulse Page, enter
your account credentials, and create a new project:

Here, you can clone a similar project: Raspi - Img Class.

Dataset

We will walk through four main steps using the EI Studio (or Studio). These steps are crucial
in preparing our model for use on the Raspi: Dataset, Impulse, Tests, and Deploy (on the Edge
Device, in this case, the Raspi).

Regarding the Dataset, it is essential to point out that our Original Dataset,
captured with the Raspi, will be split into Training, Validation, and Test. The Test
Set will be separated from the beginning and reserved for use only in the Test phase
after training. The Validation Set will be used during training.

On Studio, follow the steps to upload the captured data:

1. Go to the Data acquisition tab, and in the UPLOAD DATA section, upload the files from
your computer in the chosen categories.
2. Leave to the Studio the splitting of the original dataset into train and test and choose
the label about
3. Repeat the procedure for all three classes. At the end, you should see your “raw data” in
the Studio:

115
The Studio allows you to explore your data, showing a complete view of all the data in your
project. You can clear, inspect, or change labels by clicking on individual data items. In our
case, a straightforward project, the data seems OK.

The Impulse Design

In this phase, we should define how to:

116
• Pre-process our data, which consists of resizing the individual images and determining
the color depth to use (be it RGB or Grayscale) and
• Specify a Model. In this case, it will be the Transfer Learning (Images) to fine-tune a
pre-trained MobileNet V2 image classification model on our data. This method performs
well even with relatively small image datasets (around 180 images in our case).

Transfer Learning with MobileNet offers a streamlined approach to model training, which is
especially beneficial for resource-constrained environments and projects with limited labeled
data. MobileNet, known for its lightweight architecture, is a pre-trained model that has already
learned valuable features from a large dataset (ImageNet).

By leveraging these learned features, we can train a new model for your specific task with fewer
data and computational resources and achieve competitive accuracy.

117
This approach significantly reduces training time and computational cost, making it ideal for
quick prototyping and deployment on embedded devices where efficiency is paramount.
Go to the Impulse Design Tab and create the impulse, defining an image size of 160 × 160 and
squashing them (squared form, without cropping). Select Image and Transfer Learning blocks.
Save the Impulse.

Image Pre-Processing

All the input QVGA/RGB565 images will be converted to 76,800 features (160 × 160 × 3).

118
Press Save parameters and select Generate features in the next tab.

Model Design

MobileNet is a family of efficient convolutional neural networks designed for mobile and
embedded vision applications. The key features of MobileNet are:

1. Lightweight: Optimized for mobile devices and embedded systems with limited computa-
tional resources.
2. Speed: Fast inference times, suitable for real-time applications.

119
3. Accuracy: Maintains good accuracy despite its compact size.

MobileNetV2, introduced in 2018, improves the original MobileNet architecture. Key features
include:

1. Inverted Residuals: Inverted residual structures are used where shortcut connections are
made between thin bottleneck layers.
2. Linear Bottlenecks: Removes non-linearities in the narrow layers to prevent the destruction
of information.
3. Depth-wise Separable Convolutions: Continues to use this efficient operation from Mo-
bileNetV1.

In our project, we will do a Transfer Learning with the MobileNetV2 160x160 1.0, which
means that the images used for training (and future inference) should have an input Size of
160 × 160 pixels and a Width Multiplier of 1.0 (full width, not reduced). This configuration
balances between model size, speed, and accuracy.

Model Training

Another valuable deep learning technique is Data Augmentation. Data augmentation


improves the accuracy of machine learning models by creating additional artificial data. A data
augmentation system makes small, random changes to the training data during the training
process (such as flipping, cropping, or rotating the images).
Looking under the hood, here you can see how Edge Impulse implements a data Augmentation
policy on your data:

# Implements the data augmentation policy


def augment_image(image, label):
# Flips the image randomly
image = [Link].random_flip_left_right(image)

# Increase the image size, then randomly crop it down to


# the original dimensions
resize_factor = [Link](1, 1.2)
new_height = [Link](resize_factor * INPUT_SHAPE[0])
new_width = [Link](resize_factor * INPUT_SHAPE[1])
image = [Link].resize_with_crop_or_pad(image, new_height,
new_width)
image = [Link].random_crop(image, size=INPUT_SHAPE)

# Vary the brightness of the image


image = [Link].random_brightness(image, max_delta=0.2)

120
return image, label

Exposure to these variations during training can help prevent your model from taking shortcuts
by “memorizing” superficial clues in your training data, meaning it may better reflect the deep
underlying patterns in your dataset.
The final dense layer of our model will have 0 neurons with a 10% dropout for overfitting
prevention. Here is the Training result:

The result is excellent, with a reasonable 35 ms of latency (for a Raspi-4), which should result
in around 30 fps (frames per second) during inference. A Raspi-Zero should be slower, and the
Raspi-5, faster.

Trading off: Accuracy versus speed

If faster inference is needed, we should train the model using smaller alphas (0.35, 0.5, and
0.75) or even reduce the image input size, trading with accuracy. However, reducing the input
image size and decreasing the alpha (width multiplier) can speed up inference for MobileNet
V2, but they have different trade-offs. Let’s compare:

1. Reducing Image Input Size:

Pros:

121
• Significantly reduces the computational cost across all layers.
• Decreases memory usage.
• It often provides a substantial speed boost.

Cons:

• It may reduce the model’s ability to detect small features or fine details.
• It can significantly impact accuracy, especially for tasks requiring fine-grained recognition.

2. Reducing Alpha (Width Multiplier):

Pros:

• Reduces the number of parameters and computations in the model.


• Maintains the original input resolution, potentially preserving more detail.
• It can provide a good balance between speed and accuracy.

Cons:

• It may not speed up inference as dramatically as reducing input size.


• It can reduce the model’s capacity to learn complex features.

Comparison:

1. Speed Impact:
• Reducing input size often provides a more substantial speed boost because it reduces
computations quadratically (halving both width and height reduces computations
by about 75%).
• Reducing alpha provides a more linear reduction in computations.
2. Accuracy Impact:
• Reducing input size can severely impact accuracy, especially when detecting small
objects or fine details.
• Reducing alpha tends to have a more gradual impact on accuracy.
3. Model Architecture:
• Changing input size doesn’t alter the model’s architecture.
• Changing alpha modifies the model’s structure by reducing the number of channels
in each layer.

Recommendation:

1. If our application doesn’t require detecting tiny details and can tolerate some loss in
accuracy, reducing the input size is often the most effective way to speed up inference.

122
2. Reducing alpha might be preferable if maintaining the ability to detect fine details is
crucial or if you need a more balanced trade-off between speed and accuracy.
3. For best results, you might want to experiment with both:
• Try MobileNet V2 with input sizes like 160 × 160 or 92 × 92
• Experiment with alpha values like 1.0, 0.75, 0.5 or 0.35.
4. Always benchmark the different configurations on your specific hardware and with your
particular dataset to find the optimal balance for your use case.

Remember, the best choice depends on your specific requirements for accuracy, speed,
and the nature of the images you’re working with. It’s often worth experimenting
with combinations to find the optimal configuration for your particular use case.

Model Testing

Now, you should take the data set aside at the start of the project and run the trained model
using it as input. Again, the result is excellent (92.22%).

Deploying the model

As we did in the previous section, we can deploy the trained model as .tflite and use Raspi to
run it using Python.
On the Dashboard tab, go to Transfer learning model (int8 quantized) and click on the download
icon:

123
Let’s also download the float32 version for comparison

Transfer the models from your computer to the Raspi (./models), for example, using FileZilla.
Also, capture some images for inference and save them in (./images), or use the images in the
./dataset folder.
Let’s remember what we did in the last chapter:
Activate the environment:

source ~/tflite_env/bin/activate

Run a Jupyter Notebook, using the command:

jupyter notebook --ip=[Link] --no-browser

Change the IP address for yours

Open a new notebook and enter with the code below:


Import the needed libraries:

124
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter

Define the paths and labels:

img_path = "./images/[Link]"
model_path = "./models/ei-raspi-img-class-int8-quantized-\
[Link]"
labels = ['background', 'periquito', 'robot']

Note that the models trained on the Edge Impulse Studio will output values with
index 0, 1, 2, etc., where the actual labels will follow an alphabetic order.

Load the model, allocate the tensors, and get the input and output tensor details:

# Load the LiteRT model


interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()

# Get input and output tensors


input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

One important difference to note is that the dtype of the input details of the model is now
int8, which means that the input values go from –128 to +127, while each pixel of our image
goes from 0 to 255. This means that we should pre-process the image to match it. We can
check here:

input_dtype = input_details[0]['dtype']
input_dtype

numpy.int8

So, let’s open the image and show it:

125
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
[Link]()

And perform the pre-processing:

scale, zero_point = input_details[0]['quantization']


img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (
(img_array / scale + zero_point)
.clip(-128, 127)
.astype(np.int8)
)
input_data = np.expand_dims(img_array, axis=0)

Checking the input data, we can verify that the input tensor is compatible with what is expected
by the model:

126
input_data.shape, input_data.dtype

((1, 160, 160, 3), dtype('int8'))

Now, it is time to perform the inference. Let’s also calculate the latency of the model:

# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert
# to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))

The model will take around 125ms to perform the inference in the Raspi-Zero, which is 3 to 4
times longer than a Raspi-5.
Now, we can get the output labels and probabilities. It is also important to note that the
model trained on the Edge Impulse Studio has a softmax activation function in its output
(different from the original Movilenet V2), and we can use the model’s raw output as the
“probabilities.”

# Obtain results and map them to the classes


predictions = interpreter.get_tensor(output_details[0]
['index'])[0]

# Get indices of the top k results


top_k_results=3
top_k_indices = [Link](predictions)[::-1][:top_k_results]

# Get quantization parameters


scale, zero_point = output_details[0]['quantization']

# Dequantize the output


dequantized_output = ([Link](np.float32) -
zero_point) * scale
probabilities = dequantized_output

print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):

127
print("\t{:20}: {:.2f}%".format(
labels[top_k_indices[i]],
probabilities[top_k_indices[i]] * 100))

Let’s modify the function created before so that we can handle different type of models:

def image_classification(img_path, model_path, labels,


top_k_results=3, apply_softmax=False):
# Load the image
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')

# Load the LiteRT model


interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()

# Get input and output tensors


input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

# Preprocess
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))

input_dtype = input_details[0]['dtype']

if input_dtype == np.uint8:
input_data = np.expand_dims([Link](img), axis=0)
elif input_dtype == np.int8:
scale, zero_point = input_details[0]['quantization']

128
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (
img_array / scale
+ zero_point
).clip(-128, 127).astype(np.int8)
input_data = np.expand_dims(img_array, axis=0)
else: # float32
input_data = np.expand_dims(
[Link](img, dtype=np.float32),
axis=0
) / 255.0

# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time -
start_time
) * 1000 # Convert to milliseconds

# Obtain results
predictions = interpreter.get_tensor(output_details[0]
['index'])[0]

# Get indices of the top k results


top_k_indices = [Link](predictions)[::-1][:top_k_results]

# Handle output based on type


output_dtype = output_details[0]['dtype']
if output_dtype in [np.int8, np.uint8]:
# Dequantize the output
scale, zero_point = output_details[0]['quantization']
predictions = ([Link](np.float32) -
zero_point) * scale

if apply_softmax:
# Apply softmax
exp_preds = [Link](predictions - [Link](predictions))
probabilities = exp_preds / [Link](exp_preds)
else:
probabilities = predictions

129
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {:.1f}%".format(
labels[top_k_indices[i]],
probabilities[top_k_indices[i]] * 100))
print ("\n\tInference time: {:.1f}ms".format(inference_time))

And test it with different images and the int8 quantized model (160x160 alpha =1.0).

Let’s download a smaller model, such as the one trained for the Nicla Vision Lab (int8 quantized
model, 96x96, alpha = 0.1), as a test. We can use the same function:

The model lost some accuracy, but it is still OK once our model does not look for many details.
Regarding latency, we are aboutt ten times faster on the Raspi-Zero.

With the Raspberry Pi 5, the latencies are respectively 8 ms and 0.8 ms

130
Live Image Classification

Let’s develop an app that captures images with the camera in real-time and displays their
classification.
Using the nano on the terminal, save the code below, such as img_class_live_infer.py.

from flask import Flask, Response, render_template_string,


request, jsonify
from picamera2 import Picamera2
import io
import threading
import time
import numpy as np
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
from queue import Queue

app = Flask(__name__)

# Global variables
picam2 = None
frame = None
frame_lock = [Link]()
is_classifying = False
confidence_threshold = 0.8
model_path = "./models/ei-raspi-img-class-int8-quantized-\
[Link]"
labels = ['background', 'periquito', 'robot']
interpreter = None
classification_queue = Queue(maxsize=1)

def initialize_camera():
global picam2
picam2 = Picamera2()
config = picam2.create_preview_configuration(
main={"size": (320, 240)}
)
[Link](config)
[Link]()
[Link](2) # Wait for camera to warm up

def get_frame():

131
global frame
while True:
stream = [Link]()
picam2.capture_file(stream, format='jpeg')
with frame_lock:
frame = [Link]()
[Link](0.1) # Capture frames more frequently

def generate_frames():
while True:
with frame_lock:
if frame is not None:
yield (
b'--frame\r\n'
b'Content-Type: image/jpeg\r\n\r\n'
+ frame + b'\r\n'
)
[Link](0.1)

def load_model():
global interpreter
if interpreter is None:
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
return interpreter

def classify_image(img, interpreter):


input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims([Link](img), axis=0)\
.astype(input_details[0]['dtype'])

interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()

predictions = interpreter.get_tensor(output_details[0]
['index'])[0]
# Handle output based on type
output_dtype = output_details[0]['dtype']

132
if output_dtype in [np.int8, np.uint8]:
# Dequantize the output
scale, zero_point = output_details[0]['quantization']
predictions = ([Link](np.float32) -
zero_point) * scale
return predictions

def classification_worker():
interpreter = load_model()
while True:
if is_classifying:
with frame_lock:
if frame is not None:
img = [Link]([Link](frame))
predictions = classify_image(img, interpreter)
max_prob = [Link](predictions)
if max_prob >= confidence_threshold:
label = labels[[Link](predictions)]
else:
label = 'Uncertain'
classification_queue.put({
'label': label,
'probability': float(max_prob)
})
[Link](0.1) # Adjust based on your needs

@[Link]('/')
def index():
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Image Classification</title>
<script
src="[Link]
</script>
<script>
function startClassification() {
$.post('/start');
$('#startBtn').prop('disabled', true);
$('#stopBtn').prop('disabled', false);
}

133
function stopClassification() {
$.post('/stop');
$('#startBtn').prop('disabled', false);
$('#stopBtn').prop('disabled', true);
}
function updateConfidence() {
var confidence = $('#confidence').val();
$.post('/update_confidence',
{confidence: confidence}
);
}
function updateClassification() {
$.get('/get_classification', function(data) {
$('#classification').text([Link] + ': '
+ [Link](2));
});
}
$(document).ready(function() {
setInterval(updateClassification, 100);
// Update every 100ms
});
</script>
</head>
<body>
<h1>Image Classification</h1>
<img src="{{ url_for('video_feed') }}"
width="640"
height="480" />

<br>
<button id="startBtn"
onclick="startClassification()">
Start Classification
</button>

<button id="stopBtn"
onclick="stopClassification()"
disabled>
Stop Classification
</button>

<br>

134
<label for="confidence">Confidence Threshold:</label>
<input type="number"
id="confidence"
name="confidence"
min="0" max="1"
step="0.1"
value="0.8"
onchange="updateConfidence()" />

<br>
<div id="classification">
Waiting for classification...
</div>

</body>
</html>
''')

@[Link]('/video_feed')
def video_feed():
return Response(
generate_frames(),
mimetype='multipart/x-mixed-replace; boundary=frame'
)

@[Link]('/start', methods=['POST'])
def start_classification():
global is_classifying
is_classifying = True
return '', 204

@[Link]('/stop', methods=['POST'])
def stop_classification():
global is_classifying
is_classifying = False
return '', 204

@[Link]('/update_confidence', methods=['POST'])
def update_confidence():
global confidence_threshold
confidence_threshold = float([Link]['confidence'])
return '', 204

135
@[Link]('/get_classification')
def get_classification():
if not is_classifying:
return jsonify({'label': 'Not classifying',
'probability': 0})
try:
result = classification_queue.get_nowait()
except [Link]:
result = {'label': 'Processing', 'probability': 0}
return jsonify(result)

if __name__ == '__main__':
initialize_camera()
[Link](target=get_frame, daemon=True).start()
[Link](target=classification_worker,
daemon=True).start()
[Link](host='[Link]', port=5000, threaded=True)

On the terminal, run:

python img_class_live_infer.py

And access the web interface:

• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Raspberry
Pi’s IP address). For example: [Link]

Here are some screenshots of the app running on an external desktop

136
Here, you can see the app running on YouTube.
The code creates a web application for real-time image classification using a Raspberry Pi,
its camera module, and a TensorFlow Lite model. The application uses Flask to serve a web
interface where is possible to view the camera feed and see live classification results.

Key Components:

1. Flask Web Application: Serves the user interface and handles requests.
2. PiCamera2: Captures images from the Raspberry Pi camera module.
3. LiteRT: Runs the image classification model.
4. Threading: Manages concurrent operations for smooth performance.

Main Features:

• Live camera feed display


• Real-time image classification
• Adjustable confidence threshold
• Start/Stop classification on demand

Code Structure:

1. Imports and Setup:


• Flask for web application
• PiCamera2 for camera control
• TensorFlow Lite for inference
• Threading and Queue for concurrent operations
2. Global Variables:

137
• Camera and frame management
• Classification control
• Model and label information
3. Camera Functions:
• initialize_camera(): Sets up the PiCamera2
• get_frame(): Continuously captures frames
• generate_frames(): Yields frames for the web feed
4. Model Functions:
• load_model(): Loads the TFLite model
• classify_image(): Performs inference on a single image
5. Classification Worker:
• Runs in a separate thread
• Continuously classifies frames when active
• Updates a queue with the latest results
6. Flask Routes:
• /: Serves the main HTML page
• /video_feed: Streams the camera feed
• /start and /stop: Controls classification
• /update_confidence: Adjusts the confidence threshold
• /get_classification: Returns the latest classification result
7. HTML Template:
• Displays camera feed and classification results
• Provides controls for starting/stopping and adjusting settings
8. Main Execution:
• Initializes camera and starts necessary threads
• Runs the Flask application

Key Concepts:

1. Concurrent Operations: Using threads to handle camera capture and classification


separately from the web server.
2. Real-time Updates: Frequent updates to the classification results without page reloads.
3. Model Reuse: Loading the TFLite model once and reusing it for efficiency.
4. Flexible Configuration: Allowing users to adjust the confidence threshold on the fly.

138
Usage:

1. Ensure all dependencies are installed.


2. Run the script on a Raspberry Pi with a camera module.
3. Access the web interface from a browser using the Raspberry Pi’s IP address.
4. Start classification and adjust settings as needed.

Summary:

Image classification has emerged as a powerful and versatile application of machine learning,
with significant implications for various fields, from healthcare to environmental monitoring.
This chapter has demonstrated how to implement a robust image classification system on
edge devices like the Raspi-Zero and Raspi-5, showcasing the potential for real-time, on-device
intelligence.
We’ve explored the entire pipeline of an image classification project, from data collection and
model training using Edge Impulse Studio to deploying and running inferences on a Raspi.
The process highlighted several key points:

1. The importance of proper data collection and preprocessing for training effective models.
2. The power of transfer learning, allowing us to leverage pre-trained models like MobileNet
V2 for efficient training with limited data.
3. The trade-offs between model accuracy and inference speed, especially crucial for edge
devices.
4. The implementation of real-time classification using a web-based interface, demonstrating
practical applications.

The ability to run these models on edge devices like the Raspi opens up numerous possibilities
for IoT applications, autonomous systems, and real-time monitoring solutions. It allows for
reduced latency, improved privacy, and operation in environments with limited connectivity.
As we’ve seen, even with the computational constraints of edge devices, it’s possible to achieve
impressive results in terms of both accuracy and speed. The flexibility to adjust model
parameters, such as input size and alpha values, allows for fine-tuning to meet specific project
requirements.
Looking forward, the field of edge AI and image classification continues to evolve rapidly.
Advances in model compression techniques, hardware acceleration, and more efficient neural
network architectures promise to further expand the capabilities of edge devices in computer
vision tasks.
This project serves as a foundation for more complex computer vision applications and encour-
ages further exploration into the exciting world of edge AI and IoT. Whether it’s for industrial

139
automation, smart home applications, or environmental monitoring, the skills and concepts
covered here provide a solid starting point for a wide range of innovative projects.

Resources

• Dataset Example
• Python Scripts
• Edge Impulse Project
• Image Classification Project - Edge Impulse Notebook

140
141
Object Detection: Fundamentals

Figure 4: DALL·E prompt - A cover image for an ‘Object Detection’ chapter in a Raspberry Pi
tutorial, designed in the same vintage 1950s electronics lab style as previous covers.
The scene should prominently feature wheels and cubes, similar to those provided by
the user, placed on a workbench in the foreground. A Raspberry Pi with a connected
camera module should be capturing an image of these objects. Surround the scene
with classic lab tools like soldering irons, resistors, and wires. The lab background
should include vintage equipment like oscilloscopes and tube radios, maintaining the
detailed and nostalgic feel of the era. No text or logos should be included.
142
Introduction

Building on our exploration of image classification, we now turn to a more advanced computer
vision task: object detection. While image classification assigns a single label to an entire
image, object detection goes further by identifying and locating multiple objects within a single
image. This capability opens up many new applications and challenges, particularly in edge
computing and IoT devices like the Raspberry Pi.
Object detection combines classification and localization. It not only determines which objects
are present in an image but also pinpoints their locations, for example, by drawing bounding
boxes around them. This added complexity makes object detection a more powerful tool
for understanding visual scenes, but it also requires more sophisticated models and training
techniques.
In edge AI, where computational resources are constrained, implementing efficient object
detection models is crucial. The challenges we faced with image classification—balancing model
size, inference speed, and accuracy—are even more pronounced in object detection. However,
the rewards are also more significant, as object detection enables more nuanced and detailed
analysis of visual data.
Some applications of object detection on edge devices include:

1. Surveillance and security systems


2. Autonomous vehicles and drones
3. Industrial quality control
4. Wildlife monitoring
5. Augmented reality applications

As we put our hands into object detection, we’ll build on the concepts and techniques we
explored in image classification. We’ll examine popular object detection architectures designed
for efficiency, such as:

• Single Stage Detectors, such as MobileNet and EfficientDet,


• FOMO (Faster Objects, More Objects), and
• YOLO (You Only Look Once).

To learn more about object detection models, follow the tutorial A Gentle Intro-
duction to Object Recognition With Deep Learning.

We will explore those object detection models using:

• TensorFlow Lite Runtime (now changed to LiteRT),


• Edge Impulse Linux Python SDK and
• Ultralitics

143
Throughout this lab, we’ll cover the fundamentals of object detection and how it differs from
image classification. We’ll also learn how to train, fine-tune, test, optimize, and deploy popular
object detection architectures using a dataset created from scratch.

Object Detection Fundamentals

Object detection builds upon the foundations of image classification but extends its capabilities
significantly. To understand object detection, it’s crucial first to recognize its key differences
from image classification:

Image Classification vs. Object Detection

Image Classification:

• Assigns a single label to an entire image


• Answers the question: “What is this image’s primary object or scene?”
• Outputs a single class prediction for the whole image

Object Detection:

• Identifies and locates multiple objects within an image


• Answers the questions: “What objects are in this image, and where are they located?”
• Outputs multiple predictions, each consisting of a class label and a bounding box

144
To visualize this difference, let’s consider an example:

This diagram illustrates the critical difference: image classification provides a single label for
the entire image, while object detection identifies multiple objects, their classes, and their
locations within the image.

Key Components of Object Detection

Object detection systems typically consist of two main components:

1. Object Localization: This component identifies the location of objects within the image.
It typically outputs bounding boxes, rectangular regions encompassing each detected
object.
2. Object Classification: This component determines the class or category of each detected
object, similar to image classification but applied to each localized region.

Challenges in Object Detection

Object detection presents several challenges beyond those of image classification:

• Multiple objects: An image may contain multiple objects of various classes, sizes, and
positions.
• Varying scales: Objects can appear at different sizes within the image.
• Occlusion: Objects may be partially hidden or overlapping.

145
• Background clutter: Distinguishing objects from complex backgrounds can be challenging.
• Real-time performance: Many applications require fast inference times, especially on edge
devices.

Approaches to Object Detection

There are two main approaches to object detection:

1. Two-stage detectors: These first propose regions of interest and then classify each region.
Examples include R-CNN and its variants (Fast R-CNN, Faster R-CNN).
2. Single-stage detectors: These predict bounding boxes (or centroids) and class probabilities
in a single forward pass through the network. Examples include YOLO (You Only Look
Once), EfficientDet, SSD (Single Shot Detector), and FOMO (Faster Objects, More
Objects). These are often faster and better suited to edge devices, such as the Raspberry
Pi.

Evaluation Metrics

Object detection uses different metrics compared to image classification:

• Intersection over Union (IoU) is a metric used to evaluate the accuracy of an object
detector. It measures the overlap between two bounding boxes: the Ground Truth
box (the manually labeled correct box) and the Predicted box (the box generated by
the object detection model). The IoU value is calculated by dividing the area of the
Intersection (the overlapping area) by the area of the Union (the total area covered by
both boxes). A higher IoU value indicates a better prediction.

146
• Mean Average Precision (mAP) is a widely used metric for evaluating the perfor-
mance of object detection models. It provides a single number that reflects a model’s
ability to accurately both classify and localize objects. The “mean” in mAP refers to
the average taken over all object classes in the dataset. The “average precision” (AP) is
calculated for each class, and then these AP values are averaged to get the final mAP
score. A high mAP score indicates that the model is excellent at identifying all objects
and placing a tight-fitting, accurate bounding box around them.

147
• Frames Per Second (FPS): Measures detection speed, crucial for real-time applications
on edge devices.

Pre-Trained Object Detection Models Overview

As we saw in the introduction, given an image or a video stream, an object detection model
can identify which of a known set of objects might be present and provide information about
their positions within the image.

You can test some common models online by visiting Object Detection - MediaPipe
Studio

On Kaggle, we can find the most common pre-trained TFLite models to use with the Raspberry
Pi, ssd_mobilenet_v1, and efficiendet. Those models were trained on the COCO (Common
Objects in Context) dataset, which contains over 200,000 labeled images across 91 categories.
Download the models and upload them to the ./models folder on the Raspberry Pi.

Alternatively, you can find the models and the COCO labels on GitHub.

148
For the first part of this lab, we will focus on a pre-trained 300x300 SSD-Mobilenet V1 model
and compare it with the 320x320 EfficientDet-lite0, also trained using the COCO 2017 dataset.
Both models were converted to a TensorFlow Lite format (4.2MB for the SSD Mobilenet and
4.6MB for the EfficientDet).

SSD-Mobilenet V2 or V3 is recommended for transfer learning projects, but once


the V1 TFLite model is publicly available, we will use it for this overview.

The model outputs up to ten detections per image, including bounding boxes,
class IDs, and confidence scores.

Setting Up the TFLite Environment

We should confirm the steps done on the last Hands-On Lab, Image Classification, as follows:

• Updating the Raspberry Pi


• Installing Required Libraries
• Setting up a Virtual Environment (Optional but Recommended)

source ~/tflite/bin/activate

• Installing TensorFlow Lite Runtime


• Installing Additional Python Libraries (inside the environment)

149
Creating a Working Directory:

Considering that we have created the Documents/TFLITE folder in the last Lab, let’s now create
the specific folders for this object detection lab:

cd Documents/TFLITE/
mkdir OBJ_DETECT
cd OBJ_DETECT
mkdir images
mkdir models
cd models

Inference and Post-Processing

Let’s start a new notebook to follow all the steps to detect objects in an image:
Import the needed libraries:

import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter

Download the model and labels from the folder models and save them in the models folder
under OBJ_DETECT.
Load the model and allocate tensors:

model_path = "./models/[Link]"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()

150
Get input and output tensors.

input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

Input details will inform us how the model should be fed with an image. The shape of (1,
300, 300, 3) with a dtype of uint8 tells us that a non-normalized (pixel value range from
0 to 255) image with dimensions (300x300x3) should be input one by one (Batch Dimension:
1).

The output details include not only the labels (“classes”) and probabilities (“scores”), but
also the relative window positions of the bounding boxes (“boxes”), indicating where the object
is located in the image, and the number of detected objects (”num_detections”). The output
details also indicate that the model can detect up to 10 objects in the image.

151
So, for the above example, using the same cat image used with the Image Classification
Lab, looking for the output, we have a 76% probability of having found an object with a
class ID of 16 on an area delimited by a bounding box of [0.028011084, 0.020121813,
0.9886069, 0.802299]. Those four numbers are related to ymin, xmin, ymax, and xmax, the
box coordinates.

Considering that y ranges from the top (ymin) to the bottom (ymax) and x ranges from
left (xmin) to right (xmax), we have, in fact, the coordinates of the top-left corner and the
bottom-right one. With both edges and knowing the shape of the picture, it is possible to draw
a rectangle around the object, as shown in the figure below:

152
Next, we should find what class ID 16 means. Opening the file coco_labels.txt, we see that
each element has an associated index; inspecting index 16, we get, as expected, cat. The
probability is the value returned from the score.
Let’s now upload some images with multiple objects on them for testing.

img_path = "./images/cat_dog.jpeg"
orig_img = [Link](img_path)

# Display the image


[Link](figsize=(8, 8))
[Link](orig_img)

153
[Link]("Original Image")
[Link]()

Based on the input details, let’s pre-process the image, changing its shape and expanding its
dimensions:

img = orig_img.resize((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims(img, axis=0)
input_data.shape, input_data.dtype

The new input_data shape is(1, 300, 300, 3) with a dtype of uint8, which is compatible
with what the model expects.
Using the input_data, let’s run the interpreter, measure the latency, and get the output:

154
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))

With a latency of around 800 ms on the Raspi-Zero and 100 ms n the Raspi-5 , we can get four
distinct outputs:

boxes = interpreter.get_tensor(output_details[0]['index'])[0]
classes = interpreter.get_tensor(output_details[1]['index'])[0]
scores = interpreter.get_tensor(output_details[2]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[3]['index'])[0])

On a quick inspection, we can see that the model detected two objects with a score over 0.5:

for i in range(num_detections):
if scores[i] > 0.5: # Confidence threshold
print(f"Object {i}:")
print(f" Bounding Box: {boxes[i]}")
print(f" Confidence: {scores[i]}")
print(f" Class: {classes[i]}")

And we can also visualize the results:

[Link](figsize=(12, 8))
[Link](orig_img)
for i in range(num_detections):
if scores[i] > 0.5: # Adjust threshold as needed
ymin, xmin, ymax, xmax = boxes[i]

155
(left, right, top, bottom) = (xmin * orig_img.width,
xmax * orig_img.width,
ymin * orig_img.height,
ymax * orig_img.height)
rect = [Link]((left, top), right-left, bottom-top,
fill=False, color='red', linewidth=2)
[Link]().add_patch(rect)
class_id = int(classes[i])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {scores[i]:.2f}',
color='red', fontsize=12, backgroundcolor='white')

The choice of the confidence threshold is crucial. For example, setting it to 0.2
will show false positives. A proper code should handle it.

156
EfficientDet

EfficientDet is not technically an SSD (Single Shot Detector) model, but it shares some
similarities and builds upon ideas from SSD and other object detection architectures:

1. EfficientDet:
• Developed by Google researchers in 2019
• Uses EfficientNet as the backbone network
• Employs a novel bi-directional feature pyramid network (BiFPN)
• It uses compound scaling to efficiently scale the backbone network and object
detection components.
2. Similarities to SSD:
• Both are single-stage detectors, meaning they perform object localization and
classification in a single forward pass.
• Both use multi-scale feature maps to detect objects at different scales.
3. Key differences:
• Backbone: SSD typically uses VGG or MobileNet, while EfficientDet uses Efficient-
Net.
• Feature fusion: SSD uses a simple feature pyramid, while EfficientDet uses the more
advanced BiFPN.
• Scaling method: EfficientDet introduces compound scaling for all components of the
network
4. Advantages of EfficientDet:
• Generally achieves better trade-offs between accuracy and efficiency than SSD and
many other object detection models.
• More flexible scaling enables a family of models with varying size-performance
trade-offs.

While EfficientDet is not an SSD model, it can be seen as an evolution of single-stage detection
architectures, incorporating more advanced techniques to improve efficiency and accuracy.
When using EfficientDet, we can expect outputs similar to those of SSD (e.g., bounding boxes
and class scores).

On GitHub, you can find another notebook exploring the EfficientDet model that
we did with SSD MobileNet.

157
Object Detection on a live stream

Object detection models can also detect objects in real-time using a camera. The captured
image should be the input for the trained and converted model. On the Raspberry Pi 4 or 5,
OpenCV can capture frames and display inference results on a desktop.
However, even without a desktop, it’s possible to create a live stream with a webcam to
detect objects in real time. For example, let’s start with the script developed for the Image
Classification app and adapt it for a Real-Time Object Detection Web Application Using
TensorFlow Lite and Flask.
Download the Python script object_detection_app.py from GitHub.
This app version should work for any TFLite/LiteRT models.

Verify if the model is in its correct folder, for example:

model_path = "./models/[Link]"

Check the model and labels:

And on the terminal, run:

python object_detection_app.py

After starting, you should receive the message on the terminal (the IP is from my Raspberry):

158
* Running on [Link]
Press CTRL+C to quit

And access the web interface:

• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Raspberry
Pi’s IP address). For example: [Link]

Here is a screenshot of the app running on an external desktop

159
160
Let’s see a technical description of the key modules used in the object detection application:

1. LiteRT:
• Purpose: Efficient inference of machine learning models on edge devices.
• Why: LiteRT offers a smaller model size and optimized performance compared to
full TensorFlow, which is crucial for resource-constrained devices like the Raspberry
Pi. It supports hardware acceleration and quantization, further improving efficiency.
• Key functions: Interpreter for loading and running the model, get_input_details(),
and get_output_details() for interfacing with the model.
2. Flask:
• Purpose: Lightweight web framework for building backend servers.
• Why: Flask’s simplicity and flexibility make it ideal for rapidly developing and
deploying web applications. It’s less resource-intensive than larger frameworks
suitable for edge devices.
• Key components: route decorators for defining API endpoints, Response objects for
streaming video, render_template_string for serving dynamic HTML.
3. Picamera2:
• Purpose: Interface with the Raspberry Pi camera module.
• Why: Picamera2 is the latest library for controlling Raspberry Pi cameras, offering
improved performance and features over the original Picamera library.
• Key functions: create_preview_configuration() for setting up the camera,
capture_file() for capturing frames.
4. PIL (Python Imaging Library):
• Purpose: Image processing and manipulation.
• Why: PIL provides a wide range of image processing capabilities. It’s used here to
resize images, draw bounding boxes, and convert between image formats.
• Key classes: Image for loading and manipulating images, ImageDraw for drawing
shapes and text on images.
5. NumPy:
• Purpose: Efficient array operations and numerical computing.
• Why: NumPy’s array operations are much faster than pure Python lists, which is
crucial for efficiently processing image data and model inputs/outputs.
• Key functions: array() for creating arrays, expand_dims() for adding dimensions
to arrays.
6. Threading:
• Purpose: Concurrent execution of tasks.

161
• Why: Threading enables simultaneous frame capture, object detection, and web
server operation, which is crucial for maintaining real-time performance.
• Key components: Thread class creates separate execution threads, and Lock is used
for thread synchronization.
7. [Link]:
• Purpose: In-memory binary streams.
• Why: Allows efficient handling of image data in memory without needing temporary
files, improving speed and reducing I/O operations.
8. time:
• Purpose: Time-related functions.
• Why: Used for adding delays ([Link]()) to control frame rate and for perfor-
mance measurements.
9. jQuery (client-side):
• Purpose: Simplified DOM manipulation and AJAX requests.
• Why: It makes it easy to update the web interface dynamically and communicate
with the server without page reloads.
• Key functions: .get() and .post() for AJAX requests, DOM manipulation methods
for updating the UI.

Regarding the main app system architecture:

1. Main Thread: Runs the Flask server, handling HTTP requests and serving the web
interface.
2. Camera Thread: Continuously captures frames from the camera.
3. Detection Thread: Processes frames using the LiteRT object detection model.
4. Frame Buffer: Shared memory space (protected by locks) storing the latest frame and
detection results.

And the app data flow, we can describe in short:

1. Camera captures frame → Frame Buffer


2. Detection thread reads from Frame Buffer → Processes through TFLite/LiteRT model
→ Updates detection results in Frame Buffer
3. Flask routes access the Frame Buffer to serve the latest frame and detection results
4. Web client receives updates via AJAX and updates the UI

This architecture enables efficient, real-time object detection while maintaining a responsive web
interface on a resource-constrained edge device, such as a Raspberry Pi. Threading and efficient
libraries, such as LiteRT and PIL, enable the system to process video frames in real-time, while
Flask and jQuery provide a user-friendly way to interact with them.

162
You can test the app with another pre-processed model, such as the EfficientDet, by changing
the app line:

model_path = "./models/lite-model_efficientdet_lite0_detection_metadata_1.tflite"

If we want to use the app with the SSD-MobileNetV2 model, trained in Edge
Impulse Studio with the “Box versus Wheel” dataset, the code should also be
adapted to the input details, as we explored in its notebook.

Conclusion

This lab has explored implementing object detection on edge devices such as the Raspberry
Pi, demonstrating the power and potential of running advanced computer vision tasks on
resource-constrained hardware. We examined the object detection models SSD-MobileNet and
EfficientDet, comparing their performance and trade-offs on edge devices.
The lab demonstrated a real-time object-detection web application, showing how these models
can be integrated into practical, interactive systems.
The ability to perform object detection on edge devices opens up numerous possibilities across
domains such as precision agriculture, industrial automation, quality control, smart home
applications, and environmental monitoring. By processing data locally, these systems can offer
reduced latency, improved privacy, and operation in environments with limited connectivity.
Looking ahead, potential areas for further exploration include: - Using a custom dataset
(labeled on Roboflow), walking through the process of training models using Edge Impulse
Studio and Ultralytics, and deploying them on Raspberry Pi. - To improve inference speed on
edge devices, explore various optimization methods, such as model quantization (TFLite int8)
and format conversion (e.g., to NCNN). - Implementing multi-model pipelines for more complex
tasks - Exploring hardware acceleration options for Raspberry Pi - Integrating object detection
with other sensors for more comprehensive edge AI systems - Developing edge-to-cloud solutions
that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the power
of AI directly into the physical world, opening up new frontiers in how we interact with and
understand our environment.

Resources

• SSD-MobileNet Notebook on a Raspberry Pi


• EfficientDet Notebook on a Raspberry Pi

163
• Python Scripts
• Models

164
Custom Object Detection Project

165
Object Detection Project

In this chapter, we will develop a complete Object Detection project from data collection,
labelling, training, and deployment. As we did with the Image Classification project, the
trained and converted model will be used for inference.

We will use the same dataset to train 3 models: SSD-MobileNet V2, FOMO, and YOLO.

The Goal

All Machine Learning projects need to start with a goal. Let’s assume we are in an industrial
facility and must sort and count wheels and special boxes.

In other words, we should perform a multi-label classification, where each image can have three
classes:

• Background (no objects)


• Box

166
• Wheel

Raw Data Collection

Once we have defined our Machine Learning project goal, the next and most crucial step is
collecting the dataset. We can use a phone, the Raspi, or a mix to create the raw dataset (with
no labels). Let’s use the simple web app on our Raspberry Pi to view the QVGA (320 x 240)
captured images in a browser.
From GitHub, get the Python script get_img_data.py and open it in the terminal:

python3 get_img_data.py

Access the web interface:

• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Raspberry
Pi’s IP address). For example: [Link]

The
Python script creates a web-based interface for capturing and organizing image datasets using
a Raspberry Pi and its camera. It’s handy for machine learning projects that require labeled
image data, or not, as in our case here.
Access the web interface from a browser, enter a generic label for the images you want to
capture, and press Start Capture.

167
Note that the images to be captured will have multiple labels that should be defined
later.

Use the live preview to position the camera, then click Capture Image to save the images under
the current label (in this case, box-wheel).

168
When we have enough images, we can press Stop Capture. The captured images are saved in
the folder dataset/box-wheel:

169
Get around 60 images. Try to capture different angles, backgrounds, and light
conditions. FileZilla can transfer the raw dataset you created to your main computer.

Labeling Data

The next step in an Object Detect project is to create a labeled dataset. We should label the
raw dataset images, creating bounding boxes around each picture’s objects (box and wheel).
We can use labeling tools such as LabelImg, CVAT, Roboflow, or even the Edge Impulse Studio.
Once we have explored the Edge Impulse tool in other labs, let’s use Roboflow here.

We are using Roboflow (free version) here for two main reasons. 1) We can have an
auto-labeler, and 2) The annotated dataset is available in several formats and can
be used both on Edge Impulse Studio (we will use it for MobileNet V2 and FOMO
train) and on CoLab (YOLOv8 or YOLOv11 train), for example. An annotated
dataset created on Edge Impulse (Free account) cannot be used for training on
other platforms.

We should upload the raw dataset to Roboflow. Create a free account there and start a new
project, for example, (“box-versus-wheel”).

170
We will not go into great detail about the Roboflow process, as many tutorials are
already available.

Annotate

Once the project is created and the dataset is uploaded, you can use the “Auto-Label” tool to
generate annotations, or do it manually.

171
The Label Assist tool can be handy for the labeling process.

172
Note that you should also upload images with only a background, which should be saved w/o
any annotations using the Null Tool option.

173
Once all images are annotated, split them into training, validation, and test sets.

174
Data Pre-Processing

The last step in the dataset is preprocessing to generate a final training version. Let’s resize all
images to 320x320 and generate augmented versions of each image (augmentation) to create
new training examples from which our model can learn.
For augmentation, we will rotate the images (+/-15o ), crop, and vary the brightness and
exposure.

175
At the end of the process, we will have 153 images.

176
Now, you should export the annotated dataset in a format that Edge Impulse, Ultralitics, and
other frameworks/tools understand, for example, YOLOv8 (or v11). Let’s download a zipped
version of the dataset to our desktop.

177
Here, it is possible to review how the dataset was structured

178
There are 3 separate folders, one for each split (train/test/valid). For each of them,
there are 2 subfolders, images, and labels. The pictures are stored as image_id.jpg and
images_id.txt, where “image_id” is unique for every picture.
The labels file format will be class_id bounding box coordinates, where in our case, class_id
will be 0 for box and 1 for wheel. The numerical id (o, 1, 2…) will follow the alphabetical order
of the class name.
The [Link] file contains information about the dataset, such as the classes’ names (names:
['box', 'wheel']) following the YOLO format.
And that’s it! We are ready to start training using Edge Impulse Studio (as we will in the next
step), Ultralytics (as we will when discussing YOLO), or even training from scratch on CoLab
(as we did with the Cifar-10 dataset in the Image Classification lab).

The pre-processed dataset can be found at the Roboflow site, or here:

179
Training an SSD MobileNet Model on Edge Impulse Studio

Go to Edge Impulse Studio, enter your credentials at Login (or create an account), and start
a new project.

Here, you can clone the project developed for this hands-on lab: Raspi - Object
Detection.

On the Project Dashboard tab, go down to Project info, and for Labeling method select
Bounding boxes (object detection)

Uploading the annotated data

In Studio, go to the Data acquisition tab, and in the UPLOAD DATA section, upload the raw
dataset from your computer.
We can use the Select a folder option, choosing, for example, the train folder on your
computer, which contains two sub-folders: images and labels. Select the Image label
format, “YOLO TXT”, upload it into the category Training, and press Upload data.

180
Repeat the process for the test data (upload both folders, test, and validation). At the end
of the upload process, you should end with the annotated dataset of 153 images split in the
train/test (84%/16%).

Note that labels will be stored at the labels files 0 and 1 , which are equivalent to
box and wheel.

The Impulse Design

The first thing to define when we enter the Create impulse step is to describe the target device
for deployment. A pop-up window will appear. We will select Raspberry 4, an intermediary
device between the Raspi-Zero and the Raspi-5.

This choice will not interfere with the training; it will only give us an idea about
the latency of the model on that specific target.

181
In this phase, you should define how to:

• Pre-processing consists of resizing the individual images. In our case, the images were
pre-processed on Roboflow, to 320x320 , so let’s keep it. The resize will not matter here
because the images are already squared. If you upload a rectangular image, squash it
(squared form, without cropping). Afterward, you could define if the images are converted
from RGB to Grayscale or not.
• Design a Model, in this case, “Object Detection.”

182
Preprocessing all dataset

In the section Image, select Color depth as RGB, and press Save parameters.

183
The Studio automatically moves to the next section, Generate features, where all samples will
be preprocessed, resulting in 480 objects: 207 boxes and 273 wheels.

184
The feature explorer shows that all samples exhibit a good separation after the feature
generation.

Model Design, Training, and Test

For training, we should select a pre-trained model. Let’s use the MobileNetV2 SSD FPN-
Lite (320x320 only).

• Base Network (MobileNetV2)


• Detection Network (Single Shot Detector or SSD)
• Feature Extractor (FPN-Lite)

185
It is a pre-trained object detection model that locates up to 10 objects in an image and outputs
a bounding box for each. The model is approximately 3.7 MB. It supports an RGB input at
320x320px.

Regarding the training hyper-parameters, the model will be trained with:

• Epochs: 25
• Batch size: 32
• Learning Rate: 0.15.

For validation during training, 20% of the dataset (validation_dataset) will be spared.

186
As a result, the model achieves an overall precision score (based on COCO mAP) of 88.8%,
higher than the score on the test data (83.3%).

Deploying the model

We have two ways to deploy our model:

• TFLite model, which lets deploy the trained model as .tflite for the Raspberry Pi to
run it using Python.
• Linux (AARCH64), a binary for Linux (AARCH64), implements the Edge Impulse
Linux protocol, which lets us run our models on any Linux-based development board,
with SDKs such as Python. See the documentation for more information and setup
instructions.

Let’s deploy the TFLite model. On the Dashboard tab, go to Transfer learning model (int8
quantized) and click on the download icon:

187
Transfer the model from your computer to the Raspi folder./models and capture or get some
images for inference and save them in the folder ./images.

188
Inference and Post-Processing

The inference can be made as discussed in the Pre-Trained Object Detection Models Overview.
Let’s start a new notebook to follow all the steps to detect cubes and wheels in an image.
Import the needed libraries:

import time
import numpy as np
import [Link] as plt
import [Link] as patches
from PIL import Image
from ai_edge_litert.interpreter import Interpreter

Define the model path and labels:

model_path = "./models/ei-raspi-object-detection-SSD-MobileNetv2-320x0320-\
[Link]"
labels = ['box', 'wheel']

Remember that the model will output the class ID as values (0 and 1), following an
alphabetic order regarding the class names.

Load the model, allocate the tensors, and get the input and output tensor details:

# Create a LiteRT Interpreter


interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()

# Get input and output tensors


input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

One crucial difference to note is that the dtype of the input details of the model is now int8,
which means that the input values go from -128 to +127, while each pixel of our raw image
goes from 0 to 256. This means that we should pre-process the image to match it. We can
check here:

input_dtype = input_details[0]['dtype']
input_dtype

numpy.int8

189
So, let’s open the image and show it:

# Load the image


img_path = "./images/box_2_wheel_2.jpg"
orig_img = [Link](img_path)

# Display the image


[Link](figsize=(6, 6))
[Link](orig_img)
[Link]("Original Image")
[Link]()

190
And perform the pre-processing:

scale, zero_point = input_details[0]['quantization']


img = orig_img.resize((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (img_array / scale + zero_point).clip(-128, 127).astype(np.int8)
input_data = np.expand_dims(img_array, axis=0)

Checking the input data, we can verify that the input tensor is compatible with what is expected

191
by the model:

input_data.shape, input_data.dtype

((1, 320, 320, 3), dtype('int8'))

Now, it is time to perform the inference. Let’s also calculate the latency of the model:

# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))

The model will take around 600ms to perform the inference in the Raspi-Zero, which is around
5 times longer than a Raspi-5.
Now, we can get the output classes of objects detected, its bounding boxes coordinates, and
probabilities.

boxes = interpreter.get_tensor(output_details[1]['index'])[0]
classes = interpreter.get_tensor(output_details[3]['index'])[0]
scores = interpreter.get_tensor(output_details[0]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[2]['index'])[0])

for i in range(num_detections):
if scores[i] > 0.5: # Confidence threshold
print(f"Object {i}:")
print(f" Bounding Box: {boxes[i]}")
print(f" Confidence: {scores[i]}")
print(f" Class: {classes[i]}")

192
From the results, we can see that 4 objects were detected: two with class ID 0 (box)and two
with class ID 1 (wheel), what is correct!
Let’s visualize the result for a threshold of 0.5

threshold = 0.5
[Link](figsize=(6,6))
[Link](orig_img)
for i in range(num_detections):
if scores[i] > threshold:
ymin, xmin, ymax, xmax = boxes[i]
(left, right, top, bottom) = (xmin * orig_img.width,
xmax * orig_img.width,
ymin * orig_img.height,
ymax * orig_img.height)
rect = [Link]((left, top), right-left, bottom-top,
fill=False, color='red', linewidth=2)
[Link]().add_patch(rect)
class_id = int(classes[i])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {scores[i]:.2f}',
color='red', fontsize=12, backgroundcolor='white')

193
But what happens if we reduce the threshold to 0.3, for example?

194
We start to see false positives and multiple detections, where the model detects the same
object multiple times with different confidence levels and slightly different bounding boxes.
Commonly, sometimes, we need to adjust the threshold to smaller values to capture all objects,
avoiding false negatives, which would lead to multiple detections.
To improve the detection results, we should implement Non-Maximum Suppression (NMS),
which helps eliminate overlapping bounding boxes and keeps only the most confident detection.
For that, let’s create a general function named non_max_suppression(), with the role of
refining object detection results by eliminating redundant and overlapping bounding boxes.

195
It achieves this by iteratively selecting the detection with the highest confidence score and
removing other significantly overlapping detections based on an Intersection over Union (IoU)
threshold.

def non_max_suppression(boxes, scores, threshold):


# Convert to corner coordinates
x1 = boxes[:, 0]
y1 = boxes[:, 1]
x2 = boxes[:, 2]
y2 = boxes[:, 3]

areas = (x2 - x1 + 1) * (y2 - y1 + 1)


order = [Link]()[::-1]

keep = []
while [Link] > 0:
i = order[0]
[Link](i)
xx1 = [Link](x1[i], x1[order[1:]])
yy1 = [Link](y1[i], y1[order[1:]])
xx2 = [Link](x2[i], x2[order[1:]])
yy2 = [Link](y2[i], y2[order[1:]])

w = [Link](0.0, xx2 - xx1 + 1)


h = [Link](0.0, yy2 - yy1 + 1)
inter = w * h
ovr = inter / (areas[i] + areas[order[1:]] - inter)

inds = [Link](ovr <= threshold)[0]


order = order[inds + 1]

return keep

How it works:

1. Sorting: It starts by sorting all detections by their confidence scores, highest to lowest.
2. Selection: It selects the highest-scoring box and adds it to the final list of detections.
3. Comparison: This selected box is compared with all remaining lower-scoring boxes.
4. Elimination: Any box that overlaps significantly (above the IoU threshold) with the
selected box is eliminated.

196
5. Iteration: This process repeats with the next highest-scoring box until all boxes are
processed.

Now, we can define a more precise visualization function that will take into consideration an
IoU threshold, detecting only the objects that were selected by the non_max_suppression
function:

def visualize_detections(image, boxes, classes, scores,


labels, threshold, iou_threshold):
if isinstance(image, [Link]):
image_np = [Link](image)
else:
image_np = image

height, width = image_np.shape[:2]

# Convert normalized coordinates to pixel coordinates


boxes_pixel = boxes * [Link]([height, width, height, width])

# Apply NMS
keep = non_max_suppression(boxes_pixel, scores, iou_threshold)

# Set the figure size to 12x8 inches


fig, ax = [Link](1, figsize=(12, 8))

[Link](image_np)

for i in keep:
if scores[i] > threshold:
ymin, xmin, ymax, xmax = boxes[i]
rect = [Link]((xmin * width, ymin * height),
(xmax - xmin) * width,
(ymax - ymin) * height,
linewidth=2, edgecolor='r', facecolor='none')
ax.add_patch(rect)
class_name = labels[int(classes[i])]
[Link](xmin * width, ymin * height - 10,
f'{class_name}: {scores[i]:.2f}', color='red',
fontsize=12, backgroundcolor='white')

[Link]()

Now we can create a function that will call the others, performing inference on any image:

197
def detect_objects(img_path, conf=0.5, iou=0.5):
orig_img = [Link](img_path)
scale, zero_point = input_details[0]['quantization']
img = orig_img.resize((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (img_array / scale + zero_point).clip(-128, 127).\
astype(np.int8)
input_data = np.expand_dims(img_array, axis=0)

# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to ms
print ("Inference time: {:.1f}ms".format(inference_time))

# Extract the outputs


boxes = interpreter.get_tensor(output_details[1]['index'])[0]
classes = interpreter.get_tensor(output_details[3]['index'])[0]
scores = interpreter.get_tensor(output_details[0]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[2]['index'])[0])

visualize_detections(orig_img, boxes, classes, scores, labels,


threshold=conf,
iou_threshold=iou)

Now, running the code, having the same image again with a confidence threshold of 0.3, but
with a small IoU:

img_path = "./images/box_2_wheel_2.jpg"
detect_objects(img_path, conf=0.3,iou=0.05)

198
Training a FOMO Model at Edge Impulse Studio

The inference with the SSD MobileNet model worked well, but the latency was significantly
high. The inference varied from 0.5 to 1.3 seconds on a Raspi-Zero, which means around or
less than 1 FPS (1 frame per second). One alternative to speed up the process is to use FOMO
(Faster Objects, More Objects).
This novel machine learning algorithm lets us count multiple objects and find their location in

199
an image in real-time using up to 30x less processing power and memory than MobileNet SSD
or YOLO. The main reason this is possible is that while other models calculate the object’s
size by drawing a square around it (bounding box), FOMO ignores the size of the image,
providing only the information about where the object is located in the image through its
centroid coordinates.

How FOMO works?

In a typical object detection pipeline, the first stage is extracting features from the input image.
FOMO leverages MobileNetV2 to perform this task. MobileNetV2 processes the input
image to produce a feature map that captures essential characteristics, such as textures, shapes,
and object edges, in a computationally efficient way.

Once these features are extracted, FOMO’s simpler architecture, focused on center-point
detection, interprets the feature map to determine where objects are located in the image. The

200
output is a grid of cells, where each cell represents whether or not an object center is detected.
The model outputs one or more confidence scores for each cell, indicating the likelihood of an
object being present.
Let’s see how it works on an image.
FOMO divides the image into blocks of pixels using a factor of 8. For the input of 96x96,
the grid would be 12x12 (96/8=12). For a 160x160, the grid will be 20x20, and so on. Next,
FOMO will run a classifier through each pixel block to calculate the probability that there is a
box or a wheel in each of them and, subsequently, determine the regions that have the highest
probability of containing the object (If a pixel block has no objects, it will be classified as
background). From the overlap of the final region, the FOMO provides the coordinates (related
to the image dimensions) of the centroid of this region.

Trade-off Between Speed and Precision:

• Grid Resolution: FOMO uses a grid of fixed resolution, meaning each cell can detect if
an object is present in that part of the image. While it doesn’t provide high localization

201
accuracy, it makes a trade-off by being fast and computationally light, which is crucial
for edge devices.
• Multi-Object Detection: Since each cell is independent, FOMO can detect multiple
objects simultaneously in an image by identifying multiple centers.

Impulse Design, new Training and Testing

Return to Edge Impulse Studio, and in the Experiments tab, create another impulse. Now,
the input images should be 160x160 (this is the expected input size for MobilenetV2).

On the Image tab, generate the features and go to the Object detection tab.
We should select a pre-trained model for training. Let’s use the FOMO (Faster Objects,
More Objects) MobileNetV2 0.35.

202
Regarding the training hyper-parameters, the model will be trained with:

• Epochs: 30
• Batch size: 32
• Learning Rate: 0.001.

For validation during training, 20% of the dataset (validation_dataset) will be spared. We will
not apply Data Augmentation for the remaining 80% (train_dataset) because our dataset was
already augmented during the labeling phase at Roboflow.
As a result, the model ends with an overall F1 score of 93.3% with an impressive latency of
8ms (Raspi-4), around 60X less than we got with the SSD MovileNetV2.

203
Note that FOMO automatically added a third label background to the two previously
defined boxes (0) and wheels (1).

On the Model testing tab, we can see that the accuracy was 94%. Here is one of the test
sample results:

204
In object detection tasks, accuracy is generally not the primary evaluation metric.
Object detection involves classifying objects and providing bounding boxes around
them, making it a more complex problem than simple classification. The issue is
that we do not have the bounding box, only the centroids. In short, using accuracy
as a metric could be misleading and may not provide a complete understanding of
how well the model is performing.

Deploying the model

As we did in the previous section, we can deploy the trained model as TFLite or Linux
(AARCH64). Let’s do it now as Linux (AARCH64), a binary that implements the Edge
Impulse Linux protocol.
Edge Impulse for Linux models is delivered in .eim format. This executable contains our “full
impulse” created in Edge Impulse Studio. The impulse consists of the signal processing block(s)
and any learning and anomaly block(s) we added and trained. It is compiled with optimizations
for our processor or GPU (e.g., NEON instructions on ARM cores), plus a straightforward IPC
layer (over a Unix socket).
At the Deploy tab, select the option Linux (AARCH64), the int8model and press Build.

205
The model will be automatically downloaded to your computer.

206
On our Raspi, let’s create a new working area:

cd ~
cd Documents
mkdir EI_Linux
cd EI_Linux
mkdir models
mkdir images

Rename the model for easy identification:


For example, [Link] and transfer it to
the new Raspi folder./models and capture or get some images for inference and save them in
the folder ./images.

Inference and Post-Processing

The inference will be made using the Linux Python SDK. This library lets us run machine
learning models and collect sensor data on Linux machines using Python. The SDK is open
source and available on GitHub at edgeimpulse/linux-sdk-python.
Let’s set up a Virtual Environment for working with the Linux Python SDK

python3 -m venv ~/eilinux


source ~/eilinux/bin/activate

And Install the all the libraries needed:

sudo apt-get update


sudo apt-get install libatlas-base-dev libportaudio0 libportaudio2
sudo apt-get install libportaudiocpp0 portaudio19-dev

pip3 install edge_impulse_linux -i [Link]


pip3 install Pillow matplotlib pyaudio opencv-contrib-python

sudo apt-get install portaudio19-dev


pip3 install pyaudio
pip3 install opencv-contrib-python

Permit our model to be executable.

207
chmod +x [Link]

Install the Jupiter Notebook on the new environment

pip3 install jupyter

Run a notebook locally (on the Raspi-4 or 5 with desktop)

jupyter notebook

or on the browser on your computer:

jupyter notebook --ip=[Link] --no-browser

Let’s start a new notebook by following all the steps to detect cubes and wheels on an image
using the FOMO model and the Edge Impulse Linux Python SDK.
Import the needed libraries:

import sys, time


import numpy as np
import [Link] as plt
import [Link] as patches
from PIL import Image
import cv2
from edge_impulse_linux.image import ImageImpulseRunner

Define the model path and labels:

model_file = "[Link]"
model_path = "models/"+ model_file # Trained ML model from Edge Impulse
labels = ['box', 'wheel']

Remember that the model will output the class ID as values (0 and 1), following an
alphabetic order regarding the class names.

Load and initialize the model:

# Load the model file


runner = ImageImpulseRunner(model_path)

# Initialize model
model_info = [Link]()

208
The model_info will contain critical information about our model. However, unlike the TFLite
interpreter, the EI Linux Python SDK library will now prepare the model for inference.
So, let’s open the image and show it (Now, for compatibility, we will use OpenCV, the CV
Library used internally by EI. OpenCV reads the image as BGR, so we will need to convert it
to RGB :

# Load the image


img_path = "./images/1_box_1_wheel.jpg"
orig_img = [Link](img_path)
img_rgb = [Link](orig_img, cv2.COLOR_BGR2RGB)

# Display the image


[Link](img_rgb)
[Link]("Original Image")
[Link]()

209
Now we will get the features and the preprocessed image (cropped) using the runner:

features, cropped = runner.get_features_from_image_auto_studio_setings(img_rgb)

And perform the inference. Let’s also calculate the latency of the model:

res = [Link](features)

Let’s get the output classes of objects detected, their bounding boxes centroids, and probabili-
ties.

print('Found %d bounding boxes (%d ms.)' % (


len(res["result"]["bounding_boxes"]),
res['timing']['dsp'] + res['timing']['classification']))
for bb in res["result"]["bounding_boxes"]:
print('\t%s (%.2f): x=%d y=%d w=%d h=%d' % (
bb['label'], bb['value'], bb['x'],
bb['y'], bb['width'], bb['height']))

Found 2 bounding boxes (29 ms.)


1 (0.91): x=112 y=40 w=16 h=16
0 (0.75): x=48 y=56 w=8 h=8

The results show that two objects were detected: one with class ID 0 (box) and one with class
ID 1 (wheel), which is correct!
Let’s visualize the result (The threshold is 0.5, the default value set during the model testing
on the Edge Impulse Studio).

print('\tFound %d bounding boxes (latency: %d ms)' % (


len(res["result"]["bounding_boxes"]),
res['timing']['dsp'] + res['timing']['classification']))
[Link](figsize=(5,5))
[Link](cropped)

# Go through each of the returned bounding boxes


bboxes = res['result']['bounding_boxes']
for bbox in bboxes:

# Get the corners of the bounding box


left = bbox['x']

210
top = bbox['y']
width = bbox['width']
height = bbox['height']

# Draw a circle centered on the detection


circ = [Link]((left+width//2, top+height//2), 5,
fill=False, color='red', linewidth=3)
[Link]().add_patch(circ)
class_id = int(bbox['label'])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {bbox["value"]:.2f}',
color='red', fontsize=12, backgroundcolor='white')
[Link]()

211
Conclusion

This chapter has explored the implementation of a custom object detector on edge devices, such
as the Raspberry Pi, demonstrating the power and potential of running advanced computer
vision tasks on resource-constrained hardware. We’ve covered several vital aspects:

1. Model Comparison: We examined object detection models, including SSD-MobileNet

212
and FOMO, and compared their performance and trade-offs on edge devices.
2. Training and Deployment: Using a custom dataset of boxes and wheels (labeled on
Roboflow), we walked through the process of training models with Edge Impulse Studio
and Ultralytics and deploying them on a Raspberry Pi.
3. Optimization Techniques: To improve inference speed on edge devices, we explored
various optimization methods, such as model quantization (int8).
4. Performance Considerations: Throughout the lab, we discussed the balance between
model accuracy and inference speed, a critical consideration for edge AI applications.

As discussed earlier, the ability to perform object detection on edge devices opens up numerous
possibilities across domains, such as precision agriculture, industrial automation, quality control,
smart home applications, and environmental monitoring. By processing data locally, these
systems can offer reduced latency, improved privacy, and operation in environments with limited
connectivity.
Looking ahead, potential areas for further exploration include: - Implementing multi-model
pipelines for more complex tasks - Exploring hardware acceleration options for Raspberry Pi
- Integrating object detection with other sensors for more comprehensive edge AI systems -
Developing edge-to-cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the power
of AI directly into the physical world, opening up new frontiers in how we interact with and
understand our environment.

Resources

• Roboflow Annotated Dataset (“Box versus Wheel”)


• Edge Impulse Project - SSD MobileNet and FOMO
• Notebooksi
• Python Scripts
• Models

213
Computer Vision Applications with YOLO

Exploring a YOLO Models using Ultralitics

In this chapter, we will explore YOLOv8 and v11. Ultralytics YOLO (v8 and v11) are versions
of the acclaimed real-time object detection and image segmentation model, YOLO. YOLOv8
and v11 are built on cutting-edge advances in deep learning and computer vision, offering
unparalleled speed and accuracy. Its streamlined design makes it suitable for a wide range
of applications and easily adaptable across hardware platforms, from edge devices to cloud
APIs.

214
Talking about the YOLO Model

The YOLO (You Only Look Once) model is a highly efficient, widely used object detection
algorithm known for its real-time performance. Unlike traditional object detection systems that
repurpose classifiers or localizers to perform detection, YOLO frames the detection problem as
a single regression task. This innovative approach enables YOLO to simultaneously predict
multiple bounding boxes and their class probabilities from full images during a single evaluation,
significantly boosting its speed.

Key Features:

1. Single Network Architecture:

• YOLO employs a single neural network to process the entire image. This network
divides the image into a grid and, for each grid cell, directly predicts bounding boxes
and associated class probabilities. This end-to-end training improves speed and
simplifies the model architecture.

2. Real-Time Processing:

• One of YOLO’s standout features is its ability to perform object detection in real-
time. Depending on the version and hardware, YOLO can process images at high
frames per second (FPS). This makes it ideal for applications requiring quick and
accurate object detection, such as video surveillance, autonomous driving, and live
sports analysis.

3. Evolution of Versions:

• Over the years, YOLO has undergone significant improvements, from YOLOv1
to the latest YOLOv12. Each iteration has introduced enhancements in accuracy,
speed, and efficiency. YOLOv8, for instance, incorporates advancements in net-
work architecture, improved training methodologies, and better support for various
hardware, ensuring a more robust performance.
• YOLOv11 offers substantial improvements in accuracy, speed, and parameter effi-
ciency compared to prior versions such as YOLOv8 and YOLOv10, making it one
of the most versatile and powerful real-time object detection models available as of
2025

215
4. Accuracy and Efficiency:

• While early versions of YOLO traded off some accuracy for speed, recent versions
have made substantial strides in balancing both. The newer models are faster and
more accurate, detecting small objects (such as bees) and performing well on complex
datasets.

5. Wide Range of Applications:

• YOLO’s versatility has led to its adoption in numerous fields. It is used in traffic
monitoring systems to detect and count vehicles, security applications to identify
potential threats and agricultural technology to monitor crops and livestock. Its
application extends to any domain requiring efficient and accurate object detection.

6. Community and Development:

• YOLO continues to evolve and is supported by a strong community of developers and


researchers (with YOLOv8 being very strong). Open-source implementations and
extensive documentation have made it accessible for customization and integration
into various projects. Popular deep learning frameworks like Darknet, TensorFlow,
and PyTorch support YOLO, further broadening its applicability.

7. Model Capabilities
YOLO models support multiple computer vision tasks:

• Object Detection: Identifying and localizing objects with bounding boxes


• Instance Segmentation: Pixel-level object segmentation
• Pose Estimation: Human pose keypoint detection
• Classification: Image classification tasks

216
Ultralitics YOLO Detect, Segment, and Pose models pre-trained on the COCO
dataset, and Classify on the ImageNet dataset.

Track mode is available for all Detect, Segment, and Pose models. The latest versions of
YOLO can also perform OBB, which stands for Oriented Bounding Box, a rectangular
box in computer vision that can rotate to match the orientation of an object within
an image, providing a much tighter and more precise fit than traditional axis-aligned
bounding boxes.

Available Model Sizes

YOLO offers several model variants optimized for different use cases, for example. The
YOLOv8:

• YOLOv8n (Nano): Smallest model, fastest inference, lowest accuracy


• YOLOv8s (Small): Balanced performance for edge devices
• YOLOv8m (Medium): Higher accuracy, moderate computational requirements
• YOLOv8l (Large): High accuracy, requires more computational resources
• YOLOv8x (Extra Large): Highest accuracy, most computational intensive

By 2025, for Raspberry Pi applications, YOLOv8n or YOLO11n are typically the


best choice due to their optimized size and speed. In 2026, Ultralytics launched
YOLO 26, not covered in this lab.

Installation

First, let’s confirm the System Python version:

python --version

If we use the latest Raspberry Pi OS (based on Debian Trixie), it should be:


3.13.5

217
As of today (January 2026), Ultralytics officially supports only Python 3.9-3.12; Python
3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships with
Python 3.13 by default, we’ll need to install a compatible Python version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python.

If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.

Install pyenv Dependencies

sudo apt update


sudo apt install -y build-essential libssl-dev zlib1g-dev \
libbz2-dev libreadline-dev libsqlite3-dev curl git \
libncursesw5-dev xz-utils tk-dev libxml2-dev \
libxmlsec1-dev libffi-dev liblzma-dev \
libopenblas-dev libjpeg-dev libpng-dev cmake

Install pyenv

# Download and install pyenv


curl [Link] | bash

Configure Shell

Add pyenv to ~/.bashrc:

cat >> ~/.bashrc << 'EOF'

# pyenv configuration
export PYENV_ROOT="$HOME/.pyenv"
[[ -d $PYENV_ROOT/bin ]] && export PATH="$PYENV_ROOT/bin:$PATH"
eval "$(pyenv init -)"
EOF

Reload the shell:

218
source ~/.bashrc

Verify if pyenv is installed:

pyenv --version

Install Python 3.11 (or 3.12)

# See available versions


pyenv install --list | grep " 3.11"

# Install Python 3.11.14 (latest 3.11 stable)


pyenv install 3.11.14

# Or install Python 3.12.3 if you prefer


# pyenv install 3.12.12

This will take a few minutes to compile.

Create YOLO Workspace

cd Documents
mkdir YOLO
cd YOLO

# Set Python 3.11.14 for this directory


pyenv local 3.11.14

# Verify
python --version # Should show Python 3.11.14

Create Virtual Environment

python -m venv yolo-env


source yolo-env/bin/activate

219
# Verify if we're using the correct Python
which python
python --version

To exit the virtual environment later:

deactivate

Installing Ultralytics/Yolo

On our Raspi, let’s create new folders on the working area:

cd Documents/
cd YOLO
mkdir models
mkdir images

And install the Ultralytics packages for local inference on the Raspberry Pi (inside the env)

1. Update the packages list, install and/or upgrade PIP to the latest:

sudo apt update


pip install -U pip

2. Install the ultralytics pip package with optional dependencies:

pip install ultralytics[export]

3. Reboot the device:

sudo reboot

220
Testing the YOLO

After the Raspi booting, let’s activate the yolo env, go to the working directory,

cd /Documents/YOLO
source ~/yolo_env/bin/activate

And run inference on an image that will be downloaded from the Ultralytics website, using, for
example, the YOLOV8n model (the smallest in the family) at the Terminal (CLI):

yolo predict model='yolov8n' source='[Link]

Note that the first time we invoke a model, it will automatically be downloaded to
the current directory.

The inference result will appear in the terminal. In the image ([Link]), 4 persons, 1 bus,
and 1 stop signal were detected:

Also, we got a message that Results saved to runs/detect/predict. Inspecting that


directory, we can see a new image saved ([Link]). Let’s download it from the Raspi to our
desktop for inspection:

221
222
So, the Ultrayitics YOLO is correctly installed on our Raspberry Pi. Note that on the Raspberry
Pi Zero, an issue is the high latency for this inference, which takes several seconds, even with
the most compact model in the family (YOLOv8n).

Testing with the YOLOv11

The procedure is the same as we did with version v8. As a comparison, we can see that the
YOLOv11 is faster than the v8, but seems a little less precise, as it does not detect the “stop
sign” as the v8.

Export Models to NCNN format

Deploying computer vision models on edge devices with limited computational power, such as
the Raspberry Pi Zero, can cause latency issues. One alternative is to use a format optimized
for optimal performance. This ensures that even devices with limited processing power can
handle advanced computer vision tasks well.
Of all the model export formats supported by Ultralytics, the NCNN is a high-performance
neural network inference computing framework optimized for mobile platforms. From the
beginning of the design, NCNN was deeply considerate of deployment and use on mobile phones,
and it did not have third-party dependencies. It is cross-platform and runs faster than all
known open-source frameworks (such as TFLite).
NCNN delivers the best inference performance when working with Raspberry Pi devices. NCNN
is highly optimized for mobile embedded platforms (such as ARM architecture).
Let’s move the downloaded YOLO models to the ./models folder and [Link] to
./images.

223
And convert our models and rerun the inferences:

1. Export the YOLO PyTorch models to NCNN format, creating: yolov8n_ncnn_model


and yolo11n_ncnn_model

yolo export model=./models/[Link] format=ncnn


yolo export model=./models/[Link] format=ncnn

2. Run inference with the exported models:

yolo predict task=detect model='./models/yolov8n_ncnn_model' source='./images/[Link]'


yolo predict task=detect model='./models/yolo11n_ncnn_model' source='./images/[Link]'

The first inference, when the model is loaded, typically has a high latency; however,
from the second inference, it is possible to note that the inference time decreases.

We can now realize that neither model detects the “Stop Signal”, with YOLOv11 being the
fastest. The optimized models are more rapid but also less accurate.

Exploring YOLO with Python

To start, let’s call the Python Interpreter so we can explore how the YOLO model works, line
by line:

python

Now, we should call the YOLO library from Ultralitics and load the model:

224
from ultralytics import YOLO
model = YOLO('./models/yolov8n_ncnn_model')

Run inference over an image (let’s use again [Link]):

img = './images/[Link]'
result = [Link](img, save=True, imgsz=640, conf=0.5, iou=0.3)

We can verify that the result is almost identical to the one we get running the inference at
the terminal level (CLI), except that the bus stop was not detected with the reduced NCNN
model. Note that the latency was reduced.
Let’s analyze the “result” content.
For example, we can see result[0].[Link], showing us the main inference result, which
is a tensor with a shape of (4, 6). Each line is one of the objects detected, being the first four
columns, the bounding boxes coordinates, the 5th, the confidence, and the 6th, the class (in
this case, 0: person and 5: bus):

We can access several inference results separately, as the inference time, and have it printed in
a better format:

225
inference_time = int(result[0].speed['inference'])
print(f"Inference Time: {inference_time} ms")

Or we can have the total number of objects detected:

print(f'Number of objects: {len (result[0].[Link])}')

With Python, we can create a detailed output that meets our needs (See Model Prediction with
Ultralytics YOLO for more details). Let’s run a Python script instead of manually entering it
line by line in the interpreter, as shown below. Let’s use nano as our text editor. First, we
should create an empty Python script named, for example, yolov8_tests.py:

nano yolov8_tests.py

Enter the code lines:

from ultralytics import YOLO

# Load the YOLOv8 model


model = YOLO('./models/yolov8n_ncnn_model')

# Run inference
img = './images/[Link]'
result = [Link](img, save=False, imgsz=640, conf=0.5, iou=0.3)

# print the results


inference_time = int(result[0].speed['inference'])
print(f"Inference Time: {inference_time} ms")
print(f'Number of objects: {len (result[0].[Link])}')

226
And enter with the commands: [CTRL+O] + [ENTER] +[CTRL+X] to save the Python script.
Run the script:

python yolov8_tests.py

The result is the same as running the inference at the terminal level (CLI) and with the built-in
Python interpreter.

Calling the YOLO library and loading the model for inference for the first time
takes a long time, but the inferences after that will be much faster. For example,
the first single inference can take several seconds, but after that, the inference time
should be reduced to less than 1 second.

Inference Arguments

[Link]() accepts multiple arguments that can be passed at inference time to override
defaults:
Inference arguments:

227
Argument Type Default Description
source str Specifies the data source for inference. Can be
'ultralytics/assets'
an image path, video file, directory, URL, or
device ID for live feeds. Supports a wide
range of formats and sources, enabling flexible
application across different types of input.
conf float 0.25 Sets the minimum confidence threshold for
detections. Objects detected with confidence
below this threshold will be disregarded.
Adjusting this value can help reduce false
positives.
iou float 0.7 Intersection Over Union (IoU) threshold for
Non-Maximum Suppression (NMS). Lower
values result in fewer detections by
eliminating overlapping boxes, useful for
reducing duplicates.
imgsz int or 640 Defines the image size for inference. Can be a
tuple single integer 640 for square resizing or a
(height, width) tuple. Proper sizing can
improve detection accuracy and processing
speed.
rect bool True If enabled, minimally pads the shorter side of
the image until it’s divisible by stride to
improve inference speed. If disabled, pads the
image to a square during inference.
half bool False Enables half-precision (FP16) inference, which
can speed up model inference on supported
GPUs with minimal impact on accuracy.
device str None Specifies the device for inference (e.g., cpu,
cuda:0 or 0). Allows users to select between
CPU, a specific GPU, or other compute
devices for model execution.
batch int 1 Specifies the batch size for inference (only
works when the source is a directory, video file
or .txt file). A larger batch size can provide
higher throughput, shortening the total
amount of time required for inference.
max_det int 300 Maximum number of detections allowed per
image. Limits the total number of objects the
model can detect in a single inference,
preventing excessive outputs in dense scenes.

228
Argument Type Default Description
vid_stride int 1 Frame stride for video inputs. Allows skipping
frames in videos to speed up processing at the
cost of temporal resolution. A value of 1
processes every frame, higher values skip
frames.
stream_buffer
bool False Determines whether to queue incoming frames
for video streams. If False, old frames get
dropped to accommodate new frames
(optimized for real-time applications). If True,
queues new frames in a buffer, ensuring no
frames get skipped, but will cause latency if
inference FPS is lower than stream FPS.
visualize bool False Activates visualization of model features
during inference, providing insights into what
the model is “seeing”. Useful for debugging
and model interpretation.
augment bool False Enables test-time augmentation (TTA) for
predictions, potentially improving detection
robustness at the cost of inference speed.
agnostic_nmsbool False Enables class-agnostic Non-Maximum
Suppression (NMS), which merges overlapping
boxes of different classes. Useful in multi-class
detection scenarios where class overlap is
common.
classes list[int] None Filters predictions to a set of class IDs. Only
detections belonging to the specified classes
will be returned. Useful for focusing on
relevant objects in multi-class detection tasks.
retina_masksbool False Returns high-resolution segmentation masks.
The returned masks ([Link]) will match
the original image size if enabled. If disabled,
they have the image size used during
inference.
embed list[int] None Specifies the layers from which to extract
feature vectors or embeddings. Useful for
downstream tasks like clustering or similarity
search.
project str None Name of the project directory where
prediction outputs are saved if save is
enabled.

229
Argument Type Default Description
name str None Name of the prediction run. Used for creating
a subdirectory within the project folder,
where prediction outputs are stored if save is
enabled.
stream bool False Enables memory-efficient processing for long
videos or numerous images by returning a
generator of Results objects instead of loading
all frames into memory at once.
verbose bool True Controls whether to display detailed inference
logs in the terminal, providing real-time
feedback on the prediction process.

Visualization arguments:

Argument Type Default Description


show bool False If True, displays the annotated images or videos in
a window. Useful for immediate visual feedback
during development or testing.
save bool False or Enables saving of the annotated images or videos to
True file. Useful for documentation, further analysis, or
sharing results. Defaults to True when using CLI &
False when used in Python.
save_frames bool False When processing videos, saves individual frames as
images. Useful for extracting specific frames or for
detailed frame-by-frame analysis.
save_txt bool False Saves detection results in a text file, following the
format [class] [x_center] [y_center]
[width] [height] [confidence]. Useful for
integration with other analysis tools.
save_conf bool False Includes confidence scores in the saved text files.
Enhances the detail available for post-processing
and analysis.
save_crop bool False Saves cropped images of detections. Useful for
dataset augmentation, analysis, or creating focused
datasets for specific objects.
show_labels bool True Displays labels for each detection in the visual
output. Provides immediate understanding of
detected objects.

230
Argument Type Default Description
show_conf bool True Displays the confidence score for each detection
alongside the label. Gives insight into the model’s
certainty for each detection.
show_boxes bool True Draws bounding boxes around detected objects.
Essential for visual identification and location of
objects in images or video frames.
line_width None or None Specifies the line width of bounding boxes. If None,
int the line width is automatically adjusted based on
the image size. Provides visual customization for
clarity.

Exploring other Computer Vision Applications

Let’s set up Jupyter Notebook optimized for headless Raspberry Pi camera work and develop-
ment:

pip install jupyter jupyterlab notebook


jupyter notebook --generate-config

To run Jupyter Notebook, run the command (change the IP address for yours):

jupyter notebook --ip=[Link] --no-browser

On the terminal, you can see the local URL address and its Token to open the notebook. Copy
and paste it into the Browser.

Environment Setup and Dependencies

import time
import numpy as np
from PIL import Image
from ultralytics import YOLO
import [Link] as plt

Here we have all the necessary libraries, which we installed automatically when we installed
Ultralytics.

231
• Time: Performance measurement and benchmarking
• NumPy: Numerical computations and array operations
• PIL (Python Imaging Library): Image loading and manipulation
• Ultralytics YOLO: Core YOLO functionality
• Matplotlib: Visualization and plotting results

Model Configuration and Loading

model_path= "./models/[Link]"
task = "detect"
verbose = False

model = YOLO(model_path, task, verbose)

• Model Selection: YOLOv11n (nano) is chosen for its balance of speed and accuracy
• Task Specification: We will select detect, which in fact is the default for the model.
But remember that YOLO supports multiple computer vision tasks, which will be explored
later.
• Verbose Control: output model information during model initialization

Performance Characteristics

Let’s open the previous bus image using PIL

source = [Link]("./images/[Link]")

And run an inference in the source:

results = [Link](source, save=False, imgsz=640, conf=0.5, iou=0.3)

From the inference results info, we can see that the first time an inference is run, the latency is
greater.

# First inference
0: 640x480 4 persons, 1 bus, 7528.3ms

# Second inference
0: 640x480 4 persons, 1 bus, 2822.1ms

232
The dramatic difference between the first inference (7.5s) and subsequent inferences (2.8s)
illustrates:

• Model Loading Overhead: Initial inference includes model loading time


• Optimization Effects: Subsequent inferences benefit from cached optimizations

Results Object Structure

Let’s explore the YOLO’s output structure:

result = results[0]
# - boxes, keypoints, masks, names
# - orig_img, orig_shape, path
# - speed metrics

Bounding Box Analysis

[Link] # Class IDs: tensor([5., 0., 0., 0., 0.])


[Link] # Confidence scores
[Link] # Bounding box coordinates

• Coordinate Systems: On [Link], we can get different bounding box formats


(xyxy, xywh, normalized):
– xywh: Tensor with bounding box coordinates in center_x, center_y, width, height
format, in pixels.
– xywhn: Normalized center_x, center_y, width, height, scaled to the image dimen-
sions, values in .
– xyxy: Tensor of boxes as x1, y1, x2, y2 in pixels, representing the top-left and
bottom-right corners.
– xyxyn: Normalized x1, y1, x2, y2, scaled by image width and height, values in .

8. Visualization and Customization

The Ultralytics plot() can be customized to show as the detection result, for example, only
the bounding boxes:

im_bgr = [Link](boxes=True, labels=False, conf=False)

233
img = [Link](im_bgr[..., ::-1])

[Link](figsize=(6, 6))
[Link](img)
#[Link]('off') # This turns off the axis numbers
[Link]("YOLO Result")
[Link]()

234
235
Customization Options:
The plot() method in Ultralytics YOLO Results object accepts several arguments to control
what is visualized on the image, including boxes, masks, keypoints, confidences, labels, and
more. Common Arguments for plot()

• boxes (bool): Show/hide bounding boxes. Default is True.


• conf (bool): Show/hide confidence scores. Default is True.
• labels (bool): Show/hide class labels. Default is True.
• masks (bool): Show/hide segmentation masks (when available, e.g. in segment tasks).
• kpt_line (bool): Draw lines connecting pose keypoints (skeleton diagram). Default is
True in pose tasks.
• line_width (int): Set annotation line thickness.
• font_size (int): Set font size for text annotations.
• show (bool): If True, immediately display the image (interactive environments).

Exploring Other Computer Vision Tasks

Image Classification

As explored in previous chapters, the output of an image classifier is a single class label and a
confidence score. Image classification is useful when we need to know only what class an image
belongs to and don’t need to know where objects of that class are located or what their exact
shape is.

model_path= "./models/[Link]"
task = "clasification"

model = YOLO(model_path, task, verbose)

Note that a specific variation of the model, for image classification, will be downloaded. Now,
let’s do an inference, using the same bus image:

results = [Link](source, save=False)


result = results[0]

The model inference will return:

0: 224x224 minibus 0.57, police_van 0.34, trolleybus 0.04, recreational_vehicle 0.01, streetc

Speed: 5233.9ms preprocess, 3355.1ms inference, 28.2ms postprocess per image at shape (1, 3,

236
We can check the top5 inference results using Python:

classes = [Link].top5
classes

And we will get:

[654, 734, 874, 757, 829]

What is related to:

for id in classes:
print([Link][id])

minibus
police_van
trolleybus
recreational_vehicle
streetcar

And the top 5 probabilities:

probs = [Link]()
probs

[0.5710113048553467,
0.33745330572128296,
0.04209813103079796,
0.014150412753224373,
0.005880324635654688]

The Top 1 result can be extract directly from the result:

print([Link][[Link].top1],
round([Link](), 2))

minibus 0.57

Instance Segmentation

237
model_path= "./models/[Link]"
task = "segment"

model = YOLO(model_path, task, verbose)

Note that a specific variation of the model, for instance segmentation, will be downloaded.
Now, lt’s use another image for testing:

source = [Link]("./images/[Link]")

Display the image

[Link](figsize=(6, 6))
[Link](source)
#[Link]('off') # This turns off the axis numbers
[Link]("Original Image")
[Link]()

And run the inference:

238
results = [Link](source, save=False)
result = results[0]

Display the result:

im_bgr = [Link](boxes=False, conf=False, masks=True)


img = [Link](im_bgr[..., ::-1])
[Link](figsize=(6, 6))
[Link](img)
[Link]('off') # This turns off the axis numbers
[Link]("YOLO Segmentation Result")
[Link]()

Pose Estimation

Download the model:

model_path= "./models/[Link]"
task = "pose"

239
model = YOLO(model_path, task, verbose)

Running the Inference on the beatles image:

source = [Link]("./images/[Link]")
results = [Link](source, save=False)
result = results[0]

Showing the human pose keypoint detection and skeleton visualization.

im_bgr = [Link](boxes=False, conf=False, kpt_line=True)


img = [Link](im_bgr[..., ::-1])
[Link](figsize=(6, 6))
[Link](img)
[Link]('off') # This turns off the axis numbers
[Link]("YOLO Pose Estimation Result")
[Link]()

240
Training YOLO on a Customized Dataset

Object Detection Project

We will now develop a customized object detection project from the data collected and labelled
with Roboflow. The training and deployment will be done in Python using a CoLab and
Ultralytics functions.

We will use with YOLO, the same dataset previously used to train the SSD-
MobileNet V2 and FOMO models.

As a reminder, we are assuming we are in an industrial facility that must sort and count wheels
and special boxes.

241
Each image can have three classes:

• Background (no objects)


• Box
• Wheel

The Dataset

Return to our “Boxe versus Wheel” dataset, labeled on Roboflow. On the Download Dataset,
instead of Download a zip to computer option done for training on Edge Impulse Studio, we
will opt for Show download code. This option will open a pop-up window with a code snippet
that should be pasted into our training notebook.

242
For training, let’s choose one model (let’s say YOLOv8) and adapt one of the publicly available
examples from Ultralytics, then run it on Google Colab. Below, you can find my adaptation:

• YOLOv8 Box versus Wheel Dataset Training [Open In Colab]

Critical points on the Notebook:

1. Run it with GPU (the NVidia T4 is free)


2. Install Ultralytics using PIP.

243
3. Now, you can import the YOLO and upload your dataset to the CoLab, pasting the
Download code that we get from Roboflow. Note that our dataset will be mounted under
/content/datasets/:

4. It is essential to verify and change the file [Link] with the correct path for the images
(copy the path on each images folder).

names:
- box
- wheel
nc: 2
roboflow:
license: CC BY 4.0
project: box-versus-wheel-auto-dataset
url: [Link]

244
version: 5
workspace: marcelo-rovai-riila
test: /content/datasets/Box-versus-Wheel-auto-dataset-5/test/images
train: /content/datasets/Box-versus-Wheel-auto-dataset-5/train/images
val: /content/datasets/Box-versus-Wheel-auto-dataset-5/valid/images

5. Define the main hyperparameters that you want to change from default, for example:
MODEL = '[Link]'
IMG_SIZE = 640
EPOCHS = 25 # For a final project, you should consider at least 100 epochs

6. Run the training (using CLI):


!yolo task=detect mode=train model={MODEL} data={[Link]}/[Link] epochs={EPO

Figure 5: image-20240910111319804

The model took a few minutes to be trained and has an excellent result (mAP50 of 0.995). At the
end of the training, all results are saved in the folder listed, for example: /runs/detect/train/.
There, you can find, for example, the confusion matrix.

245
7. Note that the trained model ([Link]) is saved in the folder /runs/detect/train/weights/.
Now, you should validate the trained model with the valid/images.

!yolo task=detect mode=val model={HOME}/runs/detect/train/weights/[Link] data={[Link]

The results were similar to training.

8. Now, we should perform inference on the images left aside for testing

246
!yolo task=detect mode=predict model={HOME}/runs/detect/train/weights/[Link] conf=0.25 sourc

The inference results are saved in the folder runs/detect/predict. Let’s see some of them:

9. It is advised to export the train, validation, and test results for a Drive at Google. To do
so, we should mount the drive.
from [Link] import drive
[Link]('/content/gdrive')

and copy the content of /runs folder to a folder that you should create in your Drive, for
example:
!scp -r /content/runs '/content/gdrive/MyDrive/10_UNIFEI/Box_vs_Wheel_Project'

Inference with the trained model, using the Raspi

Download the trained model /runs/detect/train/weights/[Link] to your computer. Using


the FileZilla FTP, let’s transfer the [Link] to the Raspi models folder (before the transfer,
you may change the model name, for example, box_wheel_320_yolo.pt).
Using the FileZilla FTP, let’s transfer a few images from the test dataset to .\YOLO\images:
Let’s return to the YOLO folder and use the Python Interpreter:

247
cd ..
python

We will import the YOLO library and define the model to use::

>>> from ultralytics import YOLO


>>> model = YOLO('./models/box_wheel_320_yolo.pt')

Now, let’s define an image and call the inference (we will save the image result this time to
external verification):

>>> img = './images/1_box_1_wheel.jpg'


>>> result = [Link](img, save=True, imgsz=320, conf=0.5, iou=0.3)

Let’s repeat for several images. The inference result is saved on the variable result, and the
processed image on runs/detect/predict8

Using FileZilla FTP, we can send the inference result to our Desktop for verification:

248
We can see that the inference result is excellent! The model was trained based on the smaller
base model of the YOLOv8 family (YOLOv8n). The issue is the latency, around 1 second (or
1 FPS on the Raspi-Zero). We can reduce this latency and convert the model to TFLite or
NCNN.

The model trained with YOLO11 has a latency of around 800 ms, similar to the
result of v8 with ncnn.

In the last section of the notebook, we can find inferences made with the trained YOLO11n
model on a Raspberry 5, which took around 400ms:

The same model, when exported to NCNN, took around 80 ms.

249
Conclusion

This chapter has explored the YOLO model and the implementation of a custom object detector
on a Raspberry Pi, demonstrating the power and potential of running advanced computer
vision tasks on resource-constrained hardware. We’ve covered several vital aspects:

1. Model Comparison: We examined different object detection models, including SSD-


MobileNet, FOMO, and YOLO, comparing their performance and trade-offs on edge
devices.
2. Training and Deployment: Using a custom dataset of boxes and wheels (labeled
on Roboflow), we walked through the process of training models with Ultralytics and
deploying them on a Raspberry Pi.
3. Optimization Techniques: To improve inference speed on edge devices, we explored
various optimization methods, such as format conversion (e.g., to NCNN).
4. Performance Considerations: Throughout the lab, we discussed the balance between
model accuracy and inference speed, a critical consideration for edge AI applications.

250
AS discussed before, the ability to perform object detection on edge devices opens up numerous
possibilities across various domains, including precision agriculture, industrial automation,
quality control, smart home applications, and environmental monitoring. By processing data
locally, these systems can offer reduced latency, improved privacy, and operation in environments
with limited connectivity.
Looking ahead, potential areas for further exploration include: - Implementing multi-model
pipelines for more complex tasks - Exploring hardware acceleration options for Raspberry Pi
- Integrating object detection with other sensors for more comprehensive edge AI systems -
Developing edge-to-cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the power
of AI directly into the physical world, opening up new frontiers in how we interact with and
understand our environment.

Resources

• Dataset (“Box versus Wheel”)


• YOLOv8 Box versus Wheel Dataset Training on CoLab
• YOLO11 Box versus Wheel Dataset Training on CoLab
• Model Predictions with Ultralytics YOLO
• Python Scripts
• Models

251
Counting objects with YOLO

Deploying YOLOv8 on Raspberry Pi Zero 2W for Real-Time Bee Counting at the


Hive Entrance.”

252
Introduction

At the Federal University of Itajuba in Brazil, with the master’s student José Anderson Reis
and Professor José Alberto Ferreira Filho, we are exploring a project that delves into the
intersection of technology and nature. This tutorial will review our first steps and share our
observations on deploying YOLOv8, a cutting-edge machine learning model, on the compact

253
and efficient Raspberry Pi Zero 2W (Raspi-Zero). We aim to estimate the number of bees
entering and exiting their hive—a task crucial for beekeeping and ecological studies.
Why is this important? Bee populations are vital indicators of environmental health, and their
monitoring can provide essential data for ecological research and conservation efforts. However,
manual counting is labor-intensive and prone to errors. By leveraging the power of embedded
machine learning, or tinyML, we automate this process, enhancing accuracy and efficiency.

Figure 6: img

This tutorial will cover setting up the Raspberry Pi, integrating a camera module, optimizing
and deploying YOLOv8 for real-time image processing, and analyzing the data gathered.

Estimating the number of Bees

For our project at the university, we are preparing to collect a dataset of bees at the entrance
of a beehive using the same camera connected to the Raspberry Pi. The images should be
collected every 10 seconds. With the Arducam OV5647, the horizontal Field of View (FoV)
is 53.5o , which means that a camera positioned at the top of a standard Hive (46 cm) will
capture all of its entrance (about 47 cm).

254
Dataset

The dataset collection is the most critical phase of the project and should take several weeks or
months. For this tutorial, we will use a public dataset: “Sledevic, Tomyslav (2023), “[Labeled
dataset for bee detection and direction estimation on beehive landing boards,” Mendeley Data,
V5, doi: 10.17632/8gb9r2yhfc.5”
The original dataset contains 6,762 images (1920 x 1080), and around 8% (518) of them have
no bees (only background). This is very important in Object Detection, where we should keep
around 10% of the dataset with only background (no objects to be detected).

255
The images contain from zero to up to 61 bees:

We downloaded the dataset (images and annotations) and uploaded it to Roboflow. There,
you should create a free account and start a new project, for example, (“Bees_on_Hive_land-
ing_boards”):

256
We will not enter details about the Roboflow process once many tutorials are
available.

Once the project is created and the dataset is uploaded, you should review the annotations
using the “Auto-Label” Tool. Note that all images with only a background should be saved
w/o any annotations. At this step, you can also add additional images.

257
Once all images are annotated, you should split them into training, validation, and testing.

258
Pre-Processing

The last step with the dataset is preprocessing to generate a final version for training. The
Yolov8 model can be trained with 640 x 640 pixels (RGB) images. Let’s resize all images and
generate augmented versions of each image (augmentation) to create new training examples
from which our model can learn.
For augmentation, we will rotate the images (+/-15o ) and vary the brightness and exposure.

259
This will create a final dataset of 16,228 images.

260
Now, you should export the annotateddataset in a YOLOv8 format. You can download a
zipped version of the dataset to your desktop or get a downloaded code to be used with a
Jupyter Notebook:

261
And that is it! We are prepared to start our training using Google Colab.

The pre-processed dataset can be found at the Roboflow site.

Training YOLOv8 on a Customized Dataset

For training, let’s adapt one of the public examples available from Ultralitytics and run it on
Google Colab:

• yolov8_bees_on_hive_landing_board.ipynb [Open In Colab]

262
Critical points on the Notebook:

1. Run it with GPU (the NVidia T4 is free)


2. Install Ultralytics using PIP.

3. Now, you can import the YOLO and upload your dataset to the CoLab, pasting the
Download code that you get from Roboflow. Note that your dataset will be mounted
under /content/datasets/:

263
4. It is important to verify and change, if needed, the file [Link] with the correct path
for the images:

names:
- bee
nc: 1
roboflow:
license: CC BY 4.0
project: bees_on_hive_landing_boards
url: [Link]
version: 1
workspace: marcelo-rovai-riila
test: /content/datasets/Bees_on_Hive_landing_boards-1test/images
train: /content/datasets/Bees_on_Hive_landing_boards-1/train/images
val: /content/datasets/Bees_on_Hive_landing_boards-1/valid/images

5. Define the main hyperparameters that you want to change from default, for example:
MODEL = '[Link]'
IMG_SIZE = 640
EPOCHS = 25 # For a final project, you should consider at least 100 epochs

6. Run the training (using CLI):


!yolo task=detect mode=train model={MODEL} data={[Link]}/[Link] epochs={EPO

The model took 2.7 hours to train and has an excellent result (mAP50 of 0.984). At the end
of the training, all results are saved in the folder listed, for example: /runs/detect/train3/.
There, you can find, for example, the confusion matrix and the metrics curves per epoch.

264
7. Note that the trained model ([Link]) is saved in the folder /runs/detect/train3/weights/.
Now, you should validade the trained model with the valid/images.

!yolo task=detect mode=val model={HOME}/runs/detect/train3/weights/[Link] data={[Link]

The results were similar to training.

8. Now, we should perform inference on the images left aside for testing

!yolo task=detect mode=predict model={HOME}/runs/detect/train3/weights/[Link] conf=0.25 sour

The inference results are saved in the folder runs/detect/predict. Let’s see some of them:

We can also perform inference with a completely new and complex image from another beehive
with a different background (the beehive of Professor Maurilio of our University). The results
were great (but not perfect and with a lower confidence score). The model found 41 bees.

265
9. The last thing to do is export the train, validation, and test results for your Drive at
Google. To do so, you should mount your drive.
from [Link] import drive
[Link]('/content/gdrive')

and copy the content of /runs folder to a folder that you should create in your Drive, for
example:
!scp -r /content/runs '/content/gdrive/MyDrive/10_UNIFEI/Bee_Project/YOLO/bees_on_hive_l

Inference with the trained model, using the Rasp-Zero

Using the FileZilla FTP, let’s transfer the [Link] to our Rasp-Zero (before the transfer, you
may change the model name, for example, bee_landing_640_best.pt).
The first thing to do is convert the model to an NCNN format:

yolo export model=bee_landing_640_best.pt format=ncnn

266
As a result, a new converted model, bee_landing_640_best_ncnn_model is created in the
same directory.
Let’s create a folder to receive some test images (under Documents/YOLO/:

mkdir test_images

Using the FileZilla FTP, let’s transfer a few images from the test dataset to our Rasp-Zero:

Let’s use the Python Interpreter:

python

As before, we will import the YOLO library and define our converted model to detect bees:

from ultralytics import YOLO


model = YOLO('bee_landing_640_best_ncnn_model')

Now, let’s define an image and call the inference (we will save the image result this time to
external verification):

img = 'test_images/15_bees.jpg'
result = [Link](img, save=True, imgsz=640, conf=0.2, iou=0.3)

267
The inference result is saved on the variable result, and the processed image on
runs/detect/predict9

Using FileZilla FTP, we can send the inference result to our Desktop for verification:

268
let’s go over the other images, analyzing the number of objects (bees) found:

269
Depending on the confidence level, we may see some false positives or negatives. But in general,
with a model trained on a smaller base model in the YOLOv8 family (YOLOv8n) and converted
to NCNN, the results are pretty good, running on an Edge device such as the Rasp-Zero. Also,
note that the inference latency is around 730ms.
For example, by running the inference on [Link], we can find 40 bees. During
the test phase on Colab, 41 bees were found (we only missed one here).

Considerations about the Post-Processing

Our final project should be very simple in terms of code. We will use the camera to capture an
image every 10 seconds. As we did in the previous section, the captured image should serve as
the input to the trained and converted model. We should count the bees in each image and
store the counts in a database (e.g., timestamp: number of bees).
We can do it with a single Python script, or use a Linux system timer, such as cron, to
periodically capture images every 10 seconds, and have a separate Python script process them
as they are saved. This method can be particularly efficient at managing system resources and
is more robust against potential delays in image processing.

270
Setting Up the Image Capture with cron

First, we should set up a cron job to use the rpicam-jpeg command to capture an image
every 10 seconds.

1. Edit the crontab:

• Open the terminal and type crontab -e to edit the cron jobs.
• cron normally doesn’t support sub-minute intervals directly, so we should use a
workaround, such as a loop or a file watcher.

2. Create a Bash Script ([Link]):

• Image Capture: This bash script captures images every 10 seconds using rpicam-
jpeg, a command in the raspijpeg tool. This command lets us control the camera
and capture JPEG images directly from the command line. This is especially useful
because we are looking for a lightweight, straightforward method to capture images
without requiring additional libraries like Picamera or external software. The script
also saves the captured image with a timestamp.
#!/bin/bash
# Script to capture an image every 10 seconds

while true
do
DATE=$(date +"%Y-%m-%d_%H%M%S")
rpicam-jpeg --output test_images/$[Link] --width 640 --height 640
sleep 10
done

• We should make the script executable with chmod +x [Link].


• The script must start at boot or use a @reboot entry in cron to start it automatically.

Setting Up the Python Script for Inference

Image Processing: The Python script continuously monitors the designated directory for
new images, processes each new image using the YOLOv8 model, updates the database with
the count of detected bees, and optionally deletes the image to conserve disk space.
Database Updates: The results, along with the timestamps, are saved in an SQLite database.
For that, a simple option is to use sqlite3.
In short, we need to write a script that continuously monitors the directory for new images,
processes them using a YOLO model, and then saves the results to a SQLite database. Here’s
how we can create and make the script executable:

271
#!/usr/bin/env python3
import os
import time
import sqlite3
from datetime import datetime
from ultralytics import YOLO

# Constants and paths


IMAGES_DIR = 'test_images/'
MODEL_PATH = 'bee_landing_640_best_ncnn_model'
DB_PATH = 'bee_count.db'

def setup_database():
"""
Establishes a database connection and creates the table
if it doesn't exist.
"""
conn = [Link](DB_PATH)
cursor = [Link]()
[Link]('''
CREATE TABLE IF NOT EXISTS bee_counts
(timestamp TEXT, count INTEGER)
''')
[Link]()
return conn

def process_image(image_path, model, conn):


"""
Processes an image to detect objects and logs
the count to the database.
"""
result = [Link](image_path, save=False, imgsz=640, conf=0.2, iou=0.3, verbose=Fals
num_bees = len(result[0].[Link])
timestamp = [Link]().strftime('%Y-%m-%d %H:%M:%S')
cursor = [Link]()
[Link]("INSERT INTO bee_counts (timestamp, count) VALUES (?, ?)",
(timestamp, num_bees)
)
[Link]()
print(f'Processed {image_path}: Number of bees detected = {num_bees}')

def monitor_directory(model, conn):

272
"""
Monitors the directory for new images and processes
them as they appear.
"""
processed_files = set()
while True:
try:
files = set([Link](IMAGES_DIR))
new_files = files - processed_files
for file in new_files:
if [Link]('.jpg'):
full_path = [Link](IMAGES_DIR, file)
process_image(full_path, model, conn)
processed_files.add(file)
[Link](1) # Check every second
except KeyboardInterrupt:
print("Stopping...")
break

def main():
conn = setup_database()
model = YOLO(MODEL_PATH)
monitor_directory(model, conn)
[Link]()

if __name__ == "__main__":
main()

The Python script must be executable, for that:

1. Save the script: For example, as process_images.py.


2. Change file permissions to make it executable:
chmod +x process_images.py

3. Run the script directly from the command line:


./process_images.py

We should consider keeping the script running even after closing the terminal; for that, we can
use nohup or screen:

273
nohup ./process_images.py &

or

screen -S bee_monitor
./process_images.py

Note that we capture images with their own timestamps and log a separate timestamp when the
inference results are saved to the database. This approach can be beneficial for the following
reasons:

1. Accuracy in Data Logging:


• Capture Timestamp: The timestamp associated with each image capture repre-
sents the exact moment the image was taken. This is crucial for applications where
precise timing of events (like bee activity) is important for analysis.
• Inference Timestamp: This timestamp indicates when the image was processed
and the results were recorded in the database. This can differ from the capture time
due to processing delays or if the image processing is batched or queued.
2. Performance Monitoring:
• Separate timestamps enable us to monitor the performance and efficiency of your
image processing pipeline. We can measure the delay between image capture and
result logging, which helps optimize the system for real-time processing needs.
3. Troubleshooting and Audit:
• Separate timestamps provide a better audit trail and troubleshooting data. If there
are issues with the image processing or data recording, having distinct timestamps
can help isolate whether delays or problems occurred during capture, processing, or
logging.

Script For Reading the SQLite Database

Here is an example of a code to retrieve the data from the database:

#!/usr/bin/env python3
import sqlite3

def main():
db_path = 'bee_count.db'
conn = [Link](db_path)

274
cursor = [Link]()
query = "SELECT * FROM bee_counts"
[Link](query)
data = [Link]()
for row in data:
print(f"Timestamp: {row[0]}, Number of bees: {row[1]}")
[Link]()

if __name__ == "__main__":
main()

Adding Environment data

Besides bee counting, environmental data, such as temperature and humidity, are essential
for monitoring the bee-have health. Using a Rasp-Zero, it is straightforward to add a digital
sensor such as the DHT-22 to get this data.

Environmental data will be part of our final project. If you want to know more about connecting
sensors to a Raspberry Pi and, even more, how to save the data to a local database and send

275
it to the web, follow this tutorial: From Data to Graph: A Web Journey With Flask and
SQLite.

Conclusion

In this tutorial, we have thoroughly explored integrating the YOLOv8 model with a Raspberry
Pi Zero 2W to address the practical, pressing task of counting (or, better, “estimating”) bees
at a beehive entrance. Our project underscores the robust capability of embedding advanced
machine learning technologies within compact edge computing devices, highlighting their
potential impact on environmental monitoring and ecological studies.
This tutorial provides a step-by-step guide to deploying the YOLOv8 model in practice. We
demonstrate a tangible real-world application by optimizing it for edge computing, improving
efficiency and processing speed (using the NCNN format). This not only serves as a functional
solution but also as an instructional tool for similar projects.

276
The technical insights and methodologies shared in this tutorial are the basis for the complete
work to be developed at our university in the future. We envision further development, such as
integrating additional environmental sensing capabilities and refining the model’s accuracy and
processing efficiency. Implementing alternative energy solutions, such as the proposed solar
power setup, will enhance the project’s sustainability and applicability in remote or underserved
locations.

Resources

The Dataset paper, Notebooks, and PDF version are in the Project repository.

277
Image Classification with EXECUTORCH

Implementing efficient image classification using PyTorch EXECUTORCH on edge devices

Introduction

Image classification is a fundamental computer vision task that powers countless real-world
applications—from quality control in manufacturing to wildlife monitoring, medical diagnostics,
and smart home devices. In the edge AI landscape, the ability to run these models efficiently on

278
resource-constrained devices has become increasingly critical for privacy-preserving, low-latency
applications.
In the chapter Image Classification Fundamentals, we explored image classification with
TensorFlow Lite and demonstrated how to deploy efficient neural networks on the Raspberry
Pi. That tutorial covered the complete workflow from model conversion to real-time camera
inference, achieving excellent results with the MobileNet V2 architecture and a real dataset
(CIFAR-10).
This chapter takes a parallel approach using PyTorch EXECUTORCH—Meta’s modern
solution for edge deployment. Rather than replacing our TFLite knowledge, this chapter
expands your edge AI toolkit, giving us the flexibility to choose the right framework for our
specific needs.

What is EXECUTORCH?

EXECUTORCH is PyTorch’s official solution for deploying machine learning models on edge
devices, from smartphones and embedded systems to microcontrollers and IoT devices. Released
in 2023, it represents Meta’s commitment to bringing the entire PyTorch ecosystem to edge
computing.
Core Capabilities:

• Native PyTorch Integration: Seamless workflow from model training to edge deploy-
ment without switching frameworks
• Efficient Execution: Optimized runtime designed specifically for resource-constrained
devices
• Broad Portability: Runs on diverse hardware platforms (ARM, x86, specialized
accelerators)
• Flexible Backend System: Extensible delegate architecture for hardware-specific
optimizations
• Quantization Support: Built-in integration with PyTorch’s quantization tools for
model compression

Why EXECUTORCH for Edge AI?

EXECUTORCH offers compelling advantages for edge deployment:


1. Unified Workflow If we are training models in PyTorch, EXECUTORCH provides a
natural deployment path without framework switching. This eliminates conversion errors and
maintains model fidelity from training to deployment.

279
2. Modern Architecture Built from the ground up for edge computing with contemporary
best practices, EXECUTORCH incorporates lessons learned from previous mobile deployment
frameworks.
3. Comprehensive Quantization Native support for various quantization techniques
(dynamic, static, quantization-aware training) enables significant model size reduction with
minimal accuracy loss.
4. Extensible Backend System The delegate system allows seamless integration with
hardware accelerators (XNNPACK for CPU optimization, QNN for Qualcomm chips, CoreML
for Apple devices, and more).
5. Active Development Backed by Meta with rapid iteration and strong community support,
ensuring the framework evolves with edge AI needs.
6. Growing Model Zoo Access to pretrained models specifically optimized for edge deploy-
ment, with consistent performance across devices.

Framework Comparison: EXECUTORCH vs TensorFlow Lite

Understanding when to choose each framework is crucial for effective edge deployment:

Feature EXECUTORCH TensorFlow Lite


Training PyTorch TensorFlow/Keras
Framework
Maturity Newer (2023+) Mature (2017+)
Model Format .pte .tflite (.lite)
Quantization PyTorch native quantization TF quantization-aware
training
Backend Delegate system (XNNPACK, QNN, Delegates (GPU, NNAPI,
Acceleration CoreML) Hexagon)
Community Rapidly growing Large, established
Hardware Support Expanding quickly Extensive, mature
Learning Curve Easier for PyTorch users Easier for TF/Keras users
Documentation Growing, modern Comprehensive, mature
Industry Adoption Increasing in research Widespread in production

The Reality: Both Are Excellent Choices


In practice, both frameworks achieve similar goals with different philosophies. Our choice often
comes down to:

1. Our training framework preference

280
2. Team expertise and existing infrastructure
3. Specific hardware requirements
4. Project timeline and maturity needs

This chapter demonstrates that transitioning between frameworks is straightforward, allowing


us to make informed decisions based on project needs rather than framework limitations.

Setting Up the Environment

Updating the Raspberry Pi

First, ensure that the Raspberry Pi is up to date:

sudo apt update


sudo apt upgrade -y
sudo reboot # Reboot to ensure all updates take effect

Installing Required System-Level Libraries

Install Python tools, camera libraries, and build dependencies for PyTorch:

sudo apt install -y python3-pip python3-venv python3-picamera2


sudo apt install -y libcamera-dev libcamera-tools libcamera-apps
sudo apt install -y libopenblas-dev libjpeg-dev zlib1g-dev libpng-dev

Picamera2 Installation Test


We can test the camera with:

rpicam-hello --list-cameras

281
We should see that the OV5647 cam is installed.

Now, let’s create a test script to verify everything works:


camera_capture.py

import numpy as np
from picamera2 import Picamera2
import time

print(f"NumPy version: {np.__version__}")

# Initialize camera
picam2 = Picamera2()

config = picam2.create_preview_configuration(main={"size":(640,480)})
[Link](config)
[Link]()

# Wait for camera to warm up


[Link](2)

print("Camera working in the system!")

# Capture image
picam2.capture_file("camera_capture.jpg")
print("Image captured: cam_test.jpg")

# Stop camera

282
[Link]()
[Link]()

A test image should be created in the current directory

Setting up a Virtual Environment

First, let’s confirm the System Python version:

python --version

If we use the latest Raspberry Pi OS (based on Debian Trixie), it should be:


3.13.5
As of today (January 2026), ExecuTorch officially supports only Python 3.10 to 3.12; Python
3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships with
Python 3.13 by default, we’ll need to install a compatible Python version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python.

If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.

Install pyenv Dependencies

sudo apt update


sudo apt install -y build-essential libssl-dev zlib1g-dev \
libbz2-dev libreadline-dev libsqlite3-dev curl git \
libncursesw5-dev xz-utils tk-dev libxml2-dev \
libxmlsec1-dev libffi-dev liblzma-dev \
libopenblas-dev libjpeg-dev libpng-dev cmake

Install pyenv

# Download and install pyenv


curl [Link] | bash

283
Configure Shell

Add pyenv to ~/.bashrc:

cat >> ~/.bashrc << 'EOF'

# pyenv configuration
export PYENV_ROOT="$HOME/.pyenv"
[[ -d $PYENV_ROOT/bin ]] && export PATH="$PYENV_ROOT/bin:$PATH"
eval "$(pyenv init -)"
EOF

Reload the shell:

source ~/.bashrc

Verify if pyenv is installed:

pyenv --version

Install Python 3.11 (or 3.12)

# See available versions


pyenv install --list | grep " 3.11"

# Install Python 3.11.14 (latest 3.11 stable)


pyenv install 3.11.14

# Or install Python 3.12.3 if you prefer


# pyenv install 3.12.12

This will take a few minutes to compile.

Create ExecuTorch Workspace

284
cd Documents
mkdir EXECUTORCH
cd EXECUTORCH

# Set Python 3.11.14 for this directory


pyenv local 3.11.14

# Verify
python --version # Should show Python 3.11.14

Create Virtual Environment

python -m venv executorch-venv


source executorch-venv/bin/activate

# Verify if we're using the correct Python


which python
python --version

To exit the virtual environment later:


deactivate

Install Python Packages

Ensure we’re in the virtual environment (venv)


pip install --upgrade pip
pip install numpy pillow matplotlib opencv-python

Verify installation:

pip list | grep -E "(numpy|pillow|opencv)"

285
PyTorch and EXECUTORCH Installation

Installing PyTorch for Raspberry Pi

PyTorch provides pre-built wheels for ARM64 architecture (Raspberry Pi 3/4/5).


For Raspberry Pi 4/5 (aarch64):

# Install PyTorch (CPU version for ARM64)


pip install torch torchvision --index-url \
[Link]

For the Raspberry Pi Zero 2 W (32-bit ARM), we may need to build from
source or use lighter alternatives, which are not covered here.

Verify PyTorch installation:

python -c "import torch; print(f'PyTorch version: \


{torch.__version__}')"

We will get, for example, PyTorch version: 2.9.1+cpu

Installing EXECUTORCH Runtime

EXECUTORCH can be installed via pip:

pip install executorch

Building from Source (Optional - for latest features):


If we want the absolute latest features or need to customize:

# Clone the repository


git clone [Link]
cd executorch

# Install dependencies
./install_requirements.sh

# Install EXECUTORCH in development mode


pip install -e .

286
Verifying the Setup

Let’s verify our setup with a test script. Create setup_test.py (for example, using nano):

import torch
import numpy as np
from PIL import Image
import executorch

print("=" * 50)
print("SETUP VERIFICATION")
print("=" * 50)

# Check versions
print(f"PyTorch version: {torch.__version__}")
print(f"NumPy version: {np.__version__}")
print(f"PIL version: {Image.__version__}")
print(f"EXECUTORCH available: {executorch is not None}")

# Test basic PyTorch functionality


x = [Link](3, 224, 224)
print(f"\nCreated test tensor with shape: {[Link]}")

# Test PIL
test_img = [Link]('RGB', (224, 224), color='red')
print(f"Created test PIL image: {test_img.size}")

print("\n� Setup verification complete!")


print("=" * 50)

Run it:

python setup_test.py

Expected output (the versions can be different):

==================================================
SETUP VERIFICATION
==================================================
PyTorch version: 2.9.1+cpu
NumPy version: 2.2.6
PIL version: 12.1.0

287
EXECUTORCH available: True

Created test tensor with shape: [Link]([3, 224, 224])


Created test PIL image: (224, 224)

� Setup verification complete!


==================================================

Image Classification using MobileNet V2

Working directory:

cd Documents
cd EXECUTORCH
mkdir IMG_CLASS
cd IMG_CLASS
mkdir MOBILENET
cd MOBILENET
mkdir models images notebooks

Making inference with Torch

Load an image from the internet, for example, a cat: "[Link]


And save it in the images folder as “[Link]”:

wget "[Link] \
-O ./images/[Link]

Now, let’s create a test program where we should take into consideration:

1. First run - Downloads model & labels (and saves them)


2. Preprocessing - MobileNetV2 expects 224x224 images with ImageNet normalization
3. torch.no_grad() -Disables gradient calculation for faster inference
4. Timing - Measures only inference time, not preprocessing
5. Softmax - Converts raw outputs to probabilities
6. Top-5 - Shows the 5 most likely classes

288
and save it as img_class_test_torch.py:

import torch
import [Link] as transforms
from torchvision import models
from PIL import Image
import time
import json
import [Link]
import os

# Paths
MODEL_PATH = "models/mobilenet_v2.pth"
LABELS_PATH = "models/imagenet_labels.json"
IMAGE_PATH = "images/[Link]"

# Download and save ImageNet labels (only first time)


if not [Link](LABELS_PATH):
print("Downloading ImageNet labels...")
LABELS_URL = "[Link]
imagenet-simple-labels/master/[Link]"
with [Link](LABELS_URL) as url:
labels = [Link](url)

# Save labels locally


with open(LABELS_PATH, 'w') as f:
[Link](labels, f)
print(f"Labels saved to {LABELS_PATH}")
else:
print("Loading labels from disk...")
with open(LABELS_PATH, 'r') as f:
labels = [Link](f)

# Load or download model


if not [Link](MODEL_PATH):
print("Downloading MobileNetV2 model...")
model = models.mobilenet_v2(pretrained=True)
[Link]()
[Link](model.state_dict(), MODEL_PATH)
print(f"Model saved to {MODEL_PATH}")
else:
print("Loading model from disk...")
model = models.mobilenet_v2()

289
model.load_state_dict([Link](MODEL_PATH, map_location='cpu'))
[Link]()

# Define image preprocessing


preprocess = [Link]([
[Link](256),
[Link](224),
[Link](),
[Link](mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]),
])

# Load and preprocess image


print(f"\nLoading image from {IMAGE_PATH}...")
img = [Link](IMAGE_PATH)
img_tensor = preprocess(img)
batch = img_tensor.unsqueeze(0)

# Perform inference with timing


print("Running inference...")
start_time = [Link]()

with torch.no_grad():
output = model(batch)

inference_time = ([Link]() - start_time) * 1000

# Get predictions
probabilities = [Link](output[0], dim=0)
top5_prob, top5_idx = [Link](probabilities, 5)

# Display results
print("\n" + "="*50)
print("CLASSIFICATION RESULTS")
print("="*50)
print(f"Inference Time: {inference_time:.2f} ms\n")
print("Top 5 Predictions:")
print("-"*50)

for i in range(5):
idx = top5_idx[i].item()
prob = top5_prob[i].item()

290
print(f"{i+1}. {labels[idx]:20s} - {prob*100:.2f}%")

print("="*50)

The result:

Loading image from images/[Link]...


Running inference...

==================================================
CLASSIFICATION RESULTS
==================================================
Inference Time: 86.12 ms

Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================

The inference was OK, taking 86ms (first time). We can also verify the size of the saved Torch
model

ls -lh ./models/mobilenet_v2.pth

Which has 14Mb.

Exporting Models to EXECUTORCH Format

Unlike TensorFlow Lite, where we downloaded pre-converted .tflite models, with EXECU-
TORCH, we typically export PyTorch models to the .pte (PyTorch EXECUTORCH) format
ourselves. This gives us full control over the export process.

291
Understanding the Export Process

The EXECUTORCH export process involves several steps:

1. Load a PyTorch model (pretrained or custom)


2. Trace/script the model (convert to TorchScript)
3. Export to EXECUTORCH format (.pte file)

Optional optimization steps:

• Quantization (before or during export)


• Backend delegation (XNNPACK, QNN, etc.)
• Memory planning optimization

The complete ExecuTorch pipeline:

1. export() → Captures the model graph


2. to_edge() → Converts to Edge dialect
3. to_executorch() → Lowers to ExecuTorch format
4. .buffer → Gets the binary data to save

PyTorch Model (.pt/.pth)



[Link]() # Export to ExportedProgram

to_edge() # Convert to Edge dialect

to_executorch() # Generate EXECUTORCH program

.pte file # Ready for edge deployment

Exporting MobileNet V2 to ExecuTorch

Let’s export a MobileNet V2 model to EXECUTORCH basic format. Creating a Python script
as convert_mobv2_executorch.py

import torch
from torchvision import models
from [Link] import to_edge
from [Link] import export

# Paths
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"

292
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2.pte"

print("Loading PyTorch model...")


# Load the saved model
model = models.mobilenet_v2()
model.load_state_dict([Link](PYTORCH_MODEL_PATH, map_location='cpu'))
[Link]()

# Create example input (batch_size=1, channels=3, height=224, width=224)


example_input = ([Link](1, 3, 224, 224),)

print("Exporting to ExecuTorch format...")

# Step 1: Export to EXIR (ExecuTorch Intermediate Representation)


print(" 1. Capturing model with [Link]...")
exported_program = export(model, example_input)

# Step 2: Convert to Edge dialect


print(" 2. Converting to Edge dialect...")
edge_program = to_edge(exported_program)

# Step 3: Convert to ExecuTorch program


print(" 3. Lowering to ExecuTorch...")
executorch_program = edge_program.to_executorch()

# Step 4: Save as .pte file


print(" 4. Saving to .pte file...")
with open(EXECUTORCH_MODEL_PATH, "wb") as f:
[Link](executorch_program.buffer)

print(f"\n? Model successfully exported to {EXECUTORCH_MODEL_PATH}")

# Display file sizes for comparison


import os
pytorch_size = [Link](PYTORCH_MODEL_PATH)/(1024*1024)
executorch_size = [Link](EXECUTORCH_MODEL_PATH)/(1024*1024)

print("\n" + "="*50)
print("MODEL SIZE COMPARISON")
print("="*50)
print(f"PyTorch model: {pytorch_size:.2f} MB")
print(f"ExecuTorch model: {executorch_size:.2f} MB")

293
print(f"Reduction: {((pytorch_size - executorch_size) \
/pytorch_size * 100):.1f}%")
print("="*50)

Runing the export script:

python export_mobv2_executorch.py

We will get:

Loading PyTorch model...


Exporting to ExecuTorch format...
1. Capturing model with [Link]...
2. Converting to Edge dialect...
3. Lowering to ExecuTorch...
4. Saving to .pte file...

? Model successfully exported to models/mobilenet_v2.pte

==================================================
MODEL SIZE COMPARISON
==================================================
PyTorch model: 13.60 MB
ExecuTorch model: 13.58 MB
Reduction: 0.2%
==================================================

The basic ExecuTorch conversion doesn’t compress the model much - it’s mainly for runtime
efficiency. To get real size reduction, we need quantization, which we will explore later.
But first, let’s do an inference test using the converted model.
Runing the script mobv2_executorch.py:

import torch
import [Link] as transforms
from PIL import Image
import time
import json
from [Link].portable_lib import _load_for_executorch

# Paths
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2.pte"

294
LABELS_PATH = "models/imagenet_labels.json"
IMAGE_PATH = "images/[Link]"

# Load labels
print("Loading labels...")
with open(LABELS_PATH, 'r') as f:
labels = [Link](f)

# Load ExecuTorch model


print(f"Loading ExecuTorch model from {EXECUTORCH_MODEL_PATH}...")
model = _load_for_executorch(EXECUTORCH_MODEL_PATH)

# Define image preprocessing (same as PyTorch)


preprocess = [Link]([
[Link](256),
[Link](224),
[Link](),
[Link](mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]),
])

# Load and preprocess image


print(f"Loading image from {IMAGE_PATH}...")
img = [Link](IMAGE_PATH)
img_tensor = preprocess(img)
batch = img_tensor.unsqueeze(0) # Add batch dimension

# Perform inference with timing


print("Running ExecuTorch inference...")
start_time = [Link]()

# ExecuTorch expects a tuple of inputs


output = [Link]((batch,))

inference_time = ([Link]() - start_time) * 1000 # Convert to ms

# Get predictions
output_tensor = output[0] # ExecuTorch returns a list
probabilities = [Link](output_tensor[0], dim=0)
top5_prob, top5_idx = [Link](probabilities, 5)

# Display results

295
print("\n" + "="*50)
print("EXECUTORCH CLASSIFICATION RESULTS")
print("="*50)
print(f"Inference Time: {inference_time:.2f} ms\n")
print("Top 5 Predictions:")
print("-"*50)

for i in range(5):
idx = top5_idx[i].item()
prob = top5_prob[i].item()
print(f"{i+1}. {labels[idx]:20s} - {prob*100:.2f}%")

print("="*50)

As a result, we got a similar inference result, but a much higher latency (almost 2.5 seconds),
which was unexpected.

Loading labels...
Loading ExecuTorch model from models/mobilenet_v2.pte...
Loading image from images/[Link]...
Running ExecuTorch inference...

==================================================
EXECUTORCH CLASSIFICATION RESULTS
==================================================
Inference Time: 2445.78 ms

Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================

That export path produces a generic ExecuTorch CPU graph with reference kernels and no
backend optimizations or fusions, so significantly higher latency than PyTorch is expected for
MobileNet_v2 on a Pi 5.
ExecuTorch is designed to shine when delegated to a backend (XNNPACK, OpenVINO, etc.),
where large subgraphs are lowered into highly optimized kernels. Without a delegate, most of

296
the graph runs on the generic portable path, which is known to be significantly slower than
PyTorch for many models.
So, let’s export the .pth model again with a CPU‑optimized backend (e.g., XNNPACK) and
run with that backend enabled; this alone should reduce latency when compared with the naïve
interpreter path.
Here’s the corrected conversion script with XNNPACK delegation (convert_mobv2_xn-
[Link]):

import torch
from torchvision import models
from [Link] import to_edge
from [Link] import export
from [Link].xnnpack_partitioner \
import XnnpackPartitioner

# Paths
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2_xnnpack.pte"

print("Loading PyTorch model...")


model = models.mobilenet_v2()
model.load_state_dict([Link](PYTORCH_MODEL_PATH, map_location='cpu'))
[Link]()

# Create example input


example_input = ([Link](1, 3, 224, 224),)

print("Exporting to ExecuTorch with XNNPACK backend...")

# Step 1: Export to EXIR


print(" 1. Capturing model with [Link]...")
exported_program = export(model, example_input)

# Step 2: Convert to Edge dialect with XNNPACK partitioner


print(" 2. Converting to Edge dialect with XNNPACK delegation...")
edge_program = to_edge(exported_program)

# Step 3: Partition for XNNPACK backend


print(" 3. Delegating to XNNPACK backend...")
edge_program = edge_program.to_backend(XnnpackPartitioner())

# Step 4: Convert to ExecuTorch program

297
print(" 4. Lowering to ExecuTorch...")
executorch_program = edge_program.to_executorch()

# Step 5: Save as .pte file


print(" 5. Saving to .pte file...")
with open(EXECUTORCH_MODEL_PATH, "wb") as f:
[Link](executorch_program.buffer)

print(f"\n? Model successfully exported to {EXECUTORCH_MODEL_PATH}")

# Display file size


import os
pytorch_size = [Link](PYTORCH_MODEL_PATH) / (1024 * 1024)
executorch_size = [Link](EXECUTORCH_MODEL_PATH) / (1024 * 1024)

print("\n" + "="*50)
print("MODEL SIZE COMPARISON")
print("="*50)
print(f"PyTorch model: {pytorch_size:.2f} MB")
print(f"ExecuTorch+XNNPACK: {executorch_size:.2f} MB")
print("="*50)

Runing it we get:

Loading PyTorch model...


Exporting to ExecuTorch with XNNPACK backend...
1. Capturing model with [Link]...
2. Lowering to Edge with XNNPACK delegation...
3. Converting to ExecuTorch...
4. Saving to .pte file...

? Model successfully exported to models/mobilenet_v2_xnnpack.pte

==================================================
MODEL SIZE COMPARISON
==================================================
PyTorch model: 13.60 MB
ExecuTorch+XNNPACK: 13.35 MB
==================================================

We did not gain in terms of size, but let’s run the same inference script as before, with this
new converted model, to inspect the latency:

298
the result:

Loading labels...
Loading ExecuTorch model from models/mobilenet_v2_xnnpack.pte...
Loading image from images/[Link]...
Running ExecuTorch inference...

==================================================
EXECUTORCH CLASSIFICATION RESULTS
==================================================
Inference Time: 19.95 ms

Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================

Now, the ExecuTorch runtime detects the backend automatically from the .pte file metadata.
We have achieved much faster inference: 20ms instead of 2445ms. This latency is, in fact,
several times faster than PyTorch.
Why XNNPACK is so fast:

• � ARM NEON SIMD optimizations


• � Multi-threading on Raspberry Pi’s 4 cores
• � Operator fusion and memory optimization
• � Cache-friendly memory access patterns

This demonstrates:

1. ExecuTorch (basic) without a backend = don’t use in production


2. ExecuTorch + XNNPACK = production-ready edge AI
3. Raspberry Pi 5 can do 50+ inferences/second at this speed!

Now we can add quantization to get an even smaller model size while maintaining (or even
increasing) this speed!

299
Model Quantization

Quantization reduces model size and can further improve inference speed. EXECUTORCH
supports PyTorch’s native quantization.
Quantization Overview
Quantization is a technique that reduces the precision of numbers used in a model’s computations
and stored weights—typically from 32-bit floats to 8-bit integers. This reduces the model’s
memory footprint, speeds up inference, and lowers power consumption, often with minimal loss
in accuracy.
Quantization is especially important for deploying models on edge devices such as wearables,
embedded systems, and microcontrollers, which often have limited compute, memory, and
battery capacity. By quantizing models, we can make them significantly more efficient and
better suited to these resource-constrained environments.
Quantization in ExecuTorch
ExecuTorch uses torchao as its quantization library. This integration allows ExecuTorch to
leverage PyTorch-native tools for preparing, calibrating, and converting quantized models.
Quantization in ExecuTorch is backend-specific. Each backend defines how models should be
quantized based on its hardware capabilities. Most ExecuTorch backends use the torchao PT2E
quantization flow, which works with models exported with [Link] and enables tailored
quantization for each backend.
For a quantized XNNPACK .pte we need a different pipeline: PT2E quantization (with
XNNPACKQuantizer), then lowering with XnnpackPartitioner before to_executorch(). Oth-
erwise, we will hit errors or get an undelegated model.
For the conversion, we need: (1) calibrate with real, preprocessed images, and (2) compute the
quantized .pte size after you actually write the file.
First, let us create a small calib_images/ folder (e.g., 50–100 natural images across a few
classes). A simple way is to reuse an existing dataset (e.g., CIFAR‑10) and save 50–100 images
into calib_images/ with an ImageNet‑style folder layout.
The script gen_calibr_images.py will: • Download CIFAR‑10. • Pick 10 classes × 10 images
each = 100 images. • Save them under calib_images/<class_name>/img_XXX.jpg.

import os
from pathlib import Path

import torch
from torchvision import datasets, transforms
from [Link] import save_image

300
# Where to store calibration images
OUT_ROOT = Path("calib_images")
OUT_ROOT.mkdir(parents=True, exist_ok=True)

# 1) Load a small, natural-image dataset (CIFAR-10)


transform = [Link]() # we will NOT normalize here
dataset = datasets.CIFAR10(
root="data",
train=True,
download=True,
transform=transform,
)

# 2) Map label index -> class name (CIFAR-10 has 10 classes)


classes = [Link] # ['airplane', 'automobile', ..., 'truck']

# 3) Choose how many classes and images per class


num_classes = 10
images_per_class = 10 # 10 x 10 = 100 images

# 4) Collect and save images


counts = {cls: 0 for cls in classes[:num_classes]}

for img, label in dataset:


cls_name = classes[label]
if cls_name not in counts:
continue
if counts[cls_name] >= images_per_class:
continue

# Make class subdir


class_dir = OUT_ROOT / cls_name
class_dir.mkdir(parents=True, exist_ok=True)

idx = counts[cls_name]
out_path = class_dir / f"img_{idx:04d}.jpg"
save_image(img, out_path)

counts[cls_name] += 1

# Stop when we have enough


if all(counts[c] >= images_per_class for c in counts):

301
break

print("Saved calibration images:")


for cls_name, n in [Link]():
print(f" {cls_name}: {n} images")
print(f"\nRoot folder: {OUT_ROOT.resolve()}")

Let’s use the inference script convert_mobv2_xnnpack_int8.py, which is the same inference
script as before, with this new int8 converted model to inspect the latency:

import os
import torch
import [Link] as models
import [Link] as transforms
import [Link] as datasets

from [Link] import export


from [Link].pt2e.quantize_pt2e import (
prepare_pt2e,
convert_pt2e,
)
from [Link].xnnpack_quantizer import (
get_symmetric_quantization_config,
XNNPACKQuantizer,
)
from [Link].xnnpack_partitioner import (
XnnpackPartitioner,
)
from [Link] import to_edge_transform_and_lower

PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
EXECUTORCH_QUANTIZED_PATH = "models/mobilenet_v2_quantized_xnnpack.pte"
CALIB_IMAGES_DIR = "calib_images" # <-- put some natural images here

# 1) Load FP32 model


model = models.mobilenet_v2()
model.load_state_dict([Link](PYTORCH_MODEL_PATH, map_location="cpu"))
[Link]()

# Example input only defines shapes for export


example_inputs = ([Link](1, 3, 224, 224),)

302
# 2) Configure XNNPACK quantizer (global symmetric config)
qparams = get_symmetric_quantization_config(is_per_channel=True)
quantizer = XNNPACKQuantizer()
quantizer.set_global(qparams)

# 3) Export float model for PT2E and prepare for quantization


exported = [Link](model, example_inputs)
training_ep = [Link]()
prepared = prepare_pt2e(training_ep, quantizer)

# 4) Calibration with REAL images using SAME preprocessing as inference


calib_transform = [Link]([
[Link](256),
[Link](224),
[Link](),
[Link](mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]),
])

calib_dataset = [Link](CALIB_IMAGES_DIR,
transform=calib_transform)
calib_loader = [Link](
calib_dataset, batch_size=1, shuffle=True
)

print(f"Calibrating on {len(calib_dataset)} images from {CALIB_IMAGES_DIR}...")

num_calib = min(100, len(calib_dataset)) # or adjust


with torch.no_grad():
for i, (calib_img, _) in enumerate(calib_loader):
if i >= num_calib:
break
prepared(calib_img)

# 5) Convert calibrated model to quantized model


quantized_model = convert_pt2e(prepared)

# 6) Export quantized model and lower to XNNPACK, then to ExecuTorch


exported_quant = export(quantized_model, example_inputs)

et_program = to_edge_transform_and_lower(
exported_quant,

303
partitioner=[XnnpackPartitioner()],
).to_executorch()

# 7) Save .pte and compute sizes


with open(EXECUTORCH_QUANTIZED_PATH, "wb") as f:
et_program.write_to_file(f)

pytorch_size = [Link](PYTORCH_MODEL_PATH)/(1024*1024)
quantized_size = [Link](EXECUTORCH_QUANTIZED_PATH)/(1024*1024)

print("\n" + "="*60)
print("MODEL SIZE COMPARISON")
print("="*60)
print(f"PyTorch (FP32): {pytorch_size:6.2f} MB")
print(f"ExecuTorch Quantized (INT8): {quantized_size:6.2f} MB")
print(f"Size reduction: {((pytorch_size - quantized_size) \
/ pytorch_size * 100):5.1f}%")
print(f"Savings: {pytorch_size - quantized_size:6.2f} MB")
print("="*60)

Runing the script, we get:

Calibrating on 100 images from calib_images...

============================================================
MODEL SIZE COMPARISON
============================================================
PyTorch (FP32): 13.60 MB
ExecuTorch Quantized (INT8): 3.59 MB
Size reduction: 73.6%
Savings: 10.01 MB
============================================================

The quantized (int8) model achieved 74% size reduction: ~3.5 MB (similar to TFLite). Let’s
see about the inference latency, runing mobv2_xnnpack_int8.py.

Loading labels...
Loading ExecuTorch model from models/mobilenet_v2_quantized_xnnpack.pte...
Loading image from images/[Link]...
Running ExecuTorch inference (Quantized INT8)...

304
==================================================
EXECUTORCH QUANTIZED INT8 RESULTS
==================================================
Inference Time: 13.56 ms
Output dtype: torch.float32

Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 51.01%
2. Egyptian Mau - 34.11%
3. lynx - 7.54%
4. tabby cat - 6.17%
5. plastic bag - 0.37%
==================================================

Slightly higher top‑1 probabilities in the INT8 model are normal and do not indicate
a problem by themselves. Quantization slightly changes the logits, and softmax can
become a bit “sharper” or “flatter” even when top‑1 remains correct.

Model Size/Performance Comparison

Model Configuration File Size Size Reduction Latency


Float32 (basic export) 13.58 MB Baseline 2.5 s
Float32 + XNNPACK 13.35 MB ~0% 20 ms
INT8 + XNNPACK 3.59 MB ~75% 14 ms

NOTE

• Looking at Htop, we can see that only one of the Pi’s cores is at 100%. This indicates that
the shipped Python runtime currently runs our ExecuTorch/XNNPACK model effectively
single‑threaded on Pi.
• To exploit all four cores, the next step would be to move inference into a small C++
wrapper that sets the ExecuTorch threadpool size before executing the graph. With the
pure‑Python path, there is no clean public knob to change it yet. We will not explore it
here.

Making Inferences with EXECUTORCH

Now that we have our EXECUTORCH models, let’s explore them in more detail for image
classification using a Jupyter Notebook!

305
Setting up Jupyter Notebook

Set up Jupyter Notebook for interactive development:

pip install jupyter jupyterlab notebook


jupyter notebook --generate-config

To run the Jupyter notebook on the Raspberry Pi desktop, run:

jupyter notebook

and open the URL with the token


To run Jupyter Notebook on your computer (headless), run the command below, replacing
with your Raspberry Pi’s IP address:
To get the IP Address, we can use the command: hostname -I

jupyter notebook --ip=[Link] --no-browser

Access it from another device using the provided token in your web browser.

The Project folder

We must be sure that we have this project folder structure:

EXECUTORCH/MOBILENET/
��� convert_mobv2_executorch.py
��� convert_mobv2_xnnpack.py
��� convert_mobv2_xnnpack_int8.py
��� mobv2_executorch.py
��� mobv2_xnnpack.py
��� mobv2_xnnpack_int8.py
��� calib_images/
��� data/
��� models/
� ��� mobilenet_v2.pth # Float32 pytorch model
� ��� mobilenet_v2.pte # Float32 conv model
� ��� mobilenet_v2_xnnpack.pte # Float32 conv model
� ��� mobilenet_v2_quantized_xnnpack.pte # Quantized conv model
� ��� imagenet_labels.json # Labels
��� images/ # Test images

306
� ��� [Link]
� ��� camera_capture.jpg
��� notebooks/
��� image_classification_executorch.ipynb

Loading and Running a Model

Inside the folder ‘notebooks’, on the project space IMAGE_CLASS/MOBILENET, create a new
notebook: image_classification_executorch.ipynb.

Setup and Verification

# Import required libraries


import os
import time
import json
import [Link]
import numpy as np
import [Link] as plt
from PIL import Image
import torch
from torchvision import transforms
import executorch
from [Link].portable_lib import _load_for_executorch

print("=" * 50)
print("SETUP VERIFICATION")
print("=" * 50)

# Check versions
print(f"PyTorch version: {torch.__version__}")
print(f"NumPy version: {np.__version__}")
print(f"PIL version: {Image.__version__}")
print(f"EXECUTORCH available: {executorch is not None}")

# Test basic PyTorch functionality


x = [Link](3, 224, 224)
print(f"\nCreated test tensor with shape: {[Link]}")

# Test PIL

307
test_img = [Link]('RGB', (224, 224), color='red')
print(f"Created test PIL image: {test_img.size}")

print("\n� Setup verification complete!")


print("=" * 50)

We get:

==================================================
SETUP VERIFICATION
==================================================
PyTorch version: 2.9.1+cpu
NumPy version: 2.2.6
PIL version: 12.1.0
EXECUTORCH available: True

Created test tensor with shape: [Link]([3, 224, 224])


Created test PIL image: (224, 224)

� Setup verification complete!


==================================================

Download Test Image

• Download test image for example from:


– “[Link]
– And save it on the ../images folder as “[Link]”

img_path = "../images/[Link]"

# Load and display


img = [Link](img_path)
[Link](figsize=(6, 6))
[Link](img)
[Link]("Original Image")
#[Link]('off')
[Link]()

print(f"Image size: {[Link]}")


print(f"Image mode: {[Link]}")

308
Image size: (1600, 1598)
Image mode: RGB

Load EXECUTORCH Model

Note: You need to export a model first using the export_mobv2_executorch.py


script.
If you don’t have a model yet, run the export script first:

• python export_mobv2_executorch.py

Let’s verify what the models in the folder ../models:

imagenet_labels.json mobilenet_v2_quantized_xnnpack.pte
mobilenet_v2.pte mobilenet_v2_xnnpack.pte
mobilenet_v2.pth

The conversions were performed using the Python scripts in the previous sections.

309
# Load the EXECUTORCH model
model_path = "../models/mobilenet_v2.pte"

try:
model = _load_for_executorch(model_path)
print(f"Model loaded successfully from: {model_path}")
#print(f" Available methods: {model.method_names}")

# Check file size


file_size = [Link](model_path) / (1024 * 1024) # MB
print(f"Model size: {file_size:.2f} MB")

except FileNotFoundError:
print(f"� Model not found: {model_path}")
print("\nPlease run the export script first:")
print(" python export_mobilenet.py")

Model loaded successfully from: ../models/mobilenet_v2.pte


Model size: 13.58 MB

Download ImageNet Labels

# Download and save ImageNet labels (if you do not have it)
LABELS_PATH = "../models/imagenet_labels.json"

if not [Link](LABELS_PATH):
print("Downloading ImageNet labels...")
LABELS_URL = "[Link]
imagenet-simple-labels/master/[Link]"
with [Link](LABELS_URL) as url:
labels = [Link](url)

# Save labels locally


with open(LABELS_PATH, 'w') as f:
[Link](labels, f)
print(f"Labels saved to {LABELS_PATH}")
else:
print("Loading labels from disk...")
with open(LABELS_PATH, 'r') as f:
labels = [Link](f)

310
Check the labels:

print(f"\nTotal classes: {len(labels)}")


print(f"Sample labels: {labels[280:285]}")

Total classes: 1000


Sample labels: ['grey fox', 'tabby cat', 'tiger cat', 'Persian cat', 'Siamese cat']

Image Preprocessing

A preprocessing pipeline is needed because ExecuTorch only runs the exported core network; it
does not include the input normalization logic that MobileNet v2 expects, and the model will
give incorrect predictions if the input tensor is not in the exact format it was trained on.
What MobileNet v2 expects For typical PyTorch MobileNet v2 models (ImageNet‑pretrained):
• Input shape: 3‑channel RGB tensor of size. • Value range: floating-point values, usually in
float32 after dividing by 255. • Normalization: per‑channel mean/std (ImageNet) normalization,
e.g., mean=0.485, 0.456, 0.406, std=0.229, 0.224, 0.225.
These steps (resize, convert to tensor, normalize) are not “optional decorations”; they are part
of the functional definition of the model’s expected input distribution.

Define preprocessing pipeline

preprocess = [Link]([
[Link](256), # Resize to 256
[Link](224), # Center crop to 224x224
[Link](), # Convert to tensor [0, 1]
[Link]( # Normalize with ImageNet stats
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])

Apply preprocessing

311
input_tensor = preprocess(img)
print(f" Input shape: {input_tensor.shape}")
print(f" Input dtype: {input_tensor.dtype}")

Input shape: [Link]([3, 224, 224])


Input dtype: torch.float32

Add batch dimension: [1, 3, 224, 224]

input_batch = input_tensor.unsqueeze(0)

print(f" Input shape: {input_batch.shape}")


print(f" Input dtype: {input_batch.dtype}")
print(f" Value range: [{input_batch.min():.3f}, {input_batch.max():.3f}]")

Input shape: [Link]([1, 3, 224, 224])


Input dtype: torch.float32
Value range: [-2.084, 2.309]

The Preprocessing is complete!

Run Inference

For inference, we should run a forward pass of the model in inference mode (torch.no_grad()),
measure the time, and print basic information about the outputs.
torch.no_grad() is a context manager that disables gradient calculation inside its block.
During inference, we do not need gradients, so disabling them:

• Saves memory (no computation graph is stored).


• Can speed up computation slightly.
• Everything computed inside this block will have requires_grad=False, so we cannot
call .backward() on it.

312
# Run inference
with torch.no_grad():
start_time = [Link]()
outputs = [Link]((input_batch,))
inference_time = [Link]() - start_time

print(f"Inference completed in {inference_time*1000:.2f} ms")


print(f"Output type: {type(outputs)}")
print(f"Output shape: {outputs[0].shape}")

Inference completed in 2478.74 ms


Output type: <class 'list'>
Output shape: [Link]([1, 1000])

type(outputs) tells us what container the model returned. Often this is a tuple or list when
working with exported/ExecuTorch‑style models, e.g., <class 'tuple'>.
That container may hold one or more tensors (e.g., logits, auxiliary outputs).

• outputs[0] accesses the first element of that container (usually the main output tensor),
and .shape prints its dimensions (For image classification, this is often batch_size,
num_classes).

Process and Display Results

Now we should take the model’s raw scores (logits) for a single image, convert them into proba-
bilities with softmax, select the top‑5 most likely classes, and print them nicely formatted.

• outputs[0][0] selects the first element in the batch, giving a 1D tensor of logits of
length num_classes.
• [Link](..., dim=0) applies the softmax function along that
1D dimension, turning logits into probabilities that sum to 1.

# Apply softmax to get probabilities


probabilities = [Link](outputs[0][0], dim=0)

# Get top 5 predictions


top5_prob, top5_indices = [Link](probabilities, 5)

# Display results
print("\n" + "="*60)

313
print("TOP 5 PREDICTIONS")
print("="*60)
print(f"{'Class':<35} {'Probability':>10}")
print("-"*60)

for i in range(5):
label = labels[top5_indices[i]]
prob = top5_prob[i].item() * 100
print(f"{label:<35} {prob:>9.2f}%")

print("="*60)

============================================================
TOP 5 PREDICTIONS
============================================================
Class Probability
------------------------------------------------------------
tiger cat 12.85%
Egyptian cat 9.75%
tabby 6.09%
lynx 1.70%
carton 0.84%
============================================================

Create Reusable Classification Function

For simplicity and reuse across other tests, let’s create a reusable function that builds on what
was done so far.

def classify_image_executorch(img_path, model_path, labels_path,


top_k=5, show_image=True):
"""
Classify an image using EXECUTORCH model

Args:
img_path: Path to input image
model_path: Path to .pte model file
labels_path: Path to labels text file
top_k: Number of top predictions to return
show_image: Whether to display the image

314
Returns:
inference_time: Inference time in ms
top_indices: Indices of top k predictions
top_probs: Probabilities of top k predictions
"""
# Load image
img = [Link](img_path).convert('RGB')

# Display image
if show_image:
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
[Link]('Input Image')
[Link]()

print(f"Image Path: {img_path}")

# Load model
print(f"Model Path {model_path}")
model_size = [Link](model_path) / (1024 * 1024)
print(f"Model size: {model_size:6.2f} MB")

model = _load_for_executorch(model_path)

# Preprocess
preprocess = [Link]([
[Link](256),
[Link](224),
[Link](),
[Link](
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])

input_tensor = preprocess(img)
input_batch = input_tensor.unsqueeze(0)

# Inference
with torch.no_grad():
start_time = [Link]()

315
outputs = [Link]((input_batch,))
inference_time = ([Link]() - start_time)*1000

# Process results
probabilities = [Link](outputs[0][0], dim=0)
top_prob, top_indices = [Link](probabilities, top_k)

# Load labels
with open(labels_path, 'r') as f:
labels = [Link](f)

# Display results
print(f"\nInference time: {inference_time:.2f} ms")
print("\n" + "="*60)
print(f"{'[PREDICTION]':<35} {'[Probability]':>15}")
print("-"*60)

for i in range(top_k):
label = labels[top_indices[i]]
prob = top_prob[i].item() * 100
print(f"{label:<35} {prob:>14.2f}%")

print("="*60)

return inference_time, top_indices, top_prob

print("� Classification function defined!")

� Classification function defined!

Classification Function Test

# Test with the cat image


inf_time, indices, probs = classify_image_executorch(
img_path="../images/[Link]",
model_path="../models/mobilenet_v2.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)

316
We can also check what is retrurned fron the function

inf_time, indices, probs

(2445.200204849243,
tensor([282, 285, 287, 281, 728]),
tensor([0.4744, 0.3761, 0.0691, 0.0622, 0.0047]))

Using the XNNPACK accelerated backend

Note: We need to export a model using the convert_mobv2_xnnpack.py script first.

317
# Test with the cat image
inf_time, indices, probs = classify_image_executorch(
img_path="../images/[Link]",
model_path="../models/mobilenet_v2_xnnpack.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)

The inference time was reduced from +2.5s to around -20ms

Quantized model - XNNPACK accelerated backend

Note: We need to export a model first using the convert_mobv2_xnnpack_int8.py script.

318
# Test with the cat image
inf_time, indices, probs = classify_image_executorch(
img_path="../images/[Link]",
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)

==> Even faster inference with a lower model in size

Slightly higher probabilities in the INT8 model are normal and do not indicate a
problem by themselves. Quantization slightly changes the logits, and softmax can
become a bit “sharper” or “flatter” even when top‑1 remains correct.

319
Camera Integration

We essentially have two different Python worlds: system Python 3.13 (where the camera stack
is wired up) and our 3.11 virtual env (where ExecuTorch is installed). To run ExecuTorch on
live frames from the Pi camera, we need to bridge those worlds.

Why the camera “only works” in 3.13

• Recent Raspberry Pi OS uses Picamera2 on top of libcamera as the recom-


mended interface.
• The Picamera2/libcamera Python bindings are usually installed into the system
Python and are not trivially pip‑installable into arbitrary venvs or other Python
versions.
• Once we create a separate 3.11 environment, it will not automatically see the
Picamera2/libcamera bindings under 3.13, so imports fail or the camera device
is not accessible from that environment.

We will use a two‑process solution: capture in 3.13, infer in 3.11. For that, we should run
a small capture service under Python 3.13 that:

• Grabs frames from the Pi camera (Picamera2 / libcamera).


• Sends frames to your ExecuTorch process (3.11) over a local channel (e.g., ZeroMQ,
TCP/UDP socket, shared memory, filesystem (write JPEG/PNG to a temp directory
and signal), or a simple HTTP server.

The 3.11 process (under venev) receives the frame, decodes it, runs the preprocessing pipeline
(resize, normalize), then calls ExecuTorch for inference..

Image Capture

Outside of the ExecuTorch env and folder, we will create a folder (CAMERA).

Documents/
��� EXECUTORCH/MOBILENET/ # Python 3.11
��� CAMERA/ # Python 3.13
��� camera_capture.py
��� camera_capture.jpg

There we will run the script camera_capture.py):

320
import numpy as np
from picamera2 import Picamera2
import time

print(f"NumPy version: {np.__version__}")

# Initialize camera

picam2 = Picamera2()

config = picam2.create_preview_configuration(main={"size":(640,480)})
[Link](config)
[Link]()

# Wait for camera to warm up

[Link](2)

print("Camera working in isolated venv!")

# Capture image

picam2.capture_file("camera_capture.jpg")
print("Image captured: camera_capture.jpg")

# Stop camera

[Link]()
[Link]()

Runing the script, we um get an image that will be stored on:

• /Documents/CAMERA/camera_capture.jpg

Looking from the notebook folder, the image path will be:

../../../../CAMERA/camera_capture.jpg

Let’s run the same function used with the test image:

321
# Test the quantized model with the captured image
inf_time, indices, probs = classify_image_executorch(
img_path="../../../../CAMERA/camera_capture.jpg",
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)

322
Performance Benchmarking

Let’s now define a function to run inference several times for each model and compare their
performance.

323
def benchmark_inference(model_path, num_runs=50):
"""
Benchmark model inference speed
"""
print(f"Benchmarking model: {model_path}")
print(f"Number of runs: {num_runs}\n")

# Load model
model = _load_for_executorch(model_path)

# Create dummy input


dummy_input = [Link](1, 3, 224, 224)

# Warmup (10 runs)


print("Warming up...")
for _ in range(10):
with torch.no_grad():
_ = [Link]((dummy_input,))

# Benchmark
print(f"Running benchmark...")
times = []
for i in range(num_runs):
start = [Link]()
with torch.no_grad():
_ = [Link]((dummy_input,))
[Link]([Link]() - start)

times = [Link](times) * 1000 # Convert to ms

# Print statistics
print("\n" + "="*50)
print("BENCHMARK RESULTS")
print("="*50)
print(f" Mean: {[Link]():.2f} ms")
print(f" Median: {[Link](times):.2f} ms")
print(f" Std: {[Link]():.2f} ms")
print(f" Min: {[Link]():.2f} ms")
print(f" Max: {[Link]():.2f} ms")
print("="*50)

# Plot distribution

324
[Link](figsize=(12, 4))

# Histogram
[Link](1, 2, 1)
[Link](times, bins=20, edgecolor='black', alpha=0.7)
[Link]([Link](), color='red', linestyle='--',
label=f'Mean: {[Link]():.2f} ms')
[Link]('Inference Time (ms)')
[Link]('Frequency')
[Link]('Inference Time Distribution')
[Link]()
[Link](alpha=0.3)

# Time series
[Link](1, 2, 2)
[Link](times, marker='o', markersize=3, alpha=0.6)
[Link]([Link](), color='red', linestyle='--',
label=f'Mean: {[Link]():.2f} ms')
[Link]('Run Number')
[Link]('Inference Time (ms)')
[Link]('Inference Time Over Runs')
[Link]()
[Link](alpha=0.3)

plt.tight_layout()
[Link]()

return times

To recal, we have the folowing converted models:

mobilenet_v2.pte
mobilenet_v2_xnnpack.pte
mobilenet_v2_quantized_xnnpack.pte

Basic (Float32): mobilenet_v2.pte

# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2.pte",

325
num_runs=50
)

XNNPACK Backend (Flot32): mobilenet_v2_xnnpack.pte

# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2_xnnpack.pte",
num_runs=50
)

326
Quantization (INT8): mobilenet_v2_quantized_xnnpack.pte

# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
num_runs=50
)

327
Performance Comparison Table

Based on actual benchmarking results on Raspberry Pi 5:

Mean Median
Model Configuration (ms) (ms) Std Dev (ms) File Size (MB) Latency
Float32 (basic) 2440 2440 2.17 13.58 +600×
Float32 + 11.24 10.84 1.67 13.35 ~3×
XNNPACK
INT8 + XNNPACK 3.91 3.69 0.55 3.59 1×

Key Observations:

1. XNNPACK Impact: Backend delegation provides an important speedup even without


quantization

328
2. Quantization Benefit: INT8 quantization, besides size reduction, adds additional
speedup beyond XNNPACK
3. Variability: Quantized model shows lower standard deviation, indicating more stable
performance
4. Size-Speed Tradeoff: 75% size reduction (14MB → 3.5MB) with 3× speed improvement

Exploring Custom Models

CIFAR-10 Dataset:

• 10 classes: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck

Figure 7: cifar10

329
• The images in CIFAR-10 are of size 3x32x32 (3-channel color images of 32x32 pixels in
size).

Exporting a Custom Trained Model

Let’s create a Project folder structure as below (some files are shown as they will appear
later)

EXECUTORCH/CIFAR-10/
��� export_cifar10_xnnpack.py
��� inference_cifar10_xnnpack.py
��� models/
� ��� cifar10_model_jit.pt # Float32 pytorch model
� ��� cifar10_xnnpack.pte # Float32 conv model
��� images/ # Test images
� ��� [Link]
��� notebooks/
��� CIFAR-10_Inference_RPI.ipynb

Let’s train a model from scratch on CIFAR-10. For that, we can run the Notebook below on
Google Colab:
cifar10_colab_training.ipynb
From the training, we will have the trained model:
cifar10_model_jit.pt, which should be saved on /models folder
Next, as we did before, we should export the PyTorch model to ExecuTorch, and let’s use
XNNPACK. Run the script: export_cifar10_xnnpack.py, as a result, we have:

330
Runing it, a converted model cifar10_xnnpack.pte will be saved in ./models/ folder.

Running Custom Models on Raspberry Pi

Runing the script inference_cifar10_xnnpack.py, over the “cat” image, we can see that the
converted model is working fine:
python inference_cifar10_xnnpack.py ./images/[Link]

331
And runing 20 times….

332
Despite the exported model being OK, when we make an inference with the original PyTorch
model, in this case (a small model), we will find even lower latencies.

333
In short, our export script is conceptually the right pattern for ExecuTorch+XNNPACK on
Arm, but for this specific small CIFAR‑10 CNN, the overhead of ExecuTorch and partial
XNNPACK delegation on a Pi‑class device can easily make it slower than a well‑optimized
plain PyTorch JIT model.
Optionally, it is possible to explore those models with the notebook:
CIFAR-10_Inference_RPI_Updated.ipynb

Conclusion

This chapter adapted our image classification workflow from TensorFlow Lite to PyTorch
EXECUTORCH, demonstrating that the PyTorch ecosystem provides a powerful and modern
alternative for edge AI deployment on Raspberry Pi devices.
EXECUTORCH represents a significant evolution in edge AI deployment, bringing PyTorch’s
research-friendly ecosystem to production edge devices. While TensorFlow Lite remains
excellent and mature, having EXECUTORCH in your toolkit makes you a more versatile edge
AI practitioner.
The future of edge AI is multi-framework, multi-platform, and rapidly evolving. By mastering
both EXECUTORCH and TensorFlow Lite, you’re positioned to make informed technical
decisions and adapt as the landscape changes.

Remember: The best framework is the one that serves your specific needs. This
tutorial empowers you to make that choice confidently.

334
Key Takeaways

Technical Achievements:

• Successfully set up PyTorch and EXECUTORCH on Raspberry Pi (4/5)


• Learned the complete model export pipeline from PyTorch to .pte format
• Implemented quantization for reduced model size (~3.5MB vs ~14MB)
• Created reusable inference functions for both standard and custom models
• Integrated camera capture with EXECUTORCH inference

EXECUTORCH Advantages:

• Unified ecosystem: Training and deployment in the same framework


• Modern architecture: Built for contemporary edge computing needs
• Flexibility: Easy export of any PyTorch model
• Quantization: Native PyTorch quantization support
• Active development: Continuous improvements from Meta and the community

Comparison with TFLite: Both frameworks achieve similar goals with different philoso-
phies:

• EXECUTORCH: Better for PyTorch users, newer technology, growing ecosystem


• TFLite: More mature, broader hardware support, larger community

The choice between them often comes down to your training framework and specific require-
ments.

Performance Considerations

On Raspberry Pi 4/5, you can expect: - Float32 models: 10-20ms per inference (MobileNet
V2)

• Quantized models: 3-5ms per inference


• Memory usage: 4-15MB, depending on model size

335
Resources

Code Repository

• Tutorial Code Repository


• Model Export and Inference Scripts
• Notebooks

Official Documentation

PyTorch & EXECUTORCH:

• PyTorch Official Website


• EXECUTORCH Documentation
• EXECUTORCH GitHub Repository
• PyTorch Mobile
• Deep Learning with PyTorch: A 60 Minute Blitz

Quantization:

• PyTorch Quantization
• Quantization API Tutorial

Models:

• Torchvision Models
• Pretrained Model Deployment Guide

Hardware Resources

• Raspberry Pi Official Documentation


• Picamera2 Library
• ARM Architecture Optimization

Books

• Edge AI Engineering e-book- by Prof. Marcelo Rovai, UNIFEI


• Machine Learning Systems - by Prof. Vijay Janapa Reddi, Harvard University
• AI and ML for Coders in PyTorch by Laurence Moroney

336
Beyond CPU - Hardware Acceleration for Edge
AI

Hardware Acceleration with MemryX MX3 on Raspberry Pi 5.

337
Introduction

Throughout this course, we’ve explored various approaches to deploying AI models at the edge.
We started with TensorFlow Lite running on the Raspberry Pi’s CPU, then moved to YOLO
and ExecuTorch with optimized backends like XNNPACK. While these software optimizations
significantly improve performance, they still rely on the general-purpose CPU to execute neural
network operations.
In this chapter, we’ll take the next step: dedicated hardware acceleration. We’ll use the
MemryX MX3 M.2 AI Accelerator Module—a specialized processor designed specifically for
neural network inference. The MX3 module contains four AI accelerator chips that can run
deep learning models with dramatically lower latency and power consumption compared to
CPU execution.

Why Hardware Acceleration?

Consider the requirements for real-time edge AI applications: - Latency: Autonomous


systems need predictions in milliseconds - Power efficiency: Battery-powered devices must
conserve energy - Throughput: Multi-camera systems may need to process several streams
simultaneously - Cost: System designs often cannot afford high-end GPUs
The MX3 addresses these challenges with a unique architecture: - At-memory computing:
All memory is integrated on the accelerator, eliminating bandwidth bottlenecks - Pipelined
dataflow: Optimized for streaming inputs with a batch size of 1 - Floating-point accuracy:
No quantization required (though supported) - Low power: Maximum 10W for four accelerator
chips

338
For learning more about AI Acceleration, please refer to MLSys book and how the
MemryX module works, read the Architecture Overview.

Our goal

By the end of this lab, we will have installed and configured the MX3 hardware on a Raspberry
Pi 5, set up the MemryX SDK and development environment, and gained a clear understanding
of the MX3 compilation and deployment workflow. We will also compile neural network models
for execution on the MX3 accelerator, compare their performance against CPU-based inference
while analyzing the trade-offs, and finally build a complete end-to-end inference pipeline using
the MemryX Python API.

Hardware Installation and Verification

Prerequisites

Before starting this lab, we should have:


Raspberry Pi 5 with M.2 HAT+ adapter (or similar)

IMPORTANT NOTE: MemryX recommends the GeeekPi N04 M.2 2280 HAT
as an excellent choice for the Raspberry Pi 5. It delivers solid power and fits the
2280 MX3 M.2 form factor. Some hats can lead to instabilities, mainly due to PCIe
speed (Gen3). The Raspberry Pi 5 can have stability issues on Gen3.

339
The Raspberry Pi M.2 HAT+ is a good option. It works very well, despite the fact that we
should adapt the MX3 board to it (The MX3 is longer than the hat).
MemryX MX3 M.2 module with the heatsink installed
For heatsink installation, follow the video instructions: [Link]

Installation and Cooling Considerations

It is essential to ensure we have sufficient cooling for the MemryX MX3 M.2 module, or we
may experience thermal throttling and reduced performance. The chips will throttle their
performance if they hit 100 °C.
During normal operation, the current MemryX MX3 temperature and throttle status can be
viewed at any time with:

cat /sys/memx0/temperature

Or measurd continuay every 1 second, for example with the command:

watch -n 1 cat /sys/memx0/temperature

340
Verification

After installing the hardware, turn on the Raspberry Pi and verify the system setup.

ls /dev/memx*

It should return: /dev/memx0


If the device is not detected, see the Troubleshooting section below.
Let’s also check the initial temperature:

The lab temperature at the time of the above measurement was 25 °C.

Software Installation

Create Project Directory

First, we should create a project directory:

cd Documents
mkdir MEMRYX
cd MEMRYX

Python Version Management with Pyenv

Verify your Python version:

python --version

341
If using the latest Raspberry Pi OS (based on Debian Trixie), it should be:
Python 3.13.5
Or, if the OS is the Legacy version:
Python 3.11.2
Important: As of January 2026, MemryX officially supports only Python 3.09 to 3.12.
Python 3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships
with Python 3.13 by default, we’ll need to install a compatible Python version alongside
it.
One solution is to install Pyenv, which allows us to easily manage multiple Python versions for
different projects without affecting the system Python.

If the Raspberry Pi OS is the legacy version, the Python version should be 3.11,
and it is not necessary to install Pyenv.

Installing Pyenv on Debian Trixie

If you need to install Pyenv on Debian Trixie, follow these steps:

# Install dependencies
sudo apt install -y make build-essential libssl-dev zlib1g-dev \
libbz2-dev libreadline-dev libsqlite3-dev wget curl llvm \
libncursesw5-dev xz-utils tk-dev libxml2-dev libxmlsec1-dev \
libffi-dev liblzma-dev

# Install pyenv
curl [Link] | bash

# Add to ~/.bashrc
echo 'export PYENV_ROOT="$HOME/.pyenv"' >> ~/.bashrc
echo 'command -v pyenv >/dev/null || export PATH="$PYENV_ROOT/bin:$PATH"' \
>> ~/.bashrc
echo 'eval "$(pyenv init -)"' >> ~/.bashrc

# Reload shell
source ~/.bashrc

# Install Python 3.11.14


pyenv install 3.11.14

This process takes several minutes as it compiles Python from source.

342
Set Python Version for Project

Once Pyenv and the selected Python version are installed, define it for the project direc-
tory:

pyenv local 3.11.14

Checking the Python version again, we should see: Python 3.11.14.

MemryX Drivers and SDK Installation

The MemryX software stack consists of two main components:

• Drivers (memx-drivers): Kernel-level drivers for PCIe communication with the acceler-
ator hardware
• SDK (memx-accl): Python libraries, neural compiler, runtime, and benchmarking tools

Prepare the System

Install the Linux kernel headers required for driver compilation:

sudo apt install linux-headers-$(uname -r)

Add MemryX Repository and Key

This command downloads the repository’s GPG key for package verification and adds the
MemryX package repository:

wget -qO- [Link] | \


sudo tee /etc/apt/[Link].d/[Link] >/dev/null

echo 'deb [Link] stable main' | \


sudo tee /etc/apt/[Link].d/[Link] >/dev/null

Update and Install Drivers and SDK

sudo apt update


sudo apt install memx-drivers memx-accl

343
Configure Platform Settings

Run the ARM setup utility to configure platform-specific settings. This opens a menu to
select the platform and apply the necessary configurations (e.g., enabling PCIe Gen 3.0 on the
Raspberry Pi 5):

sudo mx_arm_setup

Select the appropriate option for your hardware, and press <OK> in the next page:

After configuration, reboot the system:

sudo reboot

Verify Driver Installation

After rebooting, verify that the MemryX driver is installed by checking its version:

344
apt policy memx-drivers

Install Utilities

Install additional utilities including GUI tools and plugins:

sudo apt install memx-accl-plugins memx-utils-gui

Prepare System Dependencies

Install system libraries required for the Python SDK:

sudo apt update


sudo apt install libhdf5-dev python3-dev cmake python3-venv build-essential

Install Tools (Inside Virtual Environment)

It’s best practice to use a virtual environment to avoid conflicts with system packages.
Create and activate a virtual environment:

python -m venv mx-env


source mx-env/bin/activate

Inside the environment, install the MemryX Python package:

pip3 install --upgrade pip wheel


pip3 install --extra-index-url [Link] memryx

345
Verify the neural compiler is installed:

mx_nc --version

Verification

Verify the complete installation by running the built-in “hello world” benchmark:

mx_bench --hello

With the benchmark results, our MemryX MX3 is properly installed and ready to use.

346
Our First Accelerated Model

Understanding the MX3 Workflow

Working with the MemryX MX3 follows a straightforward four-step workflow that differs from
traditional CPU-based inference:

347
Step 1: Select or Train a Model

Start with a pre-trained model or train your own. MemryX supports models from major
frameworks:

• TensorFlow/Keras (.h5, SavedModel)


• ONNX (.onnx)
• PyTorch (.pt, .pth) (Should be converted to ONNX first)
• TensorFlow Lite (.tflite)

The model remains in its original format—no framework-specific conversions needed yet. For
this lab, we’re using MobileNetV2 from Keras Applications, but we could equally use a custom
model we have trained for a specific task, as we have seen before.
Supported Operations: The MX3 supports most common deep learning operators (convolu-
tions, pooling, activations, etc.). Check the supported operators if using custom architectures.
Unsupported operations will fall back to CPU, though this is rare for standard vision models.

Step 2: Compile with Neural Compiler

The MemryX Neural Compiler (mx_nc) transforms the model into a DFP (Dataflow Package):

mx_nc [options] -m <model_file>

The MemryX Neural Compiler, mx_nc, is a command‑line tool that takes one or more
neural‑network models (Keras, TensorFlow, TFLite, ONNX, etc.) and compiles them into a
MemryX Dataflow Program (DFP) that can run on MemryX accelerators (MXA). Internally,
it does framework import, graph optimization (fusion/splitting, operator expansion, activation
approximation), resource mapping on the MXA cores, and finally emits the DFP used by the
runtime or simulator.
What mx_nc does:

• Compiles models into a single DFP file per compilation, then loads it onto one or more
MXA chips to run inference.
• Supports multi‑model, multi‑stream, and multi‑chip mapping, automatically distributing
models and layers across available MX3 devices for higher throughput.
• Handles mixed‑precision weights (per‑channel 4/8/16‑bit) while keeping activations in
floating point on the accelerator. By default, MemryX quantizes weights to INT8 precision
and activations to BFloat16.
• Can crop pre/post‑processing parts of the graph so the MXA focuses on the core CNN/ML
operators while the host CPU runs the cropped sections.

348
What happens during compilation?

1. Model parsing: Loads the model and extracts the computational graph
2. Graph optimization: Fuses operations, eliminates redundancies
3. Operator mapping: Maps each layer to MX3 hardware instructions
4. Dataflow scheduling: Determines optimal execution order for pipelined processing
5. Memory allocation: Assigns on-chip memory for all intermediate activations
6. Multi-chip distribution: If using multiple chips, partitions the workload

The compiler is surprisingly tolerant—most models compile without any modifications. If a


layer isn’t supported, you’ll get a clear error message indicating which operation failed.
Key command‑line options (high level)

• Model specification:
– -m / --models – input model file(s) (e.g. .h5, .pb, .onnx, TFLite).

– Multi‑model example: mx_nc -v -m model_0.h5 model_1.onnx.


• Input shapes:
– -is, --input_shapes – specify input shape(s) when they cannot be inferred, always
in NHWC order (e.g. "224,224,3").
– Supports "auto" for models where shapes can be inferred: -is "auto"
"300,300,3".
• Cropping / pre‑ and post‑processing control:
– --autocrop – experimental automatic cropping of pre/post‑processing layers. See
the YOLOv8 example further in this chapter.
– --inputs, --outputs – manually set which graph nodes are treated as the MXA
inputs/outputs; everything outside is cropped to run on the host.
– For multi‑model graphs, inputs/outputs of different models are separated with |
(vertical bar).
– --model_in_out – JSON file describing model inputs/outputs for more complex
single or multi‑model cases.
• Multi‑chip / system sizing:
– -c – number of MXA chips; -c 2 means compile for two chips and distribute
workload across them.
• Diagnostics / verbosity:
– -v, -vv, etc. – increase verbosity, useful to inspect graph transformations and
cropping decisions.
• Extensions / unsupported patterns:

349
– --extensions – load Neural Compiler Extensions (.nce files or builtin names) to
add or patch graph handling (e.g., complex transformer subgraphs or unsupported
ops) without a new SDK release.

For the complete option list (including less common flags), run mx_nc -h or consult
the Neural Compiler page in the MemryX Developer Hub, which documents all
arguments and includes usage examples for single‑model, multi‑model, cropping,
and mixed‑precision flows.

Compilation time varies by model complexity:

• Small models (MobileNet): ~30 seconds


• Medium models (ResNet50): ~2 minutes

• Large models (EfficientNet): ~5 minutes

Once compiled, the DFP file is portable across all MX3 hardware.

Step 3: Deploy and Benchmark

Before integrating into our application, we can verify performance with the benchmarking
tool:

mx_bench -d <dfp_file> -f <num_frames>

The benchmarker: - Generates synthetic input data matching the model’s input shape - Runs
warm-up inferences to stabilize performance - Measures throughput (FPS), latency, and chip
utilization - Reports first-inference latency (includes loading overhead)
Why benchmark separately? Real-world applications involve preprocessing (image loading
and resizing) and postprocessing (parsing outputs). Benchmarking isolates pure inference
performance, letting to identify bottlenecks in our full pipeline.

Step 4: Integrate into the Application

Finally, integrate the accelerator into our Python application using the MemryX API:

from memryx import SyncAccl # or AsyncAccl for concurrent processing

# Initialize accelerator with our DFP


accl = SyncAccl(dfp="[Link]")

350
# Run inference
output = [Link](input_data)

# Process results
# ...

# Clean up
[Link]()

Synchronous vs. Asynchronous APIs:

• SyncAccl: Blocking calls, simple to use, good for single-stream processing


• AsyncAccl: Non-blocking, better for multi-stream or real-time applications

The API handles all hardware communication, memory transfers, and scheduling. Our code
just provides input tensors and receives output tensors—the complexity is abstracted away.

Complete Workflow Example

Let’s see the four steps in action with MobileNetV2:

# Step 1: Get a model (already trained)


python3 -c "import tensorflow as tf; \
[Link].MobileNetV2().save('mobilenet_v2.h5');"

# Step 2: Compile to DFP


mx_nc -v -m mobilenet_v2.h5

# Step 3: Benchmark
mx_bench -d mobilenet_v2.dfp -f 1000

# Step 4: Integrate (see full Python script in next section)


python run_inference_mobilenetv2.py

This workflow is remarkably consistent across models and use cases. Once we’ve done it for
one model, adapting to others is straightforward.

Key Takeaway: The MX3 workflow separates compilation (done once) from
inference (done repeatedly). This “compile-once, run-many” approach means the
optimization overhead is amortized over thousands or millions of inferences in
production.

351
Download and Compile MobileNetV2

In Keras Applications, we can find deep learning models that are provided with pre-trained
weights. These models can be used for prediction, feature extraction, and fine-tuning.
Let’s download MobileNetV2, which was used in previous labs:

python3 -c "import tensorflow as tf; \


[Link].MobileNetV2().save('mobilenet_v2.h5');"

The model is saved in the current directory as mobilenet_v2.h5.


Next, we will compile the MobileNetV2 model using the MemryX Neural Compiler. This step
verifies that both the compiler and the SDK tools are installed and functioning as expected:

mx_nc -v -m mobilenet_v2.h5

The compiled model, mobilenet_v2.dfp, is saved in the current folder.

352
What is a DFP file?

The .dfp (Dataflow Package) file is MemryX’s proprietary compiled format. Unlike standard
model formats (H5, ONNX, etc.) that describe the network architecture, a DFP file contains:

• Optimized operator graph: The network restructured for dataflow execution


• Memory layout: Pre-calculated memory allocations for at-memory computing
• Chip mappings: Instructions for distributing work across the four MX3 accelerators
• Quantization parameters: If applicable, the bit-width and scaling factors

The neural compiler (mx_nc) performs this transformation automatically, with no manual
tuning required. The compilation process: 1. Parses the input model (H5, ONNX, TFLite,
etc.) 2. Maps operators to MX3-supported operations 3. Optimizes the dataflow graph 4.
Allocates memory on-chip 5. Generates the DFP binary
This is why compilation takes a few minutes, but inference is blazingly fast—all the optimization
work happens once, upfront.

Benchmarking Performance

Now that the model is compiled, it’s time to deploy it and run a benchmark to test its
performance on the MXA hardware. We will run 1000 frames of random data through the
accelerator to measure performance metrics:

mx_bench -v -d mobilenet_v2.dfp -f 1000

353
Let’s understand what these metrics mean:

• FPS (Frames Per Second): How many images the accelerator can process per second
(~1,200 FPS for MobileNetV2)
• Latency: Time for a single inference (shown as “Avg” in the output)
– Subsequent inferences: True steady-state performance (~2ms)
• Throughput: Total data processed per second

The benchmark runs with random input data, which is why we see consistent performance.
Real-world performance with actual images should be similar once the preprocessing pipeline
is optimized, but we have found bigger latency.

In true dataflow architecture, latency and FPS are not coupled in the traditional
sense; latency does not equal 1/FPS. Note that even though latency is ~2ms in
the above benchmarking results, FPS is not measured to be 1000 ms / 2 ms = 500
FPS; rather, the FPS from the benchmarking results is ~1160.

In MemryX’s dataflow architecture, the “usual” rule (latency = 1/FPS) only applies to
frame‑to‑frame latency, not to end‑to‑end in‑to‑out latency for a single frame. That is why we
see ~2 ms latency per frame, yet still measure around 1160 FPS in MX3 benchmarks.
Two different latencies MemryX explicitly distinguishes two metrics.

• Latency 1 (frame‑to‑frame latency): Time between consecutive outputs once the pipeline
is full. Its reciprocal is FPS.

354
• Latency 2 (full in‑to‑out latency): Time from when the first input frame enters the system
(host + MX3 pipeline) until its output appears. This can be larger, but it does not set
the FPS.

In a streaming, pipelined accelerator like MX3, multiple frames are in‑flight simultaneously, so
the pipeline “fills” once and then produces results at a steady cadence.

Building an Inference Application

Now let’s build a complete inference application that processes real images and compares CPU
vs. MX3 performance.

Prepare Directory Structure

Let’s create subdirectories for organization:

mkdir models
mkdir images

Download Test Image

Load an image from the internet, for example, a cat (for comparison, it is the same as used on
previous chapters):

wget "[Link] \
-O ./images/[Link]

Here is the image:

355
Understanding Input Requirements

All neural networks expect input data in a specific format, determined during training. For
MobileNetV2 trained on ImageNet:

• Input shape: (224, 224, 3) - RGB images at 224x224 pixels


• Batch dimension: Models expect batch inputs, so (1, 224, 224, 3) for single images
• Preprocessing: MobileNetV2 uses specific normalization (scaling pixel values to [-1, 1])
• Color channels: RGB order (not BGR)

The preprocessing must match exactly what was used during training, or accuracy will suffer.

Getting the Labels

For inference, we will need the ImageNet labels. The following function checks if the file exists,
and if not, downloads it:

356
import os, json
from pathlib import Path
import requests

MODELS_DIR = Path("./models")
IMAGENET_JSON = MODELS_DIR / "imagenet_class_index.json"
IMAGENET_JSON_URL = (
"[Link]
imagenet_class_index.json"
)

# ---- one-time label download ----


def ensure_imagenet_labels():
MODELS_DIR.mkdir(parents=True, exist_ok=True)
if IMAGENET_JSON.exists():
return
print("Downloading ImageNet class index...")
resp = [Link](IMAGENET_JSON_URL, timeout=30)
resp.raise_for_status()
IMAGENET_JSON.write_bytes([Link])
print("Saved:", IMAGENET_JSON)

The function load_idx2label() loads the labels into a list:

def load_idx2label():
with open(IMAGENET_JSON, "r") as f:
class_idx = [Link](f)
idx2label = [class_idx[str(k)][1] for k in range(len(class_idx))]
return idx2label

Image Preprocessing

The image used for inference should be preprocessed in the same way as during model training.
[Link].mobilenet_v2.preprocess_input() takes an image of shape (224, 224)
and converts it to (1, 224, 224, 3):

import numpy as np
from PIL import Image
import tensorflow as tf
from tensorflow import keras

357
def load_and_preprocess_image(image_path):
img = [Link](image_path).convert("RGB").resize((224, 224))
arr = [Link](img).astype(np.float32)
arr = [Link].mobilenet_v2.preprocess_input(arr)
arr = np.expand_dims(arr, 0) # Add batch dimension
return arr

Prepare Input Tensor

The processed image will serve as the model’s input tensor (x):

ensure_imagenet_labels()
idx2label_full = load_idx2label() # length 1000 for ImageNet

IMAGE_PATH = Path("./images/[Link]")
x = load_and_preprocess_image(IMAGE_PATH)

Run Inference on MemryX Accelerator (MXA)

Move the models (the original and compiled) to the models folder and set up the paths:

MODELS_DIR = Path("./models")
DFP_PATH = MODELS_DIR / "mobilenet_v2.dfp"
KERAS_PATH = MODELS_DIR / "mobilenet_v2.h5"

Run inference on the compiled model using the MemryX accelerator:

from memryx import SyncAccl

accl = SyncAccl(dfp=str(DFP_PATH))
mxa_outputs = [Link](x)

We get a list/array of outputs. In this case, with a shape of (1, 1000) and a dtype of float32.
This output should be normalized to a NumPy array:

mxa_outputs = [Link](mxa_outputs)
if mxa_outputs.ndim == 3:
mxa_outputs = mxa_outputs[0]

358
Decode the MXA Results

Now, using helper functions to extract top-k predictions:

def topk_from_probs(probs, k=5):


"""
probs: (1, num_classes) or (num_classes,)
Returns [(index, prob)] sorted by prob desc.
"""
probs = [Link](probs)
if [Link] == 2:
probs = probs[0]
# If outputs are logits, uncomment this:
# probs = [Link](probs).numpy()
s = [Link]()
if s > 0:
probs = probs / s
idxs = [Link](probs)[::-1][:k]
return [(int(i), float(probs[i])) for i in idxs]

def label_for(idx, idx2label):


if idx2label is not None and idx < len(idx2label):
return idx2label[idx]
return f"class_{idx}"

We can decode and print the results:

mxa_top5 = topk_from_probs(mxa_outputs, k=5)


print("\nMXA top-5:")
for idx, prob in mxa_top5:
name = label_for(idx, idx2label_full)
print(f" #{idx:4d}: {name:20s} ({prob*100:.1f}%)")

Expected output:

MXA top-5:
# 282: tiger_cat (38.6%)
# 281: tabby (18.3%)
# 285: Egyptian_cat (15.2%)
# 287: lynx (3.9%)
# 478: carton (1.7%)

359
Comparing CPU vs. MXA Performance

We can also run the unconverted model (mobilenet_v2.h5) on the CPU, applying the code to
the same input tensor:

cpu_model = [Link].load_model(KERAS_PATH)
cpu_outputs = cpu_model.predict(x)

num_classes = cpu_outputs.shape[-1]
idx2label = idx2label_full if num_classes == len(idx2label_full) else None

cpu_top5 = topk_from_probs(cpu_outputs, k=5)


print("\nCPU top-5:")
for idx, prob in cpu_top5:
name = label_for(idx, idx2label)
print(f" #{idx:4d}: {name:20s} ({prob*100:.1f}%)")

Expected output:

CPU top-5:
# 282: tiger_cat (58.4%)
# 285: Egyptian_cat (12.9%)
# 281: tabby (11.6%)
# 287: lynx (3.4%)
# 588: hamper (1.3%)

Despite the probabilities not being identical, both models reach the same top prediction.
The slight differences are due to numerical precision variations between CPU and accelerator
implementations.

Measuring Latency

Let’s create a complete Python script (run_inference_mobilenetv2.py) that also measures


and compares latency for both CPU and MXA.

Note: The following sections break down the complete inference script into logical
components. The full working script is available separately and integrates all these
pieces together.

To measure latency accurately, we’ll add timing code:

360
import time

# Warm-up run
_ = [Link](x)

# Timed inference
start = [Link]()
mxa_outputs = [Link](x)
mxa_latency = [Link]() - start

print(f"\nMXA latency: {mxa_latency*1000:.2f} ms")

Run the complete script run_inference_comp_mobilenetv2.py in the terminal:

python run_inference_comp_mobilenetv2.py

Expected results:

The Accelerator runs 11 times faster than the CPU!

361
The MobileNet V2 running with ExecuTorch/XNNPACK backend on a CPU has
around 20 ms of latency.

Testing with Larger Models

We can also test a larger model like ResNet50:

# Download ResNet50
python3 -c "import tensorflow as tf; \
[Link].ResNet50().save('resnet50.h5');"

# Compile
mx_nc -v -m resnet50.h5

# Run the script in the terminal:


python run_inference_comp_resnet50.py

The script can be found in the lab repo: run_inference_comp_resnet50.py in the


terminal:

362
The performance improvements are even more dramatic with larger models!

Clean Shutdown

Always properly shut down the accelerator when done:

[Link]()

BTW, to shut down the Raspberry Pi via SSH, we can use

sudo shutdown -h now

Folders Structure

Documents/MEMRYX/
��� run_inference_comp_mobilenetv2.py # MobileNetV2 script
��� run_inference_comp_resnet50.py # ResNet50 script
��� images/
� ��� [Link] # Test image
��� models/
� ��� mobilenet_v2.h5 # Original model
� ��� mobilenet_v2.dfp # Compiled model
� ��� resnet50.h5 # Original model
� ��� [Link] # Compiled model
� ��� imagenet_class_index.json # Labels (auto-downloaded)
��� mx-env/

Performance Comparison Summary

Here’s how the MX3 compares across different deployment approaches we’ve covered in this
course:

Approach Hardware MobileNetV2 Latency ResNet50 Latency Power (Active)


TFLite Raspberry ~110 ms ~266 ms ~
(CPU) Pi 5

363
Approach Hardware MobileNetV2 Latency ResNet50 Latency Power (Active)
Execu- Raspberry ~20 ms NA ~
Torch/XN- Pi 5
NPACK
MemryX Dedi- ~10 ms ~10 ms ~
MX3 cated
accelera-
tor

Key Observations

• 25x faster than unoptimized TFLite on CPU (Resnet50)


• 2x faster than highly optimized ExecuTorch with XNNPACK (MobileNetV2)
• Minimal CPU load: The host CPU is free for preprocessing, postprocessing, and
application logic
• Consistent latency: Hardware acceleration provides deterministic performance
• Power efficiency: Not measured

When to Use the MX3?

The MemryX MX3 is ideal for:

• � Real-time applications requiring <20ms latency


• � Multi-stream processing (multiple cameras, sensors)
• � Power-constrained environments where CPU load matters
• � Production deployments requiring consistent, predictable performance
• � Complex models where CPU inference is too slow

The MX3 may be overkill for:

• � Simple models that run fast enough on CPU


• � Non-latency-critical batch processing
• � Prototyping where development speed matters more than performance
• � Very cost-sensitive applications

YOLOv8 Object Detection with MX3 Hardware Acceleration

In this part of the lab, we’ll deploy YOLOv8n (nano) for real-time object detection on the
Raspberry Pi 5 using the MemryX MX3 AI accelerator. We’ll cover the complete workflow
from model export to inference optimization.

364
But first, we should install ULTRALYTICS
The MemoryX team suggested:

pip install ultralytics==8.3.161

If you face issues, try it with:

pip install ultralytics[export]

If you still have issues, reinstall memryx

pip3 install --extra-index-url [Link] memryx

Now, download the [Link] model and test it:

yolo predict model='yolov8n' source='[Link]

The model and the image [Link] will be download and tested with the YOLOV8n:

4 persons, 1 bus, and one stop signal were detected in 522 ms.

Under ./runs/detect, the output processed image can be analysed:

365
Model Export and Compilation

Step 1: Export YOLOv8 to ONNX

YOLOv8 must be converted to ONNX format before compilation for the MX3:

from ultralytics import YOLO

#### Load pretrained YOLOv8n model


model = YOLO('[Link]')

366
#### Export to ONNX format
[Link](format='onnx', simplify=True)

This creates [Link] with the complete model graph.


We can also use CLI for the conversion:

yolo export model=[Link] format=onnx

Step 2: Compile for MX3

We can use the MemryX Neural Compiler to generate the DFP file:

mx_nc -v --autocrop -m [Link]

Key flags:

• -v: Verbose output for debugging


• -m: Input model path
• --autocrop: Automatically split model for optimal MX3 execution

Output files:

• [Link] - Accelerator executable (runs on MX3 chips)


• yolov8n_post.onnx - Post-processing model (runs on CPU)

Why Two Files?


The MX3 compiler splits the model:

1. Feature extraction ([Link]): Neural network backbone on MX3 hardware

367
2. Detection head (yolov8n_post.onnx): Bounding box decoding on CPU

This hybrid approach optimizes performance by running computationally intensive operations


on the accelerator while keeping final post-processing flexible.

Now, let’s check the model with a benchmark

mx_bench -d [Link] -f 1000

368
Understanding YOLOv8 Output Format

YOLOv8 Output Structure

The post-processing model outputs predictions in shape (1, 84, 8400):

• Dimension 0 (1): Batch size


• Dimension 1 (84):
– First 4 values: Bounding box [x_center, y_center, width, height]
– Next 80 values: Class probabilities (COCO dataset)
• Dimension 2 (8400): Anchor points across three detection scales:
– 80×80 = 6400 points (small objects)
– 40×40 = 1600 points (medium objects)
– 20×20 = 400 points (large objects)

Decoding Process

1. Transpose: Convert from (1, 84, 8400) → (8400, 84)


2. Extract: Separate boxes and class scores
3. Filter: Keep predictions above confidence threshold
4. Convert coordinates: Transform from xywh to xyxy
5. NMS: Remove overlapping detections
6. Scale: Map to original image coordinates

Complete Inference Pipeline

We should now create a script to run an object detector (YOLOv8 with a pre/post-processing
pipeline), print each detection (label, confidence, bounding box), and save a copy of the image
with the boxes drawn.

Configuration section

At the top of __main__ we define the configuration values:

DFP_PATH = "./models/[Link]"
POST_MODEL_PATH = "./models/yolov8n_post.onnx"
IMAGE_PATH = "./images/[Link]"
CONF_THRESHOLD = 0.25

369
• DFP_PATH: path to the compiled model used for inference.

• POST_MODEL_PATH: path to a post-processing model or graph, usually converting raw


outputs into boxes, scores, and class IDs.

• IMAGE_PATH: image file we want to run detection on.

• CONF_THRESHOLD: minimum confidence score; detections below this are filtered out.

Running detection

detections, annotated_image, inference_time = detect_objects(


DFP_PATH,
POST_MODEL_PATH,
IMAGE_PATH,
CONF_THRESHOLD
)

Here we call a helper function detect_objects that encapsulates the heavy lifting:

• Loads the model(s).

• Loads and preprocesses the image.

• Runs inference.

• Applies post-processing (NMS, thresholding, etc.).

• Returns:
– detections: list/array where each element is [x1, y1, x2, y2, conf,
class_id].

– annotated_image: a PIL image with bounding boxes and labels drawn.

– inference_time: time spent doing the detection (in milliseconds).

Printing results

370
print(f"\n{'='*60}")
print("Detection Results:")
print(f"{'='*60}")
for i, det in enumerate(detections):
x1, y1, x2, y2, conf, class_id = det
print(f" {i+1}. {COCO_CLASSES[int(class_id)]}: {conf:.3f}")
print(f" Box: [{int(x1)}, {int(y1)}, {int(x2)}, {int(y2)}]")

• The loop goes over each detection, unpacks the bounding box coordinates, confidence,
and class ID.

• COCO_CLASSES[int(class_id)] converts the numeric class index into a human-readable


label (e.g., “person”, “bus”).

• Coordinates are cast to int for cleaner printing.

Saving the annotated image

if len(detections) > 0:
output_path = IMAGE_PATH.rsplit('.', 1)[0] + '_detected.jpg'
annotated_image.save(output_path)
print(f"\nSaved: {output_path}")

• Only saves an output file if at least one object was detected.

• The output filename is built by taking the original name and appending _detected
before the extension (e.g., bus_detected.jpg).

• annotated_image.save(...) writes the image with drawn boxes and labels to disk.

Final summary output

print(f"\n{'='*60}")
print(f"Total: {len(detections)} objects")
print(f"Time: {inference_time:.2f} ms")
print(f"{'='*60}")

• Prints how many objects were found in total.

371
• Prints the inference time, which is useful to talk about performance (e.g., model size
vs. speed, hardware differences).

Run the script yolov8_m3_detect.py

python yolov8_m3_detect.py

As a result, we can see that the models found 4 persons and 1 bus, missing only the stop signal.
Regarding latency, the MX3 runs inference about 11 times faster than a CPU-only
system.

Basically, the same accuracy result that we got on the YOLO chapter running
yolov11

Going deeper in the functions

1. Image preprocessing (preprocess_image)

original image → resized → padded → tensor

372
• Load and normalize:
– Open the image with PIL, convert to RGB, and get its original size.
– Compute a scale ratio so the image fits into 640×640 without distortion (preserving
aspect ratio).
• Letterboxing:
– Resize the image to (new_w, new_h) = (int(w * ratio), int(h * ratio)).
– Paste it onto a 640×640 canvas filled with color (114, 114, 114) (same as
Ultralytics).

– Compute the padding offsets (pad_w, pad_h) so we can undo this later.
• Tensor conversion:
– Convert to numpy, normalize to [0,1], permute from HWC to CHW, and add a batch
dimension to get shape [1, 3, 640, 640], which matches YOLOv8’s expected
input.[1][2]

2. Decoding YOLOv8 output (decode_predictions)

How to turn raw model numbers into human-readable detections.


The raw output format:

• For COCO YOLOv8n ONNX, the detection head outputs a tensor of shape (1, 84,
8400).
• 84 = 4 (bbox) + 80 (class scores). Each of the 8400 positions corresponds to one
candidate box.

Function Walkthrough:

• Transpose:
– From (1, 84, 8400) to (8400, 84) so each row is: [x_center, y_center,
width, height, class_0_score, ..., class_79_score].
• Best class per box:
– Take max_scores = [Link](class_scores, axis=1) and class_ids =
[Link](class_scores, axis=1) to select the most likely class and its
score for each of the 8400 candidates.
• Confidence filtering:
– Drop boxes whose max class score is below conf_threshold.
• Coordinate conversion:

373
– Convert from YOLO’s center-format (x, y, w, h) to corner-format (x1, y1, x2,
y2) to make drawing and IoU calculation simpler.
• NMS:
– Call apply_nms to remove overlapping boxes and keep only the best ones.

3. IoU and NMS (compute_iou_batch and apply_nms)

IoU:

• IoU (Intersection over Union) measures overlap between two boxes:

area of intersection
IoU =
area of union
• compute_iou_batch does this between one box and many boxes at once using vectorized
Numpy operations.

NMS:

• apply_nms:
– Sort boxes by score descending.
– Repeatedly pick the highest-score box, compute its IoU with the remaining boxes,
and discard those whose IoU is above iou_threshold.
• The result is a list of indices for boxes that don’t overlap too much and represent unique
objects.

4. Mapping back to the original image (scale_boxes_to_original)

Everything after the model must undo what preprocessing did.

• During preprocessing we:


– Rescaled the image by ratio.
– Padded by (pad_w, pad_h).
• The model’s boxes live in that padded, resized 640×640 space.
• scale_boxes_to_original:
– Subtracts the padding.
– Divides by ratio to go back to the original resolution.
– Clips coordinates so they stay inside the original image bounds.

374
5. Drawing results (draw_detections)

model output → decoded boxes → drawn on the original image

• Make a copy of the original image and get an ImageDraw context.


• For each detection:
– Choose a color deterministically using [Link](class_id) so the same
class always has the same color.
– Draw the rectangle [x1, y1, x2, y2].
– Build a label string "class_name: confidence", measure text size using textbbox,
draw a filled rectangle for the label background, and render the text.

6. The Memryx pipeline (detect_objects and AsyncAccl usage)

How to integrate Memryx’s async accelerator into a typical vision pipeline

• Preprocess once: call preprocess_image to get the model-ready tensor and the info
needed for rescaling.
• Create the accelerator:
– accl = AsyncAccl(dfp_path) loads the compiled Memryx DFP model.
– accl.set_postprocessing_model(post_model_path, model_idx=0) attaches
the ONNX post-processing graph.
• Streaming-style design:
– frame_queue is a queue of inputs; you put your tensor in it.
– generate_frame is a generator feeding frames into the accelerator.
– process_output is a callback that collects outputs into results.
– The code wires them with connect_input and connect_output, then waits for
completion with [Link]().
• Post-processing:
– Grab the first output, call decode_predictions, rescale boxes, and draw.

Making Inferences

Let’s change the script to easily handle different images and confidence threshold
(yolov8_m3_detect_v2.py). We should replace the hardcoded IMAGE_PATH with a
command-line argument:

375
import argparse

if __name__ == "__main__":
parser = [Link]()
parser.add_argument(
"-i", "--image",
type=str,
required=True,
help="Path to input image"
)
parser.add_argument(
"-c", "--conf",
type=float,
default=0.25,
help="Confidence threshold"
)
args = parser.parse_args()

# Configuration
DFP_PATH = "./models/[Link]"
POST_MODEL_PATH = "./models/yolov8n_post.onnx"
IMAGE_PATH = [Link]
CONF_THRESHOLD = [Link]

# Run detection
detections, annotated_image, inference_time = detect_objects(
DFP_PATH,
POST_MODEL_PATH,
IMAGE_PATH,
CONF_THRESHOLD
)

# Print results
print(f"\n{'='*60}")
print("Detection Results:")
print(f"{'='*60}")
for i, det in enumerate(detections):
x1, y1, x2, y2, conf, class_id = det
print(f" {i+1}. {COCO_CLASSES[int(class_id)]}: {conf:.3f}")
print(f" Box: [{int(x1)}, {int(y1)}, {int(x2)}, {int(y2)}]")

# Save annotated image

376
if len(detections) > 0:
output_path = IMAGE_PATH.rsplit('.', 1)[0] + '_detected.jpg'
annotated_image.save(output_path)
print(f"\nSaved: {output_path}")

print(f"\n{'='*60}")
print(f"Total: {len(detections)} objects")
print(f"Time: {inference_time:.2f} ms")
print(f"{'='*60}")

We can run it as:

python yolov8_mx3_detect_v2.py --image ./images/[Link] -c 0.2

Here are some results with other images:

Inference with a custom model

As we saw in the YOLO chapter, we are assuming we are in an industrial facility that must
sort and count wheels and special boxes.

377
Each image can have three classes:

• Background (no objects)


• Box
• Wheel

We have captured a raw dataset using the Raspberry Pi Camera and labeled it with the
ROBOFLOW. The Yolo model was trained on a Google Colab using Ultralytics.

After training, we download the trained model from /runs/detect/train/weights/[Link]


to our computer, renaming it to box_wheel_320_yolo.pt.

Using the FileZilla FTP, transfer a few images from the test dataset to .\images:

378
Let’s return to the ./MEMRYX/YOLO folder and using the Python Interpreter, to quickly do some
inferences:

python

We will import the YOLO library and define the model to use:

>>> from ultralytics import YOLO


>>> model = YOLO('./models/box_wheel_320_yolo.pt')

Now, let’s define an image and call the inference (we will save the image result this time to
external verification):

>>> img = './images/box_3_wheel_4.jpg'


>>> result = [Link](img, save=True, imgsz=320, conf=0.5, iou=0.3)

379
We can see that the model is working and that the latency was 168 ms.
Let’s now export the model first to ONNX and after to FFPls
, to run it in the MX3 device:

cd ./models
yolo export model=box_wheel_320_yolo.pt format=onnx
mx_nc -v --autocrop -m box_wheel_320_yolo.onnx
cd ..

In the models folder, we will have box_wheel_320_yolo.dfp and box_wheel_320_yolo_post.onnx


Let’s adapt the previous script to be more generic in terms of models (box_wheel_mx3_de-
tect_v2.py):

Naturally we should enter with the new models ’names and instead of COCO_LA-
BELS, the script was changed to:

380
# dataset class names
CLASSES = [
'Box', 'Wheel'
]

Thant’s all!

Run it with:

python box_wheel_mx3_detect_v2.py --image ./images/box_3_wheel_4.jpg

The Result was great! And the latency (~38 ms) was 4 times lower than with the
CPU-only approach (even smaller than the model exported to NCNN, runing 100% at CPU -
80 ms).

Adjusting Confidence Threshold

Lower confidence for more detections (may include false positives):

python box_wheel_mx3_detect_v2.py --image ./images/box_3_wheel_4.jpg -c 0.15

We should experiment with the right confidence threshold.

381
Advanced Topics

Batch Processing (Optimization)

For multiple images, reuse the accelerator instance:

accl = AsyncAccl(dfp_path)
accl.set_postprocessing_model(post_model_path)

for image_path in image_list:


detections = detect_single_image(accl, image_path)

Thermal Management

Always monitor temperature during operation:

watch -n 1 cat /sys/memx0/temperature

Confidence Threshold Tuning

• 0.15-0.20: Maximum recall (catch everything)


• 0.25-0.40: Balanced (default)
• 0.45-0.50: High precision (only confident detections)

Model Selection

• yolov8n: Fastest, 3.2M parameters


• yolov8s: Balanced, 11.2M parameters
• yolov8m: Accurate, 25.9M parameters

Exploring MemryX eXamples

MemryX eXamples is a collection of end-to-end AI applications and tasks powered by


MemryX hardware and software solutions. These examples provide practical, hands-on use
cases to help leverage MemryX technology.

382
Clone the MemryX eXamples Repository

Clone this repository plus any linked submodules:

git clone --recursive [Link]


cd memryx_examples

After cloning the repository, you’ll find several subdirectories with different categories of
applications:

• image_inference - Single image classification and detection


• video_inference - Real-time video processing
• multistream_video_inference - Multi-camera scenarios
• audio_inference - Audio processing and speech recognition
• open_vocabulary - Open-set classification tasks
• accuracy_calculation - Model accuracy verification
• multi_dfp_application - Running multiple models
• optimized_multistream_apps - Production-ready multi-stream examples
• fun_projects - Creative applications and demos

These examples demonstrate best practices for:

• Preprocessing pipelines
• Multi-threaded inference
• Output visualization
• Performance optimization
• Multi-model orchestration

Exploring these examples is an excellent way to learn production-ready patterns for deploying
MemryX applications.

Troubleshooting Common Issues

Device Not Detected

Symptom: ls /dev/memx* returns “No such file or directory”


Solutions:

• Verify physical connection: Reseat the M.2 module in its slot


• Check PCIe settings:

383
sudo raspi-config
# Navigate to: Advanced Options → PCIe Speed → Enable PCIe Gen 3
sudo reboot

• Verify in kernel logs:

dmesg | grep -i memryx


lspci | grep -i memryx

• Ensure sufficient power: Use the official Raspberry Pi 27W power supplyCheck HAT
installation: Ensure the M.2 HAT is properly seated.
• Lower Frequency: Try running sudo mx_set_powermode with a lower frequency, such
as 200 or 300 MHz. Then restart mxa-manager for good measure with sudo service
mxa-manager restart

sudo mx_set_powermode

384
sudo service mxa-manager restart

Check the frequency with the following command:

cat /etc/memryx/[Link]

We can check it by the first field of the file, FREQ4. If the Raspberry Pi is set to the module’s
default operating frequency (500 MHz), we should see FREQ4C=500, indicating that the module
is set to a 500 MHz clock speed for 4-chip DFPs.

If decreasing the frequency solves the issue, then you can either keep the default
frequency for all DFPs at 300 MHz (or 400, 450, etc.), or you can raise it back to
500 MHz and use the C++ API’s set_operating_frequency function to change the
clock speed on a per-DFP basis.

Compilation Errors

Symptom: mx_nc fails with “Unsupported operator” error


Solutions: - Check the supported operators

• Some custom layers may need reformulation


• Try exporting to ONNX first for better compatibility:

import tensorflow as tf
import tf2onnx

model = [Link].load_model('model.h5')
onnx_model, _ = [Link].from_keras(model)
with open("[Link]", "wb") as f:
[Link](onnx_model.SerializeToString())

385
• Check the compilation log (-v flag) to identify which specific layer is causing issues

Thermal Throttling

Symptom: Performance degrades over time, temperature > 90°C


Solutions:

• Verify heatsink installation: Ensure thermal paste is properly applied and heatsink is
firmly attached
• Improve airflow: Position the Raspberry Pi for better air circulation
• Check ambient temperature: Ensure the room temperature is reasonable (<30°C)
• Monitor continuously:

watch -n 1 cat /sys/memx0/temperature

= Consider additional cooling: Add a small fan directed at the heatsink

Python Version Conflicts

Symptom: pip install memryx fails with compatibility errors


Solutions:

• Verify Python version:

python --version # Must show 3.11.x or 3.12.x


pip --version # Should match the Python version

• Ensure you’re in the virtual environment:

which python # Should point to mx-env/bin/python

• Try reinstalling in a fresh virtual environment:

deactivate
rm -rf mx-env
python -m venv mx-env
source mx-env/bin/activate
pip install --upgrade pip wheel
pip install --extra-index-url [Link] memryx

386
Low FPS / Poor Performance

Symptom: Benchmark shows much lower FPS than expected


Solutions:

• Check for thermal throttling:

cat /sys/memx0/temperature # Should be <80°C

• Verify PCIe Gen 3 is enabled (not Gen 2):

sudo raspi-config
# Advanced Options → PCIe Speed

• Close other PCIe-intensive applications: Ensure no other devices are saturating


the PCIe bus
• Check for background CPU load:

htop

• Verify driver version: Ensure you have the latest drivers

apt policy memx-drivers

• Verify Frequency

By default, the frequency should be at 500 MHz. Smaller frequencies will reduce the FPS
(increase the latency)

cat /etc/memryx/[Link]

Import Errors

Symptom: ImportError: cannot import name 'SyncAccl' from 'memryx'


Solutions:

• Ensure memryx is installed in the current environment:

pip list | grep memryx

• Reinstall if necessary:

387
pip install --force-reinstall --extra-index-url \
[Link] memryx

• Check Python path conflicts:

import sys
print([Link])

Model Accuracy Issues

Symptom: Inference results are incorrect or significantly different from CPU


Solutions: - Verify preprocessing: Ensure the same preprocessing is applied as during
training

• Check input normalization: Confirm the value range matches training (e.g., [0, 1] vs
[-1, 1])
• Test with known inputs: Use the validation dataset to verify accuracy
• Compare outputs numerically: Print raw logits/probabilities to identify differences
• Check for quantization effects: If using -q flag, try without quantization first

Next Steps and Extensions

Project Ideas

1. Real-time Object Detection with Camera


• Integrate picamera2 with YOLO
• Display bounding boxes in real-time
• Measure end-to-end latency (capture → inference → display)
2. Multi-Model Pipeline
• Use detection + classification cascade
• Leverage multiple MX3 chips for parallel inference
• Build a smart surveillance system
3. Custom Model Deployment
• Train your own model for a specific task
• Optimize and compile for MX3

388
• Compare against the models from previous labs
4. Power Efficiency Study
• Measure power consumption with a USB meter
• Compare CPU vs. MX3 energy per inference
• Calculate battery life for mobile applications
5. Multi-Stream Processing
• Process multiple camera streams simultaneously
• Demonstrate chip utilization across streams
• Build a multi-camera monitoring system

Advanced Topics to Explore

• Quantization: Experiment with 8-bit and 4-bit quantization for even better performance
• Model Zoo: Explore pre-optimized models in the MemryX Model Explorer
• Async API: Use AsyncAccl for non-blocking, concurrent processing
• Custom Operators: Learn to handle models with custom layers
• Multi-chip Scaling: Understand how workload distributes across the four accelerators

Conclusion

In this lab, we’ve explored hardware acceleration for edge AI using the MemryX MX3 accelerator.
We’ve learned:

1. � How to install and configure the MX3 hardware


2. � The MX3 compilation and deployment workflow
3. � How to benchmark and measure performance
4. � Building complete inference applications
5. � Comparing CPU vs. dedicated accelerator performance

The MX3 demonstrates that dedicated AI accelerators can deliver significant performance
improvements for edge applications, achieving FPS several times higher (up to 25x for ResNet-
50) than CPU inference while maintaining accuracy and providing deterministic latency.
As edge AI continues to evolve, hardware acceleration will become increasingly important for
real-time, power-efficient deployments. The skills we’ve developed in this lab—understanding
the compilation workflow, benchmarking methodologies, and performance optimization—will
transfer to other accelerator platforms as well.

389
References and Further Reading

Official Documentation

1. MemryX Developer Hub


2. MX3 Product Brief
3. Architecture Overview
4. Supported Operators

Code and Examples

5. MemryX eXamples Repository


6. MemryX GitHub Organization
7. Model eXplorer
8. Lab Code Repo

Background Reading

9. MLSys Book - Hardware Acceleration


10. Raspberry Pi PCIe Documentation

Community and Support

11. MemryX YouTube Channel


12. MemryX Support Portal

#
Generative AI (Proactive)

390
Text Generation with RNNs

The Jules Verne Bot

Introduction

In this chapter, we will explore how to build a character-level text generation model using
Recurrent Neural Networks (RNNs), specifically inspired by the works of Jules Verne.
The “Jules Verne Bot” project will help to show us the fundamental concepts of sequence
modeling and text generation using deep learning techniques, as a preview of how the modern
LLMs work.
Project Overview:

391
• Goal: Create an AI model that generates text in the style of Jules Verne
• Architecture: RNN with GRU (Gated Recurrent Unit) layers
• Approach: Character-level text prediction
• Framework: TensorFlow/Keras
• Platform: Google Colab with Tesla T4 GPU

What Are We Actually Building?

Imagine we could teach a computer to write like Jules Verne, the famous author of “Twenty
Thousand Leagues Under the Sea” and “Around the World in Eighty Days.” That’s precisely
what we’re doing with the Jules Verne Bot. This project creates an artificial intelligence system
that learns the patterns, style, and vocabulary from Jules Verne’s novels, then generates new
text that sounds like it could have come from his pen.
Think of it like this: if we read enough of someone’s writing, we start to recognize their style.
We notice they use certain phrases, prefer specific sentence structures, or have favorite topics.
Our neural network does something similar, but with mathematical precision. It analyzes
millions of characters from Verne’s works and learns to predict what character should come
next in any given sequence.

Neural Network Architectures Background

Before we dive into the technical details, let’s understand why we use neural networks for this
task and why we chose the specific type we did.

The Human Brain Analogy

When you read a sentence like “The submarine descended into the dark…” your brain auto-
matically starts predicting what might come next. Maybe “depths” or “ocean” or “waters.”
Your brain does this because it has learned patterns from all the text you’ve ever read. Neural
networks work similarly, but they learn these patterns through mathematical calculations
rather than biological processes.

Recurrent Neural Networks (RNN)

Before diving into our RNN implementation, let’s understand where RNNs fit in the neural
network ecosystem:
Key Neural Network Architectures:

392
• MLP (Multi-Layer Perceptron): Basic feedforward networks for general tasks, for
example, vibration analysis
• CNN (Convolutional Neural Networks): Specialized for image processing as Image
Classification tasks and spatial data
• RNN (Recurrent Neural Networks): Designed for sequential data like text and time
series
• GAN (Generative Adversarial Networks): Two networks competing for realistic
data generation, as images
• Transformers (Attention Networks): Modern architecture using attention mecha-
nisms, as in LLMs (Large Language Models, such as GPT)

We chose a Recurrent Neural Network (RNN) for this project because text has a crucial
property: order matters tremendously. The sequence “The cat sat on the mat” means something
completely different from “Mat the on sat cat the.” Regular neural networks process all inputs
simultaneously, like looking at a photograph. But for text, we need a network that processes
information sequentially, remembering what came before to understand what should go next.

In text generation, we aim to predict the most probable word to follow a sentence.

393
Think of reading a book. You don’t just look at all the words on a page simultaneously. You
read word by word, sentence by sentence, and your understanding builds as you progress. Each
new word is interpreted in the context of everything you’ve read before in that chapter. RNNs
work the same way.

The Memory Problem and GRU Solution

Early RNNs had a significant problem: they couldn’t remember information for very long.
Imagine trying to understand a story where you could only remember the last few words you
read. You’d lose track of characters, plot points, and context very quickly.
This is where the Gated Recurrent Unit (GRU) comes in. Think of GRU as an improved
memory system with two special abilities:
Reset Gate: This decides when to “forget” old information. If the story switches to a new
scene or character, the reset gate helps the network forget irrelevant details from the previous
context.
Update Gate: This decides how much new information to incorporate. When encountering
important plot points or character names, the update gate helps the network remember these
crucial details for longer.
It’s like having a smart note-taking system that automatically decides what’s worth remembering
and what can be forgotten.
Why RNNs for Text Generation?
Recurrent Neural Networks are designed explicitly for sequential data processing. Key
characteristics:

• Memory: RNNs maintain an internal state (memory) to remember previous inputs

394
• Sequential Processing: Process data one element at a time, making them ideal for text
• Variable Length Input: Can handle sequences of different lengths
• Parameter Sharing: Same weights applied across different time steps

RNN Architecture Flow:

Input Sequence: x(t-1) → x(t) → x(t+1) → ...


Hidden State: h(t-1) → h(t) → h(t+1) → ...
Output: o(t-1) → o(t) → o(t+1) → ...

Dataset Preparation

395
Our model is trained on a curated collection of 10 classic Jules Verne novels, downloaded from
public domain texts of the Gutenberg Project:

1. “A Journey to the Centre of the Earth”


2. “In Search of the Castaways”
3. “An Antarctic Mystery”
4. “In the year 2889”
5. “Around the World in Eighty Days”
6. “Michael Strogoff”
7. “Five Weeks in a Balloon”
8. “The Mysterious Island”
9. “From the Earth to the Moon”
10. “Twenty Thousand Leagues under the Sea”

“A Journey to the Centre of the Earth” teaches the model about geological descriptions and
underground adventures. “Twenty Thousand Leagues Under the Sea” provides vocabulary
about marine life and submarine technology. “Around the World in Eighty Days” offers
geographical references and travel descriptions. Each book contributes unique vocabulary and
stylistic elements while maintaining Verne’s consistent voice.
The complete dataset contains 5,768,791 characters, of which 123 are unique. To put this
in perspective, that’s roughly equivalent to 3,000 pages of text. This provides our neural
network with ample material to learn from, enabling it to capture both common patterns and
unique expressions in Verne’s writing.

Data Preprocessing Steps

# Example preprocessing workflow


def preprocess_text(text):
# Convert to lowercase for consistency
text = [Link]()

# Remove unwanted characters (optional)


# Keep punctuation for realistic text generation

return text

# Load and combine all books


all_text = ""
for book in book_list:
with open(book, 'r') as f:
all_text += preprocess_text([Link]())

396
Tokenization and Vocabulary

Character-Level Tokenization

Here’s where our approach differs from how humans typically think about text. While we
naturally think in words and sentences, our model processes text character by character. This
means it learns that certain letters frequently follow others, that spaces separate words, and
that punctuation marks signal sentence boundaries.
Why choose character-level processing? Consider the word “extraordinary,” which appears
frequently in Verne’s work. A word-level model would need to have seen this exact word during
training to use it. But a character-level model can generate this word by learning that ‘e’ often
starts words, ‘x’ can follow ‘e’, ‘t’ often follows ‘x’, and so on. This allows our model to create
new words or handle misspellings gracefully.
The downside is that character-level processing requires more computational steps to generate
the same amount of text. Generating “Hello world” requires 11 prediction steps instead of just
2. However, for our educational purposes, this trade-off provides valuable insights into how
language generation works at its most fundamental level.

Unlike word-level tokenization, character-level tokenization treats each character as


a token.

Advantages of Character-Level Tokenization:

• No Out-of-Vocabulary Issues: Every possible character sequence can be generated


• Smaller Vocabulary: Only 123 unique characters vs thousands of words
• Language Agnostic: Works with any language or symbol system
• Handles Rare Words: Can generate new words character by character

Please see the following site for a great general visual explanation, from Andrej
Karpathy, The Unreasonable Effectiveness of Recurrent Neural Networks.

Vocabulary Building Process

Computers work with numbers, not letters, so we need to convert our text into a numerical
representation. We start by finding every unique character in our dataset. This includes
not just letters A-Z and a-z, but also numbers, punctuation marks, spaces, and even special
characters that might appear in the original texts.
Our Jules Verne collection contains 123 unique characters. These include obvious ones like
letters and common punctuation, but also less common characters like accented letters from
French names or special typography marks from the original publications.

397
Creating the Character Dictionary

We create two dictionaries: one that converts characters to numbers (encoding) and another
that converts numbers back to characters (decoding). For example:
‘a’ might become 47, ‘b’ becomes 48, ‘c’ becomes 49, and so on. The space character might be
1, and the period might be 72. These assignments are arbitrary but consistent throughout our
project.
When we want to process the phrase “The sea”, we convert it to something like [84, 72, 69, 1,
83, 69, 47]. When the model generates numbers like [84, 72, 69, 1, 87, 47, 83], we convert them
back to “The was” (as an example).

# Create character-to-index mapping


text = "Your complete dataset text here..."
vocab = sorted(set(text))
char_to_idx = {char: idx for idx, char in enumerate(vocab)}
idx_to_char = {idx: char for idx, char in enumerate(vocab)}

print(f"Vocabulary size: {len(vocab)}")


print(f"Unique characters: {vocab}")

398
It is possible to experiment with (sub-word level) tokenization using OpenAI’s
tokenizer tool at:
[Link]

Training Sequences

The Sliding Window Approach

Our model learns by playing a sophisticated prediction game. We show it sequences of 120
characters and ask it to predict what the 121st character should be. Think of it like a
fill-in-the-blank exercise, but instead of missing words, we’re missing the next character.
For example, if our text contains “The submarine descended into the dark depths of the ocean”,
we might show the model “The submarine descended into the dark depths of the ocea” and ask

399
it to predict “n”. Then we slide our window forward by one character and show it “he submarine
descended into the dark depths of the ocean” and ask it to predict the next character.

Training Configuration

• Sequence Length: 120 characters (approximately one paragraph)


• Input-Output Relationship: Predict the next character given the previous characters

Why 120 Characters?

We chose 120 characters as our context window because it roughly corresponds to one
paragraph of text in English. This gives the model enough context to understand local patterns
(like completing words and phrases) while remaining computationally manageable. In practical
terms, 120 characters might look like:
“The Nautilus had been cruising in these waters for some time. Captain Nemo stood on the
bridge, observing the vast exp”
From this context, the model might predict “a” to complete “expanse” or “l” to form “explore”.

The longer the context window, the better the model can maintain coherence, but
the more computer memory and processing time it requires.

Training Example

• Input Sequence: “Hello my nam”


• Target Sequence: “ello my name”

The model learns:

• Given “H”, predict “e”


• Given “He”, predict “l”
• Given “Hel”, predict “l”
• And so on…

This means our dataset of 5.8 million characters becomes millions of individual training
examples, each teaching the model about character sequence patterns.

Creating Training Data

400
def create_training_sequences(text, seq_length):
sequences = []
targets = []

for i in range(len(text) - seq_length):


# Input sequence
seq = text[i:i + seq_length]
# Target (next character)
target = text[i + 1:i + seq_length + 1]

[Link]([char_to_idx[char] for char in seq])


[Link]([char_to_idx[char] for char in target])

return sequences, targets

Character Embeddings

From Sparse to Dense Representation

Initially, each character is represented as a one-hot vector, which is mostly zeros with a single
one indicating which character it is. For 123 characters, this means each character is represented
by a vector with 123 elements, where 122 are zero and 1 is one. This is wasteful and doesn’t
capture any relationships between characters.
Character embeddings solve this problem by representing each character as a dense vector
of real numbers. Instead of 123 mostly-zero values, each character becomes 256 meaningful
numbers. These numbers are learned during training and end up encoding relationships between
characters.

Learning Character Relationships

Something fascinating happens during training: characters that behave similarly end up with
similar embedding vectors. Vowels tend to cluster together because they can often substitute
for each other in similar contexts. Consonants that frequently appear together (like ‘th’ or ‘ch’)
develop related embeddings.
The model learns that uppercase and lowercase versions of the same letter are related but
distinct. It discovers that digits form their own cluster since they appear in similar contexts
(dates, measurements, chapter numbers). Punctuation marks develop embeddings based on
their grammatical functions.

401
Visualization and Understanding

When we project these 256-dimensional embeddings down to 3D space for visualization, we


can see these learned relationships. The embedding space becomes a map where distance
represents similarity. Characters that can substitute for each other in many contexts end up
close together, while characters with completely different roles end up far apart.
This learned representation becomes the foundation for everything else the model does. The
quality of these embeddings directly affects the model’s ability to generate coherent text.

402
You can play with Word2Vec - Embedding Projector

Model Architecture

RNN Architecture Components

Our Jules Verne Bot consists of three main components, each serving a specific purpose in the
text generation pipeline.
Embedding Layer: This is our translation layer. It takes character indices (numbers like
47, 83, 72) and converts them into dense 256-dimensional vectors that capture character
relationships. Think of this as converting raw symbols into a format that captures meaning
and relationships.
GRU Layer: This is the brain of our operation. With 1024 hidden units, this layer processes
sequences and maintains memory about what it has seen. When processing the sequence “The
submarine descended”, the GRU maintains a hidden state that encodes information about the
submarine, the action of descending, and the overall maritime context.
Dense Output Layer: This is our decision-making layer. It takes the GRU’s 1024-dimensional
hidden state and converts it into 123 probabilities, one for each character in our vocabulary.
These probabilities represent the model’s confidence about what character should come next.

Model Summary

Model: "sequential_4"
_________________________________________________________________
Layer (type) Output Shape Param #

403
=================================================================
embedding_4 (Embedding) (1, 120, 256) 31,488
_________________________________________________________________
gru_3 (GRU) (1, 120, 1024) 3,938,304
_________________________________________________________________
dense_3 (Dense) (1, 120, 123) 126,075
=================================================================
Total params: 4,095,867 (15.62 MB)
Trainable params: 4,095,867 (15.62 MB)
Non-trainable params: 0 (0.00 B)

Our model has 4,095,867 parameters. These are the individual numbers that the model adjusts
during training to improve its predictions. To put this in perspective, each parameter is like a
tiny dial that affects how the model processes information.

Training involves adjusting all 4 million dials to work together harmoniously.

The GRU layer contains most of these parameters (about 3.9 million) because it needs to learn
complex patterns about how characters relate to each other across different time steps. The
embedding layer has about 31,000 parameters (123 characters × 256 dimensions), and the
output layer has about 126,000 parameters.

Memory and Processing Flow

When processing text, information flows through the model like this:
A character index enters the embedding layer and becomes a 256-dimensional vector. This
vector enters the GRU, which combines it with its current memory state to produce a new
1024-dimensional hidden state. This hidden state captures everything the model “knows” at
this point in the sequence.
The hidden state goes to the dense layer, which produces probability scores for each of the
123 possible next characters. The character with the highest probability becomes the model’s
prediction.
Crucially, the GRU’s hidden state becomes its memory for the next character prediction.
This creates a chain of memory that allows the model to maintain context across the entire
sequence.

404
Why GRU over Basic RNN?

GRU Advantages:

• Solves Vanishing Gradient: Better information flow through long sequences


• Selective Memory: Can choose what to remember and forget
• Computational Efficiency: Fewer parameters than LSTM
• Better Performance: More stable training than basic RNNs

Training Process: Teaching the Model to Write

The Learning Objective

Training a neural network means adjusting its millions of parameters so it makes better
predictions. We use a loss function called sparse categorical crossentropy, which measures how
far off the model’s predictions are from the correct answers.
Think of it like teaching someone to play darts. Each throw (prediction) has a target (the
correct next character). The loss function measures how far each dart lands from the bullseye.
Training adjusts the player’s technique (the model’s parameters) to improve accuracy over
time.

Hardware and Time Requirements

We trained our model on a Tesla T4 GPU, which can perform thousands of calculations
simultaneously. This parallelization is crucial because each training step involves matrix
multiplications with millions of numbers. The training took 33 minutes for 30 complete passes
through the entire dataset.
To understand why we need a GPU, consider that training involves calculating gradients for all
4 million parameters, potentially thousands of times per second. A regular CPU would take
many hours to complete the same training that a GPU accomplishes in minutes.

Monitoring Progress

During training, we watch the loss decrease from about 1.9 to 0.9. This represents the model’s
improving ability to predict the next character. Early in training, the model makes essentially
random predictions. By the end, it has learned sophisticated patterns about English spelling,
grammar, and Jules Verne’s writing style.
The learning curve typically shows rapid improvement in the first few epochs as the model
learns basic patterns like common letter combinations. Later epochs show slower but steady

405
improvement as the model refines its understanding of more complex patterns like narrative
structure and thematic elements.

Preventing Overfitting

One challenge in training is overfitting, in which the model memorizes the training data rather
than learning generalizable patterns. We use techniques like monitoring validation loss and
potentially stopping training early if the model stops improving on unseen text.

Training Configuration

Hardware Setup:

• GPU: Tesla T4 (Google Colab)


• GPU RAM: 15.0 GB
• Training Time: 33 minutes for 30 epochs

Training Parameters:

• Loss Function: Categorical Sparse Crossentropy


• Optimizer: Adam (adaptive learning rate)
• Epochs: 30
• Batch Size: 128
• Buffer Size: 10,000 (for dataset shuffling)

406
Training Implementation

# Model compilation
[Link](
optimizer='adam',
loss='sparse_categorical_crossentropy',
metrics=['accuracy']
)

# Training with callbacks


history = [Link](
dataset,
epochs=30,
batch_size=128,
validation_split=0.1,
callbacks=[
[Link](patience=3),
[Link](patience=5)
]
)

Text Generation

The Generation Process

Once trained, our model becomes a text generation engine. We start with a seed phrase like:
“THE FLYING SUBMARINE”
and ask the model to continue the story. The process works character by character:
The model receives “THE FLYING SUBMARINE” and predicts the most likely next character
based on everything it learned from Jules Verne’s works. Maybe it predicts a space, starting a
new word. Then we feed “THE FLYING SUBMARINE” (with the space) back to the model
and ask for the next character.
This process continues indefinitely, with each new character becoming part of the context
for predicting the next one. The model might generate “THE FLYING SUBMARINE de-
scended into the mysterious depths…” as it draws upon patterns learned from Verne’s nautical
adventures.

407
Temperature Control

Here’s where we can control the model’s creativity through a parameter called temperature.
Temperature affects how the model chooses between different possible next characters.
With temperature set to 0.1, the model almost always picks the most probable next character.
This produces very predictable, conservative text that closely mimics the training data but
might be repetitive or boring.
With temperature set to 1.0, the model considers all possible next characters according to
their learned probabilities. This produces more varied and creative text, but sometimes makes
unusual choices that lead to interesting narrative directions.
With temperature above 1.5, the model becomes quite random, often producing text that starts
coherently but gradually becomes nonsensical as unlikely character combinations accumulate.
In short:

• Temperature = 0.5: More predictable, conservative text


• Temperature = 1.0: More creative, diverse text
• Temperature = 1.5: Very random, potentially nonsensical text

Implementation

def generate_text(model, start_string, num_generate=1000, temperature=1.0):


# Convert start string to numbers
input_eval = [char_to_idx[s] for s in start_string]
input_eval = tf.expand_dims(input_eval, 0)

text_generated = []
model.reset_states()

for i in range(num_generate):
predictions = model(input_eval)
predictions = [Link](predictions, 0)

# Apply temperature
predictions = predictions / temperature
predicted_id = [Link](predictions, num_samples=1)[-1,0].numpy()

# Add predicted character to input


input_eval = tf.expand_dims([predicted_id], 0)
text_generated.append(idx_to_char[predicted_id])

408
return start_string + ''.join(text_generated)

Generation Example (Temperature = 0.5)

Seed: “THE FLYING SUBMARINE”


Generated Text:

THE FLYING SUBMARINE


CHAPTER 100 VENTANTILE

This eBook is for the use of anyone anywhere in the United States and most
other parts of the earth and miserable eruptions. The solar rays should be
entirely under the shock of the intensity of the sea. We were all sorts. Are
we to prepare for our feelings?"

"I can never see them a good geographer," said Mary.

"Well, then, John, for I get to the Pampas, that we ought to obey the same
time. In the country of this latitude changed my brother, and the
_Nautilus_ floated in a sea which contained the rudder and
lower colour visibly. The loiter was a fatalint region the two
scientific discoverers. Several times turning toward the river, the cry
of doors and over an inclined plains of the Angara, with a threatening
water and disappeared in the midst of the solar rays.

The weather was spread and strewn with closed bottoms which soon appeared
that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives, and the prisoners
were successively assuming the sides of this agreement for fifteen days
with a threatening voice.
...

Example Output Analysis

Let’s examine some generated text: “The weather was spread and strewn with closed bottoms
which soon appeared that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives…”
This excerpt shows both the model’s strengths and limitations. It successfully captures Verne’s
descriptive style and maritime vocabulary (“seas,” “wind,” “natives”). The sentence structure

409
feels appropriately Victorian and elaborate. However, the meaning becomes confused with
phrases like “closed bottoms” and “sheets of wind was soon and linen.”
This illustrates the fundamental challenge of character-level generation: the model learns local
patterns (how words are spelled, common phrases) much better than global coherence (logical
narrative flow, consistent meaning).

Challenges and Limitations

Context Window Constraints

Our 120-character context window creates a fundamental limitation. The model can only “see”
about one paragraph of previous text when making predictions. This means it might introduce
a character named Captain Smith, then 200 characters later introduce another character with
the same name, having “forgotten” the first introduction.
Humans writing stories maintain mental models of characters, plot lines, and world-building
details across entire novels. Our model’s memory effectively resets every 120 characters, making
long-term narrative consistency nearly impossible.

Character vs. Word Level Trade-offs

Character-level generation requires many more prediction steps than word-level generation.
Generating the phrase “extraordinary adventure” requires 22 character predictions instead
of just 2 word predictions. This makes character-level generation much slower and more
computationally expensive.
However, character-level generation offers unique advantages. The model can generate new
words it has never seen before by combining character patterns. It can handle misspellings,
made-up words, or technical terms more gracefully than word-level models that have fixed
vocabularies.

Coherence Challenges

Perhaps the biggest limitation is maintaining semantic coherence. The model might generate
grammatically correct text that makes no logical sense. It can describe “The submarine floating
in the air above the mountain peaks” because it has learned that submarines float and that
Verne often described mountains, but it hasn’t learned the physical constraint that submarines
float in water, not air.

410
This happens because the model learns statistical patterns without understanding meaning.
It knows that certain word combinations are common without understanding why they make
sense.

Summary

1. Limited Context Window


• Issue: Only 120 characters of context
• Impact: Cannot maintain coherence over long passages
• Example: May forget characters or plot points mentioned earlier
2. Character vs Word Level
• Issue: Character-level generation is slower and less efficient
• Impact: Requires more computation for equivalent output
• Trade-off: Better handling of rare words vs efficiency
3. Coherence Problems
• Issue: May generate grammatically correct but semantically inconsistent text
• Cause: Limited understanding of story structure and plot consistency
4. Repetitive Patterns
• Issue: May fall into repetitive loops
• Cause: Model overfitting to common patterns in training data

Potential Improvements

1. Longer Context Windows: Increase sequence length for better coherence


2. Hierarchical Models: Separate models for different text levels (word, sentence, para-
graph)
3. Fine-tuning: Additional training on specific styles or topics
4. Beam Search: Better text generation algorithms instead of greedy sampling

What will lead us to modern Language Models based on Transformers arquiteture .

Connecting to Modern Language Models

Scale Comparison

To appreciate how far language modeling has advanced, consider the scale differences between
our Jules Verne Bot and modern language models:

411
Our model has 4 million parameters and was trained on about 5.8 million characters (10 books).
GPT-3 has 175 billion parameters and was trained on 45 terabytes of text (roughly equivalent
to millions of books). That’s a difference of over 40,000 times more parameters and millions of
times more training data.
Modern small language models (SLMs) like Phi-3-mini still dwarf our model with 3.8 billion
parameters, but they represent more efficient designs that achieve impressive performance with
“only” 1,000 times more parameters than our model.

Aspect Jules Verne Bot GPT-3 (2020) Phi-3-mini (2024)


Architecture RNN (GRU) Transformer Transformer
Parameters 4 million 175 billion 3.8 billion
Training Data 5.8M characters 45TB text 3.3T tokens
Context Length 120 characters 2,048 tokens 128,000 tokens
Tokenization Character-level Subword (BPE) Subword
Training Time 33 minutes Weeks 7 days
GPU Requirements 1 Tesla T4 10,000 V100 GPUs 512 H100 GPUs

Architectural Evolution

The biggest advancement since RNNs is the Transformer architecture, which uses attention
mechanisms instead of recurrent processing. While RNNs process text sequentially (like
reading word by word), Transformers can examine all parts of a text simultaneously and learn
relationships between any two words, regardless of how far apart they are.
This solves the long-term memory problem that limits our RNN model. A Transformer can
maintain awareness of a character introduced in a hypothetical “chapter 1” while writing
“chapter 10”, something our 120-character context window makes impossible.

Training Efficiency

Modern models also benefit from more sophisticated training techniques. They’re pre-trained
on massive, diverse datasets to learn general language patterns, then fine-tuned on specific tasks.
They use techniques like instruction tuning, where they learn to follow human commands, and
reinforcement learning from human feedback, where they learn to generate text that humans
find helpful and appropriate.

Summary: Why Modern Models Perform Better?

1. Transformer Architecture

412
• Attention Mechanism: Can look at any part of the input sequence
• Parallel Processing: Much faster training and inference
• Better Long-range Dependencies: Maintains context over thousands of tokens
2. Scale
• More Data: Trained on vastly more diverse text
• More Parameters: Can memorize and generalize better
• More Compute: Allows for more sophisticated training techniques
3. Advanced Techniques
• Pre-training + Fine-tuning: Learn general language then specialize
• Instruction Tuning: Trained to follow human instructions
• RLHF: Reinforcement Learning from Human Feedback

Conclusion:

Building the Jules Verne Bot teaches us that creating artificial intelligence systems capable
of generating human-like text requires careful consideration of multiple components working
together. The embedding layer learns to represent characters meaningfully, the RNN layer
processes sequences and maintains memory, and the output layer makes predictions based on
learned patterns.
The project also illustrates the fundamental trade-offs in machine learning: between model
complexity and training speed, between creativity and coherence, between local accuracy and
global consistency. These trade-offs appear in every AI system, from simple character-level
generators to the most sophisticated language models.
Most importantly, this project demonstrates that impressive AI capabilities emerge from
relatively simple components combined thoughtfully. Our 4-million parameter model, while
limited compared to modern systems, genuinely learns to write in Jules Verne’s style through
nothing more than statistical pattern recognition and mathematical optimization.
The techniques we’ve explored, sequence processing, embedding learning, and generation
strategies, form the foundation for understanding any language model. Whether you encounter
RNNs, Transformers, or future architectures yet to be invented, the core concepts remain
consistent: learn patterns from data, encode meaning in mathematical representations, and
generate new content by predicting what should come next.
Understanding these fundamentals provides the foundation for working with, improving, or
creating the next generation of language models that will shape how humans and computers
communicate in the future.

413
Resourses

• Gutenberg Project
• The Unreasonable Effectiveness of Recurrent Neural Networks
• Word2Vec - Embedding Projector
• OpenAI’s tokenizer tool
• Generating Text with RNNs: The Jules Verne Bot - CoLab

414
Knowledge Distillation in Practice

From MNIST to LLMs

415
Figure 8: Image created by DALLE-3

Introduction to Knowledge Distillation

Knowledge distillation is a powerful technique in machine learning that enables the transfer
of knowledge from a large, complex model (the “teacher”) to a smaller, more efficient model
(the “student”). This process allows us to create compact models that maintain much of the

416
performance of their larger counterparts while being significantly faster and requiring fewer
computational resources.

Why Knowledge Distillation Matters

In today’s AI landscape, models are becoming increasingly large and complex. While these
models achieve remarkable performance, they often require substantial computational resources,
making deployment challenging in resource-constrained environments such as mobile devices,
edge computing systems, or real-time applications. Knowledge distillation addresses this
challenge by:

1. Reducing Model Size: Creating smaller models with fewer parameters


2. Improving Inference Speed: Enabling faster predictions in production
3. Lowering Computational Costs: Reducing memory and processing requirements
4. Maintaining Performance: Preserving much of the original model’s accuracy
5. Enabling Edge Deployment: Making AI accessible on resource-limited devices

The Teacher-Student Paradigm

The core concept of knowledge distillation revolves around the teacher-student relationship:

• Teacher Model: A large, well-trained model with high capacity and performance
• Student Model: A smaller, more efficient model trained to mimic the teacher’s behavior
• Knowledge Transfer: The process of transferring the teacher’s “dark knowledge” to
the student

417
Figure 9: Figure from “Knowledge Distillation: A Survey, Jianping Gou Baosheng Yu Stephen
J. Maybank Dacheng Tao, 2021

See paper: “Knowledge Distillation: A Survey, Jianping Gou, Baosheng Yu, Stephen
J. Maybank, Dacheng Tao, 2021

Theoretical Foundations

Soft Targets vs. Hard Targets

Traditional supervised learning uses “hard targets” - one-hot encoded labels that provide limited
information. For example, in MNIST digit classification, the label for the digit “5” would be
represented as [0, 0, 0, 0, 0, 1, 0, 0, 0, 0].
Knowledge distillation leverages “soft targets” - the probability distributions produced by
the teacher model. These soft targets contain richer information about the relationships
between classes. For instance, the teacher might output [0.01, 0.02, 0.01, 0.05, 0.1,
0.78, 0.02, 0.01, 0.0, 0.0] for a “5”, indicating that it’s most confident about “5” but
also considers “4” somewhat similar.

The Temperature Parameter

The temperature parameter (�) is crucial in knowledge distillation. It controls the “softness” of
the probability distribution by modifying the softmax function:

418
P(class_i) = exp(z_i / �) / Σ_j exp(z_j / �)

Where:

• z_i are the logits (pre-softmax outputs)


• � is the temperature parameter

Effects of Temperature:

• � = 1: Standard softmax (normal sharpness)


• � > 1: Softer distribution (more uniform, reveals class relationships)
• � < 1: Sharper distribution (more peaked, less informative)

In our implementation, we use a temperature of 5.0, which creates softer distributions


that effectively transfer the teacher’s “dark knowledge” to the student.

Distillation Loss Function

The distillation loss combines two components:

1. Distillation Loss (L_KD): Measures how well the student mimics the teacher’s soft
targets
2. Student Loss (L_CE): Traditional cross-entropy loss with hard targets

L_total = � * L_CE + (1 - �) * L_KD

Where � is a weighting parameter that balances the two objectives, in our implementation, we
use � = 0.3, giving more weight to the soft targets from the teacher.
The distillation loss is typically computed using the KL divergence:

L_KD = �² * KL_divergence(Teacher_soft_targets, Student_soft_targets)

The �² factor compensates for the gradient scaling effect of temperature. This is crucial for
stable training.

419
Implementation with TensorFlow and MNIST

Dataset Overview

MNIST is an ideal dataset for learning knowledge distillation:

• 60,000 training images of handwritten digits (0-9)


• 10,000 test images
• 28x28 grayscale images
• 10 classes (digits 0-9)
• Well-established baseline performances

Teacher Model Architecture

Our teacher model has substantial capacity with multiple convolutional layers, batch normal-
ization, and dropout for regularization:

def build_teacher_model():
model = [Link]([
layers.Conv2D(64, 3,
activation='relu',
padding='same',
input_shape=(28, 28, 1)),
[Link](),
layers.Conv2D(64, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),

layers.Conv2D(128, 3, activation='relu', padding='same'),


[Link](),
layers.Conv2D(128, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),

layers.Conv2D(256, 3, activation='relu', padding='same'),


[Link](),
layers.GlobalAveragePooling2D(),

[Link](256, activation='relu'),
[Link](),
[Link](0.5),
[Link](128, activation='relu'),

420
[Link](),
[Link](0.5),
[Link](10, activation='softmax')
])
return model

421
422
This architecture achieved 99.44 % accuracy on MNIST.

Student Model Architecture

Our student model is intentionally much simpler, with fewer layers and significantly fewer
parameters:

def build_student_model():
model = [Link]([
layers.Conv2D(16, 3,
activation='relu',
padding='same',
input_shape=(28, 28, 1)),
layers.MaxPooling2D(2, 2),
layers.Conv2D(32, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),
[Link](64, activation='relu'),
[Link](10, activation='softmax')
])
return model

The student model has fewer parameters than the teacher (105K versus 658K), or
6.2 times smaller.

423
Knowledge Distillation Implementation

In our implementation, we use a custom training loop that explicitly calculates both hard and
soft losses:

# Forward pass
with [Link]() as tape:
predictions = kd_student(x_batch, training=True)

# Hard target loss (cross-entropy with true labels)


hard_loss = [Link].categorical_crossentropy(y_hard_batch, predictions)

# Soft target loss (KL divergence with teacher predictions)


soft_loss = [Link].kullback_leibler_divergence(y_soft_batch, predictions)

# Combined loss with temperature scaling factor for soft loss


loss = alpha * hard_loss + (1 - alpha) * soft_loss * (temperature ** 2)

Key components of our implementation:

1. Temperature Scaling: We apply a temperature of 5.0 to soften the teacher’s outputs


2. Alpha Weighting: We use � = 0.3 to prioritize learning from the teacher’s soft targets
3. KL Divergence: This loss function helps the student match the teacher’s probability
distributions
4. Temperature Correction: We multiply the soft loss by �² to correct for gradient scaling

Training Process

Our implementation follows these steps:

1. Train the Teacher: First, we train the complex teacher model using early stopping and
a reduced learning rate.
2. Train a Vanilla Student: We train a student model normally on the hard labels for
comparison.
3. Generate Soft Targets: We use the teacher to create softened probability distributions.
4. Train the Distilled Student: We train another student using our custom distillation
training loop.
5. Evaluate and Compare: We compare the performance, size, and speed of all three
models.

424
Results and Analysis

Our implementation typically shows these results:

• Teacher Model: 99.44% accuracy, largest size, slowest inference (0.8632 seconds)
• Vanilla Student: 98.77% accuracy, smaller size, faster inference (0.3579 seconds)
• Distilled Student: 99.32% accuracy, same size as vanilla student, but better performance
(+0.55%) and faster inference than the Teacher (0.5467 seconds)

These results demonstrate the key benefit of knowledge distillation: the distilled student
achieves performance closer to the teacher while maintaining the efficiency benefits of the
smaller architecture.

Visualizing Distillation Benefits

We analyze the models on challenging examples where the teacher succeeds but the vanilla
student fails. This reveals how knowledge distillation enables students to handle complex
cases by learning the teacher’s “dark knowledge.”

425
Advanced Techniques

Feature-Based Distillation

Beyond output-level distillation, we can transfer knowledge from intermediate layers:

def feature_distillation_loss(teacher_features, student_features):


"""Distill knowledge from intermediate feature maps"""
loss = 0
for t_feat, s_feat in zip(teacher_features, student_features):
# Align dimensions if necessary
s_feat_aligned = align_feature_dimensions(s_feat, t_feat.shape)

426
loss += [Link](t_feat, s_feat_aligned)
return loss

Attention-Based Distillation

Transfer attention patterns from teacher to student:

def attention_distillation_loss(teacher_attention, student_attention):


"""Distill attention mechanisms"""
return [Link](teacher_attention, student_attention)

Practical Tips for Advanced Distillation

1. Layer Alignment: When using feature distillation, carefully align the feature dimensions
2. Feature Selection: Not all features are equally important; focus on the most informative
ones
3. Multi-Teacher Distillation: Combine knowledge from multiple teachers for better
results
4. Online Distillation: Train teacher and student simultaneously for mutual improvement

Scaling to Large Language Models

Challenges in LLM Distillation

Scaling knowledge distillation to Large Language Models presents unique challenges:

1. Model Size: LLMs have billions of parameters, making distillation computationally


expensive
2. Sequence Generation: Unlike classification, LLMs generate sequences, requiring
sequence-level distillation
3. Vocabulary Differences: Teacher and student may have different vocabularies
4. Context Length: Handling variable-length sequences and attention patterns
5. Multi-task Learning: LLMs perform multiple tasks simultaneously

427
LLM Distillation Techniques

1. Sequence-Level Distillation

Instead of token-level predictions, match entire sequence probabilities:

def sequence_level_distillation(teacher_logits, student_logits, sequence_lengths):


"""Distill at the sequence level for better coherence"""
teacher_probs = [Link](teacher_logits / temperature)
student_probs = [Link](student_logits / temperature)

# Mask out padding tokens


mask = create_sequence_mask(sequence_lengths)

return masked_kl_divergence(teacher_probs, student_probs, mask)

2. Recent LLM Distillation Examples

• DistilBERT: 40% smaller than BERT, retains 97% of performance.


• TinyBERT: 7.5x smaller and 9.4x faster than BERT-base
• MiniLM: Uses deep self-attention distillation for efficient transfer.
• Distil-GPT2: Compressed version of GPT-2 with minimal performance loss
• Llama 3.2 3B and 1B: Distilled from Llama 3.2 models (70B/8B parameters)

3. Distilling Reasoning Abilities

Modern approaches focus on transferring reasoning abilities, not just predictions:

• Chain-of-Thought Distillation: Transfer step-by-step reasoning process


• Explanation-Based Distillation: Use teacher’s explanations to guide student learning
• Rationale Extraction: Identify key reasoning patterns for targeted transfer

From MNIST to Real-World LLMs: Llama 3.2 Case Study

1. Teacher Model: The larger Llama 3.2 models (70B/8B parameters) serve as the teachers
2. Student Model: Llama 3.2 3B and 1B are the distilled student models. They were
created using pruning techniques, which systematically remove less meaningful connections
(weights) in the neural network.

428
Key Points

1. Scale Difference: The parameter reduction from 70B to 3B (~23x) or 1B (~70x)


demonstrates industrial-scale distillation
2. Performance Preservation: Despite massive size reduction, the smaller models main-
tain impressive capabilities:
• The 3B model preserves most of the reasoning abilities of larger models.
• The 1B model remains highly functional for many tasks.
3. Practical Benefits:
• The 1B model can run on consumer laptops and even some mobile devices.
• The 3B model offers a balance of performance and accessibility.
4. Distillation Techniques Used:
• Meta likely used a combination of response-based and feature-based distillation.
• They may have employed temperature scaling and specialized loss functions similar
to what we demonstrated in our MNIST example.
• The principles we covered (soft targets, temperature, loss weighting) all apply at
this larger scale.

429
While our MNIST example demonstrates the application of knowledge distillation principles
using a simple dataset, these same principles can be applied directly to state-of-the-art language
models.
Meta’s Llama 3.2 family provides a perfect real-world example:

Parame- Size
Model ters Reduction Use Case
Llama 3.2 70 billion (Teacher) Data centers, high-performance applications
70B
Llama 3.2 8 billion ~9x Server deployment, high-end workstations
8B
Llama 3.2 3 billion ~23x Consumer laptops, desktop applications
3B
Llama 3.2 1 billion ~70x Edge devices, mobile applications, embedded
1B systems

The 3B and 1B models represent successful distillations of the larger models, preserving core
capabilities while dramatically reducing computational requirements. This demonstrates the
industrial importance of knowledge distillation techniques we’ve explored.
It is possible to note that while the underlying principles remain the same as our MNIST
example, industrial LLM distillation includes additional techniques:

1. Distillation-Specific Data: Carefully curated datasets designed to transfer specific


capabilities
2. Multi-Stage Distillation: Gradual compression through intermediate models (70B →
8B → 3B → 1B)
3. Task-Specific Fine-Tuning: Optimizing smaller models for specific use cases after
distillation
4. Custom Loss Functions: Specialized loss terms to preserve reasoning patterns and
generation quality

Despite these additional complexities, the core concept remains the same: using a larger, more
capable model to guide the training of a smaller, more efficient one.

Applications in Modern LLM Development

1. From ChatGPT to Mobile Assistants

Large models like ChatGPT (175B+ parameters) can be distilled to create mobile-friendly
assistants (1-2B parameters) that maintain core capabilities while running locally on smart-
phones.

430
2. Domain-Specific Distillation

Instead of distilling general capabilities, focus on specific domains:

• Medical LLMs: Distill medical knowledge from large models to smaller, specialized
ones
• Legal Assistants: Create compact models focused on legal reasoning and terminology
• Educational Tools: Develop small models optimized for teaching specific subjects

3. From Research to Production

The journey from research models to production deployment often involves distillation:

1. Research Phase: Develop large, state-of-the-art models (100B+ parameters)


2. Distillation Phase: Compress knowledge into deployment-ready models (1-10B param-
eters)
3. Deployment Phase: Further optimize for specific hardware and latency requirements

Practical Guidelines and Best Practices

Choosing the Right Architecture

Based on our experiments, we recommend:

1. Teacher-Student Size Ratio: Aim for a 10-20x parameter reduction for significant
efficiency gains
2. Architectural Similarity: Maintain similar architectural patterns between teacher and
student
3. Bottleneck Identification: Ensure the student has adequate capacity at critical layers

Hyperparameter Selection

Our experiments suggest these optimal settings:

1. Temperature (�):
• For MNIST: 3-5 works well
• For complex tasks: 5-10 may be better
• If outputs are already soft: Lower temperatures (2-3) may suffice
2. Alpha Weighting (�):

431
• For simpler tasks: 0.3-0.5 (balanced approach)
• For complex reasoning: 0.1-0.3 (more emphasis on teacher’s knowledge)
• When teacher is extremely accurate: Lower � values work better
3. Training Duration:
• Distilled students often benefit from longer training (1.5-2x the epochs)
• Use early stopping with patience to avoid overfitting

Common Pitfalls and Solutions

1. Teacher Performance: Ensure the teacher actually outperforms the student (we fixed
this in our implementation)
2. Temperature Selection: If knowledge transfer is poor, experiment with different
temperatures
3. Loss Weighting: If the student ignores soft targets, reduce � to emphasize distillation
loss
4. Gradient Scaling: Always apply the �² correction factor to the soft loss

Performance Evaluation

Always measure multiple dimensions of performance:

1. Accuracy: Primary performance metric (should be closer to teacher than vanilla student)
2. Model Size: Parameter count and memory footprint (should match vanilla student)
3. Inference Speed: Time per prediction (should be significantly faster than teacher)
4. Challenging Cases: Performance on difficult examples (should be better than vanilla
student)

Conclusion

Knowledge distillation provides a powerful approach to create efficient, deployable models


without sacrificing performance. Our implementation demonstrates that even with a 15-20x
reduction in model size, we can maintain performance close to the larger teacher model.

Key Takeaways

1. Efficiency Without Sacrifice: Knowledge distillation enables smaller models to achieve


performance similar to larger ones

432
2. Dark Knowledge Matters: The soft probability distributions contain valuable infor-
mation beyond just the predicted class
3. Temperature and Alpha: These hyperparameters are crucial for effective knowledge
transfer
4. Practical Benefits: Smaller size, faster inference, and lower resource requirements make
AI more accessible

From MNIST to LLMs

The principles we’ve demonstrated with MNIST directly scale to Large Language Models:

1. Same Core Concept: Transfer knowledge from larger to smaller models


2. Same Hyperparameters: Temperature and alpha weighting remain critical
3. Same Benefits: Size reduction, speed improvement, and accessibility
4. Same Challenges: Ensuring adequate student capacity and effective knowledge transfer

Final Thoughts

Knowledge distillation isn’t just an academic technique—it’s essential for practical AI deploy-
ment. The same principles that helped us compress our MNIST classifier can be scaled to
compress models with hundreds of billions of parameters. This universality makes knowledge
distillation an indispensable skill for engineering students entering the field of AI.

Resources

Notebook

433
434
Small Language Models (SLM)

Figure 10: Nano Banana prompt - A cartoon about SLMs runing SLMs as Llama and Gemma
with Ollama - Python on a Raspberry Pi

435
Introduction

In the fast-growing area of artificial intelligence, edge computing presents an opportunity to


decentralize capabilities traditionally reserved for powerful, centralized servers. This chapter
explores the practical integration of small versions of traditional large language models (LLMs)
into a Raspberry Pi 5, transforming this edge device into an AI hub capable of real-time, on-site
data processing.
As large language models grow in size and complexity, Small Language Models (SLMs) offer a
compelling alternative for edge devices, striking a balance between performance and resource
efficiency. By running these models directly on Raspberry Pi, we can create responsive,
privacy-preserving applications that operate even in environments with limited or no internet
connectivity.
This chapter guides us through setting up, optimizing, and leveraging SLMs on Raspberry Pi.
We will explore the installation and utilization of Ollama. This open-source framework allows
us to run LLMs locally on our machines (desktops or edge devices such as NVIDIA Jetson or
Raspberry Pis).
Ollama is designed to be efficient, scalable, and easy to use, making it a good option for
deploying AI models such as Microsoft Phi, Google Gemma, Meta Llama, and more. We will
integrate some of those models into projects using Python’s ecosystem, exploring their potential
in real-world scenarios (or at least point in this direction).

Setup

We could use any Raspberry Pi model in the previous labs, but here, the choice must be the
Raspberry Pi 5 (Raspi-5). It is a robust platform that substantially upgrades the last version
4, equipped with the Broadcom BCM2712, a 2.4GHz quad-core 64-bit Arm Cortex-A76 CPU
featuring Cryptographic Extension and enhanced caching capabilities. It boasts a VideoCore
VII GPU, dual 4Kp60 HDMI® outputs with HDR, and a 4Kp60 HEVC decoder. Memory
options include 4GB, 8GB, and 16GB of high-speed LPDDR4X SDRAM, with 8GB as our
choice for running SLMs. It also features expandable storage via a microSD card slot and a
PCIe 2.0 interface for fast peripherals such as M.2 SSDs (Solid State Drives).

For real SSL applications, SSDs are a better option than SD cards.

By the way, as Alasdair Allan discussed, inferencing directly on the Raspberry Pi 5 CPU—with
no GPU acceleration—is now on par with the performance of the Coral TPU.

436
For more info, please see the complete article: Benchmarking TensorFlow and TensorFlow Lite
on Raspberry Pi 5.

Raspberry Pi Active Cooler

We suggest installing an Active Cooler, a dedicated clip-on cooling solution for Raspberry Pi 5
(Raspi-5), for this lab. It combines an aluminum heatsink with a temperature-controlled blower
fan to keep the Raspi-5 operating comfortably under heavy loads, such as running SLMs.

437
The Active Cooler has pre-applied thermal pads for heat transfer and is mounted directly to
the Raspberry Pi 5 board using spring-loaded push pins. The Raspberry Pi firmware actively
manages it: at 60°C, the blower’s fan is turned on; at 67.5°C, the fan speed is increased; and
finally, at 75°C, the fan increases to full speed. The blower’s fan will spin down automatically
when the temperature drops below these limits.

To prevent overheating, all Raspberry Pi boards begin to throttle the processor


when the temperature reaches 80°Cand throttle even further when it reaches the
maximum temperature of 85°C (more detail here).

Generative AI (GenAI)

Generative AI is an artificial intelligence system capable of creating new, original content across
various media such as text, images, audio, and video. These systems learn patterns from
existing data and use that knowledge to generate novel outputs that didn’t previously exist.
Large Language Models (LLMs), Small Language Models (SLMs), and multimodal
models can all be considered types of GenAI when used for generative tasks.
GenAI provides the conceptual framework for AI-driven content creation, with LLMs serving
as powerful general-purpose text generators. SLMs adapt this technology for edge computing,
while multimodal models extend GenAI capabilities across different data types. Together, they
represent a spectrum of generative AI technologies, each with its strengths and applications,
collectively driving AI-powered content creation and understanding.

438
Large Language Models (LLMs)

Large Language Models (LLMs) are advanced artificial intelligence systems that understand,
process, and generate human-like text. These models are characterized by their massive scale
in terms of the amount of data they are trained on and the number of parameters they contain.
Critical aspects of LLMs include:

1. Size: LLMs typically contain billions of parameters. For example, GPT-3 has 175 billion
parameters, while some newer models exceed a trillion parameters.
2. Training Data: They are trained on vast amounts of text data, often including books,
websites, and other diverse sources, amounting to hundreds of gigabytes or even terabytes
of text.
3. Architecture: Most LLMs use transformer-based architectures, which allow them to
process and generate text by paying attention to different parts of the input simultaneously.
4. Capabilities: LLMs can perform a wide range of language tasks without specific fine-
tuning, including:

• Text generation
• Translation
• Summarization
• Question answering
• Code generation
• Logical reasoning

5. Few-shot Learning: They can often understand and perform new tasks with minimal
examples or instructions.
6. Resource-Intensive: Due to their size, LLMs typically require significant computational
resources to run, often needing powerful GPUs or TPUs.
7. Continual Development: The field of LLMs is rapidly evolving, with new models and
techniques constantly emerging.
8. Ethical Considerations: The use of LLMs raises important questions about bias,
misinformation, and the environmental impact of training such large models.
9. Applications: LLMs are used in various fields, including content creation, customer
service, research assistance, and software development.
10. Limitations: Despite their power, LLMs can produce incorrect or biased information
and lack true understanding or reasoning capabilities.

439
We must note that we use large models beyond text, which we call multi-modal models. These
models integrate and process information from multiple types of input simultaneously. They
are designed to understand and generate content across various data types, such as text, images,
audio, and video.
Certainly. Let’s define open and closed models in the context of AI and language models:

Closed vs Open Models:

Closed models, also called proprietary models, are AI models whose internal workings, code,
and training data are not publicly disclosed. Examples: GPT-3 and beyond (by OpenAI),
Claude (by Anthropic), Gemini (by Google).
Open models, also known as open-source models, are AI models whose code, architecture,
and, often, training data are publicly available. Examples: Gemma (by Google), LLaMA (by
Meta), and Phi (by Microsoft).
Open models are particularly relevant for running models on edge devices like Raspberry Pi as
they can be more easily adapted, optimized, and deployed in resource-constrained environments.
Still, it is crucial to verify their Licenses. Open models come with various open-source licenses
that may affect their use in commercial applications, while closed models have clear, albeit
restrictive, terms of service.

Figure 11: Adapted from [Link]

440
Small Language Models (SLMs)

In the context of edge computing on devices like Raspberry Pi, full-scale LLMs are typically
too large and resource-intensive to run directly. This limitation has driven the development of
smaller, more efficient models, such as the Small Language Models (SLMs).
SLMs are compact versions of LLMs designed to run efficiently on resource-constrained de-
vices such as smartphones, IoT devices, and single-board computers like the Raspberry Pi.
These models are significantly smaller in size and computational requirements than their
larger counterparts while still retaining impressive language understanding and generation
capabilities.
Key characteristics of SLMs include:

1. Reduced parameter count: Typically ranging from a few hundred million to a few
billion parameters, compared to two-digit billions in larger models.
2. Lower memory footprint: Requiring, at most, a few gigabytes of memory rather than
tens or hundreds of gigabytes.
3. Faster inference time: Can generate responses in milliseconds to seconds on edge
devices.
4. Energy efficiency: Consuming less power, making them suitable for battery-powered
devices.
5. Privacy-preserving: Enabling on-device processing without sending data to cloud
servers.
6. Offline functionality: Operating without an internet connection.

SLMs achieve their compact size through various techniques such as knowledge distillation,
model pruning, and quantization. While they may not match the broad capabilities of larger
models, SLMs excel in specific tasks and domains, making them ideal for targeted applications
on edge devices.

We will generally consider SLMs —language models with fewer than 5 to 8 billion
parameters — quantized to 4 bits.

441
Examples of SLMs include compressed versions of models like Meta Llama, Microsoft PHI,
and Google Gemma. These models enable a wide range of natural language processing tasks
directly on edge devices, from text classification and sentiment analysis to question answering
and limited text generation.
For more information on SLMs, the paper, LLM Pruning and Distillation in Practice: The
Minitron Approach, provides an approach applying pruning and distillation to obtain SLMs
from LLMs. And, SMALL LANGUAGE MODELS: SURVEY, MEASUREMENTS, AND
INSIGHTS, presents a comprehensive survey and analysis of Small Language Models (SLMs),
which are language models with 100 million to 5 billion parameters designed for resource-
constrained devices.

442
Ollama

Figure 12: ollama logo

The primary and most user-friendly tool for running Small Language Models (SLMs) directly
on a Raspberry Pi, especially the Pi 5, is Ollama. It is an open-source framework that allows us
to install, manage, and run various SLMs (such as TinyLlama, smollm, Microsoft Phi, Google
Gemma, Meta Llama, MoonDream, LLaVa, among others) locally on our Raspberry Pi for
tasks such as text generation, image captioning, and translation.

Exclamation Alternative options:

[Link] and the Hugging Face Transformers library are well supported for running
SLMs on Raspberry Pi. [Link] is particularly efficient for running quantized models
natively. At the same time, Hugging Face Transformers offers a broader range of models
and tasks, which are best suited to smaller architectures due to hardware limitations.
Note that Ollama runs [Link] under the hood.
For production deployments on edge devices, a good option is Google AI Edge’s LiteRT-
LM, which is covered in the chapter “LiteRT-LM: Production-Ready LLM Inference at
the Edge”.

Here are some critical points about Ollama:

443
1. Local Model Execution: Ollama enables running LMs on personal computers or edge
devices such as the Raspberry Pi 5, eliminating the need for cloud-based API calls.
2. Ease of Use: It provides a simple command-line interface for downloading, running, and
managing different language models.
3. Model Variety: Ollama supports a variety of LLMs, including Phi, Gemma, Llama,
Mistral, and other open-source models.
4. Customization: Users can create and share custom models tailored to specific needs or
domains.
5. Lightweight: Designed to be efficient and run on consumer-grade hardware.
6. API Integration: Offers an API that allows integration with other applications and
services.
7. Privacy-Focused: By running models locally, it addresses privacy concerns associated
with sending data to external servers.
8. Cross-Platform: Available for macOS, Windows, and Linux systems (our case, here).
9. Active Development: Regularly updated with new features and model support.
10. Community-Driven: Benefits from community contributions and model sharing.

To learn more about what Ollama is and how it works under the hood, you should see this
short video from Matt Williams, one of the founders of Ollama:
[Link]

Matt has an entirely free course about Ollama that we recommend: [Link]
be/9KEUFe4KQAI?si=D_-q3CMbHiT-twuy

Installing Ollama

Let’s set up and activate a Virtual Environment for working with Ollama:

python3 -m venv ~/ollama


source ~/ollama/bin/activate

And run the command to install Ollama:

curl -fsSL [Link] | sh

444
As a result, an API will run in the background on [Link]:11434. From now on, we can
run Ollama via the terminal. For starting, let’s verify the Ollama version, which will also tell
us that it is correctly installed:

ollama -v

On the Ollama Library page, we can find the models Ollama supports. For example, by filtering
by Most popular, we can see Meta Llama, Google Gemma, Microsoft Phi, LLaVa, etc.

445
Meta Llama 3.2 1B/3B

Let’s install and run our first small language model, Llama 3.2 1B (and 3B). The Meta Llama,
3.2 collections of multilingual large language models (LLMs), is a collection of pre-trained and
instruction-tuned generative models in 1B and 3B sizes (text in/text out). The Llama 3.2
instruction-tuned text-only models are optimized for multilingual dialogue use cases, including
agentic retrieval and summarization tasks.
The 1B and 3B models were pruned from the Llama 8B, and then logits from the 8B and 70B
models were used as token-level targets (token-level distillation). Knowledge distillation was
used to recover performance (they were trained with 9 trillion tokens). The 1B model has 1,24B,
quantized to integer (Q8_0), and the 3B, 3.12B parameters, with a Q4_0 quantization, which
ends with a size of 1.3 GB and 2GB, respectively. Its context window is 131,072 tokens.

446
Install and run the Model

ollama run llama3.2:1b

Running the model with the command before, we should have the Ollama prompt available for
us to input a question and start chatting with the LLM model; for example,
>>> What is the capital of France?
Almost immediately, we get the correct answer:
The capital of France is Paris.
Using the option --verbose when calling the model will generate several statistics about its
performance (The model will be polling only the first time we run the command).

447
Each metric gives insights into how the model processes inputs and generates outputs. Here’s
a breakdown of what each metric means:

• Total Duration (2.620170326s): This is the complete time taken from the start of
the command to the completion of the response. It encompasses loading the model,
processing the input prompt, and generating the response.
• Load Duration (39.947908ms): This duration indicates the time to load the model
or necessary components into memory. If this value is minimal, it can suggest that the
model was preloaded or that only a minimal setup was required.
• Prompt Eval Count (32 tokens): The number of tokens in the input prompt. In
NLP, tokens are typically words or subwords, so this count includes all the tokens that
the model evaluated to understand and respond to the query.
• Prompt Eval Duration (1.644773s): This measures the model’s time to evaluate or
process the input prompt. It accounts for the bulk of the total duration, implying that
understanding the query and preparing a response is the most time-consuming part of
the process.
• Prompt Eval Rate (19.46 tokens/s): This rate indicates how quickly the model
processes tokens from the input prompt. It reflects the model’s speed in terms of natural

448
language comprehension.
• Eval Count (8 token(s)): This is the number of tokens in the model’s response, which
in this case was, “The capital of France is Paris.”
• Eval Duration (889.941ms): This is the time taken to generate the output based
on the evaluated input. It’s much shorter than the prompt evaluation, suggesting that
generating the response is less complex or computationally intensive than understanding
the prompt.
• Eval Rate (8.99 tokens/s): Similar to the prompt eval rate, this indicates the speed
at which the model generates output tokens. It’s a crucial metric for understanding the
model’s efficiency in output generation.

This detailed breakdown can help understand the computational demands and performance
characteristics of running SLMs like Llama on edge devices like the Raspberry Pi 5. It shows
that while prompt evaluation is more time-consuming, the actual generation of responses is
relatively quicker. This analysis is crucial for optimizing performance and diagnosing potential
bottlenecks in real-time applications.
Loading and running the 3B model, we can see the difference in performance for the same
prompt;

The eval rate is lower, 5.3 tokens/s versus 9 tokens/s with the smaller model.
When question about
>>> What is the distance between Paris and Santiago, Chile?
The 1B model answered 9,841 kilometers (6,093 miles), which is inaccurate, and the 3B
model answered 7,300 miles (11,700 km), which is close to the correct (11,642 km).
Let’s ask for the Paris’s coordinates:

449
>>> what is the latitude and longitude of Paris?

The latitude and longitude of Paris are 48.8567° N (48°55'


42" N) and 2.3510° E (2°22' 8" E), respectively.

Both 1B and 3B models gave correct answers.

Google Gemma

Google Gemma, is a collection of lightweight, state-of-the-art open models built from the same
technology that powers our Gemini models. Today, the Gemma family has the Gemma3 and
Gemma3n models.

450
Gemma3

We can, for example, install gemma3:latest, using ollama run gemma3:latest. This model
has 4.3B parameters, with a context length of 8,192 and an embedding length of 2,560. A
typical quantization schema is the Q4_K_M. This model has vision capabilities. Besides the
4B, we can also install the 1B parameter model.
Install and run the Model

ollama run gemma3:4b --verbose

Running the model with the command before, we should have the Ollama prompt available for
us to input a question and start chatting with the LLM model; for example,
>>> What is the capital of France?
Almost immediately, we get the correct answer:
The capital of France is **Paris**. It's a global center for art, fashion,
gastronomy, and culture. � Do you want to know anything more about Paris?
And its statistics.

We can see that Gemma 3:4B has roughly the same performance as Lama 3.2:3B,
despite having more parameters.

451
Other examples:

Also, a good response but less accurate than Llama3.2:3B.

A good and accurate answer (a little more verbose than the Llama answers).
An advantage of the Gemma3 models are their vision capability, for example, we can ask it to
caption an image:

452
This is a very accurate model, but it still has high latency, as we can see in the example above
(more than 3 minutes to caption the image).

The Gemma 3 1B size models are text only and don’t support image input.

Gemma 3n

Gemma3 is a powerful, efficient open-source model that runs locally on phones, tablets, and
laptops. The models are listed with parameter counts, such as E2B and E4B, that are lower than
the total number of parameters contained in the models. The E prefix indicates these models
can operate with a reduced set of Effective parameters. This reduced-parameter operation can
be achieved using the flexible parameter technology built into Gemma 3n models, which helps
them run efficiently on lower-resource devices.
The parameters in Gemma 3n models are divided into four main groups: text, visual, audio,
and per-layer embedding (PLE) parameters. In the standard execution of the E2B model, over
5 billion parameters are loaded. However, by using parameter skipping and PLE caching, this
model can be operated with an effective memory load of just under 2 billion (1.91B) parameters,
as illustrated below:

453
Figure 13: Immage from Google: Gemma 3n diagram of parameter usage

Once installed, (using ollama run gemma3n:e2b), inspecting the model we get:

• Architecture: gemma3n

• Parameters: 4.5B

• Context length 32768

• Embedding length 2048

• Quantization Q4_K_M
• Capabilities: completion only

When we run it, we can see that, despite having 4.5B parameters, it is faster than Llama3.3:3B.

454
Gemma 4

Gemma 4 models are designed to deliver frontier-level performance across all sizes. They are
well-suited for reasoning, agentic workflows, coding, and multimodal understanding. Gemma 4
is released under a commercially permissive Apache 2.0 license, which is a huge milestone for
open generative AI models at the edge.

E2B and E4B models: A new level of intelligence for mobile and IoT devices

Engineered from the ground up for maximum compute and memory efficiency, these models
operate with an effective 2- and 4-billion-parameter footprint during inference to preserve RAM
and battery life. In close collaboration with our Google Pixel team and mobile hardware leaders
such as Qualcomm Technologies and MediaTek, these multimodal models run completely offline
with near-zero latency on edge devices such as phones, Raspberry Pi, and NVIDIA Jetson Orin
Nano. Android developers can now prototype agentic flows in the AICore Developer Preview
today for forward compatibility with Gemini Nano 4.

Gemma 4 E2B and E4B are too heavy to be deployed on a Raspberry Pi 5 with
Ollama, but it works very well with Google AI Edge’s LiteRT-LM, which is covered
in the chapter “LiteRT-LM: Production-Ready LLM Inference at the Edge”.

455
Microsoft Phi3.5 3.8B

Let’s pull now the PHI3.5, a 3.8B lightweight state-of-the-art open model by Microsoft.
The model belongs to the Phi-3 model family and supports 128K token context length and
the languages: Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German,
Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish,
Swedish, Thai, Turkish, and Ukrainian.
The model size, in terms of bytes, will depend on the specific quantization format used. The
size can go from 2-bit quantization (q2_k) of 1.4 GB (higher performance/lower quality) to
16-bit quantization (fp-16) of 7.6 GB (lower performance/higher quality).
Let’s run the 4-bit quantization (Q4_0), which will need 2.2 GB of RAM, with an intermediary
trade-off regarding output quality and performance.

ollama run phi3.5:3.8b --verbose

You can use run or pull to download the model. What happens is that Ollama
keeps note of the pulled models, and once the PHI3 does not exist, before running
it, Ollama pulls it.

Let’s enter with the same prompt used before:

>>> What is the capital of France?

The capital of France is Paris. It' extradites significant


historical, cultural, and political importance to the country as
well as being a major European city known for its art, fashion,
gastronomy, and culture. Its influence extends beyond national
borders, with millions of tourists visiting each year from around
the globe. The Seine River flows through Paris before it reaches
the broader English Channel at Le Havre. Moreover, France is one

456
of Europe's leading economies with its capital playing a key role

...

The answer was very “verbose”, let’s specify a better prompt:

In this case, the answer was still longer than we expected, with an eval rate of 2.25 tokens/s,
more than double that of Gemma and Llama.

Choosing the most appropriate prompt is one of the most important skills to be
used with LLMs, no matter its size.

When we asked the same questions about distance and Latitude/Longitude, we did not get
a good answer for a distance of 13,507 kilometers (8,429 miles), but it was OK for
coordinates. Again, it could have been less verbose (more than 200 tokens for each answer).
We can use any model as an assistant since their speed is relatively decent, but in October 2024,
the Llama2:3B or Gemma 3 are better choices. Try other models, depending on your needs.
� Open LLM Leaderboard can give you an idea about the best models in size, benchmark,
license, etc.

457
The best model to use is the one fit for your specific necessity. Also, take into
consideration that this field evolves with new models everyday.

MoonDream

Moondream is an open-source visual language model that understands images using simple
text prompts. It’s fast and wildly capable. It is 3 to 4 times faster than the Gemma3, for
example.

Moondream is a compact and efficient open-source vision-language model (VLM) designed to


analyze images and generate descriptive text, answer questions, detect objects, count, point,
and perform OCR—all through natural-language instructions, even on resource-limited devices
like the Raspberry Pi or edge hardware.
Below is a diagram showing the flux from an image and its description:

458
Moondream can interpret images much like a human would, enabling tasks like: • Captioning
and describing images in detail • Visual question answering (VQA) such as “What is the
color of the shirt?” • Object detection and pointing out locations • Counting and contextual
reasoning about visual scenes

459
Llava-Phi-3

Another multimodal model is the LLaVA-Phi-3, a fine-tuned LLaVA model from Phi 3 Mini
4k. It has strong performance benchmarks that are on par with the original LLaVA (Large
Language and Vision Assistant) model.

In terms of latency, it is a pair with Gemma3 and much slower than MoonDream

The LLaVA-Phi-3 is an end-to-end trained large multimodal model designed to understand


and generate content based on visual inputs (images) and textual instructions. It combines the
capabilities of a visual encoder and a language model to process and respond to multimodal
inputs.
Let’s install the model:

ollama run llava-phi3:3.8b --verbose

Let’s start with a text input:

>>> You are a helpful AI assistant. What is the capital of France?

As an AI language model, I can tell you that the capital of France


is Paris. It's not only the largest city in the country but also
serves as its political and administrative center. Paris is known
for its iconic landmarks such as the Eiffel Tower, Notre-Dame
Cathedral, and the Louvre Museum. The city has a rich history,
beautiful architecture, and is widely considered to be one of the
most romantic cities in the world.

The response took around 30s, with an eval rate of 3.93 tokens/s! Not bad!
But let us know to enter with an image as input. For that, let’s create a directory for working:

cd Documents/
mkdir OLLAMA
cd OLLAMA

460
Let’s download a 640x320 image from the internet, for example (Wikipedia: Paris, France):

Using FileZilla, for example, let’s upload the image to the OLLAMA folder on the Raspi-5 and
name it image_test_1.jpg. We should have the whole image path (we can use pwd to get
it).
/home/mjrovai/Documents/OLLAMA/image_test_1.jpg
If you use a desktop, you can copy the image path by right-clicking the image.

461
Let’s enter with this prompt:

>>> Describe the image /home/mjrovai/Documents/OLLAMA/image_test_1.jpg

The result was great, but the overall latency was significant; almost 4 minutes to perform the
inference.

462
Ollama Commands

Ollama has several commands available, which can be accessed using/, such as /bye to exit,
or /setto toggle section variables as verbose/quiet to see variables as token/s or latency,
and think/nothink to turn on/off the thinking mode in models that “reason” as Qwen and
Gemma.

463
Inspecting Model parameters

It is possible to know more about each model downloaded by Ollama, using the command
ollama show <model>. For example:

ollama show llama3.2:3b

464
We’re looking at a 3.2B‑parameter Llama‑family instruction model, quantized for efficient local
use. Let’s walk through each field and what it implies in practice.

High‑level identity

• Architecture (Family): llama


This is an LLaMA‑style transformer (same general architecture as Meta’s Llama 3.x
family): decoder‑only transformer, causal (left‑to‑right) language model, multi‑head
attention, MLP blocks, rotary position embeddings, etc.

• Parameter Size: 3.2B


The base model has about 3.2 billion trainable parameters. This puts it in the
“small to mid‑size” range: much larger than tiny 1B models, but far below 7B/8B/70B.
On a modern CPU/GPU it’s suitable for real‑time-ish local inference, especially after
quantization.

Capabilities

• Supported Capabilities: ['completion', 'tools']


– completion: classic text generation—answering questions, writing code, explana-
tions, etc.

– tools: the model has been trained/fine‑tuned to call tools/functions (e.g., JSON
function calls) when the host runtime exposes them, so it can decide “I should call
a function to get data” vs just answering from its own weights.

This doesn’t mean the tools are built into the model; it means it knows how to
format tool calls when the surrounding system supports them.

Quantization: Q4_K_M

• Quantization Level: Q4_K_M


This is a 4‑bit K‑quantization variant, typically from GGUF/GGML families:
– Weights are stored at about 4 bits per parameter instead of 16/32, drastically
reducing RAM and disk size.
– K_M is a specific quantization scheme/variant that balances compression with accu-
racy; it’s usually better than older naive 4‑bit schemes (higher effective precision
per parameter block).
• Practical impact:

465
– Much smaller VRAM/RAM requirement, so a 3.2B model can run comfortably
on a Pi 5 + NPU/GPU or a modest desktop GPU.
– Slight drop in raw quality vs the original FP16/FP32 model, but for many tasks
(chat, simple coding, small reasoning tasks) it’s still very usable.
– Ideal for edge or low‑power devices.

Context and embedding sizes

• llama.context_length: 131072
Context length is 131,072 tokens (~131k):
– This is a very large context window (comparable to “128k context” models).
– It can ingest long documents, multiple files, or extended conversations without
immediately forgetting earlier content.
– Under the hood, this usually implies some form of advanced attention scaling (e.g.,
RoPE scaling, attention optimizations, maybe sparse/flash attention), but the key
user effect is: we can feed a lot of text at once.
• llama.embedding_length: 3072
The hidden/embedding dimension is 3072:
– Each token is represented as a 3072‑dimensional vector internally.
– This roughly sets the “width” of the model: wider models can represent richer
patterns but cost more per token.
– Often, this is also the size of embeddings we’d extract if using the model as an
encoder (assuming our runtime supports that mode).

License and usage

• License Short: LLAMA 3.2 COMMUNITY LICENSE AGREEMENT


– It’s under a Llama community license, not pure open source.
– Typically allows broad research and many commercial uses, but with conditions
(e.g., restrictions around competition, misuse, or deployment size).

– We should read the exact license text bundled with the model before using it in
commercial products, especially at scale.

Parameters

The only parameters explicitly set in this Modelfile are three stop strings:

• <|start_header_id|>

466
• <|end_header_id|>
• <|eot_id|>

No temperature, top_k, top_p, etc., are defined in the Model file. Those, therefore, use
Ollama’s defaults for this model at runtime.

• temperature 0.3,
• top_p 0.9and
• top_k 40

We will learn more about those parameters in the next section.

Inspecting local resources

Using htop, we can monitor the resources running on our device.

htop

During the time that the model is running, we can inspect the resources:

467
All four CPUs run at almost 100% of their capacity, and the memory used with the model
loaded is 3.24GB. Exiting Ollama, the memory goes down to around 377MB (with no desktop).
It is also essential to monitor the temperature. When running the Raspberry with a desktop,
you can have the temperature shown on the taskbar:

If you are “headless”, the temperature can be monitored continuously (for example, every 2
seconds) with the command:

watch -n 2 vcgencmd measure_temp

If you are doing nothing, the temperature is around 50°C for CPUs running at 1%. During
inference, with the CPUs at 100%, the temperature can rise to almost 70°C. This is OK and
means the active cooler is working, keeping the temperature below 80°C / 85°C (its limit).

Ollama Python Library

So far, we have explored SLMs’ chat capability using the command line on a terminal. However,
we want to integrate those models into our projects, so Python seems to be the right path. The
good news is that Ollama has such a library.
The Ollama Python library simplifies interaction with advanced LLM models, enabling more
sophisticated responses and capabilities, besides providing the easiest way to integrate Python
3.8+ projects with Ollama.
For a better understanding of how to create apps using Ollama with Python, we can follow
Matt Williams’s videos, as the one below:
[Link]
Installation:
In the terminal (and inside the virtual environment), run the command:

pip install ollama

Let’s test the installation on terminal using the Python interpetrer:

468
For better using, we will need a text editor or an IDE to create a Python script. If we
run Raspberry Pi OS on a desktop, several options, such as Thonny and Geany, are already
installed by default (accessed via [Menu][Programming]). You can download other IDEs, such
as Visual Studio Code, from [Menu][Recommended Software]. When the window pops up,
go to [Programming], select your option, and press [Apply].

469
A good option to run Python scripts during development is to use the Jupyter Notebook:

pip install jupyter


jupyter notebook --generate-config

We can use the Jupyter Notebook using SSH on our host computer. Run the command below,
changing the IP address for the Raspi:

jupyter notebook --ip=[Link] --no-browser

On the terminal, you can see the local URL address to open the notebook:

We can access it from another computer by entering the Raspberry Pi’s IP address and the
provided token in a web browser (we should copy it from the terminal).
In our working directory (Documents/OLLAMA/) on the Raspberry Pi, we will create a new
Python 3 notebook.

470
Let’s enter with a very simple script to verify the installed models:

import ollama
models = [Link]()
for model in models['models']:
print(model['model'])

gemma3:4b
riven/smolvlm:latest
gemma3n:e2b
moondream:latest
llama3.2:3b

Same as on the terminal using ollama show llama3.2: 3b, we can get model information
with the Python command [Link]('llama3.2:3b'):

info = [Link]('llama3.2:3b')

print("Model/Format :", getattr(info, 'model', None))


print("Parameter Size :", getattr([Link], 'parameter_size', None))
print("Quantization Level :", getattr([Link], 'quantization_level', None))
print("Family :", getattr([Link], 'family', None))

471
print("Supported Capabilities:", getattr(info, 'capabilities', None))
print("Date Modified (Local) :", getattr(info, 'modified_at', None))
print("License Short :", (str(getattr(info, 'license', None)).split('\n')[0]))

print("Key Architecture Details:")

for key in [
'[Link]', '[Link]',
'llama.context_length', 'llama.embedding_length', 'llama.block_count'
]:
print(f" - {key}: {[Link](key)}")

As a result, we have:

Model/Format : None
Parameter Size : 3.2B
Quantization Level : Q4_K_M
Family : llama
Supported Capabilities : ['completion', 'tools']
Date Modified (Local) : 2025-09-15 15:39:48.153510-03:00
License Short : LLAMA 3.2 COMMUNITY LICENSE AGREEMENT
Key Architecture Details :
- [Link]: llama
- [Link]: Instruct
- llama.context_length: 131072
- llama.embedding_length: 3072
- llama.block_count: 28

Besides the information got on terminal, here we can also see:


[Link]: Instruct : The base Llama model has been fine‑tuned on instruc-
tion‑following data (Q&A, chat, tool‑using examples). This means:

• It expects prompts that look like “user asks → model answers”.


• It tends to produce helpful, direct responses instead of random continuation.
• It may follow some chat‑style formatting conventions (system/user/assistant turns)
depending on the template used in our runtime.

llama.block_count: 28 The transformer has 28 decoder blocks (layers):

• More layers typically improve the model’s ability to perform multi‑step reasoning and to
build hierarchical representations.

472
• For 3.2B parameters, 28 layers with width 3072 is a sensible trade‑off (moderately deep,
not extremely wide).
• Compared with a 7B model (~32–40 layers, 4k+ width), this will be less capable on very
complex reasoning or coding, but noticeably more capable than 1B‑class models.

Model/Format: None : This usually means the frontend (e.g., our UI or manager) didn’t
detect a specific “format template” (like “chat”, “instruct”, “plain”) beyond knowing it’s
Llama:

• The underlying file is probably something like a GGUF or similar binary weight file.
• Prompt formatting will depend on our runtime’s default “llama‑instruct” template; if
nothing is configured, we may need to wrap instructions manually (system/user style or
clear “You are an assistant…” preamble).

Ollama Generate

Let’s repeat one of the questions that we did before, but now using [Link]() from
the Ollama Python library. This API generates a response for the given prompt using the
provided model. This is a streaming endpoint, so there will be a series of responses. The final
response object will include statistics and additional data from the request.

response = [Link](
model="llama3.2:3b",
prompt="What is the capital of Brazil"
)
print(response['response'])

We get the response: The capital of Brazil is Brasília.

If you are running the code as a Python script, you should save it as, for example,
test_ollama.py. You can run it in the IDE or run it directly in the terminal. Also,
remember always to call the model and define it when running a stand-alone script.
python test_ollama.py

Let’s create a function to work better the model:

def simple_query(prompt, model=MODEL):


response = [Link](
model=MODEL,
prompt=prompt
)
return response

473
And run a new query:

MODEL="llama3.2:3b"
response = simple_query ("What is the capital of Peru?")
response

Let’s print the full response now. As a result, we will have the model response in a JSON
format:

GenerateResponse(model='llama3.2:3b', created_at='2026-02-26T18:58:30.730414614Z', done=True,

We can print the full answer better:

import json
print([Link](response.__dict__, indent=2))

{
"model": "llama3.2:3b",
"created_at": "2026-02-26T18:58:30.730414614Z",
"done": true,
"done_reason": "stop",
"total_duration": 2307368739,
"load_duration": 343767495,
"prompt_eval_count": 32,
"prompt_eval_duration": 722764406,
"eval_count": 8,
"eval_duration": 1229567370,
"response": "The capital of Peru is Lima.",
"thinking": null,
"context": [
128006,
9125,
...
13
],
"logprobs": null
}

As we can see, several pieces of information are generated, such as:

• response: the main output text generated by the model in response to our prompt.

474
– The capital of Peru is Lima.

• context: the token IDs representing the input and context used by the model. Tokens
are numerical representations of text that the language model uses to process it.

The Performance Metrics:

• total_duration: The total time taken for the operation in nanoseconds.


• load_duration: The time taken to load the model or components in nanoseconds.
• prompt_eval_duration: The time taken to evaluate the prompt in nanoseconds.
• eval_count: The number of tokens evaluated during the generation. Here, 9 tokens.
• eval_duration: The time taken for the model to generate the response in nanoseconds.

But what we want is the plain ‘response’ and, perhaps for analysis, the total duration of the
inference, so let’s change the code to extract it from the dictionary:

print(f"\n{response['response']}")
print(f"\nTotal Duration: {(response['total_duration']/1e9):.2f} seconds")

Now, we got:

The capital of Peru is Lima.

Total Duration: 2.31 seconds

We can also get the other available metrics:

print(f"eval_count: {response['eval_count']}")
print(f"eval_duration: {(response['eval_duration']/1e9):.2f} s")
print(f"eval_rate: {response['eval_count']/(response['eval_duration']/1e9):.2f} tokens/s")

eval_count: 8
eval_duration: 1.23 s
eval_rate: 6.51 tokens/s

475
Streaming with [Link]()

To stream the output from [Link]() in Python, set stream=True and iterate over
the generator to print each chunk as it’s produced. This enables real-time response streaming,
similar to chat models.

stream = [Link](
model='llama3.2:3b',
prompt='Tell me an interesting fact about Brazil',
stream=True
)

for chunk in stream:


print(chunk['response'], end='', flush=True)

• Each chunk['response'] is a part of the generated text, streamed as it’s created


• This allows responsive, real-time interaction in the terminal or UI.

This approach is ideal for long or complex generations, making the user experience feel faster
and more interactive.

System Prompt

Add a system parameter to set overall instructions or behavior for the model (useful for role
assignment and tone control):

response = [Link](model='llama3.2:3b',
prompt='Tell about industry',
system='You are an expert on Brazil.',
stream=False)
print(response['response'])

476
Temperature and Sampling

Control creativity/randomness via temperature, and customize output style with extra settings
like top_p and num_ctx (context window size)

response = [Link](
model='llama3.2:3b',
prompt='Why the sky is blue?',
options={'temperature':0.1},
stream=True)
for chunk in response:
print(chunk['response'], end='', flush=True)

477
response = [Link](
messages=[
{"role": "user",
"content": "Poetically describe Paris in one short sentence"},
],
model='llama3.2:3b',
options={"temperature": 1.0} # Set. temp. to 1.0 for more creativity
)
print(response['message']['content'])

Like a velvet-draped secret, Paris whispers ancient mysteries to the moonlit


Seine, her sighing shadows weaving an eternal waltz of love and dreams.

More Options

Besides temperature, we can control a lot via the options={…} parameter in [Link].
Key ones:
Sampling/style

• top_k: limit candidates to top K tokens each step, e.g., top_k=40


• top_p: nucleus sampling cutoff, e.g., top_p=0.9
• repeat_penalty: discourage repetition, e.g., repeat_penalty=1.1
• seed: fixed seed for reproducible outputs

Length/context

• num_predict: max new tokens, e.g., num_predict=256


• num_ctx: context window (tokens to keep in RAM), e.g., num_ctx=4096

478
Putting it together in your code:

response = [Link](
model='llama3.2:3b',
prompt='Why is the sky blue?',
options={
'temperature': 0.1,
'top_p': 0.9,
'top_k': 40,
'repeat_penalty': 1.1,
'num_predict': 50, # Limit the answer
'num_ctx': 4096,
'seed': 42,
},
stream=True,
)
for chunk in response:
print(chunk['response'], end='', flush=True)

Here we will limit the response on 50 tokens:

The sky appears blue because of a phenomenon called Rayleigh scattering, named after the
British physicist Lord Rayleigh, who first described it in the late 19th century.

Here's what happens:

1. When sunlight enters Earth's atmosphere, it encounters

How top_k and top_p affect generation

top_k and top_p both control how random and diverse the model’s next tokens can be, but
they do it in different ways.
top_k: fixed shortlist size

• At each step, the model ranks all possible next tokens by probability.

• top_k = K means: keep only the K most likely tokens, set all others to probability
0, then sample from those K.

• Effects:

479
– Small k (1–20) → very focused, deterministic, “on‑rails”; less creative, fewer weird
tokens.
– Large k (50–200+) → more variety and creativity; higher chance of unusual or
off‑topic words.
• Extreme:
– top_k = 1 → greedy decoding: always pick the single most likely token, almost no
randomness.

top_p: dynamic shortlist by probability mass

• Tokens are sorted by probability.

• Starting from the top, you add tokens until their cumulative probability � p. Then
you sample only from that set.

• top_p = 0.9 means: “consider just enough top tokens to cover 90% of the probability
mass”.
• Effects:
– In confident situations (one token is clearly best), the shortlist may be very small
→ behavior similar to low k.
– In uncertain situations (probabilities spread out), more tokens enter the shortlist
→ more exploration.
• Typical:
– top_p � 0.9–0.95 is a common sweet spot for natural but not too wild text.

How they feel in practice

• Lower top_k / top_p → safer, more repetitive, less surprising.

• Higher top_k / top_p → more diverse, more creative, but also more risk of nonsense.

We can use them together—for example, top_k=40, top_p=0.9—where top_k gives a hard
cap and top_p then trims that set by probability mass.

[Link]()

Another way to get our response is to use [Link](), which generates the next message
in a chat with a provided model. This is a streaming endpoint, so a series of responses will
occur. Streaming can be disabled using "stream": false. The final response object will also
include statistics and additional data from the request.

480
PROMPT_1 = 'What is the capital of France?'

response = [Link](model=MODEL, messages=[


{'role': 'user','content': PROMPT_1,},])
resp_1 = response['message']['content']
print(f"\n{resp_1}")
print(f"\n [INFO] Total Duration: {(res['total_duration']/1e9):.2f} seconds")

The answer is the same as before.


An important consideration is that by using [Link](), the response is “clear” from
the model’s “memory” after the end of inference (only used once), but If we want to keep a
conversation, we must use [Link](). Let’s see it in action:

PROMPT_1 = 'What is the capital of France?'


response = [Link](model=MODEL, messages=[
{'role': 'user','content': PROMPT_1,},])
resp_1 = response['message']['content']
print(f"\n{resp_1}")
print(f"\n [INFO] Total Duration: {(response['total_duration']/1e9):.2f} seconds")

PROMPT_2 = 'and of Italy?'


response = [Link](model=MODEL, messages=[
{'role': 'user','content': PROMPT_1,},
{'role': 'assistant','content': resp_1,},
{'role': 'user','content': PROMPT_2,},])
resp_2 = response['message']['content']
print(f"\n{resp_2}")
print(f"\n [INFO] Total Duration: {(response['total_duration']/1e9):.2f} seconds")

In the above code, we are running two queries, and the second prompt considers the result of
the first one.
Here is how the model responded:

The capital of France is **Paris**. ��

[INFO] Total Duration: 2.82 seconds

The capital of Italy is **Rome**. ��

[INFO] Total Duration: 4.46 seconds

481
The above code works with two prompts. Let’s include a conversation variable to really provide
the chat with a memory:

# Initialize conversation history


conversation = []

# Function to chat with memory


def chat_with_memory(prompt, model=MODEL):
global conversation

# Add user message to conversation


[Link]({"role": "user", "content": prompt})

# Generate response with conversation history


response = [Link](
model=MODEL,
messages=conversation
)

# Add assistant's response to conversation history


[Link](response["message"])

# Return just the response text


return response["message"]["content"]

# Question
prompt = "What is the capital of Brazil"
print(chat_with_memory(prompt))

# Ask a follow-up question that relies on memory


follow_up = "And Peru?"
print(chat_with_memory(follow_up))

The capital of Brazil is Brasília.


The capital of Peru is Lima.

Image Description:

As we did with the visual models and the command line to analyze an image, the same
can be done here with Python. Let’s use the same image of Paris, but now with the
[Link]():

482
MODEL = 'llava-phi3:3.8b'
PROMPT = "Describe this picture"

with open('image_test_1.jpg', 'rb') as image_file:


img = image_file.read()

response = [Link](
model=MODEL,
prompt=PROMPT,
images= [img]
)
print(f"\n{response['response']}")
print(f"\n [INFO] Total Duration: {(res['total_duration']/1e9):.2f} seconds")

Here is the result:

This image captures the iconic cityscape of Paris, France. The vantage point
is high, providing a panoramic view of the Seine River that meanders through
the heart of the city. Several bridges arch gracefully over the river,
connecting different parts of the city. The Eiffel Tower, an iron lattice
structure with a pointed top and two antennas on its summit, stands tall in the
background, piercing the sky. It is painted in a light gray color, contrasting
against the blue sky speckled with white clouds.

The buildings that line the river are predominantly white or beige, their uniform
color palette broken occasionally by red roofs peeking through. The Seine River
itself appears calm and wide, reflecting the city's architectural beauty in its
surface. On either side of the river, trees add a touch of green to the urban
landscape.

The image is taken from an elevated perspective, looking down on the city. This
viewpoint allows for a comprehensive view of Paris's beautiful architecture and
layout. The relative positions of the buildings, bridges, and other structures
create a harmonious composition that showcases the city's charm.

In summary, this image presents a serene day in Paris, with its architectural
marvels - from the Eiffel Tower to the river-side buildings - all bathed in soft
colors under a clear sky.

[INFO] Total Duration: 256.45 seconds

The model took about 4 minutes (256.45 s) to return with a detailed image description.

483
Let’s capture an image from the Raspberry Pi camera and get the description, now using the
MoonDream model:

import time
import numpy as np
import [Link] as plt
from picamera2 import Picamera2
from PIL import Image

def capture_image(image_path):
# Initialize camera
picam2 = Picamera2() # default is index 0

# Configure the camera


config = picam2.create_still_configuration(main={"size": (520, 520)})
[Link](config)
[Link]()

# Wait for the camera to warm up


[Link](2)

# Capture image
picam2.capture_file(image_path)
print("Image captured: "+image_path)

# Stop camera
[Link]()
[Link]()

Using the above code, we can capture an image, which can be displayed with:

def show_image(image_path):
img = [Link](image_path)

# Display the image


[Link](figsize=(6, 6))
[Link](img)
[Link]("Captured Image")
[Link]()

Let’s also create a function to describe the image:

484
def image_description(img_path, model):
with open(img_path, 'rb') as file:
response = [Link](
model=model,
messages=[
{
'role': 'user',
'content': '''return the description of the image''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,
}
)
return response

Now, let’s put all togheter and capture an image from the camera:

IMG_PATH = "/home/mjrovai/Documents/OLLAMA/SST/capt_image.jpg"
MODEL = "moondream:latest"

apture_image(IMG_PATH)
show_image(IMG_PATH)
response = image_description(IMG_PATH, MODEL)
caption = response['message']['content']
print ("\n==> AI Response:", caption)
print(f"\n[INFO] ==> Total Duration: {
(response['total_duration']/1e9):.2f} seconds")

Pointing the camera at my table:

485
We got the description:

The image features a green table with various items on it. A white mug adorned with
black faces is prominently displayed, and there are several other mugs scattered around
the table as well. In addition to the mugs, there's also a microphone placed near them,
suggesting that this might be an office or workspace setting where someone could enjoy
their coffee while recording podcasts or audio content.

A computer keyboard can be seen in the background, indicating that it is likely

486
connected to a computer for work purposes. A mouse and a cell phone are also present on
the table, further emphasizing the technology-oriented nature of this scene.

[INFO] ==> Total Duration: 82.57 seconds

We can now change the image_description function to ask “Who are the faces in the mug?”.
The answer:

==> AI Response:
The mug has a picture of the Beatles on it.

In the Ollama_Python_Library Intro notebook, we can find the experiments using


the Ollama Python library.

Calling Direct API

One alternative to running an SLM in Python using Ollama is to call the API directly. Let’s
explore some advantages and disadvantages of both methods.
Python Library:

response = [Link](
model=MODEL,
prompt=QUERY)
result = response['response']

Direct API Calls:

import requests
import json

# Configuration
OLLAMA_URL = "[Link]
MODEL = MODEL

response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": QUERY,

487
"stream": False
}
)

response = [Link](
f"{OLLAMA_URL}/generate",
json={"model": MODEL,
"prompt": query,
"stream": False}
)
result = [Link]().get("response", "")

One clear advantage of the Python library is that it handles URL construction,
request formatting, and response parsing.

Error Handling

Python Library:

• Raises specific exceptions (e.g., [Link], [Link])


• Better typed error messages
• Automatically handles connection issues

Direct API Calls:

• We should manually check the response.status_code


• Generic HTTP errors
• Need to handle connection exceptions yourself

Connection Management

Python Library:

• Automatically detects Ollama instance (checks OLLAMA_HOST env variable or defaults to


localhost:11434)
• Built-in connection pooling and retry logic
• Handles timeouts gracefully

Direct API Calls:

• We should specify the URL manually


• Must implement our own retry logic
• Need to manage connection pooling yourself

488
Features & Functionality

Python Library:

• Clean access to all Ollama features (generate, chat, embeddings, list models, pull, etc.)
• Streaming is simple: for chunk in [Link](..., stream=True)
• Type hints and better IDE support

Direct API Calls:

• Access to any API endpoint, including undocumented ones


• Full control over request headers, timeouts, etc.
• Can use any HTTP library (requests, httpx, urllib, etc.)

Dependencies

Python Library:

pip install ollama

• Adds a dependency to our project


• Depends on httpx under the hood

Direct API Calls:

pip install requests

• We choose our HTTP library


• More lightweight if we only need basic functionality

Advanced Features

Python Library:

# Streaming
for chunk in [Link](model=MODEL, prompt=query, stream=True):
print(chunk['response'], end='')

# Chat history
response = [Link](
model=MODEL,
messages=[

489
{'role': 'user', 'content': 'Hello!'},
{'role': 'assistant', 'content': 'Hi there!'},
{'role': 'user', 'content': 'How are you?'}
]
)

# List models
models = [Link]()

# Pull models
[Link]('llama3.2:3b')

Direct API Calls:

• Require us to implement all these patterns manually


• More boilerplate code

When to Use Each?

Use the Python Library when:

• � Building standard applications


• � Need cleaner, more maintainable code
• � Need good error handling out of the box
• � Need streaming support
• � Okay with adding a dependency

Use Direct API Calls when:

• �Need fine-grained control over HTTP requests


• � Are working in a constrained environment
• � Need to customize timeouts, headers, or proxies
• � Need to minimize dependencies
• � Are debugging API issues
• � Need to access experimental/undocumented endpoints

490
Performance Difference

Minimal difference in practice! The Python library uses httpx which is comparable to requests.
Both make the same underlying HTTP calls to Ollama.

Bottom line: For most use cases, the Python library is the better choice due to its
simplicity and built-in features. Use direct API calls only when you need specific
control or have constraints that prevent adding the dependency.

Going Further

The small LLM models tested worked well at the edge, both with text and with images, but, of
course, the last one had high latency. A combination of specific, dedicated models can lead
to better results; for example, in real cases, an Object Detection model (such as YOLO) can
provide a general description and count of objects in an image, which, once passed to an LLM,
can help extract essential insights and actions.
According to Avi Baum, CTO at Hailo,

In the vast landscape of artificial intelligence (AI), one of the most intriguing
journeys has been the evolution of AI on the edge. This journey has taken us from
classic machine vision to the realms of discriminative AI, enhancive AI, and now,
the groundbreaking frontier of generative AI. Each step has brought us closer to a
future where intelligent systems seamlessly integrate with our daily lives, offering
an immersive experience of not just perception but also creation at the palm of our
hand.

491
Conclusion

This chapter has demonstrated how a Raspberry Pi 5 can be transformed into a potent AI
hub capable of running large language models (LLMs) for real-time, on-site data analysis and
insights using Ollama and Python. The Raspberry Pi’s versatility and power, coupled with the
capabilities of lightweight LLMs like Llama 3.2 and MoonDream, make it an excellent platform
for edge computing applications.
The potential of running LLMs on the edge extends far beyond simple data processing, as in
this lab’s examples. Here are some innovative suggestions for using this project:
1. Smart Home Automation:

• Integrate SLMs to interpret voice commands or analyze sensor data for intelligent home
automation. This could include real-time monitoring and control of home devices, security
systems, and energy management, all processed locally without relying on cloud services.

2. Field Data Collection and Analysis:

• Deploy SLMs on Raspberry Pi in remote or mobile setups for real-time data collection
and analysis. This can be used in agriculture to monitor crop health, in environmental
studies for wildlife tracking, or in disaster response for situational awareness and resource
management.

3. Educational Tools:

492
• Create interactive educational tools that leverage SLMs to provide instant feedback,
language translation, and tutoring. This can be particularly useful in developing regions
with limited access to advanced technology and internet connectivity.

4. Healthcare Applications:

• Use SLMs for medical diagnostics and patient monitoring. They can provide real-time
analysis of symptoms and suggest potential treatments. This can be integrated into
telemedicine platforms or portable health devices.

5. Local Business Intelligence:

• Implement SLMs in retail or small business environments to analyze customer behavior,


manage inventory, and optimize operations. The ability to process data locally ensures
privacy and reduces dependency on external services.

6. Industrial IoT:

• Integrate SLMs into industrial IoT systems for predictive maintenance, quality control,
and process optimization. The Raspberry Pi can serve as a localized data processing
unit, reducing latency and improving the reliability of automated systems.

7. Autonomous Vehicles:

• Use SLMs to process sensory data from autonomous vehicles, enabling real-time decision-
making and navigation. This can be applied to drones, robots, and self-driving cars for
enhanced autonomy and safety.

8. Cultural Heritage and Tourism:

• Implement SLMs to provide interactive and informative cultural heritage sites and
museum guides. Visitors can use these systems to get real-time information and insights,
enhancing their experience without internet connectivity.

9. Artistic and Creative Projects:

• Use SLMs to analyze and generate creative content, such as music, art, and literature. This
can foster innovative projects in the creative industries and allow for unique interactive
experiences in exhibitions and performances.

10. Customized Assistive Technologies:

• Develop assistive technologies for individuals with disabilities, providing personalized


and adaptive support through real-time text-to-speech, language translation, and other
accessible tools.

493
Resources

• Ollama_Python_Library Intro notebook


• Find out which AI models your machine can actually run

494
SLM: Basic Optimization Techniques

Figure 14: DALL·E prompt - I am writing a tutorial using Raspberry Pi. I am talking about
Optimization Techniques, such as Function Calling and RAG, using Ollama and
SLMs. I want a landscape-format image for the tutorial cover (without title). Should
be a cartoon styled in 5Os

Introduction

Large Language Models (LLMs) have revolutionized natural language processing, but their
deployment and optimization come with unique challenges. One significant issue is the
tendency for LLMs (and more, the SLMs) to generate plausible-sounding but factually incorrect

495
information, a phenomenon known as hallucination. This occurs when models produce
content that appears coherent but lacks grounding in truth or real-world facts.
Other challenges include the immense computational resources required to train and run these
models, the difficulty of maintaining up-to-date knowledge within the model, and the need
for domain-specific adaptations. Privacy concerns also arise when handling sensitive data
during training or inference. Additionally, ensuring consistent performance across diverse tasks
and maintaining ethical use of these powerful tools present ongoing challenges. Addressing
these issues is crucial for the effective and responsible deployment of LLMs in real-world
applications.
The fundamental and more common techniques for enhancing LLM (and SLM) performance
and efficiency are Function (or Tool) Calling, Prompt engineering, Fine-tuning, and Retrieval-
Augmented Generation (RAG).

• Function (Tool) calling allows models to perform actions beyond generating text. By
integrating with external functions or APIs, SLMs can access real-time data, automate
tasks, and perform precise calculations—addressing the reliability issues that arise from
the model’s limitations in mathematical operations.
• Prompt engineering is at the forefront of LLM optimization. By carefully crafting
input prompts, we can guide models to produce more accurate and relevant outputs. This
technique involves structuring queries that leverage the model’s pre-trained knowledge
and capabilities, often incorporating examples or specific instructions to shape the desired
response.
• Retrieval-Augmented Generation (RAG) represents a powerful approach that’s ideal
for resource-constrained edge devices. This method combines the knowledge embedded
in pre-trained models with the ability to access external, up-to-date information without
requiring fine-tuning. By retrieving relevant data from a local knowledge base, RAG
significantly enhances accuracy and reduces hallucinations—all without the computational
overhead of model retraining.
• Fine-tuning, while more resource-intensive, offers a way to specialize LLMs for specific
domains or tasks. This process involves further training the model on carefully curated
datasets, allowing it to adapt its vast general knowledge to particular applications. Fine-
tuning can lead to substantial performance improvements, especially in specialized fields
or for unique use cases.

In this chapter, we’ll start focusing on two techniques that are particularly well-suited for edge
devices like the Raspberry Pi: Function Calling and RAG.

We will learn more in detail about optimization techniques for SLMs, in the chapter:
Advancing EdgeAI: Beyond Basic SLMs

496
Function Calling Introduction

So far, we can see that, with the model’s (“response”) answer to a variable, we can efficiently
work with it and integrate it into real-world projects. However, a big problem is that the model
can respond differently to the same prompt. Let’s say, as in the last examples, that we want
the model’s response to be only the name of a given country’s capital and its coordinates,
nothing more, even with very verbose models such as the Microsoft Phi. We can use the Ollama
function's calling to guarantee the same answers, which is perfectly compatible with the
OpenAI API.

But what exactly is “function calling”?

In modern artificial intelligence, function calling with Large Language Models (LLMs) allows
these models to perform actions beyond generating text. By integrating with external functions
or APIs, LLMs can access real-time data, automate tasks, and interact with various systems.
For instance, instead of merely responding to a weather query, an LLM can call a weather API
to fetch the current conditions and provide accurate, up-to-date information. This capability
enhances the relevance and accuracy of the model’s responses, making it a powerful tool for
driving workflows and automating processes, thereby transforming it into an active participant
in real-world applications.
For more details about Function Calling, please see this video made by Marvin Prison:
[Link]
And on this link: HuggingFace Function Calling

Using the SLM for calculations

Let’s do a simple calculation in Python:

123456*123456

The result would be: 15,241,383,936. No issues on it, but let’s ask a SLM to do the same
simple task:

import ollama

response = [Link](
model='llama3.2:3B',
messages=[{
"role": "user",

497
"content": "What is 123456 multiplied by 123456? Only give me the answer"
}],
options={"temperature": 0}
)

Examining the response: [Link], we would get: 51,131,441,376, what


is completely wrong!
This is a fundamental limitation of LLMs - they’re not calculators. Here’s why the answer
is wrong:

Why LLMs Fail at Math

1. LLMs Predict Text, They Don’t Calculate

• LLMs work by predicting the next most likely token based on patterns in training data
• They don’t perform actual arithmetic operations
• They’re essentially “guessing” what a plausible answer looks like

2. Tokenization Issues
Numbers are broken into tokens in ways that don’t align with mathematical operations:

"123456" might tokenize as: ["123", "456"] or ["12", "34", "56"]

This makes it nearly impossible for the model to “see” the actual numbers properly for
computation.

3. Pattern Matching vs. Computation

• The model has seen similar multiplication problems in training


• It tries to recall patterns rather than compute
• For simple problems (2×3), it might seem to work because it memorizes common facts
• For larger numbers (123456×123456), it has no memorized pattern to fall back on

4. The Correct Answer

123456 × 123456 = 15,241,383,936

The LLM will likely give something that “looks” like a big number but is mathematically
incorrect.

498
The Solution: Function Calling / Tool Use

The pattern is:

1. Use the LLM for understanding intent (classification)


2. Use Python for actual computation (the multiply() function)

def multiply(a, b):


"""Actual computation - always correct"""
result = a * b
return f"The product of {a} and {b} is {result}."

Why Even Temperature=0 Doesn’t Help

Setting temperature=0 makes the output deterministic (same input → same output), but it
doesn’t make it correct. The model will confidently give the same wrong answer every time.

LLM/SLM Math Performance by Problem Type

Problem Type LLM Accuracy Why


2+2 ~99% Memorized in training
47 + 89 ~70% Some pattern recognition
123 × 456 ~30% Struggles with multi-digit
123456 × 123456 ~0% No chance without computation
Calculate 15% tip on $47.83 ~40% Multi-step reasoning fails

Best Practices

# � DON'T: Ask LLM/SLM to calculate directly


response = [Link](
model='llama3.2:3b',
messages=[{"role": "user", "content": "What is 123456 * 123456?"}]
)

# � DO: Use LLM to understand, Python to compute


classification = ask_ollama_for_classification("What is 123456 * 123456?")
if classification["type"] == "multiplication":
result = multiply(123456, 123456) # Python does the math

499
Bottom line: Use LLMs for natural language understanding and intent classifica-
tion, but delegate actual computations to proper tools/functions. This is the core
principle behind tool use and function calling in modern LLM applications!

Function Calling Solution for Calculations

Define the Tool (Function Schema)

multiply_tool = {
"type": "function",
"function": {
"name": "multiply_numbers",
"description": "Multiply two numbers together",
"parameters": {
"type": "object",
"required": ["a", "b"],
"properties": {
"a": {"type": "number", "description": "First number"},
"b": {"type": "number", "description": "Second number"}
}
}
}
}

Implement the Function

def multiply_numbers(a, b):


# Convert to int or float as needed
a = float(a)
b = float(b)
return {"result": a * b}

Now, let’s create a function to handle the user query and calling for the tool when needed.

def answer_query(QUERY):
response = [Link](
'llama3.2:3B',
messages=[{"role": "user", "content": QUERY}],

500
tools=[multiply_tool]
)

# Check if the model wants to call the tool


if [Link].tool_calls:
for tool in [Link].tool_calls:
if [Link] == "multiply_numbers":
# Ensure arguments are passed as numbers
result = multiply_numbers(**[Link])
print(f"Result: {result['result']:,.2f}")
else:
print(f"It is not a Multiplication")

Note the line tools=[multiply_tool], now as a part of the ollama’s calling.

And run the function, we will get the correct answer.

answer_query("What is 123456 multiplied by 123456?")

Result: 15,241,383,936.00

Great! And now, can I use the same code to answer general questions? Let’s test it:

answer_query("What is the capital of Brazil?")

Result: 1,000,000.00

The result is wrong. So, the above approach works fine for using the tool, but to answer it
correctly (even without a tool), we should implement an “agentic approach”, which is a subject
for later (See the Chapter: Advancing EdgeAI: Beyond Basic SLMs)

Project: Calculating Distances

Suppose we want an SLM to return the distance in km from the capital city of the country
specified by the user to the user’s current location. We can see that the first is not so simple:
it is not always enough to enter only the country’s name; the SLM can also give us a different
(and incorrect) answer every time.

501
OK, for trying to mitigate it, let’s create an app where the user enters a country’s name and
gets, as an output, the distance in km from the capital city of such a country and the app’s
location (for simplicity, we will use Santiago, Chile, as the app location).

Once the user enters a country name, the model should return the capital city’s name (as
a string) and its latitude and longitude (as floats). Using those coordinates, we can use
a simple Python library (haversine) to compute the great‑circle distance between the two
latitude/longitude points.

502
The idea of this project is to demonstrate a combination of language model interaction (IA)
and geospatial calculations using the Haversine formula (traditional computing).
First, let us install the Haversine library:

pip install haversine

Now, we should create a Python script designed to interact with our model (LLM) to determine
the coordinates of a country’s capital city and calculate the distance from Santiago de Chile to
that capital.
Let’s go over the code:
Importing Libraries

import time
from haversine import haversine
from ollama import chat

Basic Variables and Model

MODEL = 'llama3.2:3B' # The name of the model to be used


mylat = -33.33 # Latitude of Santiago de Chile
mylon = -70.51 # Longitude of Santiago de Chile

• MODEL: Specifies the model being used, which is, in this example, the Lhama3.2.
• mylat and mylon: Coordinates of Santiago de Chile, used as the starting point for the
distance calculation.

Defining a Python Function That Acts as a Tool

def calc_distance(lat, lon, city):


"""Compute distance and print a descriptive message."""
distance = haversine((mylat, mylon), (lat, lon), unit="km")
msg = f"\nSantiago de Chile is about {int(round(distance, -1)):,} I am running a \
few minutes late; my previous meeting is running [Link] away from {city}."
return {"city": city, "distance_km": int(round(distance, -1)), "message": msg}

This is the real Python function that Ollama will be allowed to call. It performs the following
steps: • Takes latitude, longitude, and city name as input arguments. • Uses the haversine
library to calculate the distance from Santiago to the target city. • Returns a JSON‑like
dictionary containing the computed distance and a human‑readable text summary.

503
In Ollama’s terminology, this is a tool — a callable external function that the LLM
may invoke automatically

Declaring the tool descriptor (schema)

tools = [
{
"type": "function",
"function": {
"name": "calc_distance",
"description": "Calculates the distance from Santiago, Chile to a \
given city's coordinates.",
"parameters": {
"type": "object",
"properties": {
"lat": {"type": "number", "description": "Latitude of the city"},
"lon": {"type": "number", "description": "Longitude of the city"},
"city": {"type": "string", "description": "Name of the city"}
},
"required": ["lat", "lon", "city"]
}
}
}
]

This JSON object describes the metadata and input schema of the tool so that the LLM knows:
• Name: which function to call. • Description: what purpose it serves. • Parameters: input
argument types and their descriptions. This schema mirrors the OpenAI function‑calling format
and is fully supported in Ollama � 0.4 .

Defining tools this way allows Ollama to validate arguments before sending a call
request back.

Asking the Model to Use the Tool

response = chat(
model=MODEL,
messages=[{
"role": "user",
"content": f"Find the decimal latitude and longitude of the capital of \
I am running a few minutes late; my previous meeting is running over. running a few\
minutes late; my previous meeting is running over.
{country},"

504
" then use the calc_distance tool to determine how far it is from \
Santiago de Chile."
}],
tools=tools
)

• The chat() function is called with the chosen model, a message, and the tools list.
• The prompt instructs the model first to identify the capital and its coordinates, and then
invoke the tool (calc_distance) with those values.
• Ollama returns a structured response that may include a tool_calls section, indicating
which tool to execute.

Executing the Model’s Tool Call


For example, if country = "Colombia", the model will will return as the result:

Where the arguments are the expected response in JSON format.


To get and display the result to the user, we should iterate over all tool_calls, extract
the JSON arguments (lat, lon, city), and execute the local Python function using them. To
finish it, we should print a human‑readable message (e.g., “Santiago de Chile is about X
kilometers away from “CITY”).
So, let’s first get the name of the city and the coordinates with:

for call in [Link].tool_calls:


raw_args = call["function"]["arguments"]

city = raw_args['city']
lat = float(raw_args['lat'])
lon = float(raw_args['lon'])

Now, we can calculate and print the distance, using haversine():

distance = haversine((mylat, mylon), (lat, lon), unit='km')


print(f"Santiago de Chile is about {int(round(distance, -1)):,}
kilometers away from {city}.")

505
In this case:

Santiago de Chile is about 4,240 kilometers away from Bogota.

NOTE: Sometimes the model returns parameter names that differ from what your function
expects. Specifically, Ollama occasionally returns argument objects like:
{"lat1": -33.33, "lon1": -70.51, "lat2": 48.8566, "lon2": 2.3522, "city":
"Paris"}
Instead of the schema-defined keys (lat, lon, city).
This happens because some LLMs (such as Llama 3.2 and Qwen 3) attempt to be “helpful”
by naming coordinates explicitly—lat1/lon1 for origin and lat2/lon2 for destination—even
when the schema only defines lat/lon.
To handle this, optionally a mapping-correction step can be added after decoding the tocall
arguments.

if hasattr([Link], "tool_calls") and [Link].tool_calls:


for call in [Link].tool_calls:
if call["function"]["name"] == "calc_distance":
raw_args = call["function"]["arguments"]

# Decode JSON if necessary


args = [Link](raw_args) if isinstance(raw_args, str) else raw_args

# Normalize key names


if "lat1" in args or "lat2" in args:
args["lat"] = [Link]("lat2") or [Link]("lat1")
args["lon"] = [Link]("lon2") or [Link]("lon1")
if "latitude" in args:
args["lat"] = args["latitude"]
if "longitude" in args:
args["lon"] = args["longitude"]
args = {k: v for k, v in [Link]() if k in ("lat", "lon", "city")}

# Convert numbers
args["lat"] = float(args["lat"])
args["lon"] = float(args["lon"])

result = calc_distance(**args)
print(result["message"])

506
Timing and Diagnostic Output

elapsed = time.perf_counter() - start


print(f"[INFO] ==> Model {MODEL} : {elapsed:.1f} s")

This records how long the operation took from prompt submission to tool execution, useful for
benchmarking response performance.
Example Usage
If we enter different countries, for example, France, Colombia, and the United States, We can
note that we always receive the same structured information:

ask_and_measure("France")
ask_and_measure("Colombia")
ask_and_measure("United States")

If you run the code as a script, the result will be printed on the terminal:

And the calculations are pretty good!

507
The complete script can be found at: func_call_dist_calc.py and on the 10-Ollama_Function_Calling
notebook.

Running with other models

The models that will run with the described approach are the ones that can handle tools. For
example, Gemma 3 and 3n will not work.
An alternative is to use the Pydantic library to serialize the schema using model_json_schema().
Using the Pydantic library, models as Gemma can also be used, as explored in the:
20-Ollama_Function_Calling_Pydantic notebook

Adding images

Now it is time to wrap up everything so far! Let’s modify the script using Pydantic so that
instead of entering the country name (as a text), the user enters an image, and the application
(based on SLM) returns the city in the image and its geographic location. With that data, we
can calculate the distance as before.

508
For simplicity, we will implement this new code in two steps. First, the LLM will analyze the
image and create a description (text). This text will be passed on to another instance, where
the model will extract the information needed to pass along.

We will start importing the libraries

import time
from haversine import haversine
from ollama import chat
from pydantic import BaseModel, Field

We can see the image if you run the code on the Jupyter Notebook. For that, we also need to
import:

import [Link] as plt


from PIL import Image

Those libraries are unnecessary if we run the code as a script.

Now, we define the model and the local coordinates:

MODEL = 'gemma3:4b'
mylat = -33.33
mylon = -70.51

We can download a new image, for example, Machu Picchu from Wikipedia. On the Notebook
we can see it:

# Load the image


img_path = "image_test_3.jpg"
img = [Link](img_path)

# Display the image

509
[Link](figsize=(8, 8))
[Link](img)
[Link]('off')
#[Link]("Image")
[Link]()

Now, let’s define a function that will receive the image and will return the decimal
latitude and decimal longitude of the city in the image, its name, and what
country it is located

def image_description(img_path):
with open(img_path, 'rb') as file:
response = chat(
model=MODEL,
messages=[
{
'role': 'user',
'content': '''return the decimal latitude and decimal longitude
of the city in the image, its name, and
what country it is located''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,

510
}
)
#print(response['message']['content'])
return response['message']['content']

We can print the entire response for debug purposes. In this case, we can get
something as:
'{\n "city": "Machu Picchu",\n "country": "Peru",\n "lat": -
13.1631,\n "lon": -72.5450\n}\n'

Let’s define a Pydantic model (CityCoord) that describes the expected structure of the SLM’s
response. It expects four fields: country, city (city name), lat (latitude), and lon (longitude).

class CityCoord(BaseModel):
city: str = Field(..., description="Name of the city in the image")
country: str = Field(..., description="Name of the country where the city in the\ image i
lat: float = Field(..., description="Decimal Latitude of the city in the image")
lon: float = Field(..., description="Decimal Longitude of the city in the image")

The image description generated for the function will be passed as a prompt for the model
again.

response = chat(
model=MODEL,
messages=[{
"role": "user",
"content": image_description # image_description from previous model's run
}],
format=CityCoord.model_json_schema(), # Structured JSON format
options={"temperature": 0}
)

Now we can get the required data using:

resp = CityCoord.model_validate_json([Link])

And so, we can calculate and print the distance, using haversine():

distance = haversine((mylat, mylon), ([Link], [Link]), unit='km')


print(f"\nThe image shows {[Link]}, with lat:{round([Link], 2)} and \
long: {round([Link], 2)}, located in {[Link]} and \
about {int(round(distance, -1)):,} kilometers away from Santiago, Chile.\n")

511
And we will get:

The image shows Machu Picchu, with lat:-13.16 and long: -72.55, located in Peru and about 2,2

In the 30-Function_Calling_with_images notebook, you can find experiments with multiple


images.
Let’s now download the script calc_distance_image.py from the GitHub and run it on the
terminal with the command:

python calc_distance_image.py /home/mjrovai/Documents/OLLAMA/image_test_3.jpg

Enter with the Machu Picchu image full patch as an argument. We will get the same previous
result.

Let’s change the model for the gemma3:4b:

The app is working fine with both models, with the Gemma being faster.

512
How about Paris?

Of course, there are many ways to optimize the code used here. Still, the idea is to explore
the considerable potential of function calling with SLMs at the edge, allowing those models to
integrate with external functions or APIs. Going beyond text generation, SLMs can access
real-time data, automate tasks, and interact with various systems.

Retrievel Augmentation Generation (RAG)

In a basic interaction between a user and a language model, the user asks a question, which is
sent to the model as a prompt. The model generates a response based solely on its pre-trained
knowledge.

513
In a RAG process, there’s an additional step between the user’s question and the model’s
response. The user’s question triggers a retrieval process from a knowledge base.

A simple RAG project

Here are the steps to implement a basic Retrieval Augmented Generation (RAG):

• Determine the type of documents you’ll be using: The best types are documents
from which we can get clean and unobscured text. PDFs can be problematic because
they are designed for printing, not for extracting sensible text. To work with PDFs, we
should get the source document or use tools to handle it.
• Chunk the text: We can’t store the text as one long stream because of context size
limitations and the potential for confusion. Chunking involves splitting the text into
smaller pieces. Chunk text has many ways, such as character count, tokens, words,
paragraphs, or sections. It is also possible to overlap chunks.

514
• Create embeddings: Embeddings are numerical representations of text that capture
semantic meaning. We create embeddings by passing each chunk of text through a
particular embedding model. The model outputs a vector, the length of which depends
on the embedding model used. We should pull one (or more) embedding models from
Ollama, to perform this task. Here are some examples of embedding models available at
Ollama.

Model Parameter Size Embedding Size


mxbai-embed-large 334M 1024
nomic-embed-text 137M 768
all-minilm 23M 384

Generally, larger embedding sizes capture more nuanced information about the
input. Still, they also require more computational resources to process, and a
higher number of parameters should increase the latency (but also the quality
of the response).

• Store the chunks and embeddings in a vector database: We will need a way to
efficiently find the most relevant chunks of text for a given prompt, which is where a vector
database comes in. We will use Chromadb, an AI-native open-source vector database,
which simplifies building RAGs by creating knowledge, facts, and skills pluggable for
LLMs. Both the embedding and the source text for each chunk are stored.
• Build the prompt: When we have a question, we create an embedding and query the
vector database for the most similar chunks. Then, we select the top few results and
include their text in the prompt.

The goal of RAG is to provide the model with the most relevant information from our documents,
allowing it to generate more accurate and informative responses. So, let’s implement a simple
example of an SLM incorporating a particular set of facts about bees (“Bee Facts”).
Inside the ollama env, enter the command in the terminal for Chromadb instalation:

pip install ollama chromadb

Let’s pull an intermediary embedding model, nomic-embed-text

ollama pull nomic-embed-text

And create a working directory:

515
cd Documents/OLLAMA/
mkdir RAG-simple-bee
cd RAG-simple-bee/

Let’s create a new Jupyter notebook, 40-RAG-simple-bee for some exploration:


Import the needed libraries:

import ollama
import chromadb
import time

And define aor models:

EMB_MODEL = "nomic-embed-text"
MODEL = 'llama3.2:3B'

Initially, a knowledge base about bee facts should be created. This involves collecting relevant
documents and converting them into vector embeddings. These embeddings are then stored in
a vector database, allowing for efficient similarity searches later. Enter with the “document,” a
base of “bee facts” as a list:

documents = [
"Bee-keeping, also known as apiculture, involves the maintenance of bee \
colonies, typically in hives, by humans.",
"The most commonly kept species of bees is the European honey bee (Apis \
mellifera).",

...

"There are another 20,000 different bee species in the world.",


"Brazil alone has more than 300 different bee species, and the \
vast majority, unlike western honey bees, don’t sting.",
"Reports written in 1577 by Hans Staden, mention three native bees \
used by indigenous people in Brazil.",
"The indigenous people in Brazil used bees for medicine and food purposes",
"From Hans Staden report: probable species: mandaçaia (Melipona \
quadrifasciata), mandaguari (Scaptotrigona postica) and jataí-amarela \
(Tetragonisca angustula)."
]

516
We do not need to “chunk” the document here because we will use each element of
the list as a chunk.

Now, we will create our vector embedding database bee_facts and store the document in it:

client = [Link]()
collection = client.create_collection(name="bee_facts")

# store each document in a vector embedding database


for i, d in enumerate(documents):
response = [Link](model=EMB_MODEL, prompt=d)
embedding = response["embedding"]
[Link](
ids=[str(i)],
embeddings=[embedding],
documents=[d]
)

Now that we have our “Knowledge Base” created, we can start making queries, retrieving data
from it:

User Query: The process begins when a user asks a question, such as “How many bees are in
a colony? Who lays eggs, and how much? How about common pests and diseases?”

517
prompt = "How many bees are in a colony? Who lays eggs and how much? How about\
common pests and diseases?"

Query Embedding: The user’s question is converted into a vector embedding using the
same embedding model used for the knowledge base.

response = [Link](
prompt=prompt,
model=EMB_MODEL
)

Relevant Document Retrieval: The system searches the knowledge base using the query
embedding to find the most relevant documents (in this case, the 5 more probable). This is done
using a similarity search, which compares the query embedding to the document embeddings
in the database.

results = [Link](
query_embeddings=[response["embedding"]],
n_results=5
)
data = results['documents']

Prompt Augmentation: The retrieved relevant information is combined with the original
user query to create an augmented prompt. This prompt now contains the user’s question and
pertinent facts from the knowledge base.

prompt=f"Using this data: {data}. Respond to this prompt: {prompt}",

Answer Generation: The augmented prompt is then fed into a language model, in this case,
the llama3.2:3b model. The model uses this enriched context to generate a comprehensive
answer. Parameters like temperature, top_k, and top_p are set to control the randomness and
quality of the generated response.

output = [Link](
model=MODEL,
prompt=f"Using this data: {data}. Respond to this prompt: {prompt}",
options={
"temperature": 0.0,
"top_k":10,
"top_p":0.5 }
)

518
Response Delivery: Finally, the system returns the generated answer to the user.

print(output['response'])

Based on the provided data, here are the answers to your questions:

1. How many bees are in a colony?


A typical bee colony can contain between 20,000 and 80,000 bees.

2. Who lays eggs and how much?


The queen bee lays up to 2,000 eggs per day during peak seasons.

3. What about common pests and diseases?


Common pests and diseases that affect bees include varroa mites, hive beetles,
and foulbrood.

Let’s create a function to help answer new questions:

def rag_bees(prompt, n_results=5, temp=0.0, top_k=10, top_p=0.5):


start_time = time.perf_counter() # Start timing

# generate an embedding for the prompt and retrieve the data


response = [Link](
prompt=prompt,
model=EMB_MODEL
)

results = [Link](
query_embeddings=[response["embedding"]],
n_results=n_results
)
data = results['documents']

# generate a response combining the prompt and data retrieved


output = [Link](
model=MODEL,
prompt=f"Using this data: {data}. Respond to this prompt: {prompt}",
options={
"temperature": temp,
"top_k": top_k,
"top_p": top_p }
)

519
print(output['response'])

end_time = time.perf_counter() # End timing


elapsed_time = round((end_time - start_time), 1) # Calculate elapsed time

print(f"\n[INFO] ==> The code for model: {MODEL}, took {elapsed_time}s \


to generate the answer.\n")

We can now create queries and call the function:

prompt = "Are bees in Brazil?"


rag_bees(prompt)

Yes, bees are found in Brazil. According to the data, Brazil has more than 300
different bee species, and indigenous people in Brazil used bees for medicine and
food purposes. Additionally, reports from 1577 mention three native bees used by
indigenous people in Brazil.

[INFO] ==> The code for model: llama3.2:3b, took 22.7s to generate the answer.

By the way, if the model used supports multiple languages, we can use it (for example,
Portuguese), even if the dataset was created in English:

prompt = "Existem abelhas no Brazil?"


rag_bees(prompt)

Sim, existem abelhas no Brasil! De acordo com o relato de Hans Staden, há três
espécies de abelhas nativas do Brasil que foram mencionadas: mandaçaia (Melipona
quadrifasciata), mandaguari (Scaptotrigona postica) e jataí-amarela (Tetragonisca
angustula). Além disso, o Brasil é conhecido por ter mais de 300 espécies
diferentes de abelhas, a maioria das quais não é agressiva e não põe veneno.

[INFO] ==> The code for model: llama3.2:3b, took 54.6s to generate the answer.

In the Chapter Advancing EdgeAI: Beyond Basic SLMs, we will learn how to
implement a Naive RAG System

520
Conclusion

Throughout this chapter, we’ve explored two fundamental optimization techniques that signifi-
cantly enhance the capabilities of Small Language Models (SLMs) running on edge devices like
the Raspberry Pi: Function Calling and Retrieval-Augmented Generation (RAG).
We began by addressing a critical limitation of language models—their inability to perform
accurate calculations. By implementing function calling, we demonstrated how to transform
SLMs from text generators into actionable agents that can interact with external tools and APIs.
Whether extracting structured data like geographic coordinates, fetching real-time weather
information, or performing precise mathematical operations, function calling bridges the gap
between natural language understanding and deterministic computation. The key principle
remains:

Use SLMs for intent classification and understanding, while delegating specific tasks
to specialized functions that guarantee accuracy.

The RAG implementation showcased an elegant solution to another fundamental challenge—the


static nature of model knowledge and the tendency toward hallucinations. By integrating
ChromaDB with vector embeddings, we created a system that augments the model’s responses
with relevant, factual information retrieved from a knowledge base. This approach not only
grounds the model’s answers in verifiable data but also enables the system to stay current
without the resource-intensive process of retraining or fine-tuning.
These techniques are particularly valuable for edge AI applications where computational
resources are limited, and offline operation is often required. Function calling ensures reliability
and precision, while RAG provides flexibility and accuracy without demanding continuous
internet connectivity once the knowledge base is established.
As we move forward with our edge AI projects, remember that the power of SLMs lies not just
in their standalone capabilities but in how effectively we orchestrate them with complementary
tools and techniques. The patterns demonstrated here—structured tool use and knowledge
augmentation—form the foundation for building robust, production-ready AI applications on
resource-constrained devices.
In the chapter Advancing EdgeAI: Beyond Basic SLMs, we’ll go deeper into these concepts
and explore more advanced implementations.

Resources

• 10-Ollama_Function_Calling notebook
• 20-Ollama_Function_Calling_Pydantic
• 30-Function_Calling_with_images notebook

521
• 40-RAG-simple-bee notebook
• calc_distance_image python script

522
Vision-Language Models at the Edge

We will learn Vison-Language Models across tasks such as captioning, object detection,
grounding, and segmentation on a Raspberry Pi.

Figure 15: DALL·E prompt - A Raspberry Pi setup featuring vision tasks. The image shows a
Raspberry Pi connected to a camera, with various computer vision tasks displayed
visually around it, including object detection, image captioning, segmentation, and
visual grounding. The Raspberry Pi is placed on a desk, with a display showing
bounding boxes and annotations related to these tasks. The background should be a
home workspace, with tools and devices typically used by developers and hobbyists.

In this hands-on lab, we will continuously explore AI applications at the Edge, going from the
basic setup of the Florence-2, Microsoft’s state-of-the-art vision foundation model, to advanced
implementations on devices like the Raspberry Pi.

523
Why Florence-2 at the Edge?

Florence-2 is a vision-language model open-sourced by Microsoft under the MIT license, which
significantly advances vision-language models by combining a lightweight architecture with
robust capabilities. Thanks to its training on the massive FLD-5B dataset, which contains
126 million images and 5.4 billion visual annotations, it achieves performance comparable to
larger models. This makes Florence-2 ideal for deployment at the edge, where power and
computational resources are limited.
In this tutorial, we will explore how to use Florence-2 for real-time computer vision applications,
such as:

• Image captioning
• Object detection
• Segmentation
• Visual grounding

Visual grounding involves linking textual descriptions to specific regions within


an image. This enables the model to understand where particular objects or entities
described in a prompt are in the image. For example, if the prompt is “a red
car,” the model will identify and highlight the region where the red car is found
in the image. Visual grounding is helpful for applications where precise alignment
between text and visual content is needed, such as human-computer interaction,
image annotation, and interactive AI systems.

In the tutorial, we will walk through:

• Setting up Florence-2 on the Raspberry Pi


• Running inference tasks such as object detection and captioning
• Optimizing the model to get the best performance from the edge device
• Exploring practical, real-world applications with fine-tuning.

Florence-2 Model Architecture

Florence-2 utilizes a unified, prompt-based representation to handle various vision-language


tasks. The model architecture consists of two main components: an image encoder and a
multi-modal transformer encoder-decoder.

524
• Image Encoder: The image encoder is based on the DaViT (Dual Attention Vision
Transformers) architecture. It converts input images into a series of visual token embed-
dings. These embeddings serve as the foundational representations of the visual content,
capturing both spatial and contextual information about the image.
• Multi-Modal Transformer Encoder-Decoder: Florence-2’s core is the multi-modal
transformer encoder-decoder, which combines visual token embeddings from the image
encoder with textual embeddings generated by a BERT-like model. This combination

525
allows the model to simultaneously process visual and textual inputs, enabling a unified
approach to tasks such as image captioning, object detection, and segmentation.

The model’s training on the extensive FLD-5B dataset ensures it can effectively handle diverse
vision tasks without requiring task-specific modifications. Florence-2 uses textual prompts to
activate specific tasks, making it highly flexible and capable of zero-shot generalization. For
tasks like object detection or visual grounding, the model incorporates additional location
tokens to represent regions within the image, ensuring a precise understanding of spatial
relationships.

Florence-2’s compact architecture and innovative training approach allow it to


perform computer vision tasks accurately, even on resource-constrained devices like
the Raspberry Pi.

Technical Overview

Florence-2 introduces several innovative features that set it apart:

Architecture

• Lightweight Design: Two variants available

526
– Florence-2-Base: 232 million parameters
– Florence-2-Large: 771 million parameters
• Unified Representation: Handles multiple vision tasks through a single architecture
• DaViT Vision Encoder: Converts images into visual token embeddings
• Transformer-based Multi-modal Encoder-Decoder: Processes combined visual
and text embeddings

Training Dataset (FLD-5B)

• 126 million unique images


• 5.4 billion comprehensive annotations, including:
– 500M text annotations
– 1.3B region-text annotations
– 3.6B text-phrase-region annotations

527
• Automated annotation pipeline using specialist models
• Iterative refinement process for high-quality labels

Key Capabilities

Florence-2 excels in multiple vision tasks:

Zero-shot Performance

• Image Captioning: Achieves 135.6 CIDEr score on COCO


• Visual Grounding: 84.4% recall@1 on Flickr30k
• Object Detection: 37.5 mAP on COCO val2017
• Referring Expression: 67.0% accuracy on RefCOCO

Fine-tuned Performance

• Competitive with specialist models despite the smaller size


• Outperforms larger models in specific benchmarks
• Efficient adaptation to new tasks

Practical Applications

Florence-2 can be applied across various domains:

1. Content Understanding
• Automated image captioning for accessibility
• Visual content moderation
• Media asset management
2. E-commerce
• Product image analysis
• Visual search
• Automated product tagging
3. Healthcare
• Medical image analysis
• Diagnostic assistance
• Research data processing
4. Security & Surveillance

528
• Object detection and tracking
• Anomaly detection
• Scene understanding

Comparing Florence-2 with other VLMs

Florence-2 stands out from other visual language models due to its impressive zero-shot
capabilities. Unlike models like Google PaliGemma, which rely on extensive fine-tuning to
adapt to various tasks, Florence-2 works right out of the box, as we will see in this lab. It can
also compete with larger models like GPT-4V and Flamingo, which often have many more
parameters but only sometimes match Florence-2’s performance. For example, Florence-2
achieves better zero-shot results than Kosmos-2 despite having over twice the parameters.
In benchmark tests, Florence-2 has shown remarkable performance in tasks like COCO cap-
tioning and referring expression comprehension. It outperformed models like PolyFormer and
UNINEXT in object detection and segmentation tasks on the COCO dataset. It is a highly
competitive choice for real-world applications where both performance and resource efficiency
are crucial.

Setup and Installation

Our choice of edge device is the Raspberry Pi 5 (Raspi-5). Its robust platform is equipped
with the Broadcom BCM2712, a 2.4GHz quad-core 64-bit Arm Cortex-A76 CPU featuring
Cryptographic Extension and enhanced caching capabilities. It boasts a VideoCore VII GPU,
dual 4Kp60 HDMI® outputs with HDR, and a 4Kp60 HEVC decoder. Memory options include
4GB and 8GB of high-speed LPDDR4X SDRAM, with 8GB being our choice to run Florence-2.
It also features expandable storage via a microSD card slot and a PCIe 2.0 interface for fast
peripherals such as M.2 SSDs (Solid State Drives).

For real applications, SSDs are a better option than SD cards.

We suggest installing an Active Cooler, a dedicated clip-on cooling solution for Raspberry Pi 5
(Raspi-5), for this lab. It combines an aluminum heatsink with a temperature-controlled blower
fan to keep the Raspi-5 operating comfortably under heavy loads, such as running Florense-2.

529
Environment configuration

To run Microsoft Florense-2 on the Raspberry Pi 5, we’ll need a few libraries:

1. Transformers:
• Florence-2 uses the transformers library from Hugging Face for model loading
and inference. This library provides the architecture for working with pre-trained
vision-language models, making it easy to perform tasks like image captioning, object
detection, and more. Essentially, transformers helps in interacting with the model,
processing input prompts, and obtaining outputs.
2. PyTorch:
• PyTorch is a deep learning framework that provides the infrastructure needed to
run the Florence-2 model, which includes tensor operations, GPU acceleration (if
a GPU is available), and model training/inference functionalities. The Florence-2
model is trained in PyTorch, and we need it to leverage its functions, layers, and
computation capabilities to perform inferences on the Raspberry Pi.
3. Timm (PyTorch Image Models):

530
• Florence-2 uses timm to access efficient implementations of vision models and pre-
trained weights. Specifically, the timm library is utilized for the image encoder
part of Florence-2, particularly for managing the DaViT architecture. It provides
model definitions and optimized code for common vision tasks and allows the easy
integration of different backbones that are lightweight and suitable for edge devices.
4. Einops:
• Einops is a library for flexible and powerful tensor operations. It makes it easy to
reshape and manipulate tensor dimensions, which is especially important for the
multi-modal processing done in Florence-2. Vision-language models like Florence-2
often need to rearrange image data, text embeddings, and visual embeddings to align
correctly for the transformer blocks, and einops simplifies these complex operations,
making the code more readable and concise.

In short, these libraries enable different essential components of Florence-2:

• Transformers and PyTorch are needed to load the model and run the inference.
• Timm is used to access and efficiently implement the vision encoder.
• Einops helps reshape data, facilitating the integration of visual and text features.

All these components work together to help Florence-2 run seamlessly on our Raspberry Pi,
allowing it to perform complex vision-language tasks relatively quickly.
Considering that the Raspberry Pi already has its OS installed, let’s use SSH to reach it from
another computer:

ssh mjrovai@[Link]

And check the IP allocated to it:

hostname -I

[Link]

531
Updating the Raspberry Pi
First, ensure your Raspberry Pi is up to date:

sudo apt update


sudo apt upgrade -y

Initial setup for using PIP:

sudo apt install python3-pip


sudo rm /usr/lib/python3.11/EXTERNALLY-MANAGED
pip3 install --upgrade pip

Install Dependencies

sudo apt-get install libjpeg-dev libopenblas-dev libopenmpi-dev libomp-dev

Let’s set up and activate a Virtual Environment for working with Florence-2:

python3 -m venv ~/florence


source ~/florence/bin/activate

Install PyTorch

532
pip3 install setuptools numpy Cython
pip3 install requests
pip3 install torch torchvision --index-url [Link]
pip3 install torchaudio --index-url [Link]

Let’s verify that PyTorch is correctly installed:

Install Transformers, Timm and Einops:

pip3 install transformers


pip3 install timm einops

Install the model:

pip3 install autodistill-florence-2

Jupyter Notebook and Python libraries


Installing a Jupyter Notebook to run and test our Python scripts is possible.

pip3 install jupyter


pip3 install numpy Pillow matplotlib
jupyter notebook --generate-config

Testing the installation

Running the Jupyter Notebook on the remote computer

533
jupyter notebook --ip=[Link] --no-browser

Running the above command on the SSH terminal, we can see the local URL address to open
the notebook:

The notebook with the code used on this initial test can be found on the Lab GitHub:

• 10-florence2_test.ipynb

We can access it on the remote computer by entering the Raspberry Pi’s IP address and the
provided token in a web browser ( copy the entire URL from the terminal).
From the Home page, create a new notebook [Python 3 (ipykernel) ] and copy and paste
the example code from Hugging Face Hub.
The code is designed to run Florence-2 on a given image to perform object detection. It
loads the model, processes an image and a prompt, and then generates a response to identify
and describe the objects in the image.

• The processor helps prepare text and image inputs.


• The model takes the processed inputs to generate a meaningful response.
• The post-processing step refines the generated output into a more interpretable form,
like bounding boxes for detected objects.

This workflow leverages the versatility of Florence-2 to handle vision-language


tasks and is implemented efficiently using PyTorch, Transformers, and related
image-processing tools.

534
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForCausalLM

device = "cuda:0" if [Link].is_available() else "cpu"


torch_dtype = torch.float16 if [Link].is_available() else torch.float32

model = AutoModelForCausalLM.from_pretrained("microsoft/Florence-2-base",
torch_dtype=torch_dtype,
trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base",
trust_remote_code=True)

prompt = "<OD>"

url = "[Link]
images/resolve/main/transformers/tasks/[Link]?download=true"
image = [Link]([Link](url, stream=True).raw)

inputs = processor(text=prompt, images=image, return_tensors="pt").to(


device, torch_dtype)

generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
do_sample=False,
num_beams=3,
)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=False)[0]

parsed_answer = processor.post_process_generation(generated_text, task="<OD>",


image_size=([Link],
[Link]))

print(parsed_answer)

Let’s break down the provided code step by step:

535
1. Importing Required Libraries

import requests
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForCausalLM

• requests: Used to make HTTP requests. In this case, it downloads an image from a
URL.
• PIL (Pillow): Provides tools for manipulating images. Here, it’s used to open the
downloaded image.
• torch: PyTorch is imported to handle tensor operations and determine the hardware
availability (CPU or GPU).
• transformers: This module provides easy access to Florence-2 by using AutoProcessor
and AutoModelForCausalLM to load pre-trained models and process inputs.

2. Determining the Device and Data Type

device = "cuda:0" if [Link].is_available() else "cpu"


torch_dtype = torch.float16 if [Link].is_available() else torch.float32

• Device Setup: The code checks if a CUDA-enabled GPU is available ([Link].is_available()).


The device is set to “cuda:0” if a GPU is available. Otherwise, it defaults to "cpu" (our
case here).
• Data Type Setup: If a GPU is available, torch.float16 is chosen, which uses half-
precision floats to speed up processing and reduce memory usage. On the CPU, it defaults
to torch.float32 to maintain compatibility.

3. Loading the Model and Processor

model = AutoModelForCausalLM.from_pretrained("microsoft/Florence-2-base",
torch_dtype=torch_dtype,
trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base",
trust_remote_code=True)

• Model Initialization:

536
– AutoModelForCausalLM.from_pretrained() loads the pre-trained Florence-2
model from Microsoft’s repository on Hugging Face. The torch_dtype is set
according to the available hardware (GPU/CPU), and trust_remote_code=True
allows the use of any custom code that might be provided with the model.
– .to(device) moves the model to the appropriate device (either CPU or GPU). In
our case, it will be set to CPU.
• Processor Initialization:
– AutoProcessor.from_pretrained() loads the processor for Florence-2. The pro-
cessor is responsible for transforming text and image inputs into a format the model
can work with (e.g., encoding text, normalizing images, etc.).

4. Defining the Prompt

prompt = "<OD>"

• Prompt Definition: The string "<OD>" is used as a prompt. This refers to “Object
Detection”, instructing the model to detect objects on the image.

5. Downloading and Loading the Image

url = "[Link]
images/resolve/main/transformers/tasks/[Link]?download=true"
image = [Link]([Link](url, stream=True).raw)

• Downloading the Image: The [Link]() function fetches the image from the
specified URL. The stream=True parameter ensures the image is streamed rather than
downloaded completely at once.
• Opening the Image: [Link]() opens the image so the model can process it.

6. Processing Inputs

inputs = processor(text=prompt, images=image, return_tensors="pt").to(device,


torch_dtype)

• Processing Input Data: The processor() function processes the text (prompt) and
the image (image). The return_tensors="pt" argument converts the processed data
into PyTorch tensors, which are necessary for inputting data into the model.

537
• Moving Inputs to Device: .to(device, torch_dtype) moves the inputs to the
correct device (CPU or GPU) and assigns the appropriate data type.

7. Generating the Output

generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
do_sample=False,
num_beams=3,
)

• Model Generation: [Link]() is used to generate the output based on the


input data.
– input_ids: Represents the tokenized form of the prompt.
– pixel_values: Contains the processed image data.
– max_new_tokens=1024: Specifies the maximum number of new tokens to be gener-
ated in the response. This limits the response length.
– do_sample=False: Disables sampling; instead, the generation uses deterministic
methods (beam search).
– num_beams=3: Enables beam search with three beams, which improves output quality
by considering multiple possibilities during generation.

8. Decoding the Generated Text

generated_text = processor.batch_decode(generated_ids, skip_special_tokens=False)[0]

• Batch Decode: processor.batch_decode() decodes the generated IDs (tokens) into


readable text. The skip_special_tokens=False parameter means that the output will
include any special tokens that may be part of the response.

9. Post-processing the Generation

parsed_answer = processor.post_process_generation(generated_text, task="<OD>",


image_size=([Link],
[Link]))

538
• Post-Processing: processor.post_process_generation() is called to process the
generated text further, interpreting it based on the task ("<OD>" for object detection)
and the size of the image.
• This function extracts specific information from the generated text, such as bounding
boxes for detected objects, making the output more useful for visual tasks.

10. Printing the Output

print(parsed_answer)

• Finally, print(parsed_answer) displays the output, which could include object detection
results, such as bounding box coordinates and labels for the detected objects in the image.

Result

Running the code, we get as the Parsed Answer:

{'<OD>': {'bboxes': [[34.23999786376953, 160.0800018310547, 597.4400024414062,


371.7599792480469], [272.32000732421875, 241.67999267578125, 303.67999267578125,
247.4399871826172], [454.0799865722656, 276.7200012207031, 553.9199829101562,
370.79998779296875], [96.31999969482422, 280.55999755859375, 198.0800018310547,
371.2799987792969]], 'labels': ['car', 'door handle', 'wheel', 'wheel']}}

First, Let’s inspect the image:

import [Link] as plt


[Link](figsize=(8, 8))
[Link](image)
[Link]('off')
[Link]()

539
By the Object Detection result, we can see that:

'labels': ['car', 'door handle', 'wheel', 'wheel']

It seems that at least a few objects were detected. we can also implement a code to draw the
bounding boxes in the find objects:

def plot_bbox(image, data):


# Create a figure and axes
fig, ax = [Link]()

# Display the image


[Link](image)

# Plot each bounding box


for bbox, label in zip(data['bboxes'], data['labels']):
# Unpack the bounding box coordinates

540
x1, y1, x2, y2 = bbox
# Create a Rectangle patch
rect = [Link]((x1, y1), x2-x1, y2-y1, linewidth=1,
edgecolor='r', facecolor='none')
# Add the rectangle to the Axes
ax.add_patch(rect)
# Annotate the label
[Link](x1, y1, label, color='white', fontsize=8,
bbox=dict(facecolor='red', alpha=0.5))

# Remove the axis ticks and labels


[Link]('off')

# Show the plot


[Link]()

Box (x0, y0, x1, y1): Location tokens correspond to the top-left and bottom-right
corners of a box.

And running

plot_bbox(image, parsed_answer['<OD>'])

We get:

541
Florence-2 Tasks

Florence-2 is designed to perform a variety of computer vision and vision-language tasks through
prompts. These tasks can be activated by providing a specific textual prompt to the model, as
we saw with <OD> (Object Detection).
Florence-2’s versatility comes from combining these prompts, allowing us to guide the model’s
behavior to perform specific vision tasks. Changing the prompt allows us to adapt Florence-2 to
different tasks without needing task-specific modifications in the architecture. This capability
directly results from Florence-2’s unified model architecture and large-scale multi-task training
on the FLD-5B dataset.
Here are some of the key tasks that Florence-2 can perform, along with example prompts:

542
1. Object Detection (OD)

• Prompt: "<OD>"
• Description: Identifies objects in an image and provides bounding boxes for each
detected object. This task is helpful for applications like visual inspection, surveillance,
and general object recognition.

2. Image Captioning

• Prompt: "<CAPTION>"
• Description: Generates a textual description for an input image. This task helps the
model describe what is happening in the image, providing a human-readable caption for
content understanding.

3. Detailed Captioning

• Prompt: "<DETAILED_CAPTION>"
• Description: Generates a more detailed caption with more nuanced information about
the scene, such as the objects present and their relationships.

4. Visual Grounding

• Prompt: "<CAPTION_TO_PHRASE_GROUNDING>"
• Description: Links a textual description to specific regions in an image. For example,
given a prompt like “a green car,” the model highlights where the red car is in the image.
This is useful for human-computer interaction, where you must find specific objects based
on text.

5. Segmentation

• Prompt: "<REFERRING_EXPRESSION_SEGMENTATION>"
• Description: Performs segmentation based on a referring expression, such as “the
blue cup.” The model identifies and segments the specific region containing the object
mentioned in the prompt (all related pixels).

6. Dense Region Captioning

• Prompt: "<DENSE_REGION_CAPTION>"
• Description: Provides captions for multiple regions within an image, offering a detailed
breakdown of all visible areas, including different objects and their relationships.

543
7. OCR with Region

• Prompt: "<OCR_WITH_REGION>"
• Description: Performs Optical Character Recognition (OCR) on an image and provides
bounding boxes for the detected text. This is useful for extracting and locating textual
information in images, such as reading signs, labels, or other forms of text in images.

8. Phrase Grounding for Specific Expressions

• Prompt: "<CAPTION_TO_PHRASE_GROUNDING>" along with a specific expression, such as


"a wine glass".
• Description: Locates the area in the image that corresponds to a specific textual phrase.
This task allows for identifying particular objects or elements when prompted with a
word or keyword.

9. Open Vocabulary Object Detection

• Prompt: "<OPEN_VOCABULARY_OD>"
• Description: The model can detect objects without being restricted to a predefined list
of classes, making it helpful in recognizing a broader range of items based on general
visual understanding.

Exploring computer vision and vision-language tasks

For exploration, all codes can be found on the GitHub:

• 20-florence_2.ipynb

Let’s use a couple of images created by Dall-E and upload them to the Rasp-5 (FileZilla can
be used for that). The images will be saved on a sub-folder named images :

dogs_cats = [Link]('./images/[Link]')
table = [Link]('./images/[Link]')

544
Let’s create a function to facilitate our exploration and to keep track of the latency of the
model for different tasks:

def run_example(task_prompt, text_input=None, image=None):


start_time = time.perf_counter() # Start timing
if text_input is None:
prompt = task_prompt
else:
prompt = task_prompt + text_input
inputs = processor(text=prompt, images=image,
return_tensors="pt").to(device)
generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
early_stopping=False,
do_sample=False,
num_beams=3,
)
generated_text = processor.batch_decode(generated_ids,
skip_special_tokens=False)[0]
parsed_answer = processor.post_process_generation(
generated_text,
task=task_prompt,
image_size=([Link], [Link])

545
)

end_time = time.perf_counter() # End timing


elapsed_time = end_time - start_time # Calculate elapsed time
print(f" \n[INFO] ==> Florence-2-base ({task_prompt}),
took {elapsed_time:.1f} seconds to execute.\n")

return parsed_answer

Caption

1. Dogs and Cats

run_example(task_prompt='<CAPTION>',image=dogs_cats)

[INFO] ==> Florence-2-base (<CAPTION>), took 16.1 seconds to execute.

{'<CAPTION>': 'A group of dogs and cats sitting in a garden.'}

2. Table

run_example(task_prompt='<CAPTION>',image=table)

[INFO] ==> Florence-2-base (<CAPTION>), took 16.5 seconds to execute.

{'<CAPTION>': 'A wooden table topped with a plate of fruit and a glass of wine.'}

DETAILED_CAPTION

1. Dogs and Cats

run_example(task_prompt='<DETAILED_CAPTION>',image=dogs_cats)

[INFO] ==> Florence-2-base (<DETAILED_CAPTION>), took 25.5 seconds to execute.

{'<DETAILED_CAPTION>': 'The image shows a group of cats and dogs sitting on top of a
lush green field, surrounded by plants with flowers, trees, and a house in the
background. The sky is visible above them, creating a peaceful atmosphere.'}

546
2. Table

run_example(task_prompt='<DETAILED_CAPTION>',image=table)

[INFO] ==> Florence-2-base (<DETAILED_CAPTION>), took 26.8 seconds to execute.

{'<DETAILED_CAPTION>': 'The image shows a wooden table with a bottle of wine and a
glass of wine on it, surrounded by a variety of fruits such as apples, oranges, and
grapes. In the background, there are chairs, plants, trees, and a house, all slightly
blurred.'}

MORE_DETAILED_CAPTION

1. Dogs and Cats

run_example(task_prompt='<MORE_DETAILED_CAPTION>',image=dogs_cats)

[INFO] ==> Florence-2-base (<MORE_DETAILED_CAPTION>), took 49.8 seconds to execute.

{'<MORE_DETAILED_CAPTION>': 'The image shows a group of four cats and a dog in a garden.
The garden is filled with colorful flowers and plants, and there is a pathway leading up
to a house in the background. The main focus of the image is a large German Shepherd dog
standing on the left side of the garden, with its tongue hanging out and its mouth open,
as if it is panting or panting. On the right side, there are two smaller cats, one orange
and one gray, sitting on the grass. In the background, there is another golden retriever
dog sitting and looking at the camera. The sky is blue and the sun is shining, creating a
warm and inviting atmosphere.'}

2. Table

run_example(task_prompt='< MORE_DETAILED_CAPTION>',image=table)

INFO] ==> Florence-2-base (<MORE_DETAILED_CAPTION>), took 32.4 seconds to execute.

{'<MORE_DETAILED_CAPTION>': 'The image shows a wooden table with a wooden tray on it. On
the tray, there are various fruits such as grapes, oranges, apples, and grapes. There is
also a bottle of red wine on the table. The background shows a garden with trees and a
house. The overall mood of the image is peaceful and serene.'}

We can note that the more detailed the caption task, the longer the latency and
the possibility of mistakes (like “The image shows a group of four cats and a dog in
a garden”, instead of two dogs and three cats).

547
OD - Object Detection

We can run the same previous function for object detection using the prompt <OD>.

task_prompt = '<OD>'
results = run_example(task_prompt,image=dogs_cats)
print(results)

Let’s see the result:

[INFO] ==> Florence-2-base (<OD>), took 20.9 seconds to execute.

{'<OD>': {'bboxes': [[737.7920532226562, 571.904052734375, 1022.4640502929688,


980.4800415039062], [0.5120000243186951, 593.4080200195312, 211.4560089111328,
991.7440185546875], [445.9520263671875, 721.4080200195312, 680.4480590820312,
850.4320678710938], [39.42400360107422, 91.64800262451172, 491.0080261230469,
933.3760375976562], [570.8800048828125, 184.83201599121094, 974.3360595703125,
782.8480224609375]], 'labels': ['cat', 'cat', 'cat', 'dog', 'dog']}}

Only by the labels ['cat,' 'cat,' 'cat,' 'dog,' 'dog'] is it possible to see that the main
objects in the image were captured. Let’s apply the function used before to draw the bounding
boxes:

plot_bbox(dogs_cats, results['<OD>'])

548
Let’s also do it with the Table image:

task_prompt = '<OD>'
results = run_example(task_prompt,image=table)
plot_bbox(table, results['<OD>'])

[INFO] ==> Florence-2-base (<OD>), took 40.8 seconds to execute.

549
DENSE_REGION_CAPTION

It is possible to mix the classic Object Detection with the Caption task in specific sub-regions
of the image:

task_prompt = '<DENSE_REGION_CAPTION>'

results = run_example(task_prompt,image=dogs_cats)
plot_bbox(dogs_cats, results['<DENSE_REGION_CAPTION>'])

results = run_example(task_prompt,image=table)
plot_bbox(table, results['<DENSE_REGION_CAPTION>'])

550
CAPTION_TO_PHRASE_GROUNDING

With this task, we can enter with a caption, such as “a wine glass”, “a wine bottle,” or “a half
orange,” and Florence-2 will localize the object in the image:

task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'

results = run_example(task_prompt, text_input="a wine bottle",image=table)


plot_bbox(table, results['<CAPTION_TO_PHRASE_GROUNDING>'])

results = run_example(task_prompt, text_input="a wine glass",image=table)


plot_bbox(table, results['<CAPTION_TO_PHRASE_GROUNDING>'])

results = run_example(task_prompt, text_input="a half orange",image=table)


plot_bbox(table, results['<CAPTION_TO_PHRASE_GROUNDING>'])

551
[INFO] ==> Florence-2-base (<CAPTION_TO_PHRASE_GROUNDING>), took 15.7 seconds to execute
each task.

Cascade Tasks

We can also enter the image caption as the input text to push Florence-2 to find more objects:

task_prompt = '<CAPTION>'
results = run_example(task_prompt,image=dogs_cats)
text_input = results[task_prompt]
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(task_prompt, text_input,image=dogs_cats)
plot_bbox(dogs_cats, results['<CAPTION_TO_PHRASE_GROUNDING>'])

Changing the task_prompt among <CAPTION,> <DETAILED_CAPTION> and <MORE_DETAILED_CAPTION>,


we will get more objects in the image.

552
OPEN_VOCABULARY_DETECTION

<OPEN_VOCABULARY_DETECTION> allows Florence-2 to detect recognizable objects in an im-


age without relying on a predefined list of categories, making it a versatile tool for iden-
tifying various items that may not have been explicitly labeled during training. Unlike
<CAPTION_TO_PHRASE_GROUNDING>, which requires a specific text phrase to locate and high-
light a particular object in an image, <OPEN_VOCABULARY_DETECTION> performs a broad scan
to find and classify all objects present.
This makes <OPEN_VOCABULARY_DETECTION> particularly useful for applications where you
need a comprehensive overview of everything in an image without prior knowledge of what to
expect. Enter with a text describing specific objects not previously detected, resulting in their
detection. For example:

task_prompt = '<OPEN_VOCABULARY_DETECTION>'
text = ["a house", "a tree", "a standing cat at the left",
"a sleeping cat on the ground", "a standing cat at the right",
"a yellow cat"]
for txt in text:
results = run_example(task_prompt, text_input=txt,image=dogs_cats)
bbox_results = convert_to_od_format(results['<OPEN_VOCABULARY_DETECTION>'])
plot_bbox(dogs_cats, bbox_results)

553
[INFO] ==> Florence-2-base (<OPEN_VOCABULARY_DETECTION>), took 15.1 seconds
to execute each task.

Note: Trying to use Florence-2 to find objects that were not found can leads to
mistakes (see exaamples on the Notebook).

Referring expression segmentation

We can also segment a specific object in the image and give its description (caption), such as
“a wine bottle” on the table image or “a German Sheppard” on the dogs_cats.
Referring expression segmentation results format: {'<REFERRING_EXPRESSION_SEGMENTATION>':
{'Polygons': [[[polygon]], ...], 'labels': ['', '', ...]}}, one object is repre-
sented by a list of polygons. each polygon is [x1, y1, x2, y2, ..., xn, yn].

Polygon (x1, y1, …, xn, yn): Location tokens represent the vertices of a polygon
in clockwise order.

So, let’s first create a function to plot the segmentation:

554
from PIL import Image, ImageDraw, ImageFont
import copy
import random
import numpy as np
colormap = ['blue','orange','green','purple','brown','pink','gray','olive',
'cyan','red','lime','indigo','violet','aqua','magenta','coral','gold',
'tan','skyblue']

def draw_polygons(image, prediction, fill_mask=False):


"""
Draws segmentation masks with polygons on an image.

Parameters:
- image_path: Path to the image file.
- prediction: Dictionary containing 'polygons' and 'labels' keys.
'polygons' is a list of lists, each containing vertices
of a polygon.
'labels' is a list of labels corresponding to each polygon.
- fill_mask: Boolean indicating whether to fill the polygons with color.
"""
# Load the image

draw = [Link](image)

# Set up scale factor if needed (use 1 if not scaling)


scale = 1

# Iterate over polygons and labels


for polygons, label in zip(prediction['polygons'], prediction['labels']):
color = [Link](colormap)
fill_color = [Link](colormap) if fill_mask else None

for _polygon in polygons:


_polygon = [Link](_polygon).reshape(-1, 2)
if len(_polygon) < 3:
print('Invalid polygon:', _polygon)
continue

_polygon = (_polygon * scale).reshape(-1).tolist()

# Draw the polygon

555
if fill_mask:
[Link](_polygon, outline=color, fill=fill_color)
else:
[Link](_polygon, outline=color)

# Draw the label text


[Link]((_polygon[0] + 8, _polygon[1] + 2), label, fill=color)

# Save or display the image


#[Link]() # Display the image
display(image)

Now we can run the functions:

task_prompt = '<REFERRING_EXPRESSION_SEGMENTATION>'

results = run_example(task_prompt, text_input="a wine bottle",image=table)


output_image = [Link](table)
draw_polygons(output_image,
results['<REFERRING_EXPRESSION_SEGMENTATION>'],
fill_mask=True)

results = run_example(task_prompt, text_input="a german sheppard",image=dogs_cats)


output_image = [Link](dogs_cats)
draw_polygons(output_image,
results['<REFERRING_EXPRESSION_SEGMENTATION>'],
fill_mask=True)

556
[INFO] ==> Florence-2-base (<REFERRING_EXPRESSION_SEGMENTATION>),
took 207.0 seconds to execute each task.

Region to Segmentation

With this task, it is also possible to give the object coordinates in the image to segment it. The
input format is '<loc_x1><loc_y1><loc_x2><loc_y2>', [x1, y1, x2, y2] , which is the
quantized coordinates in [0, 999].
For example, when running the code:

task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(task_prompt, text_input="a half orange",image=table)
results

The results were:

{'<CAPTION_TO_PHRASE_GROUNDING>': {'bboxes': [[343.552001953125,


689.6640625,
530.9440307617188,
873.9840698242188]],
'labels': ['a half']}}

557
Using the bboxes rounded coordinates:

task_prompt = '<REGION_TO_SEGMENTATION>'
results = run_example(task_prompt,
text_input="<loc_343><loc_690><loc_531><loc_874>",
image=table)
output_image = [Link](table)
draw_polygons(output_image, results['<REGION_TO_SEGMENTATION>'], fill_mask=True)

We got the segmentation of the object on those coordinates (Latency: 83 seconds):

Region to Texts

We can also give the region (coordinates and ask for a caption):

558
task_prompt = '<REGION_TO_CATEGORY>'
results = run_example(task_prompt, text_input="<loc_343><loc_690><loc_531>
<loc_874>",image=table)
results

[INFO] ==> Florence-2-base (<REGION_TO_CATEGORY>), took 14.3 seconds to execute.

{'<REGION_TO_CATEGORY>': 'orange<loc_343><loc_690><loc_531><loc_874>'}

The model identified an orange in that region. Let’s ask for a description:

task_prompt = '<REGION_TO_DESCRIPTION>'
results = run_example(task_prompt, text_input="<loc_343><loc_690><loc_531>
<loc_874>",image=table)
results

[INFO] ==> Florence-2-base (<REGION_TO_CATEGORY>), took 14.6 seconds to execute.

{'<REGION_TO_CATEGORY>': 'orange<loc_343><loc_690><loc_531><loc_874>'}

In this case, the description did not provide more details, but it could. Try another example.

OCR

With Florence-2, we can perform Optical Character Recognition (OCR) on an image, getting
what is written on it (task_prompt = '<OCR>' and also get the bounding boxes (location) for
the detected text (ask_prompt = '<OCR_WITH_REGION>'). Those tasks can help extract and
locate textual information in images, such as reading signs, labels, or other forms of text in
images.
Let’s upload a flyer from a talk in Brazil to Raspi. Let’s test works in another language, here
Portuguese):

flayer = [Link]('./images/[Link]')
# Display the image
[Link](figsize=(8, 8))
[Link](flayer)
[Link]('off')
#[Link]("Image")
[Link]()

559
Let’s examine the image with '<MORE_DETAILED_CAPTION>' :

[INFO] ==> Florence-2-base (<MORE_DETAILED_CAPTION>), took 85.2 seconds to execute.

{'<MORE_DETAILED_CAPTION>': 'The image is a promotional poster for an event called


"Machine Learning Embarcados" hosted by Marcelo Roval. The poster has a black
background with white text. On the left side of the poster, there is a logo of a
coffee cup with the text "Café Com Embarcados" above it. Below the logo, it says
"25 de Setembro as 17th" which translates to "25th of September as 17" in English.
\n\nOn the right side, there aretwo smaller text boxes with the names of the
participants and their names. The first text box reads "Democratizando a Inteligência
Artificial para Paises em Desenvolvimento" and the second text box says "Toda quarta-feira" w
the image, there has a photo of Marcelo, a man with a beard and glasses, smiling at
the camera. He is wearing a white hard hat and a white shirt. The text boxes are
in orange and yellow colors.'}

The description is very accurate. Let’s get to the more important words with the task OCR:

task_prompt = '<OCR>'
run_example(task_prompt,image=flayer)

[INFO] ==> Florence-2-base (<OCR>), took 37.7 seconds to execute.

560
{'<OCR>': 'Machine LearningCafécomEmbarcadoEmbarcadosDemocratizando a
InteligênciaArtificial para Paises em25 de Setembro ás 17hDesenvolvimentoToda quarta-
feiraMarcelo RovalProfessor na UNIFIEI eTransmissão viainCo-Director do TinyML4D'}

Let’s locate the words in the flyer:

task_prompt = '<OCR_WITH_REGION>'
results = run_example(task_prompt,image=flayer)

Let’s also create a function to draw bounding boxes around the detected words:

def draw_ocr_bboxes(image, prediction):


scale = 1
draw = [Link](image)
bboxes, labels = prediction['quad_boxes'], prediction['labels']
for box, label in zip(bboxes, labels):
color = [Link](colormap)
new_box = ([Link](box) * scale).tolist()
[Link](new_box, width=3, outline=color)
[Link]((new_box[0]+8, new_box[1]+2),
"{}".format(label),
align="right",

fill=color)
display(image)

output_image = [Link](flayer)
draw_ocr_bboxes(output_image, results['<OCR_WITH_REGION>'])

561
We can inspect the detected words:

results['<OCR_WITH_REGION>']['labels']

'</s>Machine Learning',
'Café',
'com',
'Embarcado',
'Embarcados',
'Democratizando a Inteligência',
'Artificial para Paises em',
'25 de Setembro ás 17h',
'Desenvolvimento',
'Toda quarta-feira',
'Marcelo Roval',
'Professor na UNIFIEI e',
'Transmissão via',
'in',
'Co-Director do TinyML4D']

562
Latency Summary

The latency observed for different tasks using Florence-2 on the Raspberry Pi (Raspi-5) varied
depending on the complexity of the task:

• Image Captioning: It took approximately 16-17 seconds to generate a caption for an


image.
• Detailed Captioning: Increased latency to around 25-27 seconds, requiring generating
more nuanced scene descriptions.
• More Detailed Captioning: It took about 32-50 seconds, and the latency increased as
the description grew more complex.
• Object Detection: It took approximately 20-41 seconds, depending on the image’s
complexity and the number of detected objects.
• Visual Grounding: Approximately 15-16 seconds to localize specific objects based on
textual prompts.
• OCR (Optical Character Recognition): Extracting text from an image took around
37-38 seconds.
• Segmentation and Region to Segmentation: Segmentation tasks took considerably
longer, with a latency of around 83-207 seconds, depending on the complexity and the
number of regions to be segmented.

These latency times highlight the resource constraints of edge devices like the Raspberry Pi
and emphasize the need to optimize the model and the environment to achieve real-time
performance.

563
Running complex tasks can use all 8GB of the Raspi-5’s memory. For example, the
above screenshot during the Florence OD task shows 4 CPUs at full speed and over
5GB of memory in use. Consider increasing the SWAP memory to 2 GB.

Checking the CPU temperature with vcgencmd measure_temp , showed that temperature can
go up to +80oC.

Fine-Tunning

As explored in this lab, Florence supports many tasks out of the box, including captioning, object
detection, OCR, and more. However, like other pre-trained foundational models, Florence-2
may need domain-specific knowledge. For example, it may need to improve with medical or
satellite imagery. In such cases, fine-tuning with a custom dataset is necessary. The Roboflow
tutorial, How to Fine-tune Florence-2 for Object Detection Tasks, shows how to fine-tune
Florence-2 on object detection datasets to improve model performance for our specific use
case.
Based on the above tutorial, it is possible to fine-tune the Florence-2 model to detect boxes
and wheels used in previous labs:

It is important to note that after fine-tuning, the model can still detect classes that don’t
belong to our custom dataset, like cats, dogs, grapes, etc, as seen before).
The complete fine-tunning project using a previously annotated dataset in Roboflow and
executed on CoLab can be found in the notebook:

• 30-Finetune_florence_2_on_detection_dataset_box_vs_wheel.ipynb

In another example, in the post, Fine-tuning Florence-2 - Microsoft’s Cutting-edge Vision


Language Models, the authors show an example of fine-tuning Florence on DocVQA. The authors
report that Florence 2 can perform visual question answering (VQA), but the released models
don’t include VQA capability.

564
Conclusion

Florence-2 offers a versatile and powerful approach to vision-language tasks at the edge,
providing performance that rivals larger, task-specific models, such as YOLO for object
detection, BERT/RoBERTa for text analysis, and specialized OCR models.
Thanks to its multi-modal transformer architecture, Florence-2 is more flexible than YOLO in
terms of the tasks it can handle. These include object detection, image captioning, and visual
grounding.
Unlike BERT, which focuses purely on language, Florence-2 integrates vision and language,
allowing it to excel in applications that require both modalities, such as image captioning and
visual grounding.
Moreover, while traditional OCR models such as Tesseract and EasyOCR are designed solely
for recognizing and extracting text from images, Florence-2’s OCR capabilities are part of a
broader framework that includes contextual understanding and visual-text alignment. This
makes it particularly useful for scenarios that require both reading text and interpreting its
context within images.
Overall, Florence-2 stands out for its ability to seamlessly integrate various vision-language
tasks into a unified model that is efficient enough to run on edge devices like the Raspberry Pi.
This makes it a compelling choice for developers and researchers exploring AI applications at
the edge.

Key Advantages of Florence-2

1. Unified Architecture
• Single model handles multiple vision tasks vs. specialized models (YOLO, BERT,
Tesseract)
• Eliminates the need for multiple model deployments and integrations
• Consistent API and interface across tasks
2. Performance Comparison
• Object Detection: Comparable to YOLOv8 (~37.5 mAP on COCO vs. YOLOv8’s
~39.7 mAP) despite being general-purpose
• Text Recognition: Handles multiple languages effectively like specialized OCR models
(Tesseract, EasyOCR)
• Language Understanding: Integrates BERT-like capabilities for text processing while
adding visual context
3. Resource Efficiency
• The Base model (232M parameters) achieves strong results despite smaller size

565
• Runs effectively on edge devices (Raspberry Pi)
• Single model deployment vs. multiple specialized models

Trade-offs

1. Performance vs. Specialized Models


• YOLO series may offer faster inference for pure object detection
• Specialized OCR models might handle complex document layouts better
• BERT/RoBERTa provide deeper language understanding for text-only tasks
2. Resource Requirements
• Higher latency on edge devices (15-200s depending on task)
• Requires careful memory management on Raspberry Pi
• It may need optimization for real-time applications
3. Deployment Considerations
• Initial setup is more complex than single-purpose models
• Requires understanding of multiple task types and prompts
• The learning curve for optimal prompt engineering

Best Use Cases

1. Resource-Constrained Environments
• Edge devices requiring multiple vision capabilities
• Systems with limited storage/deployment capacity
• Applications needing flexible vision processing
2. Multi-modal Applications
• Content moderation systems
• Accessibility tools
• Document analysis workflows
3. Rapid Prototyping
• Quick deployment of vision capabilities
• Testing multiple vision tasks without separate models
• Proof-of-concept development

566
Future Implications

Florence-2 represents a shift toward unified vision models that could eventually replace task-
specific architectures in many applications. While specialized models maintain advantages in
specific scenarios, the convenience and efficiency of unified models like Florence-2 make them
increasingly attractive for real-world deployments.
The lab demonstrates Florence-2’s viability on edge devices, suggesting future IoT, mobile
computing, and embedded systems applications where deploying multiple specialized models
would be impractical.

Resources

• 10-florence2_test.ipynb
• 20-florence_2.ipynb
• 30-Finetune_florence_2_on_detection_dataset_box_vs_wheel.ipynb

567
Audio and Vision AI Pipeline

Building Voice and Vision Interactive Edge AI Systems

Introduction

In this chapter, we extend our SLM and SVL capabilities by creating a complete audio processing
pipeline that transforms voice or image input into intelligent vocal responses. We will learn to
integrate Speech-to-Text (STT), Small Language (or Visual) Models, and Text-to-Speech (TTS)
technologies to build conversational AI systems that run entirely on Raspberry Pi hardware.
This chapter bridges the gap between our existing computer vision knowledge and multimodal
AI applications, demonstrating how different AI components work together in real-world edge
deployments.

568
The Audio to Audio AI Pipeline Architecture

We will understand how to architect and implement multimodal AI systems by building


a complete voice interaction pipeline. The goal is to gain practical experience with audio
processing on edge devices while learning to efficiently integrate multiple AI models within
resource constraints. Additionally, you will develop troubleshooting skills for complex AI
pipelines and understand the engineering trade-offs involved in edge audio processing.

Understanding Multimodal AI Systems

When we built computer vision systems earlier in the course, we processed visual data to
extract meaningful information. Audio AI systems follow a similar principle but work with
temporal audio signals instead of static images. The key insight is that speech processing
requires multiple specialized models working together, rather than a single end-to-end system.

Modern small models, such as Gemma 3n, can process audio directly and its prompt.
Today (September 2025), Gemma 3n can transcribe text from audio files using
Hugging Face Transformers, but it is not available with Ollama

Consider how humans process spoken language. We simultaneously parse the acoustic signal,
understand the linguistic content, reason about the meaning, and formulate responses. Our AI
pipeline mimics this process by breaking it into distinct, manageable components.

System Architecture Overview

Our comprehensive audio AI pipeline comprises four main components, connected in sequence.

569
[Microphone] → [STT Model] → [SLM] → [TTS Model] → [Speaker]
Audio Text Text Audio

Audio captured by the microphone is processed through a Speech-to-Text model, which converts
sound waves into text transcriptions. This text becomes input for our Small Language Model,
which generates intelligent responses. Finally, a Text-to-Speech system converts the written
response back into spoken audio.
Each component has specific requirements and limitations. The STT model must handle
various accents and noise conditions. The SLM needs sufficient context to generate coherent
responses. The TTS system must produce speech that sounds natural. Understanding these
individual requirements helps us optimize the overall system performance.

Edge AI Considerations

Running this pipeline on a Raspberry Pi presents unique challenges compared to cloud-


based solutions. We must carefully manage memory usage, processing time, and model sizes.
The benefit is complete local processing with no internet dependency and enhanced privacy
protection.
The choice of models becomes critical. We select Moonshine for STT because it’s specifically
optimized for edge devices. We utilize small language models, such as llama3.2:3b, for
reasonable performance on limited hardware. For TTS, we choose PIPER for its balance
between quality and computational efficiency.

Hardware Setup and Audio Capture

Audio Hardware Detection

Begin by identifying the audio capabilities of our system. The Raspberry Pi can work with
various audio input and output devices, but proper configuration is essential for reliable
operation.
Use the command arecord -l to list available recording devices. You should see output
showing your microphone’s card and device numbers. For USB microphones, this typically
appears as something like card 2: Microphone [USB Condenser Microphone], device 0:
USB Audio [USB Audio]. The critical information is the card number and device number,
which you’ll reference as hw:2,0 in the code.

570
Testing Basic Audio Functionality

Before writing Python code, we should verify that our audio setup works correctly at the
system level. Let’s record a short test file using:
arecord --device="plughw:2,0" --format=S16_LE --rate=16000 -c2 [Link]

571
Play back the recording with aplay [Link] to confirm that both capture and playback
work correctly (use [CTRL]+[C] to stop the recording or add a duration in seconds to the
command line).

The .WAV file can be played on another device (such as a computer) or on the
Raspberry Pi, as a speaker can be connected via USB or Bluetooth.

Python Audio Integration

Now, we should install the necessary audio processing dependencies.

source ~/ollama/bin/activate

The PyAudio library requires system-level audio libraries; therefore, install them using sudo
apt-get.

sudo apt-get update


sudo apt-get install libasound-dev libportaudio2 libportaudiocpp0 portaudio19-dev
sudo apt-get install python3-pyaudio

• libasound-dev covers ALSA development headers needed for audio libraries on Pi


• python3-pyaudio provides a prebuilt PyAudio package for most use cases

Let’s create a working directory: Documents/OLLAMA/SST and verify the USB device index,
with the below script (verify_usb_index.py:

import pyaudio
p = [Link]()
for ii in range(p.get_device_count()):
print(ii, p.get_device_info_by_index(ii).get('name'))

572
As a result, we should get:

0 USB Condenser Microphone: Audio (hw:2,0)


1 pulse
2 default

What confirms that the index is 2 (hw:2,0)

A lot of messages should appear. They are mostly ALSA and JACK warnings
about missing or undefined virtual/surround sound devices—they are common on
Raspberry Pi systems with minimal or headless sound configs and typically do
not impact basic USB microphone capture. If our USB Microphone appears as a
recording device (as it does: “hw:2,0”), we can safely ignore most of these unless
audio capture fails.

To clean the output, we can use:


python verify_usb_index.py 2>/dev/null

Test Audio Python Script

This script (test_audio_capture.py) records 10 seconds of mono audio at 16 kHz to


[Link] using the USB microphone.

import pyaudio
import wave

FORMAT = pyaudio.paInt16
CHANNELS = 1
RATE = 16000 # 16 kHz
CHUNK = 1024
RECORD_SECONDS = 10
DEVICE_INDEX = 2 # replace this with your detected USB mic's index
WAVE_OUTPUT_FILENAME = "[Link]"

573
audio = [Link]()

stream = [Link](format=FORMAT, channels=CHANNELS,


rate=RATE, input=True, input_device_index=DEVICE_INDEX,
frames_per_buffer=CHUNK)

print("Recording...")
frames = []

for i in range(0, int(RATE / CHUNK * RECORD_SECONDS)):


data = [Link](CHUNK, exception_on_overflow=False)
[Link](data)

print("Finished recording.")

stream.stop_stream()
[Link]()
[Link]()

wf = [Link](WAVE_OUTPUT_FILENAME, 'wb')
[Link](CHANNELS)
[Link](audio.get_sample_size(FORMAT))
[Link](RATE)
[Link](b''.join(frames))
[Link]()

Understanding the audio configuration parameters helps prevent common problems. We use
16-bit PCM format (pyaudio.paInt16) because it provides good quality while remaining
computationally efficient. The 16kHz sampling rate balances audio quality with processing
requirements - most speech recognition models expect this rate.
The buffer size (CHUNK = 1024) affects latency and reliability. Smaller buffers reduce latency
but may cause audio dropouts on busy systems. Larger buffers increase latency but provide
more stable recording.
Let’s Playback to verify if we get it correctly:

574
Your browser does not support the audio element.

Speech Recognition (SST) with Moonshine

Why Moonshine for Edge Deployment

Traditional speech recognition systems, such as OpenAI’s Whisper, are highly accurate but
require substantial computational resources. Moonshine is specifically designed for edge devices,
using optimized model architectures and quantization techniques to achieve good performance
on resource-constrained hardware.
The ONNX (Open Neural Network Exchange) version of Moonshine provides additional
optimization benefits. ONNX Runtime includes hardware-specific optimizations that can
significantly improve inference speed on ARM processors, such as those found in Raspberry Pi
devices.

Model Selection Strategy

Moonshine offers different model sizes with clear trade-offs between accuracy and computational
requirements. The “tiny” model processes audio quickly but may struggle with difficult audio
conditions. The “base” model provides better accuracy but requires more processing time and
memory.
For initial development, we should start with the tiny model to ensure that the pipeline works
correctly. Once the complete system is functional, we can experiment with larger models to
find the optimal balance for our specific use case and hardware capabilities.

575
Implementation and Preprocessing

Install Moonshine with

pip install useful-moonshine-onnx@git+[Link]


ai/[Link]#subdirectory=moonshine-onnx`

This specific installation method ensures compatibility with the ONNX runtime optimizations.
Let’s run the test script below (transcription_test.py):

import moonshine_onnx
text = moonshine_onnx.transcribe('[Link]', 'moonshine/tiny')
print(text[0])

As a result, we will get the corresponding text, which was recorded before:

Empty or incorrect transcriptions are common issues in speech recognition systems.


These failures can result from background noise, insufficient volume, unclear speech,
or mismatches in audio format. Implementing robust error handling prevents these
issues from crashing your entire pipeline.

SLM Integration and Response Generation

Connecting STT Output to SLMs

The text output from your speech recognition system becomes input for your Small Language
Model. However, the characteristics of spoken language differ significantly from written text,
especially when filtered through speech recognition systems.
Spoken language tends to be more informal, may contain false starts and repetitions, and might
include transcription errors. Our SLM integration should account for these characteristics. On
a final implementation, we should consider preprocessing the STT output to clean up obvious
transcription errors or providing context to the SLM about the voice interaction nature of the
input.

576
Optimizing Prompts for Voice Interaction

Voice-based interactions have different expectations than text-based chats. Responses should
be concise since users must listen to the entire output. Avoid complex formatting or long lists
that work well in text but become cumbersome when spoken aloud.
We should design our system prompts to encourage responses appropriate for voice interaction.
For example, “Provide a brief, conversational response suitable for speaking aloud” can help
guide the SLM toward more appropriate output formatting.
Unlike single-query text interactions, voice conversations often involve multiple exchanges.
Implementing conversation context memory significantly enhances the user experience. However,
context management on edge devices requires careful consideration of memory usage.
Consider implementing a sliding window approach, where you maintain the last few exchanges
in memory but discard older context to prevent memory exhaustion, balancing context length
with available system resources.
Let’s create a function to handle this. For test, run slm_test.py:

import ollama

def generate_voice_response(user_input, model="llama3.2:3b"):


"""
Generate a response optimized for voice interaction
"""

# Context-setting prompt that guides the model's behavior


system_context = """
You are a helpful AI assistant designed for voice interactions.
Your responses will be converted to speech and spoken aloud to the user.

Guidelines for your responses:


- Keep responses conversational and concise (ideally under 50 words)
- Avoid complex formatting, lists, or visual elements
- Speak naturally, as if having a friendly conversation
- If the user's input seems unclear, ask for clarification politely
- Provide direct answers rather than lengthy explanations unless specifically
requested
"""

# Combine system context with user input


full_prompt = f"{system_context}\n\nUser said: {user_input}\n\nResponse:"

577
response = [Link](
model=model,
prompt=full_prompt
)

return response['response']

# Answering the user question:


user_input = "What is the capital of Malawi?"
response = generate_voice_response(user_input)
print (response)

Text-to-Speech (TTS) with PIPER

Text-to-speech systems face different challenges than speech recognition systems. While STT
must handle various input conditions, TTS must generate consistent, natural-sounding output
across diverse text inputs. The quality of TTS has a significant impact on the user experience
in voice interaction systems.
PIPER provides an excellent balance between voice quality and computational efficiency for
edge deployments. Unlike cloud-based TTS services, PIPER runs entirely locally, ensuring
privacy and eliminating network dependencies.

Voice Model Selection and Installation

PIPER offers various voice models with different characteristics. The “low”, “medium”, and
“high” quality designations primarily refer to model size and computational requirements
rather than dramatic quality differences. For most applications, the low-quality models provide
acceptable voice output while running efficiently on Raspberry Pi hardware.

578
Install PIPER with pip install piper-tts, create a voices directory:

pip install piper-tts


mkdir -p voices

Download voice models from the Hugging Face repository. Each voice model requires both the
model file (.onnx) and a configuration file (.json). The configuration file contains model-specific
parameters essential for generating proper audio.
We should download both files for our chosen voice; for example, the English female “lessac”
voice provides clear, natural speech suitable for most applications.

wget -O voices/en_US-[Link] [Link]


voices/resolve/v1.0.0/en/en_US/lessac/low/en_US-[Link]

wget -O voices/en_US-[Link] [Link]


voices/resolve/v1.0.0/en/en_US/lessac/low/en_US-[Link]

Let’s run the code below (tts_test.py) for testing:

import subprocess
import os

def text_to_speech_piper(text, output_file="piper_output.wav"):


"""
Convert text to speech using PIPER and save to WAV file

Args:
text (str): Text to convert to speech
output_file (str): Output WAV file path
"""
# Path to your voice model
model_path = "voices/en_US-[Link]"

# Check if model exists


if not [Link](model_path):
print(f"Error: Model file not found at {model_path}")
return False

try:
# Run PIPER command
process = [Link](

579
['piper', '--model', model_path, '--output_file', output_file],
stdin=[Link],
stdout=[Link],
stderr=[Link],
text=True
)

# Send text to PIPER


stdout, stderr = [Link](input=text)

if [Link] == 0:
print(f"\nSpeech generated successfully: {output_file}")
return True
else:
print(f"Error: {stderr}")
return False

except Exception as e:
print(f"Error running PIPER: {e}")
return False

# converting text to sound:


txt = "Lilongwe is the capital of Malawi. Would you like to know more about \
Lilongwe or Malawi in general?"

if text_to_speech_piper(txt):
print("You can now play the file with: aplay piper_output.wav")
else:
print("Failed to generate speech")

Runing the script, a piper_output.wav file will be generated, which is the text converted into
speech.
Your browser does not support the audio element.
To listen to the sound, we can run: aplay piper_output.wav.

580
Handling Long Text and Special Cases

TTS systems may struggle with very long input text or special characters. Implement text
preprocessing to handle these cases gracefully. Break long responses into shorter segments,
handle abbreviations and numbers appropriately, and filter out problematic characters that
might cause TTS failures.
Consider implementing text chunking for responses longer than a reasonable speaking length.
This prevents both TTS processing issues and user fatigue from overly long audio responses.

Pipeline Integration and Optimization

Building the Complete System

Integrating all components requires careful attention to error handling and resource management.
Each stage of the pipeline can fail independently, and robust systems must handle these failures
gracefully rather than crashing.
Design our integration with modularity in mind. Test each component independently before
combining them. This approach simplifies debugging and allows you to optimize individual
components separately.
We should also implement proper logging throughout our pipeline. When complex systems
fail, detailed logs help identify whether the issue occurs in audio capture, speech recognition,
language model processing, text-to-speech conversion, or audio playback.

581
Performance Optimization Strategies

Measure the timing of each pipeline component to identify bottlenecks. Typically, the SLM
inference takes the longest time, followed by TTS generation. Understanding these timing
characteristics helps prioritize optimization efforts.
Consider implementing concurrent processing where possible. For example, you might begin
TTS processing for the first part of an extended response while the SLM is still generating the
remainder. However, be cautious about memory usage when implementing parallel processing
on resource-constrained devices.

Memory Management Considerations

Edge devices have limited RAM, and loading multiple large models simultaneously can cause
memory pressure. Implement strategies to manage memory efficiently, such as loading models
only when needed or using model swapping for infrequently used components.
Monitor system memory usage during operation and implement safeguards to prevent memory
exhaustion. Consider implementing graceful degradation where your system switches to smaller,
more efficient models if memory becomes constrained.

Designing Resilient Systems

Production-quality voice interaction systems must handle various failure modes gracefully.
Network interruptions, hardware disconnections, model loading failures, and unexpected input
conditions should not cause your system to crash.
Implement comprehensive error handling at each pipeline stage. When speech recognition
produces empty output, provide the user with meaningful feedback rather than processing
empty strings. When TTS fails, consider falling back to text display or simplified audio
feedback.
Design user feedback mechanisms that work within your voice interaction paradigm. Audio
beeps, LED indicators, or simple voice messages can communicate system status without
requiring visual displays.

Debugging Complex Pipelines

Multi-stage systems present unique debugging challenges. When the overall system fails,
identifying the specific failure point requires systematic testing approaches.

582
Implement test modes that allow you to inject known inputs at each pipeline stage. This
capability enables you to isolate problems to specific components rather than repeatedly testing
the entire system.
Create diagnostic outputs that help understand system behavior. For example, displaying
transcription confidence scores, SLM response times, or TTS processing status helps identify
performance issues or quality problems.

The full Python Script

Considering the previous points, let’s assemble all the essential components that work together:
audio capture, transcription, language model processing, and text-to-speech. We should combine
these into a complete voice pipeline that flows naturally from one step to the next.
The key insight here is that each of our developed scripts represents a stage in what’s called an
“audio processing pipeline”. Let’s walk through how we can connect these pieces.
Understanding the Pipeline Architecture
The pipeline follows a logical sequence: the voice becomes audio data, that audio is converted
into text, the text is processed into an AI response, the response is converted into speech audio,
and finally, that speech audio is transformed into sound that we can hear.
We should have a run_voice_pipeline() function in addition to the previous ones that acts
as a coordinator, ensuring each step completes successfully before proceeding to the next. If
any step fails, the entire pipeline stops gracefully rather than trying to continue with missing
data.
Key Integration Points
We should connect the scripts by ensuring the output of each function becomes the input
for the following function. For example, record_audio() creates “user_input.wav”, which
transcribe_audio() reads to produce text, which generate_response() processes to develop
an AI response, and so on.
The error handling at each step ensures that if our microphone isn’t working, or if the AI model
is busy, or if the voice model files are missing, we get clear feedback about what went wrong
rather than mysterious crashes.
Optimizations for Raspberry Pi
The Raspberry Pi has limited resources compared to a desktop computer, so we should include
several optimizations. The cleanup_temp_files() function prevents our storage from filling
up with temporary audio files. The audio configuration uses 16kHz sampling (which matches
Moonshine’s expectations) rather than CD-quality 44kHz, reducing processing overhead.

583
The continuous assistant mode includes a manual trigger (pressing Enter) rather than voice
activation detection, which saves CPU cycles that would otherwise be spent constantly
monitoring audio input.

On a final product, we could include, for example, a KWS (Keyword Spotting)


function, based on a TinyML device, where only when a trigger word is spoken, the
record_audio() function starts to work.

Understanding Voice Activity Detection


One of the main improvements from the single scripts stacked and tested separately is what
audio engineers call “voice activity detection” (VAD). When we speak to someone face-to-face,
we don’t announce, “I’m going to talk for exactly 10 seconds now.” Instead, we say our thoughts,
pause naturally, and the listener intuitively knows when you’ve finished.
Our new record_audio_with_silence_detection() function mimics this natural process by
continuously analyzing the audio signal’s amplitude—essentially measuring how “loud” each
tiny slice of audio is. When the amplitude stays below a threshold for 2 seconds, for example,
the system intelligently concludes that we have finished speaking.
The Technical Magic Behind Silence Detection
The recording process now works like a sophisticated audio surveillance system. Every fraction
of a second, it captures a small chunk of audio data (determined by your CHUNK size of 1024
samples) and immediately analyzes it. Using Python’s struct module, it converts the raw byte
data into numerical values representing sound pressure levels.
The crucial calculation happens in this line: volume = max(chunk_data) / 32768.0. This
finds the loudest moment in that tiny audio slice and normalizes it to a scale from 0 to 1, where
0 represents complete silence and 1 represents the maximum possible volume your microphone
can capture.
Calibrating the Sensitivity
The silence_threshold=0.01 parameter is our sensitivity control knob: too sensitive (closer
to 0.00) and it might stop recording when you pause to think; not sensitive enough (closer
to 0.1) and it might keep recording through long periods of quiet background noise.
For a typical indoor Raspberry Pi setup, 0.01 strikes a good balance. It’s sensitive enough to
detect when we have stopped talking, but robust enough to ignore minor background sounds,
such as air conditioning or distant traffic.

We should experiment with this value based on the specific environment.

Testing and Deployment Strategy


Download the complete script from GitHub: voice_ai_pipeline.py

584
We can start by testing individual components using detect_audio_devices() first to confirm
our USB microphone is still at index 2. Then, we can run run_voice_pipeline() with a
simple question to verify the complete flow works.
Once we are confident in single interactions, we can use continuous_voice_assistant()
for extended conversations. This mode lets us have back-and-forth exchanges with your AI
assistant, making it feel more like a natural conversation partner.

If you want to experiment with different SLM models, we need to change the model
parameter in generate_response().

The test with sound can be followed in the video: [Link]

The Image to Audio AI Pipeline Architecture

Voice interaction systems have significant potential for educational applications and accessibility
improvements by designing interfaces that adapt to different user needs and capabilities. By
integrating small visual models, such as Moondream, with existing TTS pipelines, we can create
multimodal assistants that describe images for people with visual impairments, converting
visual content into detailed spoken descriptions of scenes, objects, and spatial relationships.

585
To use the camera, we must ensure that the NumPy version is compatible with the Raspberry
Pi system (1.24.2). If this version was changed (what it should be due to the Moonshine
installation), revert it to the 1.24.2 version.

pip uninstall numpy


pip install 'numpy==1.24.2'

Now, let’s apply what we have explored in previous chapters, along with what we learn in this
one, to create a simple code that converts an image captured by a camera into its corresponding
spoken caption by the speaker.
The code is straightforward and is intended solely to test the solution’s potential.

1. The capture_image(IMG_PATH) function captures a photo and saves it to the specified


path (IMG_PATH).
2. The caption = image_description(IMG_PATH, MODEL) function will describe the im-
age using MoonDream, a powerful yet compact visual language model (VLM). The descrip-
tion in a text format will be saved in the variable caption.
3. The text_to_speech_piper(caption) function will create a .WAV file from the caption.
4. And finally, the play_audio() function will play the .WAV file generated by PIPER.

Here is the complete Python script (img_caption_speech.py):

import os
import time
import subprocess
import ollama
from picamera2 import Picamera2

586
def capture_image(img_path):
# Initialize camera
picam2 = Picamera2()
[Link]()

# Wait for camera to warm up


[Link](2)

# Capture image
picam2.capture_file(img_path)
print("\n==> Image captured: "+img_path)

# Stop camera
[Link]()
[Link]()

def image_description(img_path, model):


print ("\n==> WAIT, SVL Model working ...")
with open(img_path, 'rb') as file:
response = [Link](
model=model,
messages=[
{
'role': 'user',
'content': '''return the description of the image''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,
}
)
return response['message']['content']

def text_to_speech_piper(text, output_file="assistant_response.wav"):

# Path to your voice model


model_path = "voices/en_US-[Link]"

# Check if model exists

587
if not [Link](model_path):
print(f"Error: Model file not found at {model_path}")
return False

try:
# Run PIPER command
process = [Link](
['piper', '--model', model_path, '--output_file', output_file],
stdin=[Link],
stdout=[Link],
stderr=[Link],
text=True
)

# Send text to PIPER


stdout, stderr = [Link](input=text)

if [Link] == 0:
print(f"\nSpeech generated successfully: {output_file}")
return True
else:
print(f"Error: {stderr}")
return False

except Exception as e:
print(f"Error running PIPER: {e}")
return False

def play_audio(filename="assistant_response.wav"):
try:
# Use aplay to play the audio file
result = [Link](['aplay', filename],
capture_output=True,
text=True)

if [Link] == 0:
print("\nAudio playback completed")
return True
else:
print(f"\nPlayback error: {[Link]}")
return False

588
except Exception as e:
print(f"\nError playing audio: {e}")
return False

# Example usage and testing functions


if __name__ == "__main__":

print("\n============= Image to Speech AI Pipeline Ready ===========")


print("Press [Enter] to capture an image and voice caption it ...")

# Step 1: Wait for user to initiate recording


input("Press Enter to start ...")

IMG_PATH = "/home/mjrovai/Documents/OLLAMA/SST/capt_image.jpg"
MODEL = "moondream:latest"

capture_image(IMG_PATH)
caption = image_description(IMG_PATH, MODEL)
print ("\n==> AI Response:", caption)

text_to_speech_piper(caption)
play_audio()

And here we can see (and listen) to the result:

589
Your browser does not support the audio element.

Troubleshooting

To work simultaneously with STT and the camera in the same environment, we should ensure
that all core scientific packages are version-aligned for the Raspberry Pi environment. To use
NumPy 1.24.2 on our Raspberry Pi, we must ensure that both our SciPy and Librosa versions
are compatible with that NumPy version.
So, to avoid an eventual [Link] import error, and other possible incompatibilities,
we should downgrade scipy (and librosa) to versions that support numpy 1.24.2. According to
official compatibility tables, scipy 1.11.x works with numpy 1.24.x. Librosa versions released
after 0.9.0 also provide better support for recent NumPy releases. However, some older versions
of Librosa are not compatible with NumPy 1.24.2 due to deprecated NumPy attributes. So, to
fix it, we should run the lines below:

pip uninstall scipy librosa


pip install 'scipy>=1.11,<1.12' 'librosa>=0.10,<0.11'

Conclusion

This chapter introduced us to multimodal AI system development through an audio and vision
processing pipeline. We learned to integrate speech recognition, language models, and speech
synthesis into a cohesive system that runs efficiently on edge hardware.
We explored how to architect systems with multiple AI components, handle complex error
conditions, and optimize performance within resource constraints.
These skills are essential for the advanced topics in upcoming chapters, including RAG systems,
agent architectures, and the integration of physical computing. The system thinking approach
used here will be essential for future AI engineering work.
We should consider how the voice interaction capabilities we built might enhance other AI
systems. Many applications benefit from voice interfaces, and the foundation established
here can be adapted and extended for various use cases. We can, for example, transform our
audio pipeline into a smart home assistant by integrating physical computing elements. Voice
commands can trigger LED indicators, read sensor values, or control actuators connected to
our Raspberry Pi GPIO pins. Voice queries about environmental conditions can trigger sensor
readings, while voice commands can control connected devices.

590
This chapter extends our SLM and VLM work by adding input and/or output modalities beyond
text and images. The same language models we used previously now process voice-derived
input and generate responses for speech synthesis.
Consider how RAG systems from later chapters might integrate with voice interactions. Voice
queries could trigger document retrieval, with synthesized responses incorporating retrieved
information.

Resources

Python Scripts

591
592
Physical Computing with Raspberry Pi

From Sensors to Smart Analysis with Small Language Models

593
Introduction

Physical computing creates interactive systems that sense and respond to the analog world.
While this field has traditionally focused on direct sensor readings and programmed responses,
we’re entering an exciting new era where Large Language Models (LLMs) can add sophisticated
decision-making and natural language interaction to physical computing projects.
In the Small Language Models (SLM) chapter, we learned how to run an LLM (or, more
precisely, an SLM) on a Single Board Computer (SBC) such as the Raspberry Pi. In this
chapter, we will go through the process of setting up a Raspberry Pi for physical computing,
with an eye toward future AI integration. We’ll cover:

• Setting up the Raspberry Pi for physical computing


• Working with essential sensors and actuators
• Understanding GPIO (General Purpose Input/Output) programming
• Establishing a foundation for integrating LLMs with physical devices
• Creating interactive systems that can respond to both sensor data and natural language
commands

We will also use a Jupyter notebook (programmed in Python) to interact with sensors and
actuators—an important and necessary first step toward the goal of integrating the Raspi with
an SLM.
The combination of Raspberry Pi’s versatility and the power of SLMs opens up exciting
possibilities for creating more intelligent and responsive physical computing systems.
The diagram below gives us an overview of the project:

594
Prerequisites

• Raspberry Pi (model 4 or 5)
• DHT22 Temperature and Relative Humidity Sensor
• BMP280 Barometric Pressure, Temperature and Altitude Sensor
• Colored LEDs (3x)
• Push Button (1x)
• Resistor 4K7 ohm (2x)
• Resistor 220 or 330 ohm (3x)

Accessing the GPIOs

The Raspberry Pi’s GPIO (General Purpose Input/Output) pins allow us to connect electronic
components and control them with Python code. This opens up endless possibilities for creating
interactive projects, home automation systems, robotics, and more.
This chapter covers the modern GPIO Zero library for interactions with buttons and LEDs.
� IMPORTANT: [Link] does NOT support Raspberry Pi 5!

595
With the Raspberry Pi 5, we must use GPIO Zero or the newer lgpio library.
[Link] only works on Pi models 1-4, the Pi Zero, Pi Zero 2, and the Pi Zero
2W.

The Raspberry Pi has 40 pins on its header, but not all of them are GPIO pins. Some provide
power (3.3V and 5V), others are ground pins, and the rest are programmable GPIO pins.

Pin Numbering Systems

There are two ways to reference GPIO pins:

• BCM (Broadcom): Uses the GPIO number (e.g., GPIO17, GPIO27)


• BOARD: Uses the physical pin number on the header (e.g., Pin 11, Pin 13)

In this chapter, we’ll use BCM numbering as it’s more commonly used in Python
programming.

Safety First

Before connecting any components, please read these important safety guidelines:

• Never connect 5V directly to GPIO pins - GPIO pins are 3.3V tolerant only

596
• Always use current-limiting resistors with LEDs - Without them, you risk damaging
the LED or GPIO pin
• Double-check your connections before powering on
• Disconnect power when making circuit changes
• Respect polarity - LEDs for example, have positive (long leg) and negative (short leg)
sides

GPIO Zero Library

A modern, high-level library that makes GPIO programming much simpler and more intuitive.
It uses object-oriented programming and includes built-in features like automatic pin cleanup,
device abstraction, and event detection.

It is essential to note that the GPIO Zero Library uses Broadcom (BCM) pin
numbering for GPIO pins, rather than physical (board) numbering. Any pin
marked “GPIO ” in the previous diagram can be used as a PIN. For example, if an
LED were attached to GPIO13, we would specify the PIN as 13 rather than 33 (the
physical one).

It was created by Ben Nuttall of the Raspberry Pi Foundation, Dave Jones, and other
contributors (GitHub).
Advantages of GPIO Zero:

• Simpler, more readable code


• Object-oriented approach (LED, Button, etc.)
• Built-in device patterns and behaviors
• Automatic cleanup on program exit
• Better for beginners and rapid prototyping
• Includes composite devices (robots, traffic lights, etc.)

Installation:

• GPIO Zero comes pre-installed on Raspberry Pi OS!

597
“Hello World”: Blinking an LED

Let’s start with the classic ‘Hello World’ of physical computing - making an LED blink!
To connect our RPi to the world, let’s first connect:

• Physical Pin 6 (GND) to GND Breadboard Power Grid (Blue -), using a black jumper
• Physical Pin 1 (3.3V) to +VCC Breadboard Power Grid (Red +), using a red jumper

Now, let’s connect an LED (red) using the physical pin 33 (GPIO13) connected to the LED
cathode (longer LED leg). Connect the LED anode to the breadboard GND using a 220 ohms
resistor to reduce the current drawn from the Raspberry Pi, as shown below:

Understanding the LED Circuit

An LED (Light-Emitting Diode) requires current to flow through it to produce light. However,
without a resistor, too much current can flow, damaging the LED or your Raspberry Pi. We
use a resistor to limit the current to a safe level.

598
Why 220Ω or 330Ω Resistors?
These values limit the current to approximately 10-15mA, which is safe for most standard
LEDs and GPIO pins. The exact value isn’t critical—anything from 220 Ω to 1 kΩ will work
fine.

Testing with GPIO Zero

We can use the built-in Python interpreter to test the LED. In the terminal, enter python, and
once in the interpreter, enter with the commands below:

python
>>> from gpiozero import LED
>>> led = LED(13)
>>> [Link]()
>>> [Link]

599
On the Raspberry, start at home and go to Documents.

cd Documents

Create a directory to save the scripts and install the libraries. Move to there:

mkdir GPIO
cd GPIO

We can use any text editor (such as Nano) to create and run the script. Save the file, for
example, as led_test.py, and then execute it using the terminal:

python led_test.py

Now, let’s blink the LED (the actual “Hello world”) when talking about physical computing.
To do that, we must also import another library: time. We need it to define how long the LED
will be ON and OFF. In the case below, the LED will blink every 1 second.

from gpiozero import LED


from time import sleep
led = LED(13)
while True:
[Link]()
sleep(1)
[Link]()
sleep(1)

600
We can use any text editor (such as Nano) to create and run the script. Save the file, for
example, as [Link], and then execute it using the terminal:

python [Link]

Alternatively, we can reduce the blink code as below:

from gpiozero import LED


from signal import pause
led = LED(13)
[Link]() # 1 second by default
# [Link](on_time=0.5, off_time=0.5) # Fast blink
pause()

Installing all LEDs (the “actuators”)

The LEDs can be used as “actuators”; depending on the condition of a code running on our Pi,
we can command one of the LEDs to fire! We will install two more LEDs, in addition to the
red one already installed. Follow the diagram and install the yellow (on GPIO 19 ) and the
green (on GPIO 26).

601
For testing we can run a similar code as the used with the single red led, changing the pin
accordantly, for example.

from gpiozero import LED


from time import sleep

ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)

[Link]()
[Link]()
[Link]()

[Link]()
[Link]()
[Link]()

602
sleep(5)

[Link]()
[Link]()
[Link]()

Remember that instead of LEDs, we could have relays, motors, etc.

Sensors Installation and setup

In this section, we will setup the Raspberry Pi to capture data from several different sensors:
Sensors and Communication type:

• Button (Command via a Push-Button) ==> Digital direct connection


• DHT22 (Temperature and Humidity) ==> Digital communication
• BMP280 (Temperature and Pressure) ==> I2C Protocol

Button

Now, let’s learn how to read input from a button. This allows your Raspberry Pi to respond to
physical interactions!

603
Understanding Pull-Up and Pull-Down Resistors

When a button is not pressed, the GPIO pin is ‘floating’ - it’s not connected to anything and
can read random values. We use pull-up or pull-down resistors to set the pin’s default state.

• Pull-Down: Default state is LOW (0V), becomes HIGH when button pressed
• Pull-Up: Default state is HIGH (3.3V), becomes LOW when button pressed

The Raspberry Pi has internal pull-up and pull-down resistors!

Figure 16: img

The simplest way to run an external command is with a push button, and the GPIO Zero
Library makes it easy to include in the project. We do not need to think about Pull-up or
Pull-down resistors, etc. In terms of HW, the only thing to do is to connect one leg of our
push-button to any one of the Raspi GPIOs and the other one to GND, as shown in the
diagram:

604
• Push-Button leg1 to GPIO 20
• Push-Button leg2 to GND

A simple code for reading the button can be:

from gpiozero import Button


button = Button(20)
while True:
if button.is_pressed:
print("Button is pressed")
else:
print("Button is not pressed")

605
On a Raspberry Pi Zero 2W, for example, we could use the [Link], whose code is:

import [Link] as GPIO


import time

[Link]([Link])
[Link](False)

BUTTON_PIN = 20

# Set up with internal pull-up


[Link](BUTTON_PIN, [Link], pull_up_down=GPIO.PUD_UP)

try:
while True:
if [Link](BUTTON_PIN) == [Link]:
print('Button pressed!')
else:
print('Button not pressed')
[Link](0.1)

except KeyboardInterrupt:
[Link]()

606
Installing Adafruit CircuitPython

The GPIO Zero library is an excellent hardware interfacing library for Raspberry Pi. It’s great
for digital in/out, analog inputs, servos, basic sensors, etc. However, it doesn’t cover SPI/I2C
sensors or drivers. By using CircuitPython via adafruit_blinka, we can take advantage of all
of the drivers and example code developed by Adafruit!

Note that we will keep using GPIO Zero for pins, buttons, and LEDs.

Enable Interfaces
Run these commands to enable the various interfaces such as I2C and SPI:

sudo raspi-config nonint do_i2c 0


sudo raspi-config nonint do_spi 0
sudo raspi-config nonint do_serial_hw 0
sudo raspi-config nonint disable_raspi_config_at_boot 0

Install required dependencies

sudo apt-get install -y python3-libgpiod i2c-tools libgpiod-dev

Install Blinka
Let’s enter the Ollama environment, alheady created to install Blinka:

source ~/ollama/bin/activate

Install the library

pip install adafruit-blinka

Check I2C and SPI


The script will automatically enable I2C and SPI. You can run the following command to
verify:

ls /dev/i2c* /dev/spi*

607
Verify Blinka Version

python3 -m pip show adafruit-blinka

Blinka Installation Test


Create a new file called blinka_test.py with nano or your favorite text editor and put the
following in:

import board
import digitalio
import busio

print("Hello, blinka!")

# Try to create a Digital input

608
pin = [Link](board.D4)
print("Digital IO ok!")

# Try to create an I2C device


i2c = busio.I2C([Link], [Link])
print("I2C ok!")

# Try to create an SPI device


spi = [Link]([Link], [Link], [Link])
print("SPI ok!")

print("done!")

Save it and run it at the command line:

python blinka_test.py

DHT22 - Temperature & Humidity Sensor

The first sensor to be installed will be the DHT22 for capturing air temperature and relative
humidity data.
Overview
The low-cost DHT temperature and humidity sensors are elementary and slow, but great
for logging basic data. They consist of a capacitive humidity sensor and a thermistor. A
bare chip inside performs the analog-to-digital conversion and outputs a digital signal con-
taining the temperature and humidity. The digital signal is relatively easy to read using any
microcontroller.

609
DHT22 Main characteristics:

• Suitable for 0-100% humidity readings with 2-5% accuracy


• Suitable for -40 to 125°C temperature readings ±0.5°C accuracy
• No more than 0.5 Hz sampling rate (once every 2 seconds)
• Low cost
• 3 to 5V power and I/O
• 2.5mA max current use during conversion (while requesting data)
• Body size 15.1mm x 25mm x 7.7mm
• 4 pins with 0.1” spacing

Once we use the sensor at distances less than 20m, a 4K7 ohm resistor should be connected
between the Data and VCC pins. The DHT22 output data pin will be connected to Raspberry
GPIO 16. Check the electrical diagram, connecting the sensor to RPi pins as below:

• Pin 1 - Vcc ==> 3.3V


• Pin 2 - Data ==> GPIO 16
• Pin 3 - Not Connect
• Pin 4 - Gnd ==> Gnd

Do not forget to Install the 4K7 ohm resistor between the VCC and Data pins.

610
Once the sensor is connected, we must install its library on our Raspberry Pi. First, we
should install the Adafruit CircuitPython library, which we have already done, and the
Adafruit_CircuitPython_DHT.

pip install adafruit-circuitpython-dht

Create a new Python script as below and name it, for example, dht_test.py:

import time
import board
import adafruit_dht
dhtDevice = adafruit_dht.DHT22(board.D16)

611
while True:
try:
# Print the values to the serial port
temperature_c = [Link]
temperature_f = temperature_c * (9 / 5) + 32
humidity = [Link]
print(
"Temp: {:.1f} F / {:.1f} C Humidity: {}% ".format(
temperature_f, temperature_c, humidity
)
)
[Link](2.0)

except RuntimeError as error:


# Errors happen fairly often, DHT's are hard to read,
# just keep going
print([Link][0])
[Link](2.0)
continue
except Exception as error:
[Link]()
raise error

Placing a finger on the sensor, we can see that both temperature and humidity begin to rise.

612
613
Addendum: DHT22 on Raspberry Pi 5 with Debian Trixie

Exclamation Raspberry Pi 5 + Debian Trixie: adafruit_dht does not work

The adafruit-circuitpython-dht library fails silently or raises GPIO busy on Rasp-


berry Pi 5 running Debian Trixie (Debian 13). This is not a wiring issue — it is a
fundamental incompatibility between the library’s timing backend and the Pi 5 hardware.
The fix described below uses the Linux kernel’s built-in DHT driver instead.

Why the Adafruit library breaks on Pi 5

The DHT22 protocol is timing-sensitive: it communicates over a single wire using pulses as
short as 26 µs to encode each bit. Reading those pulses correctly requires a GPIO backend
with microsecond-level precision.
The adafruit_dht library offers two backends, both broken on Pi 5 + Trixie:

Backend Why it fails


use_pulseio=True (default) Relies on [Link], which does not support
the Pi 5’s RP1 GPIO chip
use_pulseio=False Falls back to pure-Python bit-banging via
lgpio; the Pi 5’s GPIO travels over PCIe to
the RP1 chip, adding unpredictable latency
that corrupts pulse timing

The pigpio daemon — which was the traditional workaround — was removed from Debian
Trixie’s official repositories. Compiling it from source is possible but fragile, and its DMA-based
approach does not work with the RP1 chip anyway.

The fix: use the kernel dht11 IIO driver

The Linux kernel ships a dht11 driver that works with both DHT11 and DHT22 sensors.
It runs in kernel space using GPIO interrupts, capturing edge transitions with nanosecond
timestamps — the only approach reliable enough on Pi 5.
Step 1 — Add the overlay to /boot/firmware/[Link]

sudo nano /boot/firmware/[Link]

Add the following line, replacing 16 with whichever GPIO pin your sensor data line uses:

614
dtoverlay=dht11,gpiopin=16

Save and reboot:

sudo reboot

Step 2 — Verify the driver loaded


After rebooting, confirm the IIO device is available:

cat /sys/bus/iio/devices/iio:device0/name

You should see output like dht11@16. Then do a quick read to confirm the sensor responds:

cat /sys/bus/iio/devices/iio:device0/in_temp_input
cat /sys/bus/iio/devices/iio:device0/in_humidityrelative_input

Values are returned in millidegrees and millipercent. A reading of 22900 means 22.9°C; 55000
means 55.0% RH.

INFO About the dht11@N device name

The number after @ in dht11@10 or dht11@16 is the internal device-tree unit address,
which does not always match the BCM GPIO number directly on the Pi 5 / RP1 chip.
What matters is that the device exists and returns readings — ignore the suffix.

Step 3 — Replace adafruit_dht with sysfs reads


Since the kernel driver owns the GPIO pin, any attempt by user-space code (including
adafruit_dht) to claim that pin will raise [Link]: 'GPIO busy'. The replacement is
straightforward: read the values directly from the IIO sysfs interface.
Save the following as dht_test_trixie.py (replacing the original Adafruit version):

import time

DHT_IIO = "/sys/bus/iio/devices/iio:device0/"

def read_dht22():
try:
temp_c = int(open(DHT_IIO + "in_temp_input").read()) / 1000.0
humidity = int(open(DHT_IIO + "in_humidityrelative_input").read()) / 1000.0

615
return temp_c, humidity
except Exception as e:
print(f"Read error: {e}")
return None, None

while True:
temperature_c, humidity = read_dht22()
if temperature_c is not None:
temperature_f = temperature_c * (9 / 5) + 32
print(
"Temp: {:.1f} F / {:.1f} C Humidity: {:.1f}% ".format(
temperature_f, temperature_c, humidity
)
)
[Link](3.0)

Run it:

python dht_test_trixie.py

Expected output:

Temp: 70.2 F / 21.2 C Humidity: 22.1%


Temp: 69.8 F / 21.0 C Humidity: 21.5%

Summary of changes (Pi 4 → Pi 5 + Trixie)

Raspberry Pi 4
What (Bullseye/Bookworm) Raspberry Pi 5 (Trixie)
DHT22 library adafruit-circuitpython- Kernel dht11 IIO driver
dht
DHT22 read [Link] /sys/bus/iio/devices/iio:device0/
GPIO backend [Link] or pigpio lgpio (explicit
LGPIOFactory)
pigpio Available in apt Removed from Trixie repos
[Link] no DHT entry needed dtoverlay=dht11,gpiopin=16

616
LIGHTBULB Why is the kernel driver more reliable anyway?

The dht11 kernel driver uses GPIO interrupts: every rising and falling edge on the data
line triggers a hardware interrupt that is timestamped in nanoseconds inside the kernel.
There is no scheduler jitter, no PCIe round-trip overhead, and no competition from other
user-space threads. Even on Pi 4, this approach is more robust than bit-banging from
Python — on Pi 5, it is the only approach that works.

Installing the BMP280: Barometric Pressure & Altitude Sensor

Sensor Overview:
Environmental sensing has become increasingly important in various industries, from weather
forecasting to indoor navigation and consumer electronics. At the forefront of this technological
advancement are sensors like the BMP280 and BMP180 (deprected), which excel in measuring
temperature and barometric pressure with exceptional precision and reliability.
As its predecessor, the BMP180, the BMP280 is an absolute barometric pressure sensor,
which is especially feasible for mobile applications. Its diminutive dimensions and low power
consumption allow for its implementation in battery-powered devices such as mobile phones,
GPS modules, or watches. The BMP280 is based on Bosch’s proven piezo-resistive pressure
sensor technology featuring high accuracy and linearity as well as long-term stability and high
EMC robustness. Numerous device operation options guarantee the highest flexibility. The
device is optimized for power consumption, resolution, and filter performance.
Technical data

Parameter Technical data


Operation range Pressure: 300…1100 hPa Temp.: -40…85°C
Absolute accuracy (950…1050 hPa, 0…+40°C) ~ ±1 hPa
Relative accuracy p = 700…900hPa (Temp. @ ± 0.12 hPa (typical) equivalent to ±1 m
25°C)
Average typical current consumption (1 Hz 3.4 �A @ 1 Hz
dt/rate)
Average current consumption (1 Hz dt refresh 2.74 �A, typical (ultra-low power mode)
rate)
Average current consumption in sleep mode 0.1 �A
Average measurement time 5.5 msec (ultra-low power preset)
Supply voltage VDDIO 1.2 … 3.6 V
Supply voltage VDD 1.71 … 3.6 V
Resolution of data Pressure: 0.01 hPa ( < 10 cm) Temp.:
0.01° C

617
Parameter Technical data
Temperature coefficient offset (+25°…+40°C @ 1.5 Pa/K, equiv. to 12.6 cm/K
900hPa)
Interface I²C and SPI

BMP280 Sensor Installation


Follow the diagram and make the connections:

• Vin ==> 3.3V


• GND ==> GND
• SCL ==> GPIO 3
• SDA ==> GPIO 2

618
Enabling I2C Interface
Go to RPi Configuration and confirm that the I2C interface is enabled. If not, enable it.

sudo raspi-config nonint do_i2c 0

Using the BMP280

619
If everything has been installed and connected correctly, you can turn on your Rapspi and
start interpreting the BMP180’s information about the environment.
The first thing to do is to check if the Raspi sees your BMP280. Try the following in a
terminal:

sudo i2cdetect -y 1

We should confirm that the BMP280 is on channel 77 (default) or 76.

In my case, the bus address is 0x76, so we should define it in the code.

Installing the BMP 280 Library:


Once the sensor is connected, we must install its library on our Raspi. For that, we should
install the Adafruit_CircuitPython_BMP280.

pip install adafruit-circuitpython-bmp280

Create a new Python script as below and name it, for example, bmp280_test.py:

import time
import board

import adafruit_bmp280

i2c = board.I2C()
bmp280 = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address = 0x76)

620
bmp280.sea_level_pressure = 1013.25

while True:
print("\nTemperature: %0.1f C" % [Link])
print("Pressure: %0.1f hPa" % [Link])
print("Altitude = %0.2f meters" % [Link])
[Link](2)

Execute the script:

python [Link]

The Terminal shows the result.

Note that pressure is presented in hPa. See the next section to better understand
this unit.

Below is the complete setup for our SLM tests.

621
Measuring Weather and Altitude With BMP280

Let’s take some time to understand more about what we will get with the BMP readings.

You can skip this part of the tutorial, or return later.

622
The BMP280 (and its predecessor, the BMP180) was designed to measure atmospheric pressure
accurately. Atmospheric pressure varies with both weather and altitude.
What is Atmospheric Pressure?
Atmospheric pressure is a force that the air around you exerts on everything. The weight of the
gasses in the atmosphere creates atmospheric pressure. A standard unit of pressure is “pounds
per square inch” or psi. We will use the international notation, newtons per square meter,
called pascals (Pa).

If you took 1 cm wide column of air would weigh about 1 kg

This weight, pressing down on the footprint of that column, creates the atmospheric pressure
that we can measure with sensors like the BMP280. Because that cm-wide column of air weighs
about 1 kg, the average sea level pressure is about 101,325 pascals, or better, 1013.25 hPa
(1 hPa is also known as milibar - mbar). This will drop about 4% for every 300 meters you
ascend. The higher you get, the less pressure you’ll see because the column to the top of the
atmosphere is much shorter and weighs less. This is useful because you can determine your
altitude by measuring the pressure and doing math.

The air pressure at 3,810 meters is only half that at sea level.

The BMP280 outputs absolute pressure in hPa (mbar). One pascal is a minimal amount of
pressure, approximately the amount that a sheet of paper will exert resting on a table. You will
often see measurements in hectopascals (1 hPa = 100 Pa). The library here provides outputs
of floating-point values in hPa, equaling one millibar (mbar).
Here are some conversions to other pressure units:

• 1 hPa = 100 Pa = 1 mbar = 0.001 bar


• 1 hPa = 0.75006168 Torr
• 1 hPa = 0.01450377 psi (pounds per square inch)
• 1 hPa = 0.02953337 inHg (inches of mercury)
• 1 hPa = 0.00098692 atm (standard atmospheres)

Temperature Effects
Because temperature affects the density of a gas, density affects the mass of a gas, and mass
affects the pressure (whew), atmospheric pressure will change dramatically with temperature.
Pilots know this as “density altitude”, which makes it easier to take off on a cold day than a hot
one because the air is denser and has a more significant aerodynamic effect. To compensate for
temperature, the BMP280 includes a rather good temperature sensor and a pressure sensor.
To perform a pressure reading, you first take a temperature reading, then combine that with a
raw pressure reading to come up with a final temperature-compensated pressure measurement.
(The library makes all of this very easy.)

623
Measuring Absolute Pressure
If your application requires measuring absolute pressure, all you have to do is get a temperature
reading, then perform a pressure reading (see the test script for details). The final pressure
reading will be in hPa = mbar. You can convert this to a different unit using the above
conversion factors.

Note that the absolute pressure of the atmosphere will vary with both your altitude
and the current weather patterns, both of which are useful things to measure.

Weather Observations
The atmospheric pressure at any given location on Earth (or anywhere with an atmosphere)
isn’t constant. The complex interaction between the earth’s spin, axis tilt, and many other
factors result in moving areas of higher and lower pressure, which in turn cause the variations
in weather we see every day. By watching for changes in pressure, you can predict short-term
changes in the weather. For example, dropping pressure usually means wet weather or a storm
is approaching (a low-pressure system is moving in). Rising pressure usually means clear
weather is coming (a high-pressure system is moving through). But remember that atmospheric
pressure also varies with altitude. The absolute pressure in my home, Lo Barnechea, in Chile
(altitude 960m), will always be lower than that in San Francisco (less than 2 meters, almost
sea level). If weather stations just reported their absolute pressure, it would be challenging
to compare pressure measurements from one location to another (and large-scale weather
predictions depend on measurements from as many stations as possible).
To solve this problem, weather stations continuously remove the effects of altitude from their
reported pressure readings by mathematically adding the equivalent fixed pressure to make it
appear that the reading was taken at sea level. When you do this, a higher reading in San
Francisco than in Lo Barnechea will always be because of weather patterns and not because of
altitude.
Sea Level Pressure Calculation
The See Level Pressure can be calculated with the formula:

Where,
po = SeaLevel Pressure
p = Atmospheric Pressure
L = Temperature Lapse Rate
h = Altitude
To = Sea Level Standard Temperature
g = Earth Surface Gravitational Acceleration

624
M = Molar Mass Of Dry Air
R = Universal Gas Constant
Having the absolute pressure in Pa, you check the sea level pressure using the
Calculator.

Or calculating in Python, where the altitude is the real altitude in meters where the sensor
is located.

presSeaLevel = pres / pow(1.0 - altitude/44330.0, 5.255)

Determining Altitude
Since pressure varies with altitude, you can use a pressure sensor to measure altitude (with a
few caveats). The average pressure of the atmosphere at sea level is 1013.25 hPa (or mbar).
This drops off to zero as you climb towards the vacuum of space. Because the curve of this
drop-off is well understood, you can compute the altitude difference between two pressure
measurements (p and p0) by using a specific equation. The BMP280 gives the measured
altitude using [Link].

The above explanation was based on the BMP 180 Sparkfun tutorial.

Playing with Sensors and Actuators

In this section, using the Jupyter Notebook, we will read sensors and act on actuators directly
on the Pi.
On the terminal, start the Jupyter notebook server with the command (change the IP address
with the one for your Raspi):

jupyter notebook --ip=[Link] --no-browser

625
You will need the Token; you can copy it from the terminal as shown above.

The Jupyter Notebook will be running as a server on:

http:localhost:8888

The first time you connect, you’ll need the token that appears in the Pi terminal
when you start the notebook server.

When you start your Pi and want to use Jupyter Notebook, type the “Jupyter
Notebook” command on your terminal and keep it running. This is very important!
If you need to use the terminal for another task, such as running a program, open
a new Terminal window.

To stop the server and close the “kernels” (the Jupyter notebooks), press [Ctrl] + [C].

626
Testing the Notebook setup

Let’s create a new notebook (Kernel: Python 3). Open dht_test.py, copy the code, and paste
it into the notebook. That’s it. We can see the temperature and humidity values appearing on
the cell. To interrupt the execution, go to the [stop] button at the top menu.

OK, this means we can access the physical world from our notebook! Let’s create a more
structured code for dealing with sensors and actuators.

Initialization

Import libraries, instantiate and initialize sensors/actuators

# time library
import time
import datetime

# Adafruit DHT library (Temperature/Humidity)


import board

627
import adafruit_dht
DHT22Sensor = adafruit_dht.DHT22(board.D16)

# BMP library (Pressure/Temperature)


import adafruit_bmp280
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address = 0x76)
bmp280Sensor.sea_level_pressure = 1013.25

# LEDs
from gpiozero import LED

ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)

[Link]()
[Link]()
[Link]()

# Push-Button
from gpiozero import Button
button = Button(20)

GPIO Input and Output

Create a function to get GPIO status:

# Get GPIO status data


def getGpioStatus():
global timeString
global buttonSts
global ledRedSts
global ledYlwSts
global ledGrnSts

# Get time of reading


now = [Link]()
timeString = [Link]("%Y-%m-%d %H:%M")

# Read GPIO Status

628
buttonSts = button.is_pressed
ledRedSts = ledRed.is_lit
ledYlwSts = ledYlw.is_lit
ledGrnSts = ledGrn.is_lit

And another to print the status:

# Print GPIO status data


def PrintGpioStatus():
print ("Local Station Time: ", timeString)
print ("Led Red Status: ", ledRedSts)
print ("Led Yellow Status: ", ledYlwSts)
print ("Led Green Status: ", ledGrnSts)
print ("Push-Button Status: ", buttonSts)

Now, we can, for example, turn on the LEDs:

[Link]()
[Link]()
[Link]()

And see their status:

629
If you press the push-button, its status will also be shown:

And turning off the LEDS:

[Link]()
[Link]()
[Link]()

630
We can create a function to simplify turning LEDs on and off:

# Acting on GPIOs and printing Status


def controlLeds(r, y, g):
if (r):
[Link]()
else:
[Link]()
if (y):
[Link]()
else:
[Link]()
if (g):
[Link]()
else:
[Link]()

getGpioStatus()
PrintGpioStatus()

For example, turning on the Yellow LED:

631
Getting and displaying Sensor Data

First, we should create a function to read the BMP280 and calculate the pressure value at sea
level, once the sensor only gives us the absolute pressure based on the actual altitude:

# Read data from BMP280


def bmp280GetData(real_altitude):

temp = [Link]
pres = [Link]
alt = [Link]
presSeaLevel = pres / pow(1.0 - real_altitude/44330.0, 5.255)

temp = round (temp, 1)


pres = round (pres, 2) # absolute pressure in mbar
alt = round (alt)
presSeaLevel = round (presSeaLevel, 2) # absolute pressure in mbar

return temp, pres, alt, presSeaLevel

Entering the BMP280 real altitude where it is located, run the code:

bmp280GetData(960)

As a result, we will get (26.9, 906.73, 927, 1017.29)which means:

• Temperature of 26.9 oC
• Absolute Pressure of 906.73 hPa

632
• Measured Altitude (from Pressure) of 927 m
• Sea Level converted Pressure: 1,017.29 hPa

Now, we will generate a unique function to get the BMP280 and the DHT data, including a
timestamp:

# Get data (from local sensors)


def getSensorData(altReal=0):
global timeString
global humExt
global tempLab
global tempExt
global presSL
global altLab
global presAbs
global buttonSts

# Get time of reading


now = [Link]()
timeString = [Link]("%Y-%m-%d %H:%M")

tempLab, presAbs, altLab, presSL = bmp280GetData(altReal)

tempDHT = [Link]
humDHT = [Link]

if humDHT is not None and tempDHT is not None:


tempExt = round (tempDHT)
humExt = round (humDHT)

And another function to print the values:

# Display important data on-screen


def printData():
print ("Local Station Time: ", timeString)
print ("External Air Temperature (DHT): ", tempExt, "oC")
print ("External Air Humidity (DHT): ", humExt, "%")
print ("Station Air Temperature (BMP): ", tempLab, "oC")
print ("Sea Level Air Pressure: ", presSL, "mBar")
print ("Absolute Station Air Pressure: ", presAbs, "mBar")
print ("Station Measured Altitude: ", altLab, "m")

Runing them:

633
real_altitude = 960 # real altitude of where the BMP280 is installed
getSensorData(real_altitude)
printData()

Results:

Using Python, we can command the actuators (LEDs) and read the sensors and GIPOs status
at this stage. This is important, for example, to generate a data log to be read by an SLM in
the future.

The notebook can be found on Github: Monitoring_Actuating_GPIOs.ipynb

� IMPORTANT: The Notebook Kernel should end after using to liberate GPIOs

The problem is that Jupyter Notebook is still holding onto the GPIO pins
even after our code finishes running. When we try to run it from the terminal,
those pins are already claimed by the Jupyter process. So, before running a code
that deals with GPIO in the terminal, In Jupyter:

• Click Kernel → Restart Kernel


• Or: Click Kernel → Shutdown Kernel

Widgets

pywidgets, or jupyter-widgets orwidgets, are interactive HTML widgets for Jupyter note-
books and the IPython kernel. Notebooks come alive when interactive widgets are used. We
can gain control of our data and visualize changes in them.
Widgets are eventful Python objects that have a representation in the browser, often as a
control like a slider, text box, etc. We can use widgets to build interactive GUIs for our
project.

634
In this lab, for example, we will use a slide bar to control the state of actuators in real time,
such as by turning on or off the LEDs. Widgets are great for adding more dynamic behavior
to Jupyter Notebooks.
Installation
To use Widgets, we must install the Ipywidgets library using the commands:

pip install ipywidgets

After installation, we should call the library:

# widget library
from ipywidgets import interactive
import ipywidgets as widgetsfrom
[Link] import display

And running the below line, we can control the LEDs in real-time:

f = interactive(controlLeds, r=(0,1,1), y=(0,1,1), g=(0,1,1))


display(f)

This interactive widget is very easy to implement and very powerful. You can learn
more about Interactive on this link: Interactive Widget.

Advanced GPIO Zero Features

GPIO Zero includes many advanced features that make complex projects much easier to build.

635
PWM LED Brightness Control

Use PWMLED to control LED brightness:

from gpiozero import PWMLED


from time import sleep

led = PWMLED(13)

# Fade in and out


while True:
[Link]() # Smooth fade in and out
sleep(5)

Composite Devices

Create a traffic light system with one object:

from gpiozero import TrafficLights


from time import sleep

lights = TrafficLights(red=13, amber=19, green=26)

while True:
[Link]()
sleep(2)
[Link]()
sleep(1)
[Link]()
[Link]()
[Link]()
sleep(3)
[Link]()
[Link]()
sleep(1)
[Link]()

Other Useful GPIO Zero Devices

• Buzzer: from gpiozero import Buzzer

636
• Motor: from gpiozero import Motor
• Servo: from gpiozero import Servo
• DistanceSensor: from gpiozero import DistanceSensor
• MotionSensor: from gpiozero import MotionSensor
• Robot: from gpiozero import Robot

Interacting an SLM with the Physical world

This section demonstrates in a simple way how to integrate a Small Language Model (SLM)
with the sensors and LEDs we have set up. The diagram below shows how data flows from
sensors through processing and AI analysis to control the actuators and ultimately provide
user feedback.

637
We will use the Transformers library from Hugging Face for model loading and inference. This
library provides the architecture for working with pre-trained language models, helping interact
with the model, processing input prompts, and obtaining outputs.
Installation

pip install transformers torch

Let’s create a simple SLM test in the Jupyter Notebook that checks if the model loads
and measures inference time. The model used here is the TinyLLama 1.1B. We will ask a
straightforward question:

"The weather today is"

As a result, besides the SLM answer, we will also measure the latency.
Run this script:

import time
from transformers import pipeline
import torch

# Check if CUDA is available (it won't be on our case, Raspberry Pi)


device = "cuda" if [Link].is_available() else "cpu"
print(f"Using device: {device}")

# Load the model and measure loading time


start_time = [Link]()

model='TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T'
generator = pipeline('text-generation',
model=model,
device=device)
load_time = [Link]() - start_time
print(f"Model loading time: {load_time:.2f} seconds")

# Test prompt
test_prompt = "The weather today is"

# Measure inference time


start_time = [Link]()
response = generator(test_prompt,
max_length=50,

638
num_return_sequences=1,
temperature=0.7)
inference_time = [Link]() - start_time

print(f"\nTest prompt: {test_prompt}")


print(f"Generated response: {response[0]['generated_text']}")
print(f"Inference time: {inference_time:.2f} seconds")

As we can see, the SLM works, but the latency is very high (+3 minutes). It is OK because this
particular test is on a Raspberry Pi 4. With a Raspberry Pi 5, the result would be better.:

The Raspi uses around 1GB of memory (model + process) and all four cores to process the
answer. The model alone needs around 800MB.

639
Now, let us create a code showing a basic interaction pattern where the SLM can respond to
sensor data and interact with the LEDs.
Install the Libraries:

import time
import datetime
import board
import adafruit_dht
import adafruit_bmp280
from gpiozero import LED, Button
from transformers import pipeline

Initialize sensors

DHT22Sensor = adafruit_dht.DHT22(board.D16)
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25

Initialize LEDs and Button

ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)

Initialize the SLM pipeline

# We're using a small model suitable for Raspberry Pi

model='TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T'
generator = pipeline('text-generation',
model=model,
device='cpu')

Support Functions
Now, let’s create support functions for readings from all sensors and control the LEDs:

640
def get_sensor_data():
"""Get current readings from all sensors"""
try:
temp_dht = [Link]
humidity = [Link]
temp_bmp = [Link]
pressure = [Link]

return {
'temperature_dht': round(temp_dht, 1) if temp_dht else None,
'humidity': round(humidity, 1) if humidity else None,
'temperature_bmp': round(temp_bmp, 1),
'pressure': round(pressure, 1)
}
except RuntimeError:
return None

def control_leds(red=False, yellow=False, green=False):


"""Control LED states"""
[Link] = red
[Link] = yellow
[Link] = green

def process_conditions(sensor_data):
"""Process sensor data and control LEDs based on conditions"""
if not sensor_data:
control_leds(red=True) # Error condition
return

temp = sensor_data['temperature_dht']
humidity = sensor_data['humidity']

# Example conditions for LED control


if temp > 30: # Hot
control_leds(red=True)
elif humidity > 70: # Humid
control_leds(yellow=True)
else: # Normal conditions
control_leds(green=True)

641
Generating an SLM’s response
So far, the LEDs reaction is only based on logic, but let’s also use the SLM to “analyse” the
sensors condition, generating a response based on that:

def generate_response(sensor_data):
"""Generate response based on sensor data using SLM"""
if not sensor_data:
return "Unable to read sensor data"

prompt = f"""Based on these sensor readings:


Temperature: {sensor_data['temperature_dht']}°C
Humidity: {sensor_data['humidity']}%
Pressure: {sensor_data['pressure']} hPa

Provide a brief status and recommendation in 2 sentences.


"""

# Generate response from SLM


response = generator(prompt,
max_length=100,
num_return_sequences=1,
temperature=0.7)[0]['generated_text']

return response

Main Function
And now, let’s create a main() function to wait for the user to, for example, press a button
and, capture the data generated by the sensors, delivering some observation or recommendation
from the SLM:

def main_loop():
"""Main program loop"""
print("Starting Physical Computing with SLM Integration...")
print("Press the button to get a reading and SLM response.")

try:
while True:
if button.is_pressed:
# Get sensor readings
sensor_data = get_sensor_data()

642
# Process conditions and control LEDs
process_conditions(sensor_data)

if sensor_data:
# Get SLM response
response = generate_response(sensor_data)

# Print current status


print("\nCurrent Readings:")
print(f"Temperature: {sensor_data['temperature_dht']}°C")
print(f"Humidity: {sensor_data['humidity']}%")
print(f"Pressure: {sensor_data['pressure']} hPa")
print("\nSLM Response:")
print(response)

[Link](2) # Debounce and allow time to read

[Link](0.1) # Reduce CPU usage

except KeyboardInterrupt:
print("\nShutting down...")
control_leds(False, False, False) # Turn off all LEDs

Test Result
The sensors are read after the user presses the button to trigger a reading, and LEDs are
controlled based on conditions. Sensor data is formatted into a prompt for the SLM to generate
a response analyzing the current conditions. The results are displayed in the terminal, and the
LED indicators are shown.

• Red: High temperature (>30°C) or error condition


• Yellow: High humidity (>70%)
• Green: Normal conditions

This simple code integrates a Small Language Model (TinyLlama model (1.1B parameters)
with our physical computing setup, providing raw sensor data and intelligent responses from
the SLM about the environmental conditions.

643
We can extend this first test to more sophisticated and valuable uses of the SLM integration,
for example: adding:

• Starting the process from a User Prompt.


• Receive commands from the User to switch LEDs ON or OFF
• Provide the status of LEDS, Button, or specific sensor data from the user prompt
• Log data and responses to a file. Provide historical information by user request
• Implement different types of prompts for various use cases

Other Models

We can use other SLMs in a Raspberry Pi that have distinct ways of handling them. For
example, many modern models use GGUF formats, and to use them, we need to install
llama-cpp-python, which is designed to work with GGUF models.
Also, as we saw in a previous lab, Ollama is a great way to download and test SLMs on the
Raspberry Pi.

644
Conclusion

Key Achievements

Throughout this tutorial, we’ve successfully: - Set up a complete physical computing environ-
ment using Raspberry Pi - Integrated multiple environmental sensors (DHT22 and BMP280)
- Implemented visual feedback through LED actuators - Created interactive controls using
push buttons - Integrated a Small Language Model (TinyLLama 1.1B) for intelligent analysis -
Developed a foundation for AI-enhanced environmental monitoring

Technical Insights

Hardware Integration

The combination of digital (DHT22) and I2C (BMP280) sensors demonstrated different commu-
nication protocols and their implementations. This multi-sensor approach provides redundancy
and comprehensive environmental monitoring capabilities. The LED actuators and push-button
interface created a responsive and interactive system that bridges the digital and physical
worlds.

Software Architecture

The layered software architecture we developed supports: 1. Low-level sensor communication


and actuator control 2. Data preprocessing and validation 3. SLM integration for intelligent
analysis 4. Interactive user interfaces through both hardware and software

AI Integration Learnings

The integration of TinyLLama 1.1B revealed several important insights: - Small Language
Models can effectively run on edge devices like Raspberry Pi - Natural language processing can
enhance sensor data interpretation - Real-time analysis is possible, though with some latency
considerations - The system can provide human-readable insights from complex sensor data

Practical Applications

This project serves as a foundation for numerous real-world applications: - Environmental


monitoring systems - Smart home automation - Industrial sensor networks - Educational
platforms for IoT and AI integration - Prototyping platforms for larger-scale deployments

645
Challenges and Solutions

Throughout the development, we encountered and addressed several challenges: 1. Resource


Constraints: - Optimized SLM inference for Raspberry Pi capabilities - Implemented efficient
sensor reading strategies - Managed memory usage for stable operation

2. Data Integration:
• Developed robust sensor data validation
• Created effective data preprocessing pipelines
• Implemented error handling for sensor failures
3. AI Integration:
• Designed effective prompting strategies
• Managed inference latency
• Balanced accuracy with response time

Future Enhancements

The system can be extended in several directions: 1. Hardware Expansions: - Additional


sensor types (air quality, light, motion) - Camera for IA applications - More complex actuators
(displays, motors, relays) - Wireless connectivity options as WiFI, BLE, or LoRa 2. Software
Improvements: - Advanced data logging and analysis - Web-based monitoring interface -
Real-time visualization tools 3. AI Capabilities: 1. Models for detecting and counting objects
2. RAG or Fine-tuning SLM for specific applications 3. Multi-modal AI integration via sensor
integration 4. Automated decision-making systems 5. Predictive maintenance capabilities

Final Thoughts

This chapter demonstrates that integrating physical computing with AI is feasible and practical
on readily accessible hardware such as the Raspberry Pi. Combining sensors, actuators, and
AI creates a powerful platform for developing intelligent environmental monitoring and control
systems.
While the current implementation focuses on environmental monitoring, the principles and
techniques can be adapted to various applications. The modular nature of hardware and
software components allows for customization and expansion based on specific needs.
Integrating small language models into physical computing opens new possibilities for creating
more intuitive and intelligent IoT devices. As edge AI capabilities evolve, projects like this
will become increasingly important in developing the next generation of smart devices and
systems.

646
Remember that this is just the beginning. Our foundation can be extended in countless ways
to create more sophisticated and capable systems. The key is to build on these basics while
balancing functionality, reliability, and resource usage.

Resources

• GPIOs - Scripts
• Sensors - Scripts
• Notebooks

647
Experimenting with SLMs for IoT Control

Introduction

This chapter explores the implementation of Small Language Models (SLMs) in IoT control
systems, demonstrating the possibility of creating a monitoring and control system using edge
AI. We’ll integrate these models with physical sensors and actuators, creating an intelligent
IoT system capable of natural language interaction. While this implementation shows
the potential of integrating AI with physical systems, it also highlights current limitations and
areas for improvement.

This chapter builds on the concepts introduced in “Small Language Models (SLMs)”
and “Physical Computing with Raspberry Pi.”

648
The Physical Computing chapter laid the groundwork for interfacing with hardware compo-
nents using the Raspberry Pi’s GPIO pins. We’ll revisit these concepts, focusing on connecting
and interacting with sensors (DHT22 for temperature and humidity, BMP280 for temperature
and pressure, and a push-button for digital inputs), as well as controlling actuators (LEDs) in
a more sophisticated setup.
We will progress from a simple IoT system to a more advanced platform that combines real-time
monitoring, historical data analysis, and natural language processing (NLP).

This chapter demonstrates a progressive evolution through several key stages:

1. Basic Sensor Integration


• Hardware interface with DHT22 (temperature/humidity) and BMP280 (tempera-
ture/pressure) sensors
• Digital input through a push-button
• Output control via RGB LEDs
• Foundational data collection and device control
2. SLM Basic Analysis
• Initial integration with small language models

649
• Simple observation and reporting of system state
• Demonstration of SLM’s ability to interpret sensor data
3. Active Control Implementation
• Direct LED control based on SLM decisions
• Temperature threshold monitoring
• Emergency state detection via button input
• Real-time system state analysis
4. Natural Language Interaction
• Free-form command interpretation
• Context-aware responses
• Multiple SLM model support
• Flexible query handling
5. Data Logging and Analysis
• Continuous system state recording
• Trend analysis and pattern detection
• Historical data querying
• Performance monitoring

Let’s begin by setting up our hardware and software environment, building upon the foundation
established in our previous labs.

Setup

Hardware Setup

Connection Diagram

Component GPIO Pin


DHT22 GPIO16
BMP280 - SCL GPIO03
BMP280 - SDA GPIO02
Red LED GPIO13
Yellow LED GPIO19
Green LED GPIO26
Button GPIO20

650
• Raspberry Pi 5 (with an OS installed, as detailed in previous labs)
• DHT22 temperature and humidity sensor
• BMP280 temperature and pressure sensor
• 3 LEDs (red, yellow, green)
• Push button
• 330Ω resistors (3)
• Jumper wires and breadboard

651
Software Prerequisites

1. Install required libraries:

pip install adafruit-circuitpython-dht


pip install adafruit-circuitpython-bmp280

Basic Sensor Integration

Let’s create a Python script ([Link]) to handle the sensors and actuators. This script
will contain functions to be called from other scripts later:

import time
import board
import adafruit_dht
import adafruit_bmp280
from gpiozero import LED, Button

DHT22Sensor = adafruit_dht.DHT22(board.D16)
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25

ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)

def collect_data():
try:
temperature_dht = [Link]
humidity = [Link]
temperature_bmp = [Link]
pressure = [Link]
button_pressed = button.is_pressed
return temperature_dht, humidity, temperature_bmp, pressure, button_pressed
except RuntimeError:
return None, None, None, None, None

def led_status():
ledRedSts = ledRed.is_lit

652
ledYlwSts = ledYlw.is_lit
ledGrnSts = ledGrn.is_lit
return ledRedSts, ledYlwSts, ledGrnSts

def control_leds(red, yellow, green):


[Link]() if red else [Link]()
[Link]() if yellow else [Link]()
[Link]() if green else [Link]()

We can test the functions using:

while True:
ledRedSts, ledYlwSts, ledGrnSts = led_status()
temp_dht, hum, temp_bmp, press, button_state = collect_data()

#control_leds(True, True, True)

if all(v is not None for v in [temp_dht, hum, temp_bmp, press]):


print(f"DHT22 Temp: {temp_dht:.1f}°C, Humidity: {hum:.1f}%")
print(f"BMP280 Temp: {temp_bmp:.1f}°C, Pressure: {press:.2f}hPa")
print(f"Button {'pressed' if button_state else 'not pressed'}")
print(f"Red LED {'is on' if ledRedSts else 'is off'}")
print(f"Yellow LED {'is on' if ledYlwSts else 'is off'}")
print(f"Green LED {'is on' if ledGrnSts else 'is off'}")

[Link](2)

653
IMPORTANT: Updated [Link] for Pi 5 + Trixie

When integrating the DHT22 into the full monitoring script alongside LEDs and BMP280, two
changes are needed: replace adafruit_dht with the sysfs reader, and explicitly set gpiozero’s
pin factory to LGPIOFactory (required on Pi 5).

import time
import board
import adafruit_bmp280
from gpiozero import LED, Button, Device
from [Link] import LGPIOFactory

# Required on Raspberry Pi 5
Device.pin_factory = LGPIOFactory()

# DHT22 via kernel IIO driver (replaces adafruit_dht)


DHT_IIO = "/sys/bus/iio/devices/iio:device0/"

def read_dht22():
try:
temp_c = int(open(DHT_IIO + "in_temp_input").read()) / 1000.0
humidity = int(open(DHT_IIO + "in_humidityrelative_input").read()) / 1000.0
return temp_c, humidity
except Exception:
return None, None

# BMP280 via I2C (unchanged)

654
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25

# LEDs and button (unchanged)


ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)

def collect_data():
temperature_dht, humidity = read_dht22()
try:
temperature_bmp = [Link]
pressure = [Link]
except Exception:
temperature_bmp, pressure = None, None
button_pressed = button.is_pressed
return temperature_dht, humidity, temperature_bmp, pressure, button_pressed

def led_status():
return ledRed.is_lit, ledYlw.is_lit, ledGrn.is_lit

def control_leds(red, yellow, green):


[Link]() if red else [Link]()
[Link]() if yellow else [Link]()
[Link]() if green else [Link]()

SLM Basic Analysis

Now, let’s create a new script, slm_basic_analysis.py, which will be responsible for analysing
the hardware components’ status, according to the following diagram:

655
The diagram shows the basic analysis system, which consists of:

1. Hardware Layer:
• Sensors: DHT22 (temperature/humidity), BMP280 (temperature/pressure)
• Input: Emergency button
• Output: Three LEDs (Red, Yellow, Green)
2. [Link]:
• Handles all hardware interactions
• Provides two main functions:
– collect_data(): Reads all sensor values
– led_status(): Checks current LED states
3. slm_basic_analysis.py:
• Creates a descriptive prompt using sensor data
• Sends prompt to SLM (for example, the Llama 3.2 1B)
• Displays analysis results
• In this step we will not control the LEDs (observation only)

Okay, let’s implement the code, starting for importing the Ollama library and the functions to
monitor the HW (from the previous script):

import ollama
from monitor import collect_data, led_status

Calling the monitor functions, we will get all data:

656
ledRedSts, ledYlwSts, ledGrnSts = led_status()
temp_dht, hum, temp_bmp, press, button_state = collect_data()

Now, the heart of out code, we will generate the Prompt, using the data captured on the
previous variables:

prompt = f"""
You are an experienced environmental scientist.
Analyze the information received from an IoT system:

DHT22 Temp: {temp_dht:.1f}°C and Humidity: {hum:.1f}%


BMP280 Temp: {temp_bmp:.1f}°C and Pressure: {press:.2f}hPa
Button {"pressed" if button_state else "not pressed"}
Red LED {"is on" if ledRedSts else "is off"}
Yellow LED {"is on" if ledYlwSts else "is off"}
Green LED {"is on" if ledGrnSts else "is off"}

Where,
- The button, not pressed, shows a normal operation
- The button, when pressed, shows an emergency
- Red LED when is on, indicates a problem/emergency.
- Yellow LED when is on indicates a warning situation.
- Green LED when is on, indicates system is OK.

If the temperature is over 20°C, mean a warning situation

You should answer only with: "Activate Red LED" or


"Activate Yellow LED" or "Activate Green LED"

"""

Now, the Prompt will be passed to the SLM, which will generate a response:

MODEL = 'llama3.2:3b'
PROMPT = prompt
response = [Link](
model=MODEL,
prompt=PROMPT
)

The last stage will be show the real monitored data and the SLM’s response:

657
print(f"\nSmart IoT Analyser using {MODEL} model\n")

print(f"SYSTEM REAL DATA")


print(f" - DHT22 ==> Temp: {temp_dht:.1f}°C, Humidity: {hum:.1f}%")
print(f" - BMP280 => Temp: {temp_bmp:.1f}°C, Pressure: {press:.2f}hPa")
print(f" - Button {'pressed' if button_state else 'not pressed'}")
print(f" - Red LED {'is on' if ledRedSts else 'is off'}")
print(f" - Yellow LED {'is on' if ledYlwSts else 'is off'}")
print(f" - Green LED {'is on' if ledGrnSts else 'is off'}")

print(f"\n>> {MODEL} Response: {response['response']}")

Runing the Python script, we got:

In this initial experiment, the system successfully collected sensor data (temperatures of
26.3°C and 26.1°C from DHT22 and BMP280, respectively, 40.2% humidity, and 908.84hPa
pressure) and processed this information through the SLM, which produced a coherent response
recommending the activation of the yellow LED due to elevated temperature conditions.

658
The model’s ability to interpret sensor data and provide logical, rule-based decisions shows
promise. Still, the simplistic nature of the current implementation (using basic thresholds and
binary LED outputs) suggests significant room for improvement through more sophisticated
prompting strategies, historical data integration, and the implementation of safety mechanisms.
Also, the result is probabilistic, meaning it should change after execution.

Act on Output (Actuators)

Let’s use the output generated and use it to actuate on the LEDs, our “actuators”:

We can add a new function, parse_llm_response(), to return a command to the the LEDs
based on the SLM’s response:

def parse_llm_response(response_text):
"""Parse the LLM response to extract LED control instructions."""
response_lower = response_text.lower()
red_led = 'activate red led' in response_lower
yellow_led = 'activate yellow led' in response_lower
green_led = 'activate green led' in response_lower
return (red_led, yellow_led, green_led)

The return to this function:

659
red, yellow, green = parse_llm_response(response['response'])

Would be:(False, True, False), which can be the input to the function control_leds(red,
yellow, green)'

control_leds(red, yellow, green)


led_status()

(False, True, False)

Creating new functions for the Actuation:

One for the inference:

def slm_inference(PROMPT, MODEL):


response = [Link](
model=MODEL,
prompt=PROMPT
)
return response

And another for output analysis and actuation:

660
def output_actuator(response, MODEL):
print(f"\nSmart IoT Actuator using {MODEL} model\n")

print(f"SYSTEM REAL DATA")


print(f" - DHT22 ==> Temp: {temp_dht:.1f}°C, Humidity: {hum:.1f}%")
print(f" - BMP280 => Temp: {temp_bmp:.1f}°C, Pressure: {press:.2f}hPa")
print(f" - Button {'pressed' if button_state else 'not pressed'}")

print(f"\n>> {MODEL} Response: {response['response']}")

# Control LEDs based on response


red, yellow, green = parse_llm_response(response['response'])
control_leds(red, yellow, green)

print(f"\nSYSTEM ACTUATOR STATUS")


ledRedSts, ledYlwSts, ledGrnSts = led_status()
print(f" - Red LED {'is on' if ledRedSts else 'is off'}")
print(f" - Yellow LED {'is on' if ledYlwSts else 'is off'}")
print(f" - Green LED {'is on' if ledGrnSts else 'is off'}")

By updating the system and calling the two functions in sequence, we can have the full cycle of
analysis by the SLM and the actuation on the output LEDs:

ledRedSts, ledYlwSts, ledGrnSts = led_status()


temp_dht, hum, temp_bmp, press, button_state = collect_data()

response = slm_inference(PROMPT, MODEL)


output_actuator(response, MODEL)

The new script can be found on GitHub: slm_basic_act_leds.

Let’s press the button, call the functions, and see what happens:

ledRedSts, ledYlwSts, ledGrnSts = led_status()


temp_dht, hum, temp_bmp, press, button_state = collect_data()
response = slm_inference(PROMPT, MODEL)
output_actuator(response, MODEL)

661
We can see that, despite the button being pressed, the SLM did not consider it and misinterpreted
the temperature value as ABOVE the threshold, not below. Also, despite the fact that we
asked for a straight answer about which LED to turn on, the model lost time with analysis,
This issue relates to the prompt we wrote. Let’s cover it in the next section.

662
Prompting Engineering

Looking at the answers we got with the previous implementation, which are not always correct,
we can see that the main issue is unreliable text parsing. Using JSON to parse the answer
should be much better!

Key Changes in the code:

1. Added JSON import


2. Fixed variable scope bug - The previous code tried to use temp_dht, hum, etc. in the
prompt before they were defined. The prompt is now created in a function after receiving
sensor data.

663
3. Changed to JSON format - The prompt now asks for a structured JSON response:
{"red_led": true, "yellow_led": false, "green_led": false}

4. Updated parser - Now parses JSON instead of searching for text strings. Includes error
handling and fallback to safe state (all LEDs off) if parsing fails.

Why JSON is Better:

• More reliable: No ambiguity about which LEDs to activate


• Structured: Clear true/false values instead of parsing text
• Error-resistant: The parser handles markdown code blocks (some models
wrap JSON in “‘) and provides safe fallback
• Flexible: Easy to add more fields later if needed

Let’s revise the previous code, with a better prompt and parcing funcion, now based on JSON.

System Overview: Enhanced IoT Environmental Monitoring with SLM Control

This project demonstrates how a Small Language Model (SLM) can make intelligent decisions
in an IoT system. The system monitors environmental conditions using sensors and controls
LED indicators based on the SLM’s analysis.

664
New System Architecture

The system is organized into three main layers:

1. Hardware Components Layer

The physical layer consists of:

• DHT22 Sensor: Measures temperature and humidity


• BMP280 Sensor: Measures temperature and atmospheric pressure
• Emergency Button: Manual override for emergency situations
• Three LEDs: Visual indicators
– Red LED: Emergency/Problem state
– Yellow LED: Warning state

665
– Green LED: Normal operation

2. Hardware Interface Layer ([Link])

This layer provides the bridge between hardware and software with three key functions:

• collect_data(): Reads all sensor values (temperature, humidity, pressure) and button
state
• led_status(): Checks the current state of all LEDs
• control_leds(red, yellow, green): Controls which LED is turned on based on
boolean values

3. Intelligence Layer (slm_act_leds.py)

This is where the SLM makes decisions. The workflow follows these steps:
a) Data Collection & Preparation

• slm_analyse_act(MODEL, TEMP_THRESHOLD): Main function that orchestrates the entire


process
– Calls collect_data() to get current sensor readings
– Calls led_status() to check LED states

b) Prompt Generation

• create_prompt(): Builds a detailed prompt for the SLM, including:


– Current sensor readings
– Temperature threshold value
– Decision rules with clear priorities:
1. Button pressed → Red LED (Emergency - highest priority)
2. Temperature exceeds threshold → Yellow LED (Warning)
3. Normal conditions → Green LED (All OK)
– Pre-analyzed data to help the SLM understand the current state

c) SLM Inference

• slm_inference(PROMPT, MODEL): Sends the prompt to the Ollama SLM (llama3.2:3b


or any other SLM, as Gemma or Phi)
• The SLM analyzes the situation and returns a JSON response:

666
{
"red_led": false,
"yellow_led": false,
"green_led": true
}

d) Response Processing

• parse_llm_response(): Parses the JSON response from the SLM


– Handles edge cases like markdown code blocks
– Extracts boolean values for each LED
– Returns a tuple: (red, yellow, green)

e) Actuation

• output_actuator(): Takes the SLM’s decision and executes it


– Displays all sensor data
– Shows the SLM’s JSON response
– Shows the final decision
– Calls control_leds() to physically turn LEDs on/off
– Displays final LED status for verification

How It Works: A Complete Cycle

1. Sensors continuously monitor the environment (temperature, humidity, pressure,


button state)
2. Data flows from hardware through [Link] to slm_act_leds.py
3. The SLM receives a structured prompt with current conditions and decision rules
4. The SLM analyzes the data and decides which LED should be activated
5. The decision returns as a JSON object specifying exactly which LED to turn on
6. The system executes the decision by controlling the LEDs through [Link]
7. Visual feedback is provided via the LED, and the system logs all data and decisions

Key Design Principles

JSON Communication: Using JSON format ensures reliable, structured communication


between the SLM and the actuation system, reducing parsing errors.
SLM-Driven Decisions: The SLM has complete control over LED decisions, demonstrating
how AI can autonomously manage IoT systems.

667
Configurable Threshold: The temperature threshold is a parameter that makes the system
adaptable to different environments without code changes. Can be updated by the user.
Clear Priority Rules: The system follows explicit priority rules that the SLM understands:

• Safety first (button press = emergency)


• Then environmental warnings (temperature)
• Then normal operation

Now, this architecture showcases how Small Language Models can be integrated into IoT
systems to provide intelligent, context-aware decision-making while maintaining simplicity and
reliability.

The new code

import ollama
import json
from monitor import collect_data, led_status, control_leds

def create_prompt(temp_dht, hum, temp_bmp, press, button_state,


ledRedSts, ledYlwSts, ledGrnSts, TEMP_THRESHOLD):
"""Create a prompt for the LLM with current sensor data."""
return f"""
You are controlling an IoT LED system. Analyze the sensor data and decide
which ONE LED to activate.

SENSOR DATA:
- DHT22 Temperature: {temp_dht:.1f}°C
- BMP280 Temperature: {temp_bmp:.1f}°C
- Humidity: {hum:.1f}%
- Pressure: {press:.2f}hPa
- Button: {"PRESSED" if button_state else "NOT PRESSED"}

TEMPERATURE THRESHOLD: {TEMP_THRESHOLD}°C

DECISION RULES (apply in this priority order):


1. IF button is PRESSED → Activate Red LED (EMERGENCY - highest priority)
2. IF button is NOT PRESSED AND (DHT22 temp > {TEMP_THRESHOLD}°C OR
BMP280 temp > {TEMP_THRESHOLD}°C) → Activate Yellow LED (WARNING)
3. IF button is NOT PRESSED AND (DHT22 temp � {TEMP_THRESHOLD}°C AND
BMP280 temp � {TEMP_THRESHOLD}°C) → Activate Green LED (NORMAL)

668
CURRENT ANALYSIS:
- Button status: {"PRESSED" if button_state else "NOT PRESSED"}
- DHT22 temp ({temp_dht:.1f}°C) is {"OVER" if temp_dht > TEMP_THRESHOLD
else "AT OR BELOW"} threshold ({TEMP_THRESHOLD}°C)
- BMP280 temp ({temp_bmp:.1f}°C) is {"OVER" if temp_bmp > TEMP_THRESHOLD
else "AT OR BELOW"} threshold ({TEMP_THRESHOLD}°C)

Based on these rules, respond with ONLY a JSON object (no other text):
{{"red_led": true, "yellow_led": false, "green_led": false}}

Only ONE LED should be true, the other two must be false.
"""

def parse_llm_response(response_text):
"""Parse the LLM JSON response to extract LED control instructions."""
try:
# Clean the response - remove any markdown code blocks if present
response_text = response_text.strip()
if response_text.startswith('```'):
# Extract JSON from markdown code block
lines = response_text.split('\n')
response_text = '\n'.join(lines[1:-1])
if len(lines) > 2
else response_text

# Parse JSON
data = [Link](response_text)
red_led = [Link]('red_led', False)
yellow_led = [Link]('yellow_led', False)
green_led = [Link]('green_led', False)
return (red_led, yellow_led, green_led)
except ([Link], KeyError) as e:
print(f"Error parsing JSON response: {e}")
print(f"Response was: {response_text}")
# Fallback to safe state (all LEDs off)
return (False, False, False)

def output_actuator(response, MODEL, temp_dht, hum, temp_bmp,


press, button_state):
print(f"\nSmart IoT Actuator using {MODEL} model\n")

print(f"SYSTEM REAL DATA")

669
print(f" - DHT22 ==> Temp: {temp_dht:.1f}°C, Humidity: {hum:.1f}%")
print(f" - BMP280 => Temp: {temp_bmp:.1f}°C, Pressure: {press:.2f}hPa")
print(f" - Button {'pressed' if button_state else 'not pressed'}")

print(f"\n>> {MODEL} Response: {response['response']}")

# Parse LLM response and use it directly (no validation)


red, yellow, green = parse_llm_response(response['response'])
print(f">> SLM decision: Red={red}, Yellow={yellow}, Green={green}")

# Control LEDs based on SLM decision


control_leds(red, yellow, green)

print(f"\nSYSTEM ACTUATOR STATUS")


ledRedSts, ledYlwSts, ledGrnSts = led_status()
print(f" - Red LED {'is on' if ledRedSts else 'is off'}")
print(f" - Yellow LED {'is on' if ledYlwSts else 'is off'}")
print(f" - Green LED {'is on' if ledGrnSts else 'is off'}")

def slm_analyse_act(MODEL, TEMP_THRESHOLD):


"""Main function to get sensor data, run SLM inference, and actuate LEDs."""
# Get system info
ledRedSts, ledYlwSts, ledGrnSts = led_status()
temp_dht, hum, temp_bmp, press, button_state = collect_data()

# Create prompt with current sensor data


PROMPT = create_prompt(temp_dht,
hum,
temp_bmp,
press,
button_state,
ledRedSts,
ledYlwSts,
ledGrnSts,
TEMP_THRESHOLD)

# Analyse and actuate on LEDs


response = slm_inference(PROMPT, MODEL)
output_actuator(response, MODEL, temp_dht, hum, temp_bmp, press, button_state)

670
Code Flow Diagram

TEST1: Temp above the threshold

Definitions

# Model to be used
MODEL = 'llama3.2:3b'

# Temperature threshold for warning


TEMP_THRESHOLD = 20.0

671
Calling the program

slm_analyse_act(MODEL, TEMP_THRESHOLD)

Result

Smart IoT Actuator using llama3.2:3b model

SYSTEM REAL DATA


- DHT22 ==> Temp: 21.5°C, Humidity: 28.8%
- BMP280 => Temp: 22.4°C, Pressure: 910.04hPa
- Button not pressed

>> llama3.2:3b Response: {"red_led": false, "yellow_led": true, "green_led": false}


>> SLM decision: Red=False, Yellow=True, Green=False

SYSTEM ACTUATOR STATUS


- Red LED is off
- Yellow LED is on
- Green LED is off

672
TEST2: Temp below the threshold

# Temperature threshold for warning


TEMP_THRESHOLD = 25.0

slm_analyse_act(MODEL, TEMP_THRESHOLD)

Smart IoT Actuator using llama3.2:3b model

SYSTEM REAL DATA


- DHT22 ==> Temp: 21.6°C, Humidity: 29.1%
- BMP280 => Temp: 22.4°C, Pressure: 910.07hPa
- Button not pressed

>> llama3.2:3b Response: {"red_led": false, "yellow_led": false, "green_led": true}


>> SLM decision: Red=False, Yellow=False, Green=True

SYSTEM ACTUATOR STATUS


- Red LED is off
- Yellow LED is off
- Green LED is on

673
TEST3: Alarm Button pressed

slm_analyse_act(MODEL, TEMP_THRESHOLD)

Smart IoT Actuator using llama3.2:3b model

SYSTEM REAL DATA


- DHT22 ==> Temp: 21.6°C, Humidity: 29.0%
- BMP280 => Temp: 22.5°C, Pressure: 909.99hPa
- Button pressed

>> llama3.2:3b Response: {"red_led": true, "yellow_led": false, "green_led": false}


>> SLM decision: Red=True, Yellow=False, Green=False

SYSTEM ACTUATOR STATUS


- Red LED is on
- Yellow LED is off
- Green LED is off

Interacting with IoT Systems, using Natural Language Commands

Now, let’s transform our IoT monitoring setup into an interactive assistant that accepts
natural language commands and queries. Instead of autonomously monitoring conditions, the

674
system will now respond to our requests in real-time.

What will change?

Original System (Autonomous)

• Continuously monitored sensors


• Automatically decided LED states based on pre-defined rules
• No user interaction

New System (Interactive)

• Waits for user commands


• Accepts natural language queries and commands
• Provides conversational responses
• Executes actions based on user requests
• Displays comprehensive system status

675
How It will Work

1. Interactive Loop

The system runs in a continuous loop:

User Input → Sensor Reading → SLM Analysis → LED Control → Status Display → Wait for Next Inp

2. Dual-Purpose Response

The SLM now returns a JSON object with TWO components:

{
"message": "Helpful text response to the user",
"leds": {
"red_led": false,
"yellow_led": true,
"green_led": false
}
}

• message: What the assistant tells you (conversational response)


• leds: What the assistant does (LED control)

3. Command Types

A. Information Queries
Ask questions about sensor readings - LEDs remain unchanged.
Examples:

• “What’s the current temperature?”


• “What are the actual conditions?”
• “What’s the humidity level?”
• “Is the button pressed?”

Response format:

{
"message": "The current temperature is 21.5°C from DHT22 and 22.3°C from BMP280.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false} // keeps current state
}

676
B. Direct LED Commands
Tell the system which LEDs to turn on/off.
Examples:

• “Turn on the yellow LED”


• “Turn on all LEDs”
• “Turn off all LEDs”
• “Turn on the red and green LEDs”

Response format:

{
"message": "Yellow LED turned on.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false}
}

C. Conditional Commands
Commands that depend on sensor readings or button state.
Examples:

• “If the temperature is above 20°C, turn on the yellow LED”


• “If the button is pressed, turn on the red LED”
• “If the button is not pressed, turn on the green LED”

Response format:

{
"message": "Temperature is 21.5°C, which is above 20°C. Yellow LED turned on.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false}
}

D. Toggle/Switch Commands
Commands that change LED states based on current conditions.
Examples:

• “If the button is pressed, switch the LED conditions”


• “Toggle all LEDs”

Response format:

677
{
"message": "Button is pressed. LED states switched.",
"leds": {"red_led": true, "yellow_led": false, "green_led": false} // inverted from curren
}

E. Analysis Queries
Ask the SLM to analyze sensor data.
Examples:

• “Based on the conditions, would we have rain?”


• “Is the weather getting warmer?”
• “What do the sensor readings indicate?”

Response format:

{
"message": "Based on pressure of 910.18hPa and humidity of 28.8%, conditions are dry. Rain
"leds": {"red_led": false, "yellow_led": false, "green_led": true} // keeps current state
}

Key Functions

create_interactive_prompt()

• Creates a comprehensive prompt with:


– Current sensor readings
– Current LED states
– User’s request
– Examples of how to respond
– Rules for the SLM to follow

parse_interactive_response()

• Extracts the message from the SLM


• Extracts the LED control commands
• Handles JSON parsing errors gracefully

678
display_system_status()

• Shows comprehensive system status:


– All sensor readings
– Button state
– LED states (with visual indicators � �)

interactive_mode()

• Main loop that:


– Accepts user input
– Reads current sensor data
– Sends request to SLM
– Parses response
– Controls LEDs
– Displays results

Here is the complete code: slm_act_leds_interactive.py

Running the System

python slm_act_leds_interactive.py

Tests

============================================================
IoT Environmental Monitoring System - Interactive Mode
Using Model: llama3.2:3b
============================================================

Commands you can try:


- What's the current temperature?
- What are the actual conditions?
- Turn on the yellow LED
- If temperature is above 20°C, turn on yellow LED
- If button is pressed, turn on red LED
- Turn on all LEDs
- Turn off all LEDs

679
- Will it rain based on current conditions?
- Type 'status' to see system status
- Type 'exit' or 'quit' to stop
============================================================

You: what's the current temperature?


Assistant: [Thinking...]
Assistant: The current temperature is 21.5°C from DHT22 and 22.3°C from BMP280.

LED Update: Red=OFF, Yellow=OFF, Green=ON

You: if temperature is above 20°C, turn on yellow LED


Assistant: [Thinking...]
Assistant: Temperature is 21.5°C, which is above 20°C. Yellow LED turned on.

LED Update: Red=OFF, Yellow=ON, Green=OFF

You: will it rain?


Assistant: [Thinking...]
Assistant: Based on pressure of 910.18hPa and humidity of 28.8%, the air is
relatively dry. Rain is unlikely in the immediate future. LEDs unchanged.

LED Update: Red=OFF, Yellow=ON, Green=OFF

You: status
============================================================
SYSTEM STATUS
============================================================
DHT22 Sensor: Temp = 21.5°C, Humidity = 28.8%
BMP280 Sensor: Temp = 22.3°C, Pressure = 910.18hPa
Button: NOT PRESSED

LED Status:
Red LED: � OFF
Yellow LED: � ON
Green LED: � OFF
============================================================

You: exit
Exiting interactive mode. Goodbye!

680
Special Commands

• status: Display comprehensive system status without calling the SLM


• exit or quit: Exit the interactive mode

Languages

One advantage of using SLMs is that, once multilingual models are used (as Llama 3,2 or
Gemma), the user can choose the better language for them, independent of the language used
during coding.

Advantages of This Approach

1. Natural Language Interface: Users don’t need to know programming or specific


commands
2. Flexible Control: Can handle complex conditional logic based on sensor data
3. Informational: Can answer questions about the system without changing states
4. Context-Aware: SLM understands current conditions and makes appropriate decisions
5. Conversational: Provides helpful feedback about what it’s doing

Error Handling

• If sensor data cannot be read, the system will notify you and wait for the next command
• If the SLM response cannot be parsed, the system keeps LEDs in their current state

681
Tips for Best Results

1. Be specific: “Turn on the yellow LED” works better than “turn on the light”
2. Use conditions clearly: “If temperature is above 20°C” is clearer than “when it’s hot”
3. Ask for status: Use the status command frequently to verify system state
4. One LED at a time: Unless you specifically say “all LEDs”, the system defaults to one
LED on

Flow Diagram

The development of our Interactive IoT-SLM system can also be followed using the
Jupyter Notebook: SLM_IoT.ipynb.

682
Adding Data Logging

Now, we will develop an enhanced version that adds data logging, analysis, and historical query
capabilities. The system automatically logs all sensor readings and commands, allowing it to
analyze trends and query historical data using natural language.

Key Features

1. Automatic Background Logging

• Logs sensor readings every 60 seconds


• Runs in background thread
• Stores data in CSV files

2. CSV Data Storage

• sensor_readings.csv: All sensor data + LED states


• command_history.csv: All commands and responses

683
3. Statistical Analysis

• Min/max/average calculations
• LED state change tracking
• Button press counting
• Trend analysis

4. Natural Language Queries

Ask questions like:

• “Show me temperature trends for the last 24 hours”


• “What was the average humidity today?”
• “How many times was the button pressed?”

Running the System

python slm_act_leds_with_logging.py

Example Queries

Real-time:

• “What’s the current temperature?”


• “Turn on the yellow LED”

Historical:

• “Show me temperature trends”


• “What was the average humidity today?”
• “How many times was the button pressed?”

Built-in commands:

• status - Show current system status


• stats - Show 24-hour statistics
• exit - Stop system

684
Data Files

sensor_readings.csv

timestamp,temp_dht,humidity,temp_bmp,pressure,
button_pressed,red_led,yellow_led,green_led

command_history.csv

timestamp,user_command,slm_response,red_led,yellow_led,green_led

Tips

1. Let system run 1+ hour for meaningful trends


2. Use specific time frames: “last 6 hours”
3. Use stats command for quick overview
4. CSV files can be opened in Excel

The complete scripts for the datalogger version arehere: data_logger.py and
slm_act_leds_with_logging.py

685
Flow Diagram

686
Examples:

687
Prompt Optimization and Efficiency

We can modify for example the interactive_mode(MODEL, USER_INPUT) function, to include


at its end, some statistics about the model’s performance:
...
# Display Latency
print(f"\nTotal Duration: {(response['total_duration']/1e9):.2f} seconds")
print(f"prompt_eval_duration: {(response['prompt_eval_duration']/1e9):.2f} s")
print(f"load_duration: {(response['load_duration']/1e9):.2f} s")
print(f"eval_count: {response['eval_count']}")
print(f"eval_duration: {(response['eval_duration']/1e9):.2f} s")

688
print(f"eval_rate: {response['eval_count']/(response['eval_duration']/1e9):.2f} \
tokens/s")

For example, running

USER_INPUT = "Turn off all LEDs"


interactive_mode(MODEL, USER_INPUT)

We will get:

Total Duration: 88.81 seconds


prompt_eval_duration: 80.41 s
load_duration: 1.98 s
eval_count: 32
eval_duration: 6.41 s
eval_rate: 4.99 tokens/s

Based on the prompt_eval_duration value, the PROMPT is our main bottleneck, so we must
make it as concise as possible. The SLM must process all of the context before it can generate
the first output token.

Quick Solution:

1. Drastically Shorten the Prompt

• Condense System Status: Remove unnecessary descriptive text and present the status
information in a compact, structured format.

– Example (Before):
CURRENT SYSTEM STATUS:
- DHT22: Temperature 25.5°C, Humidity 45.1%
- BMP280: Temperature 25.6°C, Pressure 1012.34hPa
- Button: NOT PRESSED
- Red LED: OFF
- Yellow LED: OFF
- Green LED: ON

– Example (After):
STATUS: DHT22=22.5°C/65.0% BMP280=22.3°C/1013.25hPa Button=OFF LEDs:R=OFF/Y=ON/G=OFF

689
• Reduce Examples: While examples are crucial for instruction-following, eliminate
redundancy. Keep only the most diverse and representative examples. The current
prompt is very long due to verbose examples and instructions. Focus on the single-shot
example that shows the required JSON output format.
• Simplify Instructions: Make the instructions as direct and short as possible. Use
keywords instead of full sentences where clarity is maintained.
• System message: It defines the assistant’s behavior and should sent once at initialization,
not at PROMPT
SYSTEM_MESSAGE = """You are an IoT assistant controlling an environmental
monitoring system with LEDs.

Respond with JSON only:


{"message": "your helpful response", "leds": {"red_led": bool, "yellow_led":
bool, "green_led": bool}}

RULES:

- Information queries: keep current LED states unchanged


- LED commands: update LEDs as requested
- Conditional commands (if/when): evaluate condition from sensor data first
- Only ONE LED should be on at a time UNLESS user explicitly says "all"
- Be concise and conversational

Always respond with valid JSON containing both "message" and "leds" fields."""

2. chat API Instead of generate

• Uses [Link]() to maintain conversation context


• System message sent only once at initialization
• Subsequent messages are much shorter (only status + user query)

Smart Context Management

• Keeps last four conversation exchanges (8 messages) + system message


• Prevents context from growing too large over time

690
Model Pre-loading

• Loads model into memory before first query


• Eliminates the 2-second load delay on first run

Let’s re-create the functions or add new as bellow:

New Code

Original Functions:

• create_interactive_prompt() - Now creates compact user messages (was ~800 tokens,


now ~100 tokens)
• slm_inference() - Now uses [Link]() API instead of [Link]()
• parse_interactive_response() - Unchanged
• display_system_status() - Unchanged
• interactive_mode() - Updated with all optimizations

New functions:

• SYSTEM_MESSAGE - Constant for system prompt (sent only once)


• preload_model() - Helper to pre-load model into memory

Key Changes Under the Hood:

1. create_interactive_prompt() now returns a compact format:


• Before: Long prompt with 8 examples (~800 tokens)
• After: Compact status line + user input (~100 tokens)
2. slm_inference() now uses chat API:
• Before: [Link](model, prompt) - regenerates context each time
• After: [Link](model, messages) - maintains conversation context
3. System prompt is sent only once at initialization, not with every query

Using the optimized code: slm_act_leds_interactive_optimized.py, we get as a response:

691
The Prompt evaluation time was reduced drastically!

Using Pydantic

Using Pydantic is a robust way to improve the reliability, efficiency, and maintainability of
our system, mainly since we rely on the LLM to output precise JSON.
Here’s how Pydantic can help and what you would need to do:

692
1. How Pydantic Reduces Latency and Improves Reliability

Pydantic doesn’t directly speed up the model’s token generation, but it can indirectly reduce
latency and eliminate error-handling overhead by enabling cleaner, faster parsing and
more robust communication.

A. Reliable and Fast JSON Parsing


The most critical benefit is moving from the general, potentially brittle [Link](response_text)
in our parse_interactive_response function to Pydantic’s fast validation engine.

• Pydantic’s JSON Parsing is C-Accelerated: Pydantic is built on top of the fast


pydantic-core library written in Rust, which is significantly faster than Python’s
standard json module, especially for large or complex data structures.
• Structured Output: Our current code manually handles potential JSON formatting
issues (such as the jsoncode block delimiters) and then uses a try/except block to
extract fields. Pydantic handles all of this automatically and throws a clean error if the
structure is wrong.

B. Stronger Instruction for the SLM


The model is more likely to generate the exact JSON required when you give it a formal schema
to follow.

• JSON Schema: Pydantic can generate a JSON Schema from your Python classes.
We can include this schema directly in our prompt, which acts as an unambiguous,
machine-readable instruction for the SLM. This often leads to fewer errors in the model’s
output, reducing the need for costly retries or complex string manipulation.

C. Cleaner Python Code


It simplifies our parse_interactive_response function immensely, making your code easier
to read and debug.

2. Implementing Pydantic in the Code

Let’s test it on the Optimized Interactive version: slm_act_leds_optimized.py

We cshould modify the workflow in three main steps:

693
Step 1: Define the Pydantic Models

We must replace the implicit structure with explicit Pydantic models:

from pydantic import BaseModel, Field

class LEDControl(BaseModel):
"""LED control configuration."""
red_led: bool = Field(description="Red LED state (on/off)")
yellow_led: bool = Field(description="Yellow LED state (on/off)")
green_led: bool = Field(description="Green LED state (on/off)")

class AssistantResponse(BaseModel):
"""Complete assistant response with message and LED control."""
message: str = Field(description="Helpful response to the user")
leds: LEDControl = Field(description="LED control configuration")

Step 2: Update the SLM inference function

We would update the create_interactive_prompt to include the schema:

def slm_inference(messages, MODEL):


"""Send chat request to Ollama using chat API with structured output
(Pydantic)."""
response = [Link](
model=MODEL,
messages=messages,
format=AssistantResponse.model_json_schema() # Using Pydantic schema
)
return response

Step 3: Update the Parsing Function

Our parse_interactive_response becomes much cleaner and more reliable:

def parse_interactive_response(response_text):
"""Parse the interactive SLM response using Pydantic (guaranteed valid)."""
try:
# Parse directly into Pydantic model - guaranteed valid JSON structure

694
data = AssistantResponse.model_validate_json(response_text)

# Extract values from Pydantic model


message = [Link]
red_led = [Link].red_led
yellow_led = [Link].yellow_led
green_led = [Link].green_led

return message, (red_led, yellow_led, green_led)

except Exception as e:
print(f"Error parsing response: {e}")
print(f"Response was: {response_text}")
return "Error: Could not parse SLM response.", (False, False, False)

With such modifications, the final latency was reduced from 90 to around 60 seconds

Note that when we sent two different commands to turn on the LEDs, the new one did not
turn off the previous one. This is due to the change in the rules. If we want, we can modify it
to match whatever we wish to.

695
The final code can be found on GitHub: slm_act_leds_interactive_pydantic.py
and in notebook SLM_IoT.ipynb

In short, using Pydantic over tradicional approuch with JSON, we have as benefits:

• No more parsing errors - Guaranteed valid JSON


• Cleaner output - No markdown code blocks
• Type safety - IDE autocomplete and type checking
• Faster generation - Constrained decoding is more efficient
• Better validation - Pydantic catches invalid values

Next Steps

This chapter involved experimenting with simple applications and verifying the feasibility of
using an SLM to control IoT devices. The final result is far from usable in the real world, but
it can serve as a starting point for more interesting applications. Below are some observations
and suggestions for improvement:

• SLM responses can be probabilistic and inconsistent. To increase reliability, consider


implementing a confidence threshold or voting system using multiple prompts/responses.
• Try to add data validation and sanity checks for sensor readings before passing them to
the SLM.
• Apply Structured Response Parsing as discussed early. Future improvements in this
approuch could include:

– Add more sophisticated validation rules


– Implement command history tracking
– Add support for compound commands
– Integrate with the logging system
– Add user permission levels
– Implement command templates for common operations

• Consider implementing a fallback mechanism when SLM responses are ambiguous or


inconsistent.
• Study using RAG and fine-tuning to increase the system’s reliability when using very
small models.
• Consider adding input validation for user commands to prevent potential issues.
• The current implementation queries the SLM for every command. We did it to study
how SLMs would behave. We should consider implementing a caching mechanism for
common queries.

696
• Some simple commands could be handled without SLM intervention. We can do it
programmatically.
• Consider implementing a proper state machine for LED control to ensure consistent
behavior.
• Implement more sophisticated trend analysis using statistical methods.
• Add support for more complex queries combining multiple data points.

697
Conclusion

This chapter has demonstrated the progressive evolution of an IoT system from basic sensor
integration to an intelligent, interactive platform powered by Small Language Models. Through
our journey, we’ve explored several key aspects of combining edge AI with physical computing:

Key Achievements

1. Progressive System Development


• Started with basic sensor integration and LED control
• Advanced to SLM-based analysis and decision making
• Implemented natural language interaction
• Added historical data logging and analysis
• Created a complete interactive system
2. SLM Integration Insights
• Demonstrated the feasibility of using SLMs for IoT control
• Explored different models and their capabilities
• Implemented various prompting strategies
• Handled both real-time and historical data analysis
3. Practical Learning Outcomes
• Hardware-software integration techniques
• Real-time sensor data processing
• Natural language command interpretation
• Data logging and trend analysis
• Error handling and system reliability

Challenges and Limitations

Our implementation revealed several important challenges:

1. SLM Reliability
• Probabilistic nature of responses
• Consistency issues in decision making

698
• Need for better validation and verification
2. System Performance
• Response time considerations
• Resource usage on edge devices
• Efficiency of data logging and analysis
3. Architectural Constraints
• Simple state management
• Basic error handling
• Limited data validation

Final Thoughts

While this implementation demonstrates the potential of combining SLMs with IoT systems,
it also highlights the exciting possibilities and challenges ahead. Though experimental, the
system we’ve built provides a solid foundation for understanding how edge AI can enhance IoT
applications. As SLMs evolve and improve, their integration with physical computing systems
will likely become more robust and practical for real-world applications.
This chapter has shown that, despite current limitations, SLMs can provide intelligent, natural-
language interfaces to IoT systems, opening new possibilities for human-machine interaction in
the physical world.
The future of IoT systems is shaped by intelligent, edge-based solutions that combine AI’s
power with the practicality of physical computing.

Resourses

• Python Scripts and notebook

699
Advancing EdgeAI: Beyond Basic SLMs

Exploring CoT Prompting, Agents, Function Calling, RAG, and more.

Figure 17: Image from author, with a Raspberry Pi from ImageFX - prompt - Create a cartoon
image with a single Raspberry Pi with a white background

Building on the foundation established in previous chapters of “Edge AI Engineering,” this


chapter explores the limitations of Small Language Models (SLMs) and advanced techniques to
enhance their capabilities at the edge. While we’ve demonstrated the feasibility of running
SLMs on devices like the Raspberry Pi, practical applications often require addressing inherent
limitations in these models.
As we explore techniques like chain-of-thought prompting, agent architectures, function call-
ing, response validation, and Retrieval-Augmented Generation (RAG), we’ll see how clever
engineering and system design can mitigate SLMs’ limitations.

700
Understanding SLM Limitations

Small Language Models, while impressive in their ability to run on edge devices, face several
key limitations:

1. Knowledge Constraints

SLMs have limited knowledge based on their training data, often outdated and incomplete.
Unlike their larger counterparts, they cannot store the vast information needed for comprehensive
expertise across all domains.
Let’s run the below example to verify this limitation.

import ollama

response = [Link](
model="llama3.2:1b",
prompt="Who won the 2024 Summer Olympics men's 100m sprint final?"
)
print(response['response'])

The output of the previous code will likely show hallucination or admission of not knowing, as
in the case below:

This constraint could be solved simply by having an Agent search the Internet for the answer
or using Retrieval-Augmented Generation (RAG), as we will see later.

701
2. Reasoning Limitations

Complex reasoning tasks often exceed the capabilities of SLMs, which struggle with multi-
step logical deductions, mathematical computations, and a nuanced understanding of context.
Agents can be used to mitigate such limitations.
For example, let’s reuse the previous code and ask to the SLM to multiply two numbers :

import ollama

response = [Link](
model="llama3.2:3b",
prompt="Multiply 123456 by 123456"
)
print(response['response'])

The response is wrong; once the multiplication result should be 15,241,383,936. This is
expected once the language models are not suitable for mathematical computations. Still, we
can use an “agent” to determine whether a user asks for multiplication or a general question.
We will learn how to create an agent later.

3. Inconsistent Outputs

SLMs may produce inconsistent responses to the same query, making them unreliable for
critical applications requiring deterministic outputs. Several enhancements, such as Function
Calling and Response Validation, can improve reliability.

4. Domain Specialization

SLMs perform worse than specialized models in domain-specific tasks like visual recognition or
time-series analysis. Fine-tuning can adapt models to specific domains or tasks, improving
performance for targeted applications.

702
Techniques for Enhancing SLM at the Edge

Small Language Models (SLMs) offer remarkable capabilities for edge devices, but various
techniques can significantly enhance their effectiveness. Here, we present a comprehensive
framework for optimizing SLMs on resource-constrained devices like the Raspberry Pi, organized
from fundamental to advanced approaches.
We will divide those technics into 3 segments:

• Fundamentals: Optimizing Prompting Strategies

– Chain-of-Thought Prompting
– Few-Shot Learning
– Task Decomposition

• Intermediate: Building Intelligence Systems

– Building Agents with SLMs


– General Knowledge Router
– Function Calling
– Response Validation

• Advanced: Extending Knowledge and Specialization

– Retrieval-Augmented Generation (RAG)


– Fine-Tuning for Domain Specialization

• Integration: Combining Techniques for Optimal Performance

The true power of these techniques emerges when they’re strategically combined:

1. Agent Architecture with RAG: Create agents that can access both tools and knowl-
edge bases
2. Validation-Enhanced RAG: Apply response validation to ensure RAG outputs are
accurate
3. Fine-Tuned Routers: Use specialized fine-tuned models to handle routing decisions
4. Chain-of-Thought with Function Calling: Combine reasoning traces with structured
outputs

For example, a comprehensive weather monitoring system, as we introduced in the chapter


“Experimenting with SLMs for IoT Control,” might use the following:

• RAG to access historical weather patterns and interpretation guides

703
• Function calling to structure sensor data analysis
• Response validation to verify recommendations
• Task decomposition to handle complex multi-part weather analysis

Optimizing Prompting Strategies

Chain-of-Thought Prompting

Chain-of-thought prompting encourages SLMs to break down complex problems into step-by-
step reasoning, leading to more accurate results:

def solve_math_problem(problem: str) -> str:


prompt = f"""You are a math expert. Solve the problem using chain-of-thought reasoning.

Follow this structure:


1. Restate the problem in your own words.
2. Identify all given quantities and what is being asked.
3. Plan the solution (what formulas or steps you will use).
4. Execute the plan step by step, showing intermediate calculations.
5. Double-check the result for reasonableness.
6. On the last line, write only: Answer: <final answer>

Problem:
{problem}
"""
response = [Link](model="llama3.2:3b", prompt=prompt)
return response["response"]

This technique significantly improves performance on reasoning tasks by emulating human


problem-solving approaches.

Few-Shot Learning

Few-shot learning provides examples within the prompt, helping SLMs understand the expected
response format and reasoning pattern:

def classify_sentiment(text):
prompt = f"""
Task: Classify the sentiment of the text as positive, negative, or neutral.
Examples:

704
Text: "I love this product, it works perfectly!"
Sentiment: positive
Text: "This is the worst experience I've ever had."
Sentiment: negative
Text: "The package arrived on time."
Sentiment: neutral
Text: "{text}"
Sentiment:
"""
response = [Link](model="llama3.2:1b", prompt=prompt)
return response['response'].strip()

This approach is particularly effective for classification tasks and standardized outputs.

Task Decomposition

For complex tasks, breaking them into smaller subtasks helps SLMs manage complexity:

def analyze_product_review(review):
# Step 1: Extract main points
points_prompt = f"Extract the main points from this product review: {review}"
points_response = [Link](model="llama3.2:1b", prompt=points_prompt)
main_points = points_response['response']

# Step 2: Determine sentiment


sentiment_prompt = f"Determine the overall sentiment of this review: {review}"
sentiment_response = [Link](model="llama3.2:1b",
prompt=sentiment_prompt)
sentiment = sentiment_response['response']

# Step 3: Identify improvement suggestions


improvements_prompt = f"What suggestions for improvement can be found in \
this review? {review}"
improvements_response = [Link](model="llama3.2:1b",
prompt=improvements_prompt)
improvements = improvements_response['response']

# Final synthesis
final_prompt = f"""
Create a concise analysis of this product review based on:
Main points: {main_points}

705
Overall sentiment: {sentiment}
Improvement suggestions: {improvements}
"""
final_response = [Link](model="llama3.2:1b",
prompt=final_prompt)
return final_response['response']

This technique distributes cognitive load across multiple simpler prompts, enabling SLMs to
handle tasks that might otherwise exceed their capabilities.

Building Agents with SLMs

To address some of these limitations, we can develop agents that leverage SLMs as part of a
more extensive system with additional capabilities.

What is an Agent? It is an AI model capable of reasoning, planning, and


interacting with its environment. It can be called an Agent because it has an
agency that can interact with the environment.

Let’s think about the multiplication problem that we faced before. A minimal agent can be
used for that.
An agent is a system that uses an AI Model as its core reasoning engine to:

• Understand natural language: (1) Interpret and respond to human instructions


meaningfully.
• Reason and plan: (2) Analyze information, make decisions, and devise problem-solving
strategies.
• Interact with its environment: (3 and 4) Gather information, take actions, and
observe the results of those actions.

706
For example, if it is a multiplication, we can use a Python function as a “tool” to calculate it,
as shown in the diagram:

707
This example shows a minimal agentic workflow: the model decides whether to use
a tool for multiplication or answer normally.

Our code works through the following steps:

1. User Input: The user types a query like “What is 7 times 8?” or “What is the capital
of France?”
2. Process Query: The process_query() function handles the input and decides what to
do with it.

708
3. Classification: The ask_ollama_for_classification() function sends the user’s
query to the SLM (using Ollama) with a prompt asking it to classify whether the query
is requesting multiplication or asking a general question.
4. Decision: Based on the SLM’s classification:

• If it’s a multiplication request, the SLM also extracts the numbers, and we use our
multiply() function.
• If it’s a general question, we send the original query to the SLM for a direct answer.

5. Response: The system returns either the multiplication result or the SLM’s answer to
the general question.

Here’s a Python script that creates a simple agent (or router) between multiplication operations
and general questions as described:

import requests
import json

# Configuration
OLLAMA_URL = "[Link]
MODEL = "llama3.2:3b" # You can change this to any model you have installed
VERBOSE = True

def ask_ollama_for_classification(user_input):
"""
Ask Ollama to classify whether the query is a multiplication request or a \
general question.
"""
classification_prompt = f"""
Analyze the following query and determine if it's asking for multiplication \
or if it's a general question.

Query: "{user_input}"

If it's asking for multiplication, respond with a JSON object in this format:
{{
"type": "multiplication",
"numbers": [number1, number2]
}}

If it's a general question, respond with a JSON object in this format:


{{
"type": "general_question"

709
}}

Respond ONLY with the JSON object, nothing else.


"""

try:
if VERBOSE:
print(f"Sending classification request to Ollama")

response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": classification_prompt,
"stream": False
}
)

if response.status_code == 200:
response_text = [Link]().get("response", "").strip()
if VERBOSE:
print(f"Classification response: {response_text}")

# Try to parse the JSON response


try:
# Find JSON content if there's any surrounding text
start_index = response_text.find('{')
end_index = response_text.rfind('}') + 1
if start_index >= 0 and end_index > start_index:
json_str = response_text[start_index:end_index]
return [Link](json_str)
return {"type": "general_question"}
except [Link]:
if VERBOSE:
print(f"Failed to parse JSON: {response_text}")
return {"type": "general_question"}
else:
if VERBOSE:
print(f"Error: Received status code {response.status_code} \
from Ollama.")
return {"type": "general_question"}

710
except Exception as e:
if VERBOSE:
print(f"Error connecting to Ollama: {str(e)}")
return {"type": "general_question"}

def multiply(a, b):


"""
Perform multiplication and return a formatted response.
"""
result = a * b
return f"The product of {a} and {b} is {result}."

def ask_ollama(query):
"""
Send a query to Ollama for general question answering.
"""
try:
if VERBOSE:
print(f"Sending query to Ollama")

response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": query,
"stream": False
}
)

if response.status_code == 200:
return [Link]().get("response", "")
else:
return f"Error: Received status code {response.status_code} \
from Ollama."

except Exception as e:
return f"Error connecting to Ollama: {str(e)}"

def process_query(user_input):
"""
Process the user input by first asking Ollama to classify it,
then either performing multiplication or sending it back as a

711
general question.
"""
# Let Ollama classify the query
classification = ask_ollama_for_classification(user_input)

if VERBOSE:
print("Ollama classification:", classification)

if [Link]("type") == "multiplication":
numbers = [Link]("numbers", [0, 0])
if len(numbers) >= 2:
return multiply(numbers[0], numbers[1])
else:
return "I understood you wanted multiplication, but couldn't \
extract the numbers properly."
else:
return ask_ollama(user_input)

def main():
"""
Main function to run the agent interactively.
"""
print("Ollama Agent (Type 'exit' to quit)")
print("-----------------------------------")

while True:
user_input = input("\nYou: ")

if user_input.lower() in ["exit", "quit", "bye"]:


print("Goodbye!")
break

response = process_query(user_input)
print(f"\nAgent: {response}")

# Example usage
if __name__ == "__main__":
# Set to True to see detailed logging
VERBOSE = True
main()

When we run the script, we can see that, first, the SLM chooses multiplication, passing the

712
numbers entered by the user to the “tool,” which, in this case, is the multiply() function. As
a result, we got 15,241,383,936, which it is correct.

Let’s now enter with another question that has no relation with arithmetic, for example: What
is the capital of Brazil? In this case, the SLM will decide that the query is a general
question and pass it on to the SLM to answer it.

This simple agent (or router) demonstrates the fundamental concept of using an SLM to make
decisions about processing different types of user inputs. It shows both the power of SLMs for
natural language understanding and their limitations in structured tasks.

713
Limitations and Considerations

This agent seems to resolve our problem, but it has several limitations that are common when
working with SLMs:

1. JSON Parsing Issues: SLMs don’t always perfectly format JSON responses as requested.
The code includes error handling for this.
2. Classification Reliability: The SLM might not always correctly classify the query,
especially with ambiguous questions.
3. Number Extraction: The SLM might extract numbers incorrectly or miss them entirely.
4. Error Handling: Robust error handling is essential when working with SLMs because
their outputs can be unpredictable.
5. Latency: Significant latency is involved in making multiple calls to the SLM. For example,
for the above simple agent, the latency was about 50s when using the llama3.2:3B on a
Raspberry Pi 5.

Here, you can see the SLM latency (simple query) per device (in tokens/s):

Model Raspi 5 (Cortex A-76) PC (i7) Mac (M1 Pro)


Gemma3:4b 3.8 8.7 39
Llama3.2:3b 5.5 12 63
Llama3.2:1b 7.5 19.5 111
Gemma3:1b 12 22.45 91

In my simple tests, the 1B models struggled to classify the tasks correctly. The the
3B and 4B models worked fine

Improvements

To create a more robust agent, we can, for example:

1. Expand Capabilities: Add support for more operations (addition, subtraction, division).
2. Better Error Handling: Improve fallback mechanisms when the SLM fails to extract
numbers or classify correctly.
3. Model Preloading: Initialize the model at startup to reduce latency.
4. Adding Regex Fallbacks: Use regular expressions as a fallback to extract numbers
when the SLM fails.
5. Context Preservation: Maintain conversation context for multi-turn interactions.

A more robust script can be used with the above improvements. The diagram shows how it
would work:

714
Figure 18: image-20250322113135882

The diagram illustrates the key components of the system:

1. Initialization:
• The system starts by initializing both models in parallel threads
• This prevents cold starts and reduces latency
2. Query Processing Flow:

715
• User input is first sent to a classification step
• A model (llama3.2:3B) determines if it’s a calculation or a general question (we can
choose a different model here).
• If it’s a calculation:
– The system extracts the operation type and numbers
– Numbers are converted from strings to floats
– The appropriate calculation is performed
– Results are formatted with comma separators (e.g., 1,234,567.89)
• If it’s a general question:
– The query is sent to the main model (llama3.2:3b) for answering (we can choose
a different model here)
3. Optimizations (highlighted in the subgraph):
• Persistent HTTP session for connection reuse
• Keep-alive parameter to prevent model unloading
• Simplified classification prompt for faster processing
• Using a smaller model for the classification task
• Rule-based fallback logic if the model classification fails

The main performance improvements come from:

1. Keeping models loaded in memory


2. Using connection pooling
3. Simplifying the classification task
4. Using a smaller model for classification
5. Initializing models in parallel

This approach maintains the intelligent classification capability while significantly reducing
execution time compared to the original implementation.
Runing the script [Link], we get correct results with reduced latency
of about 60%.

716
General Knowledge Router

Remember when we asked our SLM: Who won the 2024 Summer Olympics men's 100m
sprint final? We could not receive an answer because the modes were trained with
information previously in late 2023.
To solve this issue, let’s build a more advanced agent to classify whether it should use its
knowledge to answer a question or fetch updated information from the Internet. This addresses
a key limitation of Small Language Models: their knowledge cutoff date.
The general architecture of our agent will be similar to the calculator, but now, we will use a
web search API as a tool.

717
This agent addresses a critical limitation of SLMs - their knowledge cutoff date - by determining
when to use the model’s built-in knowledge versus when to search for up-to-date information
from the web.

How it works:

Uses SLM for Classification: Relies entirely on the SLM to determine whether a query
needs web search or can be answered from the model’s knowledge.
Provides Date Context: This section supplies the current date to help the SLM make
informed decisions about whether information is outdated.
Integrates Tavily Search: Uses Tavily’s powerful search API to find relevant information
for queries that need external data.
Handles Timeouts: Includes fallback mechanisms when the model takes too long to respond.
Maintains Source Attribution: Clearly indicates to the user whether the answer comes
from the model’s knowledge or web search.
Let’s run the script: [Link]
But first, we should install the required libraries:

pip install requests


pip install tavily-python

Replace "tvly-YOUR_API_KEY" with your actual Tavily API key.

Why Tavily is Superior for This Use Case

1. Built for RAG: Tavily is specifically designed for retrieval-augmented gener-


ation, making it perfect for our knowledge router.
2. High-Quality Results: It prioritizes reputable sources and provides context-
relevant results.
3. Built-in Summarization: The API can provide an AI-generated summary
of search results, giving an additional layer of processing before your SLM.
4. Simple Integration: Clean API with straightforward responses that are easy
to parse.
5. Generous Free Tier: 1,000 free searches is plenty for testing and personal
use.

718
Runing the script and entering with the same questions that could not be answered before,
we now have: Noah Lyles won the men's 100m sprint final at the 2024 Summer
Olympics. He set a new personal best time of 9.79 seconds. This victory
marked the United States' first win in the event since 2004.

When the user enters a common-knowledge question, the agent will send it directly to the SLM.
For example, if the user asks, "Who is Albert Einstein?", we get:

719
Improving Agent Reliability

There are several ways to enhance an agent’s reliability. One is to implement effective, approved,
structured function calling, which makes agents’ responses more consistent and predictable.

1. Function Calling with Pydantic

In the SLM chapter, we explored function calling when we created an app where the user enters
a country’s name and gets, as an output, the distance in km from the capital city of such a
country and the app’s location.

720
Once the user enters a country name, the model will return the name of its capital city (as a
string) and the latitude and longitude of such city (in float). Using those coordinates, the app
used a simple Python library (haversine) to calculate the distance between those 2 points.
The critical library used was Pydantic (and instructor), a robust data validation and settings
management library engineered by Python to enhance the robustness and reliability of our
codebase. In short, Pydantic helps ensure that the model’s response will always be consistent.
Function calling can improve an agent’s reliability by ensuring structured outputs and clear
tool selection logic. Here’s a generic template about how we can implement it :

import time
from haversine import haversine
from pydantic import BaseModel, Field
from ollama import chat

MODEL = 'llama3.2:3B' # The name of the model to be used


mylat = -33.33 # Latitude of Santiago de Chile
mylon = -70.51 # Longitude of Santiago de Chile

class CityCoord(BaseModel):
city: str = Field(..., description="Name of the city")
lat: float = Field(..., description="Decimal Latitude of the city")
lon: float = Field(..., description="Decimal Longitude of the city")

def calc_dist(country, model=MODEL):

start_time = time.perf_counter() # Start timing

# Ask Ollama for structured data


response = chat(
model=model,
messages=[{
"role": "user",
"content": f"Return the capital city of {country}, \
with its decimal latitude and longitude."
}],
format=CityCoord.model_json_schema(), # Structured JSON format
options={"temperature": 0}
)

resp = CityCoord.model_validate_json([Link])

distance = haversine((mylat, mylon), ([Link], [Link]), unit='km')

721
end_time = time.perf_counter() # End timing
elapsed_time = end_time - start_time # Calculate elapsed time

print(f"\nSantiago de Chile is about {int(round(distance, -1)):,} \


kilometers away from {[Link]}.")
print(f"[INFO] ==> {MODEL}: {elapsed_time:.1f} seconds")

# Test
calc_dist('france')
calc_dist('colombia')
calc_dist('united states')

2. Response Validation

Response validation is crucial to developing and deploying AI agents powered by language


models. Here are key points regarding LLM validation for agents: Types of Validation

• Response Relevancy: Determines if the LLM output addresses the input informatively
and concisely.
• Prompt Alignment: Check if the LLM output follows instructions from the prompt
template.
• Correctness: Assesses factual accuracy based on ground truth.
• Hallucination Detection: Identifies fake or made-up information in LLM outputs.

Adding validation prevents incorrect or harmful responses, and here, we can test it with a
simple script:

import ollama
import json

722
def validate_response(query, response):
"""Validate that the response is appropriate for the query"""
validation_prompt = f"""
User query: {query}
Generated response: {response}

Evaluate if this response:


1. Directly addresses the user's query
2. Is factually accurate to the best of your knowledge
3. Is helpful and complete

Respond in the following JSON format:


{{
"valid": true/false,
"reason": "Explanation if invalid",
"score": 0-10
}}
"""

try:
validation = [Link](
model="llama3.2:3b",
prompt=validation_prompt
)

result = [Link](validation['response'])
return result
except Exception as e:
print(f"Error during validation: {e}")
return {"valid": False, "reason": "Validation error", "score": 0}

# Test
query = "What is the Raspberry Pi 5?"
response = "It is a pie created with raspberry and cooked in an oven"
validation = validate_response(query, response)
print(validation)

723
Retrieval-Augmented Generation (RAG)

RAG systems enhance Small Language Models (SLMs) by providing relevant information from
external sources before generation. This is particularly valuable for edge devices with limited
model sizes, as it allows them to access knowledge beyond their training data without increasing
the model size.

Understanding RAG

In a basic interaction between a user and a language model, the user asks a question, which is
sent as a prompt to the model. The model generates a response based solely on its pre-trained
knowledge. In a RAG process, there’s an additional step between the user’s question and the
model’s response. The user’s question triggers a retrieval process from a knowledge base.

724
The RAG process consists of these key steps:

1. Query Processing: When a user asks a question, the system converts it into an
embedding (a numerical representation).
2. Document Retrieval: The system searches a knowledge base for documents with similar
embeddings.
3. Context Enhancement: Relevant documents are retrieved and combined with the
original query.
4. Generation: The SLM generates a response using both the query and the retrieved
context.

Implementing a Naive RAG System

We will develop two crucial components (scripts) of an RAG system:

1. Creating the Vector Database ([Link]). This script


builds a knowledge base by:
• Loading documents from PDFs and URLs
• Splitting them into manageable chunks
• Creating embeddings for each chunk
• Storing these embeddings in a vector database (Chroma)
2. Querying the Database ([Link]). This script:
• Loads the saved vector database

725
• Accepts user queries
• Retrieves relevant documents based on query similarity
• Combines documents with the query in a prompt
• Generates a response using the SLM

Instalation

Once inside an environment, install the required libraries:

pip install -U 'langchain-chroma'


pip install -U langchain
pip install -U langchain-community
pip install -U langchain-ollama
pip install -U langchain-text-splitter
pip install -U langchain-community pypdf
pip install tiktoken
pip install -U langsmith

Let’s examine how these components work together to implement a RAG system on edge
devices.

Key Components of the Naive RAG System

1. Document Processing

726
def create_vectorstore():
# Load documents from PDFs and URLs
docs_list = []
# [Document loading code]

# Split documents into chunks


text_splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
chunk_size=300, chunk_overlap=30
)
doc_splits = text_splitter.split_documents(docs_list)

# Create embeddings and store in vector database


embedding_function = OllamaEmbeddings(model="nomic-embed-text")
vectorstore = Chroma.from_documents(
documents=doc_splits,
collection_name="rag-edgeai-eng-chroma",
embedding=embedding_function,
persist_directory=PERSIST_DIRECTORY
)

# Persist to disk
[Link]()

This function processes our documents (chunk size of 300 with an overlap of 30), creating a
searchable knowledge base. Notice we’re using OllamaEmbeddings with the nomic-embed-text
model, which can run efficiently on edge devices like the Raspberry Pi.

1. Query Processing and Retrieval

def answer_question(question, retriever):


"""Generate an answer using the RAG system"""
start_time = [Link]()

print(f"Question: {question}")
print("Retrieving documents...")
docs = [Link](question)
docs_content = "\n\n".join(doc.page_content for doc in docs)
print(f"Retrieved {len(docs)} document chunks")

print("Generating answer...")

# Using new LangSmith client and pull_prompt

727
client = Client()
rag_prompt = client.pull_prompt("rlm/rag-prompt")

# Compose the RAG chain


if isinstance(rag_prompt, str):
rag_prompt = ChatPromptTemplate.from_template(rag_prompt)

rag_chain = rag_prompt | llm | StrOutputParser()


answer = rag_chain.invoke({"context": docs_content, "question": question})

end_time = [Link]()
latency = end_time - start_time
print(f"Response latency: {latency:.2f} seconds using model: {local_llm}")

return answer

This function retrieves relevant documents based on the query and combines them with a
specialized RAG prompt to generate a response. The RAG prompt is particularly important
as it tells the model how to use the context documents to answer the question.

1. SLM Integration

# Initialize the LLM


local_llm = "llama3.2:3b"
llm = ChatOllama(model=local_llm, temperature=0)

We’re using Ollama to run the SLM locally on our edge device, in this case using the 3B
parameter version of Llama 3.2.

Advantages of RAG for Edge AI

Using RAG on edge devices offers several significant advantages:

1. Knowledge Extension: RAG allows small models to access knowledge beyond their
training data, effectively extending their capabilities without increasing model size.
2. Reduced Hallucination: By providing factual context, RAG significantly reduces the
likelihood of SLMs generating incorrect information.
3. Up-to-date Information: Unlike the fixed knowledge in a model’s weights, RAG
knowledge bases can be updated regularly with new information.
4. Domain Specialization: RAG can make general SLMs perform like domain specialists
by providing domain-specific knowledge bases.

728
5. Resource Efficiency: RAG allows smaller models (which require less memory and
computation) to achieve performance comparable to much larger models.

Optimizing RAG for Edge Devices

When implementing RAG on resource-constrained edge devices like the Raspberry Pi, consider
these optimizations:

1. Chunk Size: Smaller chunks (300-500 tokens) reduce memory usage during retrieval
and generation.
2. Retrieval Limits: Limit the number of retrieved documents (k=3 to 5) to reduce context
size.
3. Embedding Model Selection: Choose lightweight embedding models like nomic-
embed-text (137M parameters) or all-minilm (23M parameters).
4. Persistent Storage: As shown in our examples, using persistent storage prevents
recomputing embeddings every time that the RAG system is initiated.
5. Query Optimization: Implement query preprocessing to improve retrieval accuracy
while reducing computational load.

def optimize_query(query):
"""Optimize the query for better retrieval results"""
# Remove filler words, focus on key terms
stop_words = {"and", "or", "the", "a", "an", "in", "on", "at", "to",
"for", "with"}
terms = [term for term in [Link]().split() if term not in stop_words]
return " ".join(terms)

Application: Enhanced Weather Station with RAG

Building on our advanced weather station (see the chapter “Experimenting with SLMs for IoT
Control”), we can, for example, integrate RAG to provide more contextual responses about
weather conditions and historical patterns:

729
def weather_station_with_rag(retriever, model="llama3.2:3b"):
# Get current sensor readings
temp_dht, humidity, temp_bmp, pressure, button_state = collect_data()

# Formulate a query for the RAG system based on current readings


query = f"Analysis of temperature {temp_dht}°C, humidity {humidity}%, \
and pressure {pressure}hPa"

# Retrieve relevant context


docs = [Link](query)
context = "\n\n".join(doc.page_content for doc in docs)

# Create a prompt that combines current readings with retrieved context

730
prompt = f"""
Current Weather Station Data:
- Temperature (DHT22): {temp_dht:.1f}°C
- Humidity: {humidity:.1f}%
- Pressure: {pressure:.2f}hPa

Reference Information:
{context}

Based on current readings and the reference information, provide:


1. An analysis of current weather conditions
2. What these conditions typically indicate
3. Recommendations for any actions needed
"""

# Generate response using SLM


llm = ChatOllama(model=model, temperature=0)
response = [Link](prompt)

return [Link]

This function enhances our weather station by providing context-aware responses incorporating
current sensor readings and relevant information from our knowledge base. This is only an
example. To use it, we should have “Weather Reference Data,” which we do not currently have.
Instead, let’s create a general RAG system specializing in Edge AI Engineering.

Using the RAG System for Edge AI Engineering

For our RAG system, we will create a database with all chapters alheady written for the
EdgeAI Engineering book (chapters as URLs) and a PDF Wevolver 2025 Edge AI Technology
Report.

# PDF documents to include


pdf_paths = ["./data/2025_Edge_AI_Technology_Report.pdf"]

# Define URLs for document sources


urls = [
"[Link]
/object_detection/object_detection.html",
"[Link]
/image_classification.html",

731
"[Link]
"[Link]
/counting_objects_yolo.html",
"[Link]
"[Link]
"[Link]
/RPi_Physical_Computing.html",
"[Link]
]

Using the RAG system is straightforward. First, ensure you’ve created the vector database:

# Run once to create the database


python [Link]

Then, interact with the system through queries:

732
# Start the interactive query interface
python [Link]

Example interactions:

Your question: What is edge AI?

Generating answer...

Question: what is EdgeAI?


Retrieving documents...
Retrieved 4 document chunks
Generating answer...
Response latency: 165.72 seconds using model: llama3.2:3b

ANSWER:
==================================================
EdgeAI refers to the application of artificial intelligence (AI) at the edge
of a network, typically in real-time applications such as IoT sensors, industrial
robots, and smart cameras. The Edge AI ecosystem includes edge devices, edge
servers, and cloud platforms that work together to enable low-latency AI
inferencing and processing of data on-site without relying on continuous cloud connectivity.
in AI by leveraging energy-efficient, affordable, and scalable solutions for machine
learning and advanced edge computing.
==================================================

733
Those responses demonstrate how RAG enhances the SLM’s response with specific information
from our knowledge base about Edge AI applications on Raspberry Pi. One issue that should
be addressed is the latency.
To reduce latency, we can use for embedding, the all-minilm model which is much smaller
(23M parameters vs. 137M for nomic-embed-text) and creates 384-dimensional embeddings
instead of 768, significantly reducing computation time.
Also, smaller chunks can be helpful but have some disadvantages. For example, let’s say that
we can use a small chunk size (100 tokens with 50 overlap). Here are some considerations:

Advantages

1. Memory Efficiency: Smaller chunks require less memory during retrieval and processing,
which is beneficial for resource-constrained devices like the Raspberry Pi.
2. More Granular Retrieval: Smaller chunks can potentially provide more precise matches
to specific questions, especially for targeted queries about very specific details.
3. Reduced Context Window Usage: SLMs have limited context windows; smaller
chunks allow you to include more distinct pieces of information while staying within these
limits.

734
Disadvantages

1. Loss of Context: 100 tokens is approximately 75-80 words, which is often insufficient to
capture complete concepts or explanations. Many paragraphs and technical descriptions
require more space to convey their full meaning.
2. Increased Vector Store Size: More chunks mean more embeddings to store, potentially
increasing the overall size of your vector database.
3. Fragmented Information: With such small chunks, related information will be split
across multiple chunks, making it harder for the model to synthesize coherent answers.

Testing Different Models and Chunk Sizes

A good practice would be to experiment with different chunk sizes and embedding models and
measure:

1. Retrieval Quality: Are the retrieved chunks relevant to our queries?


2. Answer Accuracy: Does the SLM generate correct and comprehensive answers?
3. Memory Usage: Is the system staying within the memory constraints of our device?
4. Response Time: How does chunk size affect latency?

We can create a simple benchmarking function to have one embedding model defined test the
best chunk size:

def benchmark_chunk_sizes(document_list,
query_list,
sizes=[(100, 50), (300, 30), (500, 50), (1000, 100)]):
"""Test different chunk sizes and measure performance"""
results = {}

for chunk_size, overlap in sizes:


print(f"Testing chunk_size={chunk_size}, overlap={overlap}")

# Create splitter with current settings


text_splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
chunk_size=chunk_size, chunk_overlap=overlap
)

# Split documents
start_time = [Link]()
doc_splits = text_splitter.split_documents(document_list)
split_time = [Link]() - start_time

735
# Create embeddings and store
embedding_function = OllamaEmbeddings(model="nomic-embed-text")
temp_db_path = f"temp_db_{chunk_size}_{overlap}"

start_time = [Link]()
vectorstore = Chroma.from_documents(
documents=doc_splits,
collection_name="benchmark",
embedding=embedding_function,
persist_directory=temp_db_path
)
db_time = [Link]() - start_time

# Create retriever
retriever = vectorstore.as_retriever(k=3)

# Test queries
query_times = []
for query in query_list:
start_time = [Link]()
docs = [Link](query)
query_time = [Link]() - start_time
query_times.append(query_time)

# Store results
results[(chunk_size, overlap)] = {
"num_chunks": len(doc_splits),
"splitting_time": split_time,
"db_creation_time": db_time,
"avg_query_time": sum(query_times) / len(query_times),
"max_query_time": max(query_times),
"min_query_time": min(query_times)
}

# Clean up temporary DB
[Link](temp_db_path)

return results

Regarding the query side, some optimations can also reduce the latency at the edge. Let’s
modify the previous script, with:

1. Direct Ollama API Calls: Bypasses the LangChain abstraction layer for embedding

736
and LLM generation to reduce overhead.
2. Embedding Caching: Uses lru_cache to prevent recalculating embeddings for repeated
queries.
3. Preloading Models: Initializes models at startup to avoid cold-start latency.
4. Optimized Retriever Settings: Uses minimal k-value (2) and adds a score threshold
to filter out irrelevant matches.
5. Reduced Dependency Usage: Removes unnecessary imports and simplifies the
pipeline.
6. Concurrent Processing: Uses ThreadPoolExecutor for batch document embedding
(when needed).
7. Early Termination: Checks for empty document results before running the LLM.
8. Simplified Prompt: Uses a more concise prompt template focused on getting direct
answers.
9. Fixed Seed: Uses a consistent seed for the LLM to reduce variability in response times.

This optimized version (25-optimized_RAG_query.py) significantly reduces the latency com-


pared to our original implementation while maintaining compatibility with our existing nomic-
embed-text vector database and chunk size (300/30).

The direct Ollama API approach removes several layers of abstraction in the
LangChain implementations.

We can see latency improvements from 2 minutes down to approximately 50-110 seconds,
depending on the complexity of the queries.

737
In the next section, we’ll explore how RAG can be combined with our agent architecture to
create even more powerful edge AI systems.

Advanced Agentic RAG System

We can significantly enhance traditional RAG implementations by incorporating some of the


modules discussed earlier, such as intelligent routing, validation feedback loops, and explicit
knowledge gap identification. This will provide more reliable and transparent answers for users
querying document-based knowledge bases.
For example, let’s enhance the last RAG system created on the Edge AI Engineering dataset
so that the agent can use tools, such as a calculator for arithmetic calculations.

Note that any tool could be used here; the calculator is only a simple example to
demonstrate the concept.

738
When the user asks a question, the system first determines if it needs to use a tool or the RAG
approach. For knowledge queries, the RAG system enhances the response with information
from the database. The system then validates the answer quality, and if it’s not sufficient, tries
again with an improved prompt. In cases where questions fall outside the database’s scope, the
system will clearly inform the user rather than attempting to generate potentially misleading
answers.

System Architecture

Figure 19: image-20250322113245945

739
The system functions through several key components:

1. Query Router
• Analyzes incoming queries to determine if they’re calculations or knowledge queries
• Can use the same model for the response generator or a lightweight model to reduce
overhead
• Implements rule-based fallbacks for robust classification
2. Document Retriever
• Connects to a persistent vector database (Chroma)
• Uses semantic embeddings to find relevant documents
• Returns contextually similar content for knowledge generation
3. Response Generator
• Creates answers based on retrieved documents
• Implements a two-stage approach with validation and improvement
• Adds appropriate disclaimers when information is insufficient
4. Validation Engine
• Evaluates answer quality using structured criteria
• Assigns a numerical score to each generated response
• Triggers enhancement processes when quality is insufficient
5. Interactive Interface
• Provides user-friendly interaction with clear quality indicators
• Supports model switching and verbosity control
• Offers guidance for improving query outcomes

Key Workflow

The system follows this high-level workflow:

1. User submits a query


2. Router determines the query type (calculation vs. knowledge)
3. For calculations:
• Extract operation and numbers
• Compute and return result
4. For knowledge queries:
• Retrieve relevant documents
• Generate initial answer with RAG

740
• Validate response quality
• If quality is low, attempt enhancement with improved prompt
• If still insufficient, add disclaimer about knowledge gaps
5. Return final answer with quality metrics

Important Code Sections

Query Routing

def route_query(query: str) -> Dict[str, Any]:


"""Determine if the query is a calculation, otherwise use RAG"""
if VERBOSE:
print(f"Routing query: {query}")

# Check for calculation keywords


calc_terms = ["+", "add", "plus", "sum", "-", "subtract", "minus",
"difference", "*", "×", "multiply", "times", "product",
"/", "÷", "divide",
"division", "quotient"]

# Simple rule-based detection for calculations


is_calc = any(term in [Link]() for term in calc_terms) and \
[Link](r'\d+', query)

if is_calc:
# Use smaller, faster model for operation and number extraction
# ...extraction logic here...
return route_info

# For everything else, use RAG


return {"type": "rag", "reasoning": "Non-calculation query, using RAG"}

Enhanced RAG with Feedback Loop

# First RAG attempt with standard prompt


answer = get_answer_with_rag(query, documents, llm)
processing_type = "rag_standard"

741
# Validate the response quality
validation = validate_response(llm, query, answer)
validation_score = [Link]("score", 5)

# If validation score is low, try again with enhanced prompt


if validation_score < 7:
if VERBOSE:
print(f"First RAG attempt validation score: {validation_score}/10. \
Trying enhanced prompt.")

# Second RAG attempt with enhanced prompt


enhanced_context = "\n\n".join(documents)
enhanced_prompt = f"""
I need a more detailed and accurate answer to the following question:

{query}

The previous answer wasn't satisfactory. Let me provide you with \


relevant information:

{enhanced_context}

Based strictly on this information, provide a comprehensive answer.


Focus specifically on addressing the user's question with precise \
information from the provided context.
If the information doesn't fully answer the question, clearly state \
what you can determine
from the available information and what remains unknown.
"""
# ... process enhanced response ...

Knowledge Gap Handling

# If still low quality after enhancement, add a note


if improved_score < 6:
processing_type = "rag_insufficient_info"
information_gap_note = (
"\n\nNote: The information in my knowledge base may be incomplete on this"
"topic. I've provided the best answer based on available information, but"
"there might be gaps or additional details that would provide a more "

742
"complete answer."
)
answer = answer + information_gap_note

Detailed Workflow Diagram

743
744
Examples

Run the script 30-advanced_agentic_rag.py:

Simple Calculation

First Pass Rag

745
More Complex Queries

746
Queries outside of the database scope:

Fine-Tuning SLMs for Edge Deployment

Fine-tuning can adapt models to specific domains or tasks, improving performance for targeted
applications.

747
Preparing for Fine-Tuning

# Example dataset for fine-tuning a weather response model


training_data = [
{"input": "What's the weather in London?",
"output": "I need to check London's current weather. Please use the \
weather tool."},
{"input": "Is it going to rain tomorrow in Paris?",
"output": "To answer about tomorrow's weather in Paris, \
I need to use the weather tool."},
{"input": "Will it be sunny this weekend in Tokyo?",
"output": "To predict Tokyo's weekend weather, \
I should use the weather tool."}
]

# Format data for fine-tuning


formatted_data = []
for item in training_data:
formatted_data.append({
"prompt": item["input"],
"response": item["output"]
})

# Save formatted data to a file


import json
with open("weather_finetune_data.json", "w") as f:
[Link](formatted_data, f)

Setting Up a Fine-Tuning Process

Fine-tuning on edge devices is typically impractical due to resource constraints. Instead,


fine-tune on a more powerful machine and deploy the result to the edge:

This is a conceptual example - actual implementation depends on the framework

def prepare_for_finetuning(data_path, output_path):


"""
Prepare a model for fine-tuning
(run this on a more powerful machine)
"""
# This is a conceptual example

748
print(f"Fine-tuning model using data from {data_path}")
print(f"Fine-tuned model will be saved to {output_path}")

# The process would typically involve:


# 1. Loading the base model
# 2. Loading and preprocessing the training data
# 3. Setting up training parameters (learning rate, epochs, etc.)
# 4. Running the fine-tuning process
# 5. Evaluating the fine-tuned model
# 6. Saving and optimizing for edge deployment

# For Ollama, we would create a custom model definition (Modelfile)


modelfile = f"""
FROM llama3.2:1b

# Fine-tuning settings would go here


PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER top_k 40

# Custom system prompt for the specific domain


SYSTEM You are a specialized assistant for weather-related questions.
"""

with open(output_path, "w") as f:


[Link](modelfile)

print("Model preparation complete. Next steps:")


print("1. Run fine-tuning on a powerful machine")
print("2. Optimize the resulting model for edge deployment")
print("3. Deploy to your Raspberry Pi")

Real implementation: Supervised Fine-Tuning (SFT)

Supervised fine tuning (SFT) is a method to improve and customize pre-trained LLMs.
It involves retraining base models on a smaller dataset of instructions and answers. The
main goal is to transform a basic model that predicts text into an assistant that can follow
instructions and answer questions. SFT can also enhance the model’s overall performance, add
new knowledge, or adapt it to specific tasks and domains
Before considering SFT, it is recommended to try prompt engineering techniques like few-shot
prompting or retrieval augmented generation (RAG), as discussed previously. In practice,

749
these methods can solve many problems without fine-tuning. If this approach doesn’t meet
our objectives (regarding quality, cost, latency, etc.), then SFT becomes a viable option when
instruction data is available. SFT also offers benefits like additional control and customizability
to create personalized LLMs.
However, SFT has limitations. It works best when leveraging knowledge already present in
the base model. Learning completely new information, like an unknown language, can be
challenging and lead to more frequent hallucinations. For new domains unknown to the base
model, it is recommended that it be continuously pre-trained on a raw dataset first.
On the opposite end of the spectrum, instruct models (i.e., already fine-tuned models) can be
very close to our requirements. By providing chosen and rejected samples for a small set of
instructions (between 100 and 1000 samples), we can force the LLM to behave as we need.
The easiest way to finetune an SLM is by using Unsloth. The three most popular SFT techniques
are full fine-tuning, LoRA, and QLoRA.

For details, see: Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth.

For example, using this link, it is possible to find several notebooks with the steps to finetune
SLMs—for instance, the Gemma 3:1B.
The fine-tuned model can be saved on HF Hub or locally as GGUF, and to run a GGUF
model locally, we can use Ollama, as shown below:

1. Download the finetuned GGUF model


2. Create a Modelfile in the home directory:

cd ~
nano Modelfile

3. In the Modelfile, specify the path to the GGUF file:

FROM ~/Downloads/[Link]
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 64

4. Save the Modelfile and exit the editor.


5. Create the loadable model in Ollama:

ollama create your-model-name -f Modelfile

The model name here can be anything.

6. We can now use the model through as we have done with the llama3.2:3B in this chapter..

750
Conclusion

This chapter has explored comprehensive strategies for overcoming the inherent limitations of
Small Language Models in edge computing environments. By implementing techniques ranging
from optimized prompting strategies to sophisticated agent architectures and knowledge inte-
gration systems, we’ve demonstrated that it’s possible to significantly enhance the capabilities
of edge AI systems without requiring more powerful hardware or cloud connectivity.
The techniques presented—chain-of-thought prompting, task decomposition, function calling,
response validation, and RAG—form a toolkit that edge AI engineers can apply individually
or in combination to address specific challenges. Each approach offers unique advantages:
prompting techniques improve reasoning capabilities with minimal overhead, agent architectures
enable SLMs to perform actions beyond text generation, and RAG systems dramatically expand
an SLM’s knowledge without increasing model size.
Our practical implementations on the Raspberry Pi showcase that these enhancements are not
merely theoretical but can be deployed in real-world edge scenarios. From the simple calculator
agent to the more sophisticated knowledge router and RAG-enabled question answering system,
these examples provide templates that developers can adapt to their specific application
requirements.
The true power of these techniques emerges when they’re strategically combined. An agent
architecture with RAG capabilities, enhanced by chain-of-thought reasoning and validated with
a feedback loop, creates an edge AI system that approaches the capabilities of much larger
models while maintaining the advantages of edge deployment—privacy preservation, reduced
latency, and operation without internet connectivity.
As edge AI continues to evolve, these techniques will become increasingly important in bridging
the gap between the limited resources available on edge devices and the growing expectations
for AI capabilities. By thoughtfully applying these approaches, developers can create intelli-
gent systems that process data locally, respect user privacy, and operate reliably in diverse
environments.
The future of edge AI lies not necessarily in deploying ever-larger models but in developing
more innovative systems that combine efficient models with intelligent architectures, contextual
knowledge integration, and robust validation mechanisms. By mastering these techniques, edge
AI practitioners can create solutions that are not just technologically impressive but genuinely
useful and trustworthy in addressing real-world challenges.

Resources

The scripts used in this chapter can be found here: Advancing EdgeAI Scripts

751
752
LiteRT-LM: Production-Ready LLM Inference
at the Edge

Figure 20: DALL·E prompt — A 1950s-style cartoon illustration of a Raspberry Pi 5 running


a language model entirely on-device. The board glows with a warm amber halo.
Around it, speech bubbles with gears inside float in mid-air, representing tool calling.
A tiny robot arm extends from the board to press a physical button, hinting at IoT
integration. The background shows a cozy retro-futuristic lab bench cluttered with
transistors, vacuum tubes, and blinking LED panels. Soft pastel palette with bold
outlines typical of 1950s editorial cartoons.

753
Introduction

In the previous chapters, we explored Ollama as our primary gateway to running Small Language
Models (SLMs) on the Raspberry Pi 5. Ollama is excellent for rapid experimentation: one
command downloads and serves a model, and the Python library wraps everything in a clean,
familiar API. However, Ollama was designed primarily for desktop and server machines; it
runs [Link] under the hood, and its runtime is not purpose-built for the tight memory and
latency budgets of IoT deployments.
This chapter introduces LiteRT-LM, Google AI Edge’s production-ready, open-source inference
framework for deploying Large Language Models directly on edge devices. While Ollama is great
for getting started, LiteRT-LM offers a different set of tradeoffs that make it particularly well-
suited for scenarios where startup latency, memory footprint, and cross-platform consistency
matter — or simply when you want your Raspberry Pi to run the same model binary that
powers Google’s own Pixel Watch and Chromebook Plus.

By the end of this chapter, you will be able to:

• Explain the architecture of LiteRT-LM and how it differs from Ollama / [Link].
• Install the litert-lm Python package on a Raspberry Pi 5.
• Download and run a Gemma 4 model in the .litertlm format from Hugging Face.
• Build a simple query function, an interactive chat loop, a basic tool-calling pipeline, and
a streaming inference loop with performance metrics — all using LiteRT-LM’s Python
API.

Prerequisites: A Raspberry Pi 5 (8 GB recommended) running Raspberry Pi OS


(64-bit), set up as described in the Setup chapter, and familiarity with the Ollama-
based examples from the Small Language Models and SLM: Basic Optimization
Techniques chapters.

754
What is LiteRT-LM?

From TensorFlow Lite to LiteRT

Google’s on-device ML story has a long history. TensorFlow Lite (TFLite) was released in
2017 as a stripped-down runtime for deploying neural networks on mobile and embedded devices.
It popularized quantization-aware training, delegate-based hardware acceleration (GPU, DSP,
Edge TPU), and the flat .tflite model format — all ideas that have since become standard
practice across the industry.
In 2024, Google rebranded TensorFlow Lite as LiteRT (Lite RunTime), reflecting a broader
vision: a unified on-device ML runtime that covers not just classification and detection models,
but the full spectrum of modern AI workloads, including generative models. LiteRT-LM is
the part of this ecosystem dedicated to Large Language Models.

TensorFlow Lite (2017)



LiteRT (2024) ← unified on-device ML runtime

LiteRT-LM (2025) ← LLM-specific inference framework

755
Architecture Overview

LiteRT-LM is organized as a layered stack:


Layer 1 — User-Facing APIs: Language bindings (Python, Kotlin, C++, Swift) and a
high-level Conversation API that handles the full chat loop for you.
Layer 2 — Core Runtime: An Engine object that loads model weights into RAM and holds
them for the lifetime of the context manager. Loading takes several seconds; create the Engine
once and reuse it for multiple conversations.
Layer 3 — Executors: Low-level interfaces to the actual hardware. On Raspberry Pi 5
(ARM64 CPU), the XNNPACK delegate is used—the same highly optimized CPU kernel
library that accelerates TFLite models.

756
The execution pipeline follows a two-phase pattern familiar from any autoregressive LLM:

1. Prefill — the full prompt is tokenized and processed to fill the KV cache.
2. Decode — tokens are generated one at a time, each conditioned on the cached context.

The important insight for Edge AI is that LiteRT-LM manages the KV cache for
you. You never need to manually rebuild message history turn by turn (as you do
with raw REST APIs). As long as you stay inside the same Conversation object,
the model retains full context.

The .litertlm Format

LiteRT-LM uses its own model packaging format: .litertlm. A .litertlm file is a single
archive that bundles together:

• The quantized model weights (INT4 or INT8 by default).


• A tokenizer.
• Inference metadata (context length, sampler settings, etc.).

This is analogous to Ollama’s .gguf format, but with one key difference: .litertlm files are
specifically optimized for the LiteRT executor pipeline. You cannot use GGUF files with
LiteRT-LM, and you cannot use .litertlm files with Ollama. Pre-converted models are hosted
on the LiteRT Community on Hugging Face.

Supported Models

The table below shows some of the models available at the time of writing. Check the LiteRT-LM
overview page for the latest additions.

Quantiza- Con- RAM


Model Parameters tion text (approx.) Best For
Gemma3-1B- 1B INT4 4096 ~1 GB Ultra-low-memory devices
IT
Gemma 4 2.3B eff. INT4 8192 ~2.7 GB Raspberry Pi 5 sweet
E2B spot
Gemma 4 4B eff. INT4 8192 ~3.7 GB Higher quality, more
E4B RAM
Gemma-3n- 2B eff. INT4 4096 ~3 GB Previous generation
E2B
Gemma-3n- 4B eff. INT4 4096 ~4.2 GB Previous generation
E4B
Phi-4-mini 3.9B INT4 4096 ~2.5 GB Reasoning-focused tasks

757
Quantiza- Con- RAM
Model Parameters tion text (approx.) Best For
Qwen3-0.6B 0.6B INT4 4096 ~0.5 GB Fastest option

Gemma 4 architecture note: The “E” in Gemma 4 E2B stands for Effective
parameters. Gemma 4 uses a MatMul-free attention design combined with Per-
Layer Embeddings (PLE), which achieves frontier-level quality with substantially
fewer active parameters than traditional dense transformers. The E2B model
benchmarks above Gemma 3 27B on several reasoning tasks despite being 12×
smaller.

LiteRT-LM vs Ollama: When to Use Each

Both Ollama and LiteRT-LM are excellent tools for running SLMs on Raspberry Pi. They
solve the same fundamental problem but with different priorities. Understanding the tradeoffs
will help you choose the right tool for each project.

Feature Ollama LiteRT-LM


Ease of setup ����� One-line install ���� pip install
Model variety ����� Hundreds of models ��� ~20 models
Format GGUF .litertlm
Backend engine [Link] LiteRT / XNNPACK
Python API Stable, mature In active development
Conversation memory Manual (you manage Automatic (KV-cache per session)
history)
OpenAI compatibility � Drop-in replacement � Custom API
Production Community-driven Powers Chrome, Pixel Watch,
provenance Chromebook+
Cross-platform binary � Separate runtimes � Same binary, all platforms
IoT / embedded focus Secondary Primary design goal
GPU acceleration � CPU only � CPU only
(RPi)

Choose Ollama when you need the widest model selection, the simplest API, or OpenAI-
compatible endpoints for existing code.
Choose LiteRT-LM when you are targeting a production IoT deployment, want a consistent
runtime from prototype (Raspberry Pi) to mobile (Android/iOS), or need tighter memory
control and a framework that is actively hardened for edge constraints.

758
Important benchmark note: On the Raspberry Pi 5 with the Gemma 4 E2B
model, you can expect roughly 4–8 tokens/second at INT4 precision using the
CPU backend. This is comparable to Ollama with Gemma 3 2B, confirming that
both frameworks extract similar throughput from the RPi5’s Cortex-A76 cores.
The practical difference shows up on mobile hardware with GPU/NPU backends,
where LiteRT-LM achieves massive speedups that [Link]-based tools cannot
match.

Installation on Raspberry Pi 5

Prerequisites

Confirm that your Raspberry Pi 5 is running the 64-bit (aarch64) version of Raspberry Pi OS.
LiteRT-LM requires this architecture.

uname -m
# Expected: aarch64

Also, confirm your Python version. LiteRT-LM’s Python package targets Python 3.10–3.13:

python3 --version

Update the system

sudo apt update


sudo apt upgrade -y

Creating a Virtual Environment

It is good practice to use a dedicated virtual environment for LiteRT-LM experiments, especially
since litert-lm may conflict with torch or tensorflow packages when installed in the same
environment.
Starting criating a working directory

mkdir -p ~/Documents/LITERT
cd ~/Documents/LITERT

Create the environment:

759
python3 -m venv .venv
source .venv/bin/activate

Installing the Package

pip install --upgrade litert-lm

This single command pulls the pre-built wheel for aarch64 Linux and its dependencies (including
the XNNPACK-accelerated LiteRT runtime). There is no separate server to start — LiteRT-LM
runs in-process, embedded inside your Python script.

No server, no port. Unlike Ollama, which runs as a background daemon on port


11434, LiteRT-LM is a library. When your script exits, the model is unloaded. This
is a natural fit for embedded applications where you want complete control over
the application lifecycle.

Verify the installation:

python3 -c "import litert_lm; print('LiteRT-LM OK')"

Installing the Hugging Face CLI (for model downloads)

Models are hosted on Hugging Face. The huggingface_hub CLI is the most reliable way to
download .litertlm files.

pip install huggingface_hub


huggingface-cli login # optional; required for gated models

760
Downloading a Model

We will use Gemma 4 E2B (gemma-4-E2B-it), the sweet-spot model for Raspberry Pi 5.
It runs in approximately 1.5 GB of working memory, leaving headroom for your application
code.

mkdir -p ~/Documents/LITERT/models
huggingface-cli download \
litert-community/gemma-4-E2B-it-litert-lm \
[Link] \
--local-dir ~/Documents/LITERT/models

The download is approximately 1.5 GB. Once complete:

~/Documents/LITERT/models
��� [Link]

Quick Sanity Check with the CLI

LiteRT-LM ships with a command-line runner. Before writing any Python, confirm the model
works:

litert-lm run \
~/Documents/LITERT/models/[Link] \
--prompt="What is the capital of Brazil?"

Expected output (after a few seconds of model loading):

The capital of Brazil is **Brasília**.

761
First-run latency: Loading a 1.5 GB model from an SD card takes 10–20 seconds.
Using an NVMe SSD via the PCIe slot reduces this to 2–4 seconds. For production
IoT applications, an SSD is strongly recommended.

Python API

We can create Python scripts with a text editor, such as nano, or install a Jupyter Notebook:

pip install jupyter


jupyter notebook

Now let’s explore the Python API systematically, building from a single query to a streaming
chat loop with performance metrics. All examples assume the following preamble:

import litert_lm

litert_lm.set_min_log_severity(litert_lm.[Link]) # suppress verbose logs

MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"

The set_min_log_severity call silences LiteRT-LM’s internal diagnostic messages (model-


loading progress, tokenizer initialization, etc.) so your terminal output stays clean. During
development, you may want to remove this line to see what’s happening under the hood.

The Engine and Conversation Objects

The two core objects you will use in every LiteRT-LM script are:

• litert_lm.Engine(model_path) — a heavyweight object that loads model weights into


RAM and holds them for the lifetime of the context manager. Create the Engine once
and reuse it for multiple conversations.
• engine.create_conversation() — a lightweight object representing a single conversa-
tion thread. It maintains its own KV cache, so it automatically retains the full message
history.

Both objects are Python context managers (use them with with statements), which ensures
clean resource deallocation.

762
��������������������������������������������
� Engine (model weights loaded once) �
� �������������������������������������� �
� � Conversation A (KV cache A) � �
� �������������������������������������� �
� �������������������������������������� �
� � Conversation B (KV cache B) � �
� �������������������������������������� �
��������������������������������������������

A single Engine can host multiple concurrent Conversation objects, each with an independent
state. This is useful for multi-user applications or for maintaining separate conversation threads
for different tasks.

1. Simple Single-Turn Query

The simplest pattern: open an Engine, create a Conversation, send one message, read the
response.

import litert_lm

litert_lm.set_min_log_severity(litert_lm.[Link])
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"

PROMPT = "What is the capital of Colombia?"

with litert_lm.Engine(MODEL_PATH) as engine:


with engine.create_conversation() as conversation:
response = conversation.send_message(PROMPT)
answer = response["content"][0]["text"]
print(answer)

Output:

The capital of Brazil is **Bogotá**.

The response object is a dictionary with a content list, where each item has a type field
("text") and a text field containing the generated string. This structure is intentionally similar
to the Anthropic and OpenAI messages API, making the pattern familiar.

763
Wrapping it as a Reusable Function

For quick experiments, a helper function is convenient:

def simple_query(prompt, model=MODEL_PATH):


with litert_lm.Engine(model) as engine:
with engine.create_conversation() as conversation:
response = conversation.send_message(prompt)
return response["content"][0]["text"]

Let’s use the function:

print(simple_query("What is the capital of France?"))

Output:

The capital of France is **Paris**.

If we ask the following question:

print(simple_query("And Peru"))

Output:

Please provide more context! "And Peru?" is a very open-ended question. To give you a helpful

No conversation history in simple_query: Each call creates a fresh Engine


and Conversation. A follow-up question like "and Peru?" will elicit a confused
answer because the model has no memory of the previous question about a country’s
capital. The next section solves this.

2. Interactive Chat with Persistent History

The key advantage of LiteRT-LM’s Conversation object is that it automatically manages


history. As long as you stay inside the same conversation context, every new send_message
call has access to the full prior exchange — no manual history list required.

764
import litert_lm

litert_lm.set_min_log_severity(litert_lm.[Link])
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"

with litert_lm.Engine(MODEL_PATH) as engine:


with engine.create_conversation() as conversation:
print("Chat started. Type 'quit' or 'exit' to stop.\n")
while True:
user_input = input(">>> ").strip()

if user_input.lower() in {"exit", "quit"}:


print("[Exiting chat]")
break
if not user_input:
continue

# Streaming: iterate over chunks as they are generated


for chunk in conversation.send_message_async(user_input):
print(chunk["content"][0]["text"], end="", flush=True)
print() # newline after the full response

Example session:

Chat started. Type 'quit' or 'exit' to stop.

>>> What is the capital of France?


The capital of France is **Paris**.

>>> and Peru?


The capital of Peru is **Lima**.

>>> quit
[Exiting chat]

Notice how "and Peru?" (a context-dependent follow-up) is correctly resolved to Peru’s capital.
The model understood the implicit reference because the KV cache preserved the prior turn.
With the simple_query function from the previous section, the same follow-up would fail.
The key API difference here is send_message_async vs send_message. The async variant
returns a generator that yields partial chunks as tokens are produced, enabling streaming
output that makes interactions feel responsive. Each chunk has the same structure as the full
response from send_message.

765
3. Tool Calling (Function Calling)

As we saw in the SLM: Basic Optimization Techniques chapter, LLMs are unreliable calculators.
Let’s reproduce that demonstration with LiteRT-LM:

PROMPT = "What is 123456 multiplied by 123456?"


print(simple_query(PROMPT))

Output (typical — wrong!):

...123,456² = 15,238,827,536

The correct answer is:

print(f"{123456*123456:,}") # → 15,241,383,936

LiteRT-LM supports native function-calling protocols for Gemma models. At the Python
API level (still in active development), the most robust and universally compatible approach
is JSON-based prompt engineering — the same technique we used with Ollama in the
previous chapter. The idea is identical: use a system instruction to constrain the model’s
output to a specific JSON schema, then parse and dispatch that JSON in Python.

766
Step 1 — Define the Tool Function

def multiply_numbers(a: float, b: float) -> float:


"""Multiply two numbers.

Args:
a: The first number.
b: The second number.
"""
return f'{a * b:,}'

Step 2 — Define the System Instruction

LiteRT-LM does not yet have a separate system role in its Python API. The standard
workaround is to send the system instruction as the first user message before any real query.
The conversation’s KV cache retains this instruction for all subsequent turns.

import json

SYSTEM_INSTRUCTION = """
You are a helpful assistant with access to a single tool: multiply_numbers(a, b).

When the user asks for a multiplication, DO NOT explain.


Instead, respond ONLY with a JSON object of this form:

{
"tool": "multiply_numbers",
"a": <number>,
"b": <number>
}

No extra text. No Markdown. No backticks.


If the user is not asking for a multiplication, answer normally in plain text.
"""

Step 3 — The Tool-Calling Loop

767
with litert_lm.Engine(MODEL_PATH) as engine:
with engine.create_conversation() as conversation:
# Prime the model with the system instruction
conversation.send_message(SYSTEM_INSTRUCTION)

# Send the user's actual question


user_prompt = "What is 123456 multiplied by 123456?"
response = conversation.send_message(user_prompt)
text = response["content"][0]["text"]
print("Model raw output:\n", text)

# Parse the JSON and dispatch to the tool


result = None
try:
tool_call = [Link](text)
if tool_call.get("tool") == "multiply_numbers":
a = float(tool_call["a"])
b = float(tool_call["b"])
result = multiply_numbers(a, b)
except Exception as e:
print("[Could not parse tool call]", e)

if result is not None:


print(f"\n[Tool result] multiply_numbers({a}, {b}) = {result}")
else:
print("\n[No tool call detected; treating reply as plain text]")

Output:

Model raw output:


{
"tool": "multiply_numbers",
"a": 123456,
"b": 123456
}

[Tool result] multiply_numbers(123456.0, 123456.0) = 15,241,383,936.0

The pattern is clean and reusable:

1. The LLM handles natural language understanding — extracting intent and operands.

768
2. Python handles actual computation — always correct, always deterministic.

Swapping the tool: You can replace multiply_numbers with any Python function
— a GPIO sensor read, a file lookup, a distance calculation, a REST API call. The
model doesn’t care how the function is implemented internally; it only needs to
emit the correct JSON. This makes the pattern reusable across all kinds of Edge
AI + IoT applications.

If we run the tool calling loop again with a new query, such as “What is the capital of
Guatemala?”, the answer will be a normal response from the model:

Model raw output:


The capital of Guatemala is Guatemala City.

[Could not parse tool call or invalid format] Expecting value: line 1 column 1 (char 0)

[No tool call detected; treat model reply as plain text]

4. Streaming with Performance Metrics

For production deployments, knowing your inference speed is essential. The following code
adds Time to First Token (TTFT), total latency, and tokens/second measurements to
the chat loop.
TTFT measures how long the user waits before any output appears — it is dominated by the
prefill phase. Tokens/second measures decode throughput. Together, they give a complete
picture of perceived responsiveness.

import time
import re
import litert_lm

litert_lm.set_min_log_severity(litert_lm.[Link])
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"

def rough_token_count(text: str) -> int:


"""Estimate token count using word/punctuation splitting.

This is an approximation — a proper tokenizer would be more accurate,


but this is sufficient for ballpark throughput measurements.
"""
return len([Link](r"\w+|[^\w\s]", text, [Link]))

769
def chunk_to_text(chunk) -> str:
"""Extract plain text from a LiteRT-LM response chunk."""
parts = []
if isinstance(chunk, dict):
for item in [Link]("content", []):
if isinstance(item, dict) and [Link]("type") == "text":
[Link]([Link]("text", ""))
return "".join(parts)

with litert_lm.Engine(MODEL_PATH) as engine:


with engine.create_conversation() as conversation:
print("Chat with metrics. Type 'quit' to exit.\n")
while True:
user_input = input(">>> ").strip()

if user_input.lower() in {"exit", "quit"}:


break
if not user_input:
continue

start = time.perf_counter()
first_token_time = None
pieces = []

for chunk in conversation.send_message_async(user_input):


text = chunk_to_text(chunk)
if text:
if first_token_time is None:
first_token_time = time.perf_counter()
[Link](text)
print(text, end="", flush=True)

end = time.perf_counter()
print()

output = "".join(pieces)
output_tokens = rough_token_count(output)
total_latency = end - start
ttft = (first_token_time - start) if first_token_time else None
decode_time = (end - first_token_time) if first_token_time else None

770
tps = (output_tokens / decode_time) if decode_time and decode_time > 0 else 0.0

if ttft is not None:


print(
f"[TTFT: {ttft:.3f}s | Total: {total_latency:.3f}s | "
f"OutTok(est): {output_tokens} | Tk/s(est): {tps:.2f}]"
)
else:
print(f"[Total: {total_latency:.3f}s | OutTok(est): {output_tokens}]")

Example session:

Chat with metrics. Type 'quit' to exit.

>>> what is the capital of Brazil?


The capital of Brazil is **Brasília**.
[TTFT: 1.524s | Total: 3.438s | OutTok(est): 11 | Tk/s(est): 5.75]

>>> and France?


The capital of France is **Paris**.
[TTFT: 1.468s | Total: 3.156s | OutTok(est): 11 | Tk/s(est): 6.52]

>>> y la de Bolivia?
The capital of Bolivia is **Sucre**.

However, it's important to note that **La Paz** is the administrative capital and the seat of
[TTFT: 1.427s | Total: 11.427s | OutTok(est): 54 | Tk/s(est): 5.40]

Interpreting the Metrics

Typical RPi 5
Metric Value Meaning
TTFT 1-2 s (SSD) Time until first visible output. Dominated by prefill
(prompt processing).
Total 3-11 s Full wall-clock time from send to end-of-stream.
latency
OutTok(est) varies Estimated number of output tokens (word-level
approximation).
Tk/s(est) 5-7 Decode throughput in tokens per second.

771
Why is TTFT several seconds? On the first question in a session, TTFT
includes model-loading overhead. On subsequent turns, it shrinks because the
Engine is already warm. Longer prompts increase TTFT proportionally, since
prefill scales with prompt length.

Improving performance: The single most impactful hardware upgrade is replac-


ing the SD card with an NVMe SSD. This reduces model load time from ~20 s to ~2
s. Within a session, performance is CPU-bound and not affected by storage speed.

Complete Example: Putting it All Together

#!/usr/bin/env python3
"""
litert_lm_chat.py — Interactive streaming chat with Gemma 4 E2B on RPi5.

Usage:
source ~/litert-env/bin/activate
python3 litert_lm_chat.py
"""

import re
import time
import litert_lm

MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"
litert_lm.set_min_log_severity(litert_lm.[Link])

def rough_token_count(text: str) -> int:


return len([Link](r"\w+|[^\w\s]", text, [Link]))

def chunk_to_text(chunk) -> str:


parts = []
if isinstance(chunk, dict):
for item in [Link]("content", []):
if isinstance(item, dict) and [Link]("type") == "text":
[Link]([Link]("text", ""))
return "".join(parts)

772
def main():
print(f"Loading model: {MODEL_PATH}")
print("Type 'quit' or 'exit' to stop.\n")

with litert_lm.Engine(MODEL_PATH) as engine:


with engine.create_conversation() as conversation:
while True:
try:
user_input = input("You: ").strip()
except (EOFError, KeyboardInterrupt):
print("\n[Interrupted]")
break

if user_input.lower() in {"exit", "quit"}:


print("[Goodbye]")
break
if not user_input:
continue

print("AI: ", end="", flush=True)


start = time.perf_counter()
first_token_time = None
pieces = []

for chunk in conversation.send_message_async(user_input):


text = chunk_to_text(chunk)
if text:
if first_token_time is None:
first_token_time = time.perf_counter()
[Link](text)
print(text, end="", flush=True)

end = time.perf_counter()
print("\n")

output = "".join(pieces)
output_tokens = rough_token_count(output)
total_latency = end - start
ttft = (first_token_time - start) if first_token_time else None
decode_time = (end - first_token_time) if first_token_time else None
tps = (output_tokens / decode_time) if decode_time and decode_time > 0 else 0

773
if ttft is not None:
print(
f" � [TTFT: {ttft:.2f}s | Total: {total_latency:.2f}s | "
f"~{output_tokens} tokens | {tps:.1f} tk/s]\n"
)

if __name__ == "__main__":
main()

During the inference, it is possible to see that the memory consumption is low (2.1GB of RAM
+ ~1GB of Swap):

774
Runing a simple query on the Raspi, using the same model but with Ollama, we can see that mem-

ory usage is intense (7.7GB of RAM + 4GB of Swap):


Another point is the model’s size. The Gemma4:e2b needs 7.2GB of RAM, and its LiteRT-LM
version only 2.7GB.

Going Further

Multi-modal inputs. Gemma 4 is natively multi-modal. LiteRT-LM can accept image buffers
alongside text prompts, opening the door to vision-language tasks directly on Raspberry Pi —
using the same .litertlm file.

UNFORTUNATELY, at this time, the multimodality is still not working properly


on the Raspberry Pi using LiteRT-LM, only with OLLAMA, but with high latency
(around 4 minutes).

775
Feeding tool results back into the conversation. The tool-calling example stops after the
function is called. A richer pattern sends the function’s result back as a follow-up message,
allowing the model to formulate a natural-language reply that incorporates the correct answer.
Multiple simultaneous conversations. A single Engine can host multiple Conversation
objects at the same time — useful for multi-user local assistants in, say, a classroom setting,
without loading model weights twice.
Upgrading to Gemma 4 E4B. If your Raspberry Pi 5 has 8 GB of RAM, the E4B model
(~3.7 GB working memory) offers noticeably better reasoning quality, especially for longer
contexts (with higher latency).
Cross-platform consistency. A .litertlm binary tested on a Raspberry Pi 5 will run
without modification on an NVIDIA Jetson Orin Nano or a Qualcomm-based Android device
— with the additional benefit of GPU/NPU acceleration on those targets. This is the key
architectural advantage of LiteRT-LM: one model file, one runtime, every platform.

Conclusion

LiteRT-LM occupies a distinct niche in the edge AI toolbox. It is not a replacement for
Ollama — it is a complement. Ollama excels at broad model selection, simplicity, and desktop
development. LiteRT-LM excels at production-quality, cross-platform deployment where the
same runtime binary needs to travel from prototyping (Raspberry Pi) to production (Android
phone, wearable, embedded system).
In this chapter, we covered the complete development arc: understanding LiteRT-LM’s ar-
chitecture (Engine → Conversation → XNNPACK executor); installing the Python package
and downloading a Gemma 4 .litertlm model; and building four progressively more complex
patterns — single query, persistent chat, tool calling, and streaming with metrics.
The most important mental model to take away is the Engine / Conversation separation:
weights are loaded once (expensive), and conversation state is maintained automatically per
session (free). This design makes LiteRT-LM well-suited for always-on IoT applications where
the model is loaded at boot and conversations flow naturally for as long as the device is
running.

Resources

• LiteRT-LM Official Documentation: [Link]/edge/litert-lm


• LiteRT-LM GitHub Repository: [Link]/google-ai-edge/LiteRT-LM
• LiteRT Community on Hugging Face (pre-converted models): hugging-
[Link]/litert-community

776
• Gemma 4 E2B model: [Link]/litert-community/gemma-4-E2B-it-litert-lm
• LiteRT-Optimized INT8 LLM for Raspberry Pi (ICCV Workshop 2025):
[Link]
• Google AI Edge Gallery (Android/iOS demo app): [Link]/google-ai-
edge/gallery
• Study Notebook for this chapter: [Link]

#
Weekly Labs

777
Edge AI Engineering - Weekly Labs

Week 1: Introduction and Setup

Lab 1: Raspberry Pi Configuration

Objectives:

• Install Raspberry Pi OS using Raspberry Pi Imager


• Configure basic settings (hostname, SSH, WiFi)
• Learn essential Linux commands
• Manage files between your computer and Raspberry Pi

Instructions:

1. Download Raspberry Pi Imager on your computer


2. Configure OS settings (enable SSH, set hostname, WiFi credentials)
3. Boot your Raspberry Pi and confirm connectivity
4. Learn how to use SSH for remote access
5. Transfer files using SCP or FileZilla
6. Update your Raspberry Pi OS (sudo apt update && sudo apt upgrade)
7. Practice basic Linux commands (ls, cd, mkdir, cp, mv)

Deliverable: Screenshot showing successful SSH connection to your Raspberry Pi

Lab 2: Development Environment Setup

Objectives:

• Set up Python environment for development


• Configure remote development tools
• Install essential libraries
• Test camera functionality

Instructions:

1. Install Python essentials: pip install jupyter matplotlib numpy pillow

778
2. Configure Jupyter Notebook for remote access:
pip install jupyter
jupyter notebook --generate-config
jupyter notebook --ip=[Link] --no-browser

3. Connect the camera module (USB or CSI) to your Raspberry Pi


4. Test camera functionality using command-line tools
5. Write a simple Python script to capture an image

Deliverable: A simple Python script that captures and displays an image from your camera
and a Screenshot showing a successful image capture

Week 2: Image Classification Fundamentals

Lab 3: Working with Pre-trained Models

Objectives:

• Install TensorFlow Lite runtime


• Download and run MobileNet V2 model
• Process and classify images
• Understand model inputs and outputs

Instructions:

1. Install TensorFlow Lite runtime: pip install tflite_runtime


2. Download MobileNet V2 model:

wget [Link]

3. Download labels file


4. Create a Python script that:

• Loads the TFLite model


• Processes input images to 224x224 format
• Runs inference on test images
• Displays top-5 predicted classes with confidence scores

Deliverable: Python script that successfully classifies sample images with MobileNet V2 and
a Screenshot showing a successful result

779
Lab 4: Custom Dataset Creation

Objectives:

• Create a simple custom dataset using Raspberry Pi camera


• Organize images into classes
• Prepare dataset for model training

Instructions:

1. Create a web interface for image capture:


• Use Flask to create a simple web server
• Set up camera preview and capture functionality
• Save captured images with appropriate filenames
2. Capture at least 50 images per class for 3 classes
3. Organize the dataset into an appropriate directory structure
4. Document your dataset creation process

Deliverable: Structured dataset with at least 3 classes and 50 images per class

Week 3: Custom Image Classification

Lab 5: Edge Impulse Model Training

Objectives:

• Create an Edge Impulse project


• Upload and process the dataset
• Design and train a transfer learning model
• Evaluate model performance

Instructions:

1. Create an Edge Impulse account and a new project


2. Upload your custom dataset
3. Create an impulse design:
• Set image size to 160x160
• Use Transfer Learning for feature extraction
4. Generate features for all images

780
5. Train model using MobileNet V2
6. Analyze model performance (accuracy, confusion matrix)
7. Test model on validation data

Deliverable: Edge Impulse project link and screenshot of model performance metrics

Lab 6: Model Deployment to Raspberry Pi

Objectives:

• Export trained model to TFLite format


• Deploy model to Raspberry Pi
• Create a real-time inference application
• Optimize inference speed

Instructions:

1. Export model as TensorFlow Lite (.tflite)


2. Transfer the model to Raspberry Pi
3. Create a Python application that:
• Captures live images from the camera
• Preprocesses images for the model
• Runs inference and displays results
• Shows confidence scores
4. Implement a web interface for real-time classification

Deliverable: Python script for real-time image classification with your custom model and a
Screenshot showing a successful result

Week 4: Object Detection Fundamentals

Lab 7: Pre-trained Object Detection

Objectives:

• Understand object detection architecture


• Run pre-trained SSD-MobileNet model
• Process detection outputs
• Visualize detected objects

781
Instructions:

1. Download the pre-trained SSD-MobileNet V1 model


2. Create a Python script that:
• Loads the model and labels
• Preprocesses input images
• Runs inference
• Extracts bounding boxes, classes, and scores
• Implements Non-Maximum Suppression (NMS)
• Visualizes detections with bounding boxes
3. Test on various images with multiple objects

Deliverable: Python script that performs and visualizes object detection on test images and
a Screenshot showing a successful result.

Lab 8: EfficientDet and FOMO Models

Objectives:

• Compare different object detection architectures


• Implement EfficientDet and FOMO models
• Analyze performance differences
• Understand trade-offs between models

Instructions:

1. Download the EfficientDet Lite0 model


2. Implement inference with EfficientDet
3. Compare with SSD-MobileNet implementation
4. Learn about FOMO (Faster Objects, More Objects)
5. Analyze trade-offs in accuracy vs. speed
6. Measure inference time on Raspberry Pi

Deliverable: Comparison report of SSD-MobileNet vs. EfficientDet with performance metrics


and visualized results

782
Week 5: Custom Object Detection

Lab 9: Dataset Creation and Annotation

Objectives:

• Create an object detection dataset


• Learn annotation techniques
• Prepare dataset for model training

Instructions:

1. Capture at least 100 images containing objects to detect


2. Upload images to Roboflow or a similar annotation tool
3. Create bounding box annotations for each object
4. Apply data augmentation (rotation, brightness adjustment)
5. Export dataset in YOLO format
6. Document the annotation process

Deliverable: Annotated dataset with at least 2 object classes and 100 total images

Lab 10: Training Models in Edge Impulse

Objectives:

• Upload annotated dataset to Edge Impulse


• Train SSD MobileNet object detection model
• Evaluate model performance
• Export model for deployment

Instructions:

1. Create a new Edge Impulse project for object detection


2. Upload annotated dataset (train/test splits)
3. Create object detection impulse
4. Train SSD MobileNet model
5. Evaluate model performance
6. Export model as TensorFlow Lite

Deliverable: Edge Impulse project link with trained object detection model and performance
metrics

783
Week 6: Advanced Object Detection

Lab 11: FOMO Model Training

Objectives:

• Understanding FOMO architecture benefits


• Train FOMO model on Edge Impulse
• Compare performance with SSD MobileNet
• Deploy optimized model to Raspberry Pi

Instructions:

1. Create a new impulse in Edge Impulse using the same dataset


2. Train FOMO model instead of SSD MobileNet
3. Compare inference speed and accuracy
4. Deploy both models to Raspberry Pi
5. Create an application that can switch between models
6. Measure and document performance differences

Deliverable: Python application that compares SSD MobileNet vs. FOMO performance in
real-time

Lab 12: YOLO Implementation

Objectives:

• Install and configure Ultralytics YOLO


• Convert models to optimized NCNN format
• Create real-time detection application
• Implement object counting

Instructions:

1. Install Ultralytics: pip install ultralytics


2. Download and test YOLO (n) model
3. Export model to NCNN format for optimization
4. Create Python script for real-time detection
5. Implement object counting algorithm
6. Add visualization of counts over time

Deliverable: Python application for real-time object detection and counting using YOLO

784
Week 7: Object Counting Project

Lab 13: Custom YOLO Training

Objectives:

• Train YOLO on a custom dataset


• Optimize model for edge deployment
• Create a complete application for object counting

Instructions:

1. Train the YOLO model on your custom dataset


• Use Google Colab for training if needed
• Set appropriate hyperparameters
2. Export optimized model for Raspberry Pi
3. Create a Python application that:
• Captures video feed
• Detects objects using YOLO
• Counts objects over time
• Logs results to a database

Deliverable: Complete object counting application with data logging

Lab 14: Fixed-Function AI Integration (Optional)

Objectives:

• Integrate multiple AI models into a single application


• Create a dashboard for visualization
• Optimize application for long-term deployment

Instructions:

1. Create an integration application that combines:


• Object detection capabilities
• Classification for detected objects
• Counting and tracking over time
2. Implement a simple web dashboard for visualization
3. Add performance monitoring
4. Configure application for startup at boot

785
Deliverable: Integrated application combining multiple AI capabilities with visualization
dashboard

Week 8: Introduction to Generative AI

Lab 15: Raspberry Pi Configuration for SLMs

Objectives:

• Optimize Raspberry Pi for running Small Language Models


• Install an active cooling solution
• Configure memory and swap
• Install essential libraries

Instructions:

1. Install active cooling solution on Raspberry Pi 5


2. Optimize system configuration:

• Increase swap memory: sudo dphys-swapfile swapoff, edit /etc/dphys-


swapfile
• Set CONF_SWAPSIZE to 2048
• sudo dphys-swapfile setup && sudo dphys-swapfile swapon

3. Install dependencies:
sudo apt update
sudo apt install build-essential python3-dev

Deliverable: Screenshot showing system configuration with increased swap and temperature
monitor during stress test

Lab 16: Ollama Installation and Testing

Objectives:

• Install Ollama framework


• Pull and test Small Language Models
• Benchmark model performance
• Monitor resource usage

786
Instructions:

1. Install Ollama:

curl -fsSL [Link] | sh

2. Pull different models (such as):

ollama pull llama3.2:1b


ollama pull gemma:2b
ollama pull phi3:latest

3. Run a basic inference test with each model


4. Measure and compare:

• Load time
• Inference speed (tokens/sec)
• Memory usage
• Temperature

Deliverable: Benchmark report comparing performance metrics of different SLM models on


your Raspberry Pi

Week 9: SLM Python Integration

Lab 17: Ollama Python Library

Objectives:

• Use the Ollama Python library


• Create interactive applications
• Process SLM responses programmatically
• Handle multiple conversation turns

Instructions:

1. Install Ollama Python library: pip install ollama


2. Create Python script to:
• Connect to Ollama API
• Send prompts to models
• Process and format responses

787
• Handle conversation context
3. Implement proper error handling
4. Create a simple interactive CLI application

Deliverable: Python script demonstrating Ollama library usage with conversation handling

Lab 18: Function Calling and Structured Outputs

Objectives:

• Implement function calling with SLMs


• Create applications with structured outputs
• Build validation mechanisms
• Handle image inputs

Instructions:

1. Install required libraries:

pip install pydantic instructor openai

2. Create Pydantic models for structured outputs


3. Implement function calling with an instructor:
client = [Link]( OpenAI(base_url="[Link] api_key="olla

4. Build distance calculator application using SLM for city/country recognition


5. Add image input processing

Deliverable: Python application that uses function calling for structured interaction with
SLMs

788
Week 10: Retrieval-Augmented Generation

Lab 19: RAG Fundamentals

Objectives:

• Understand RAG architecture


• Create vector database
• Implement embedding generation
• Build simple RAG system

Instructions:

1. Install required libraries:

pip install langchain chromadb

2. Create a simple dataset with text documents


3. Implement document splitting and chunking
4. Generate embeddings using Ollama
5. Store embeddings in ChromaDB
6. Create query system

Deliverable: Python implementation of a basic RAG system with simple text documents

Lab 20: Advanced RAG

Objectives:

• Optimize RAG for edge devices


• Implement more efficient retrieval
• Create a specialized knowledge base
• Build validation mechanisms

Instructions:

1. Create a specialized knowledge base (e.g., technical documentation)


2. Implement optimized embedding generation
3. Fine-tune retrieval parameters
4. Add response validation
5. Create a persistent vector store
6. Benchmark performance

789
Deliverable: Optimized RAG implementation with specialized knowledge base and perfor-
mance analysis

Week 11: Vision-Language Models

Lab 21: Florence-2 Setup

Objectives:

• Install Florence-2 model


• Configure environment
• Run basic inference tests
• Understand model capabilities

Instructions:

1. Install required dependencies:


pip install transformers torch torchvision torchaudio
pip install timm einops
pip install autodistill-florence-2

2. Download model and test basic functionality


3. Run image captioning test
4. Measure performance (memory usage, inference time)

Deliverable: Python script demonstrating basic Florence-2 functionality with performance


metrics

Lab 22: Vision Tasks with Florence-2

Objectives:

• Implement various vision tasks


• Create applications for captioning, detection, grounding
• Optimize performance
• Combine tasks

Instructions:

790
1. Implement image captioning:
• Basic caption generation
• Detailed caption generation
2. Implement object detection:
• Bounding box visualization
• Multiple object detection
3. Implement visual grounding:
• Highlight specific objects based on text prompts
4. Create segmentation application
5. Measure the performance of each task

Deliverable: Python application demonstrating multiple vision tasks with Florence-2 and
performance analysis

Week 12: Physical Computing Basics

Lab 23: Sensor and Actuator Integration

Objectives:

• Connect digital sensors


• Read environmental data
• Control LEDs and actuators
• Create a data collection system

Instructions:

1. Connect hardware components:

• DHT22 temperature/humidity sensor


• BMP280 pressure sensor
• LEDs (red, yellow, green)
• Push button

2. Install required libraries:

pip install adafruit-circuitpython-dht adafruit-circuitpython-bmp280

791
3. Create a Python script to read sensor data
4. Implement LED control based on conditions
5. Create visualization of sensor data

Deliverable: Python application for reading sensor data and controlling actuators with
visualization

Lab 24: Jupyter Notebook Integration

Objectives:

• Use Jupyter Notebook for physical computing


• Create interactive widgets
• Visualize sensor data in real-time
• Control actuators from a notebook

Instructions:

1. Install ipywidgets: pip install ipywidgets


2. Create a Jupyter Notebook for sensor data collection
3. Implement interactive widgets for control
4. Create real-time visualization
5. Build a dashboard with multiple data views

Deliverable: Jupyter Notebook with interactive widgets for sensor monitoring and actuator
control

Week 13: SLM-Physical Computing Integration

Lab 25: Basic SLM Analysis

Objectives:

• Integrate SLMs with sensor data


• Create analysis application
• Implement decision-making logic
• Control actuators based on SLM responses

Instructions:

792
1. Create a Python application that:
• Collects sensor data
• Formats data for SLM prompt
• Sends prompt to model
• Parses response
• Controls actuators based on response
2. Implement multiple analysis modes
3. Add error handling for SLM responses

Deliverable: Python application integrating SLMs with physical sensors and actuators

Lab 26: SLM-IoT Control System

Objectives:

• Create a complete IoT monitoring system


• Implement natural language interaction
• Add data logging and analysis
• Create web interface

Instructions:

1. Build a complete system with:


• Sensor data collection
• SLM-based analysis
• Natural language command processing
• Data logging to the database
• Web interface for interaction
2. Implement multiple SLM models
3. Add historical data analysis
4. Create visualization dashboard

Deliverable: Complete IoT monitoring system with SLM integration and web interface

793
Week 14: Advanced Edge AI Techniques

Lab 27: Building Agents

Objectives:

• Create agent architecture


• Implement tool usage
• Build decision-making system
• Handle complex tasks

Instructions:

1. Implement calculator agent:


• Create a query routing system
• Implement tool functions
• Build decision-making logic
2. Create knowledge router:
• Implement web search integration
• Build classification system
• Handle time-based queries
3. Measure and optimize performance

Deliverable: Python implementation of agent architecture with tool usage and decision
routing

Lab 28: Advanced Prompting and Validation

Objectives:

• Implement chain-of-thought prompting


• Create few-shot learning examples
• Build task decomposition system
• Implement response validation

Instructions:

1. Create examples for different prompting strategies


2. Implement chain-of-thought framework
3. Build few-shot learning templates
4. Create a task decomposition system

794
5. Implement validation mechanisms
6. Compare the effectiveness of different strategies

Deliverable: Python implementation demonstrating different prompting strategies with


performance comparison

Week 15: Final Project Integration

Lab 29: Agentic RAG System

Objectives:

• Combine agent architecture with RAG


• Create a complete knowledge system
• Implement advanced validation
• Build query optimization

Instructions:

1. Create a complete agentic RAG system:


• Build knowledge database
• Implement agent architecture
• Add tool functions
• Create validation mechanisms
• Optimize retrieval
2. Test with complex queries
3. Measure performance
4. Create visualization of system components

Deliverable: Complete agentic RAG system with documentation and performance analysis

Lab 30: Final Project

Objectives:

• Design and implement a comprehensive Edge AI system


• Combine multiple techniques
• Create complete documentation
• Present project

795
Instructions:

1. Design final project combining:


• Computer vision capabilities
• SLM integration
• Physical computing
• Advanced techniques (RAG, agents, etc.)
2. Implement complete system
3. Create documentation
4. Measure performance
5. Prepare presentation

Deliverable: Complete final project with documentation, code, and presentation

Hardware Requirements

Basic Setup (Weeks 1-7)

• Raspberry Pi Zero 2W or Pi 5
• MicroSD card (32GB+)
• Camera module (USB webcam or Pi camera)
• Power supply

Generative AI (Weeks 8-15)

• Raspberry Pi 5 (8GB RAM recommended)


• Active cooling solution
• MicroSD card (64GB+ recommended)

Physical Computing (Weeks 12-15)

• DHT22 temperature/humidity sensor


• BMP280 pressure/temperature sensor
• LEDs (red, yellow, green)
• Push button
• Resistors (4.7kΩ, 330Ω)
• Jumper wires

796
• Breadboard

Software Requirements

Development Environment

• Raspberry Pi OS (64-bit)
• Python 3.9+
• Jupyter Notebook
• SSH client

Computer Vision and DL

• TensorFlow Lite runtime


• OpenCV
• Edge Impulse
• Ultralytics

Generative AI

• Ollama
• Transformers
• Pytorch
• ChromaDB
• LangChain
• Pydantic
• Instructor

Physical Computing

• GPIO Zero
• Adafruit CircuitPython libraries

797
Assessment Criteria

Each lab will be evaluated based on:

1. Functionality (40%): Does the implementation work as specified?


2. Code Quality (20%): Is the code well-structured, documented, and efficient?
3. Documentation (20%): Are the process and results documented?
4. Analysis (20%): Is there a thoughtful analysis of results and performance?

The final project will be evaluated based on:

1. Integration (30%): How well different components are integrated


2. Innovation (20%): Novel approaches or applications
3. Implementation (30%): Overall quality and functionality
4. Presentation (20%): Clear explanation and demonstration

Tips for Success

1. Start Early: These labs build on each other. Falling behind makes later labs more
difficult.
2. Document As You Go: Take notes, screenshots, and document issues/solutions.
3. Optimize Resources: SLMs and VLMs require careful resource management.
4. Collaborate: Discuss approaches with classmates while ensuring individual work.
5. Backup Regularly: Create backups of your SD card after significant progress.
6. Measure Performance: Always benchmark and optimize your implementations.
7. Ask Questions: If you’re stuck, ask for help early rather than falling behind.

#
References & Author

798
References

To learn more:

Online Courses

• Harvard School of Engineering and Applied Sciences - CS249r: Tiny Machine Learning
• Professional Certificate in Tiny Machine Learning (TinyML) – edX/Harvard
• Introduction to Embedded Machine Learning - Coursera/Edge Impulse
• Computer Vision with Embedded Machine Learning - Coursera/Edge Impulse
• UNIFEI-IESTI01 TinyML: “Machine Learning for Embedding Devices”

Books

• “Python for Data Analysis” by Wes McKinney


• “Deep Learning with Python” by François Chollet - GitHub Notebooks
• “TinyML” by Pete Warden and Daniel Situnayake
• “TinyML Cookbook 2nd Edition” by Gian Marco Iodice
• “Technical Strategy for AI Engineers, In the Era of Deep Learning” by Andrew Ng
• “AI at the Edge” book by Daniel Situnayake and Jenny Plunkett
• “XIAO: Big Power, Small Board” by Lei Feng and Marcelo Rovai
• “Machine Learning Systems” by Vijay Janapa Reddi

Projects Repository

• Edge Impulse Expert Network

799
AiEng4D

Edge AI Engineering, an eBook in a series of Hands-On tutorials, is part of the AIEng4D


(formerly TinyML4D) group, an initiative to make Embedded Machine Learning education avail-
able to everyone, explicitly enabling innovative solutions for the unique challenges Developing
Countries face.

800
About the author

Marcelo Rovai, a Brazilian living in Chile, is a recognized figure in engineering and technology
education. He holds the title of Professor Honoris Causa from the Federal University of Itajubá
(UNIFEI), Brazil. His educational background includes an Engineering degree from UNIFEI
and a specialization from the Polytechnic School of São Paulo University (POLI/USP). Further
enhancing his expertise, he earned an MBA from IBMEC (INSPER) and a Master’s in Data
Science from the Universidad del Desarrollo (UDD) in Chile.

801
With a career spanning several high-profile technology companies, including AVIBRAS Airspace,
AT&T, NCR, and IGT, where he served as Vice President for Latin America, he brings industry
experience to his academic endeavors. He is a prolific writer on electronics-related topics and
shares his knowledge through open platforms like [Link].
In addition to his professional pursuits, he is dedicated to educational outreach, serving as a
volunteer professor at UNIFEI and engaging with the AIEng4D (formerly TinyML4D) group,
and the EDGE AIP– the Academia-Industry Partnership of EDGEAI Foundation as a Co-Chair,
promoting TinyML education in developing countries. His work underscores a commitment to
leveraging technology for societal advancement.
LinkedIn profile: [Link]
Lectures, books, papers, and tutorials: [Link]

802

You might also like