Edge AI Engineering
Edge AI Engineering
Marcelo Rovai
2026-02-27
Table of contents
Preface 3
Acknowledgments 5
Introduction 6
Edge AI Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
Why Edge AI Matters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
The Raspberry Pi Advantage . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
What You’ll Learn . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Who This Book Is For . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Classification of AI Applications 11
Fixed Function AI vs. Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Fixed Function AI (Reactive) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Generative AI (Proactive) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
Summary Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
The Edge AI Advantage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
Setup 15
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
Raspberry Pi Models (covered in this book) . . . . . . . . . . . . . . . . . . . . 17
Engineering Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
Hardware Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Raspberry Pi Zero 2W . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
Installing the Operating System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
The Operating System (OS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
Installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
Initial Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
2
Remote Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
SSH Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
To shut down the Raspi via terminal: . . . . . . . . . . . . . . . . . . . . . . . . 27
Transfer Files between the Raspberry Pi and a computer . . . . . . . . . . . . . 27
Increasing SWAP Memory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
Fixed‑size zram (better for Zero‑2’s SD card) . . . . . . . . . . . . . . . . . . . 34
Notes specific to Zero 2 W Trixie . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Installing a Camera . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
Installing a Camera Module on the CSI port . . . . . . . . . . . . . . . . . . . . 35
Installing a USB WebCam . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
Running the Raspi Desktop remotely . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
Updating and Installing Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
Model-Specific Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Raspberry Pi Zero (Raspi-Zero) . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Raspberry Pi 4 or 5 (Raspi-4 or Raspi-5) . . . . . . . . . . . . . . . . . . . . . . 52
Measuring Temperature and Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Check CPU Temperature . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
Understand Power Measurement on Pi 5 . . . . . . . . . . . . . . . . . . . . . . 53
Script: Average Temperature and Power . . . . . . . . . . . . . . . . . . . . . . 54
Run a measurement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
Interpreting the Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
3
Custom Image Classification Project 88
Image Classification Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
The Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
Data Collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
Training the model with Edge Impulse Studio . . . . . . . . . . . . . . . . . . . . . . 97
Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
The Impulse Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
Image Pre-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
Model Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
Model Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
Trading off: Accuracy versus speed . . . . . . . . . . . . . . . . . . . . . . . . . 104
Model Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Live Image Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
Summary: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
4
Impulse Design, new Training and Testing . . . . . . . . . . . . . . . . . . . . . 185
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 188
Inference and Post-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 190
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
5
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 260
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261
6
Camera Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 304
Image Capture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
Performance Benchmarking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 307
Basic (Float32): mobilenet_v2.pte . . . . . . . . . . . . . . . . . . . . . . . . . 310
XNNPACK Backend (Flot32): mobilenet_v2_xnnpack.pte . . . . . . . . . . . 310
Quantization (INT8): mobilenet_v2_quantized_xnnpack.pte . . . . . . . . . . 311
Performance Comparison Table . . . . . . . . . . . . . . . . . . . . . . . . . . . 312
Exploring Custom Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 313
Exporting a Custom Trained Model . . . . . . . . . . . . . . . . . . . . . . . . 314
Running Custom Models on Raspberry Pi . . . . . . . . . . . . . . . . . . . . . 315
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318
Key Takeaways . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Performance Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
Code Repository . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
Official Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
Books . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
7
Run Inference on MemryX Accelerator (MXA) . . . . . . . . . . . . . . . . . . 342
Decode the MXA Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 343
Comparing CPU vs. MXA Performance . . . . . . . . . . . . . . . . . . . . . . 344
Measuring Latency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 344
Testing with Larger Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Clean Shutdown . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
Folders Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 348
Performance Comparison Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . 348
Key Observations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 348
When to Use the MX3? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 349
YOLOv8 Object Detection with MX3 Hardware Acceleration . . . . . . . . . . . . . 349
Model Export and Compilation . . . . . . . . . . . . . . . . . . . . . . . . . . . 351
Understanding YOLOv8 Output Format . . . . . . . . . . . . . . . . . . . . . . 354
Complete Inference Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 354
Going deeper in the functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357
Making Inferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 360
Inference with a custom model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 362
Adjusting Confidence Threshold . . . . . . . . . . . . . . . . . . . . . . . . . . 366
Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
Batch Processing (Optimization) . . . . . . . . . . . . . . . . . . . . . . . . . . 367
Thermal Management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
Confidence Threshold Tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
Exploring MemryX eXamples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 367
Clone the MemryX eXamples Repository . . . . . . . . . . . . . . . . . . . . . 368
Troubleshooting Common Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 368
Device Not Detected . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 368
Compilation Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 370
Thermal Throttling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Python Version Conflicts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Low FPS / Poor Performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
Import Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Model Accuracy Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Next Steps and Extensions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
Project Ideas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
Advanced Topics to Explore . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
References and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Official Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Code and Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Background Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Community and Support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 376
8
Text Generation with RNNs 377
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 377
What Are We Actually Building? . . . . . . . . . . . . . . . . . . . . . . . . . . 378
Neural Network Architectures Background . . . . . . . . . . . . . . . . . . . . . . . . 378
The Human Brain Analogy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 378
Recurrent Neural Networks (RNN) . . . . . . . . . . . . . . . . . . . . . . . . . 378
The Memory Problem and GRU Solution . . . . . . . . . . . . . . . . . . . . . 380
Dataset Preparation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 382
Data Preprocessing Steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Tokenization and Vocabulary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Character-Level Tokenization . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Vocabulary Building Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . 384
Creating the Character Dictionary . . . . . . . . . . . . . . . . . . . . . . . . . 384
Training Sequences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 385
The Sliding Window Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . 385
Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
Creating Training Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 387
Character Embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 387
From Sparse to Dense Representation . . . . . . . . . . . . . . . . . . . . . . . 387
Learning Character Relationships . . . . . . . . . . . . . . . . . . . . . . . . . . 387
Visualization and Understanding . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
RNN Architecture Components . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
Model Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
Memory and Processing Flow . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
Why GRU over Basic RNN? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
Training Process: Teaching the Model to Write . . . . . . . . . . . . . . . . . . . . . 392
The Learning Objective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 392
Hardware and Time Requirements . . . . . . . . . . . . . . . . . . . . . . . . . 392
Monitoring Progress . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 392
Preventing Overfitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
Training Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Text Generation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
The Generation Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Temperature Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
Generation Example (Temperature = 0.5) . . . . . . . . . . . . . . . . . . . . . 396
Example Output Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 396
Challenges and Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
Context Window Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
Character vs. Word Level Trade-offs . . . . . . . . . . . . . . . . . . . . . . . . 397
Coherence Challenges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
9
Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398
Potential Improvements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398
Connecting to Modern Language Models . . . . . . . . . . . . . . . . . . . . . . . . . 398
Scale Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398
Architectural Evolution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 399
Training Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 399
Summary: Why Modern Models Perform Better? . . . . . . . . . . . . . . . . . 399
Conclusion: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 400
Resourses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 401
10
Final Thoughts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419
11
Function Calling Solution for Calculations . . . . . . . . . . . . . . . . . . . . . . . . 485
Define the Tool (Function Schema) . . . . . . . . . . . . . . . . . . . . . . . . . 485
Implement the Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 485
Project: Calculating Distances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 486
Running with other models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 493
Adding images . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 493
Retrievel Augmentation Generation (RAG) . . . . . . . . . . . . . . . . . . . . . . . 498
A simple RAG project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 499
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 506
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 506
12
Trade-offs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 551
Best Use Cases . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 551
Future Implications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 552
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 552
13
Accessing the GPIOs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 580
Pin Numbering Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 581
Safety First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 581
GPIO Zero Library . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 582
“Hello World”: Blinking an LED . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 583
Understanding the LED Circuit . . . . . . . . . . . . . . . . . . . . . . . . . . . 583
Testing with GPIO Zero . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 584
Installing all LEDs (the “actuators”) . . . . . . . . . . . . . . . . . . . . . . . . 586
Sensors Installation and setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 588
Button . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 589
Installing Adafruit CircuitPython . . . . . . . . . . . . . . . . . . . . . . . . . . 592
DHT22 - Temperature & Humidity Sensor . . . . . . . . . . . . . . . . . . . . . 594
Installing the BMP280: Barometric Pressure & Altitude Sensor . . . . . . . . . 598
Measuring Weather and Altitude With BMP280 . . . . . . . . . . . . . . . . . 603
Playing with Sensors and Actuators . . . . . . . . . . . . . . . . . . . . . . . . . . . 606
Testing the Notebook setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 608
Initialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 609
GPIO Input and Output . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
Getting and displaying Sensor Data . . . . . . . . . . . . . . . . . . . . . . . . 614
Widgets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 616
Advanced GPIO Zero Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 618
PWM LED Brightness Control . . . . . . . . . . . . . . . . . . . . . . . . . . . 618
Composite Devices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 618
Other Useful GPIO Zero Devices . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Interacting an SLM with the Physical world . . . . . . . . . . . . . . . . . . . . . . . 619
Other Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 627
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 627
Key Achievements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 627
Technical Insights . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 627
Practical Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Challenges and Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Future Enhancements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Final Thoughts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
14
Creating new functions for the Actuation: . . . . . . . . . . . . . . . . . . . . . 641
Prompting Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 644
Key Changes in the code: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 644
System Overview: Enhanced IoT Environmental Monitoring with SLM Control 645
The new code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 649
Code Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 652
TEST1: Temp above the threshold . . . . . . . . . . . . . . . . . . . . . . . . . 652
TEST2: Temp below the threshold . . . . . . . . . . . . . . . . . . . . . . . . . 654
TEST3: Alarm Button pressed . . . . . . . . . . . . . . . . . . . . . . . . . . . 655
Interacting with IoT Systems, using Natural Language Commands . . . . . . . . . . 655
What will change? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 656
How It will Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 657
Key Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 659
Running the System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 660
Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 660
Special Commands . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Languages . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Advantages of This Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Error Handling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Tips for Best Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 663
Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 663
Adding Data Logging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Running the System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665
Example Queries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665
Data Files . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
sensor_readings.csv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
command_history.csv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
Tips . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 667
Examples: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
Prompt Optimization and Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . 669
Quick Solution: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 670
New Code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 672
Key Changes Under the Hood: . . . . . . . . . . . . . . . . . . . . . . . . . . . 672
Using Pydantic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 673
2. Implementing Pydantic in the Code . . . . . . . . . . . . . . . . . . . . . . . 674
Next Steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 677
Conclusion 679
Resourses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 680
15
Advancing EdgeAI: Beyond Basic SLMs 681
Understanding SLM Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 682
1. Knowledge Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 682
2. Reasoning Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 683
3. Inconsistent Outputs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 683
4. Domain Specialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 683
Techniques for Enhancing SLM at the Edge . . . . . . . . . . . . . . . . . . . . . . . 684
Optimizing Prompting Strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
Chain-of-Thought Prompting . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
Few-Shot Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
Task Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 686
Building Agents with SLMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
General Knowledge Router . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 698
Improving Agent Reliability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 701
1. Function Calling with Pydantic . . . . . . . . . . . . . . . . . . . . . . . . . 701
2. Response Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 704
Retrieval-Augmented Generation (RAG) . . . . . . . . . . . . . . . . . . . . . . . . . 705
Understanding RAG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 705
Implementing a Naive RAG System . . . . . . . . . . . . . . . . . . . . . . . . 706
Instalation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 707
Key Components of the Naive RAG System . . . . . . . . . . . . . . . . . . . . 707
Advantages of RAG for Edge AI . . . . . . . . . . . . . . . . . . . . . . . . . . 709
Optimizing RAG for Edge Devices . . . . . . . . . . . . . . . . . . . . . . . . . 710
Application: Enhanced Weather Station with RAG . . . . . . . . . . . . . . . . 710
Using the RAG System for Edge AI Engineering . . . . . . . . . . . . . . . . . 712
Testing Different Models and Chunk Sizes . . . . . . . . . . . . . . . . . . . . . 716
Advanced Agentic RAG System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 719
System Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 720
Key Workflow . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 721
Important Code Sections . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 722
Detailed Workflow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 724
Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 726
Fine-Tuning SLMs for Edge Deployment . . . . . . . . . . . . . . . . . . . . . . . . . 728
Preparing for Fine-Tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 729
Setting Up a Fine-Tuning Process . . . . . . . . . . . . . . . . . . . . . . . . . 729
Real implementation: Supervised Fine-Tuning (SFT) . . . . . . . . . . . . . . . 730
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 732
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 733
16
Week 2: Image Classification Fundamentals . . . . . . . . . . . . . . . . . . . . . . . 735
Lab 3: Working with Pre-trained Models . . . . . . . . . . . . . . . . . . . . . . 735
Lab 4: Custom Dataset Creation . . . . . . . . . . . . . . . . . . . . . . . . . . 736
Week 3: Custom Image Classification . . . . . . . . . . . . . . . . . . . . . . . . . . 736
Lab 5: Edge Impulse Model Training . . . . . . . . . . . . . . . . . . . . . . . . 736
Lab 6: Model Deployment to Raspberry Pi . . . . . . . . . . . . . . . . . . . . 737
Week 4: Object Detection Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . 737
Lab 7: Pre-trained Object Detection . . . . . . . . . . . . . . . . . . . . . . . . 737
Lab 8: EfficientDet and FOMO Models . . . . . . . . . . . . . . . . . . . . . . 738
Week 5: Custom Object Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . . 739
Lab 9: Dataset Creation and Annotation . . . . . . . . . . . . . . . . . . . . . . 739
Lab 10: Training Models in Edge Impulse . . . . . . . . . . . . . . . . . . . . . 739
Week 6: Advanced Object Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
Lab 11: FOMO Model Training . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
Lab 12: YOLO Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
Week 7: Object Counting Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 741
Lab 13: Custom YOLO Training . . . . . . . . . . . . . . . . . . . . . . . . . . 741
Lab 14: Fixed-Function AI Integration (Optional) . . . . . . . . . . . . . . . . 741
Week 8: Introduction to Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Lab 15: Raspberry Pi Configuration for SLMs . . . . . . . . . . . . . . . . . . . 742
Lab 16: Ollama Installation and Testing . . . . . . . . . . . . . . . . . . . . . . 742
Week 9: SLM Python Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 743
Lab 17: Ollama Python Library . . . . . . . . . . . . . . . . . . . . . . . . . . . 743
Lab 18: Function Calling and Structured Outputs . . . . . . . . . . . . . . . . 744
Week 10: Retrieval-Augmented Generation . . . . . . . . . . . . . . . . . . . . . . . 745
Lab 19: RAG Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
Lab 20: Advanced RAG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
Week 11: Vision-Language Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 746
Lab 21: Florence-2 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 746
Lab 22: Vision Tasks with Florence-2 . . . . . . . . . . . . . . . . . . . . . . . 746
Week 12: Physical Computing Basics . . . . . . . . . . . . . . . . . . . . . . . . . . . 747
Lab 23: Sensor and Actuator Integration . . . . . . . . . . . . . . . . . . . . . . 747
Lab 24: Jupyter Notebook Integration . . . . . . . . . . . . . . . . . . . . . . . 748
Week 13: SLM-Physical Computing Integration . . . . . . . . . . . . . . . . . . . . . 748
Lab 25: Basic SLM Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 748
Lab 26: SLM-IoT Control System . . . . . . . . . . . . . . . . . . . . . . . . . . 749
Week 14: Advanced Edge AI Techniques . . . . . . . . . . . . . . . . . . . . . . . . . 750
Lab 27: Building Agents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 750
Lab 28: Advanced Prompting and Validation . . . . . . . . . . . . . . . . . . . 750
Week 15: Final Project Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
Lab 29: Agentic RAG System . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
Lab 30: Final Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
17
Hardware Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 752
Basic Setup (Weeks 1-7) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 752
Generative AI (Weeks 8-15) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 752
Physical Computing (Weeks 12-15) . . . . . . . . . . . . . . . . . . . . . . . . . 752
Software Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
Development Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
Computer Vision and DL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
Physical Computing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
Assessment Criteria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 754
Tips for Success . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 754
References 755
To learn more: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
Online Courses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
Books . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
Projects Repository . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
TinyML4D 756
18
Preface
In the rapidly evolving technology landscape, the convergence of artificial intelligence and edge
computing is one of the most exciting frontiers. This intersection promises to revolutionize how
we interact with the world around us, bringing intelligence and decision-making capabilities
directly to the devices we use every day. At the heart of this revolution lies the Raspberry Pi,
a powerful yet accessible single-board computer (SBC) that has democratized computing and
now stands poised to do the same for edge AI.
This book, which serves as the official textbook for IESTI05 Edge AI Engineering at the
Federal University of Itajubá (UNIFEI) in Brazil, embodies both a passion for technology
and a conviction in its capacity to address real-world problems. While developed to support
UNIFEI’s engineering curriculum, the content is valuable for all learners, whether in academic
settings or pursuing independent study.
“Edge AI Engineering: Hands-on with the Raspberry Pi” is not just about theory or abstract
concepts. It’s about getting your hands dirty, writing code, training models, and seeing your
creations come to life. Each chapter blends foundational knowledge with practical application,
focusing on what’s possible with the Raspberry Pi platform.
From the compact Raspberry Pi Zero to the more powerful Pi 5, we explore how these incred-
ible devices can become the brains of intelligent systems—recognizing images, understanding
speech, detecting objects, and even running small language models. Each project serves as a
stepping stone, building your skills and confidence as you progress.
Beyond the technical skills, this book aims to instill something more valuable – a sense of
curiosity and possibility. The field of edge AI is still in its infancy, with new applications
and techniques emerging daily. By mastering the fundamentals presented here, you’ll be well-
equipped to explore these frontiers, perhaps even pushing the boundaries of what’s possible
on edge devices.
Whether you’re a student seeking to understand AI’s practical applications, a professional
looking to expand your skill set, or an enthusiast eager to add intelligence to your projects, we
hope this book serves as both a guide and an inspiration.
As you embark on this journey, remember that every expert was once a beginner. The learning
path is filled with challenges and moments of joy and discovery. Embrace both, and let your
creativity guide you.
19
Thank you for joining us on this exciting adventure into edge machine learning. Let’s begin
exploring what’s possible when we bring AI to the edge, one Raspberry Pi at a time.
Happy coding, and may your models always converge!
Prof. Marcelo Rovai
January, 2026
20
Acknowledgments
Google Nano Banana and OpenAI’s GPT were used to generate some of the images
in the book. Claude Sonnet and Perplexity helped with code and text reviews.
21
Introduction
Edge AI Engineering
Traditional AI deployment often relies on cloud infrastructure, which requires constant con-
nectivity and introduces latency. Edge AI addresses these limitations by bringing intelligence
directly to where data is generated and actions occur. This approach offers several compelling
advantages:
The Raspberry Pi, with its combination of affordability, processing capability, and extensive
GPIO options, provides an ideal platform for exploring Edge AI concepts. From the compact
Raspberry Pi Zero 2W to the more powerful Pi 5, these devices offer:
22
• Sufficient computational power for running optimized AI models
• A complete Linux-based operating system for straightforward development
• Extensive connectivity options for integrating with sensors and actuators
• A vibrant community and ecosystem of libraries and tools
• An accessible entry point for students, hobbyists, and professionals alike
This book takes a progressive approach to Edge AI engineering, starting with foundational
concepts and building toward more advanced applications:
1. Essential setup and configuration: Prepare your Raspberry Pi for Edge AI develop-
ment
2. Computer vision applications: Implement image classification and object detection
systems
3. Small Language Models (SLMs): Run and optimize language models directly on
your Raspberry Pi
4. Vision-Language Models: Explore multimodal AI with Florence-2
5. Physical computing integration: Connect AI systems with sensors and actuators
6. Advanced optimization techniques: Enhance model performance through methods
like RAG, agents, and function calling
Each chapter includes detailed explanations, step-by-step instructions, and practical projects
demonstrating real-world applications of Edge AI concepts.
Whether you’re a student exploring AI for the first time, an educator developing a curriculum,
a maker building innovative projects, or a professional seeking to expand your skills, this book
provides the knowledge and hands-on experience needed to successfully implement Edge AI
solutions on the Raspberry Pi platform.
Join us on this journey to the edge of AI innovation, where we’ll bridge theory and prac-
tice through engaging, accessible projects that demonstrate the transformative potential of
intelligent edge computing.
23
About this Book
Several chapters (Labs) in this book also accompany the open-source book Machine Learning
Systems by Professor Vijay Janapa Reddi from Harvard, which we invite you to read.
“Edge AI Engineering: Hands-on with the Raspberry Pi” is designed as a practical, project-
based learning resource that bridges theoretical AI concepts with tangible implementations.
This book is part of the open-source Machine Learning Systems initiative, democratizing access
to AI education and applications.
Key Features
1. Progressive Learning Path: The book structure follows a natural progression from
basic to advanced concepts, beginning with foundational computer vision applications
and advancing to generative AI techniques.
24
2. Model-Specific Optimizations: Each chapter provides targeted guidance for specific
Raspberry Pi models, helping you maximize performance whether you’re using a Pi Zero
2W or a Pi 5.
3. Open-Source Foundation: We emphasize accessible tools and frameworks, including
Edge Impulse Studio, TensorFlow Lite, PyTorch, Transformers, and Ollama, ensuring
you can continue your learning journey with widely available resources.
4. Practical Problem-Solving: Rather than abstract exercises, each project addresses
real-world challenges that demonstrate Edge AI’s practical value.
5. Resource Optimization Techniques: Learn essential strategies for deploying AI on
resource-constrained devices, balancing performance needs with hardware limitations.
6. Cross-Domain Applications: Explore implementations spanning computer vision,
natural language processing, and physical computing, showcasing the versatility of Edge
AI.
Prerequisites
25
By completing this book, you’ll possess the skills to design, implement, and optimize Edge
AI applications across a wide range of use cases, leveraging the unique capabilities of the
Raspberry Pi platform to bring intelligence to the edge.
26
Classification of AI Applications
As we embark on our journey through Edge AI Engineering with the Raspberry Pi, it’s essential
to understand the fundamental classification of AI applications that form the structure of this
book. Our exploration is divided into two parts, each representing a different paradigm in
artificial intelligence implementation.
AI applications can be broadly categorized into two approaches that represent different capa-
bilities, interaction models, and implementation strategies:
Fixed Function AI, or Reactive AI, operates by analyzing specific inputs according to prede-
termined patterns and rules and then producing consistent outputs for given scenarios. These
systems:
• Respond to specific triggers: They activate only when presented with particular
inputs.
• Follow defined patterns: Their behavior is predictable and consistent.
• Excel at structured tasks: They perform exceptionally well at classification, detection,
and pattern recognition
• Operate within boundaries: Their capabilities are limited to their specific program-
ming.
In the first part of this book (Chapters 2-4), we explore fixed-function AI through computer
vision applications:
These applications demonstrate how edge devices can deliver reliable, efficient AI in constrained
environments, focusing on specific, well-defined tasks.
27
Generative AI (Proactive)
Generative AI, also known as Proactive AI, represents a fundamental shift in capability. These
systems can:
The second part of this book (Chapters 5-9) explores Generative AI at the edge:
This progression from Fixed Function to Generative AI mirrors the evolution of artificial
intelligence itself—from specialized systems designed for specific tasks to more flexible, creative
systems capable of addressing a broader range of challenges.
Summary Table
Conclusion
28
The Edge AI Advantage
Both Fixed Function and Generative AI gain unique benefits when deployed at the edge:
29
30
Setup
31
This chapter will guide you through setting up the Raspberry Pi Zero 2 W (Raspi-Zero) and the
Raspberry Pi 5 (Raspi-5) models. We’ll cover hardware setup, operating system installation,
initial configuration, and tests.
The general instructions for the Raspi-5 also apply to the older Raspberry Pi
versions, such as the Raspi-3 and Raspi-4.
Introduction
The Raspberry Pi is a powerful and versatile single-board computer that has become an essen-
tial tool for engineers across various disciplines. Developed by the Raspberry Pi Foundation,
these compact devices offer a unique combination of affordability, computational power, and
extensive GPIO (General Purpose Input/Output) capabilities, making them ideal for proto-
typing, embedded systems development, and advanced engineering projects.
Key Features
1. Computational Power: Despite their small size, Raspberry Pis offer significant pro-
cessing capabilities, with the latest models featuring multi-core ARM processors and up
to 8GB of RAM.
2. GPIO Interface: The 40-pin GPIO header enables direct interaction with sensors,
actuators, and other electronic components, facilitating hardware-software integration.
3. Extensive Connectivity: Built-in Wi-Fi, Bluetooth, Ethernet, and multiple USB ports
enable a wide range of communication and networking projects.
4. Low-Level Hardware Access: Raspberry Pis provide access to interfaces such as I2C,
SPI, and UART, enabling detailed control and communication with external devices.
5. Real-Time Capabilities: With proper configuration, Raspberry Pis can be used for
soft real-time applications, making them suitable for control systems and signal process-
ing tasks.
6. Power Efficiency: Low power consumption enables battery-powered and energy-
efficient designs, especially in models like the Pi Zero.
32
Raspberry Pi Models (covered in this book)
Engineering Applications
1. Embedded Systems Design: Develop and prototype embedded systems for real-world
applications.
2. IoT and Networked Devices: Create interconnected devices and explore protocols
like MQTT, CoAP, and HTTP/HTTPS.
3. Control Systems: Implement feedback control loops, PID controllers, and interface
with actuators.
4. Computer Vision and AI: Utilize libraries like OpenCV and TensorFlow Lite for
image processing and machine learning at the edge.
5. Data Acquisition and Analysis: Collect sensor data, perform real-time analysis, and
create data logging systems.
6. Robotics: Build robot controllers, implement motion planning algorithms, and interface
with motor drivers.
7. Signal Processing: Perform real-time signal analysis, filtering, and DSP applications.
8. Network Security: Set up VPNs, firewalls, and explore network penetration testing.
This lab will guide you through setting up the most common Raspberry Pi models, so you
can get started on your machine learning project quickly. We’ll cover hardware setup, operat-
ing system installation, and initial configuration, focusing on preparing your Pi for Machine
Learning applications.
33
Hardware Overview
Raspberry Pi Zero 2W
34
Raspberry Pi 5
• Processor:
– Pi 5: Quad-core 64-bit Arm Cortex-A76 CPU @ 2.4GHz
– Pi 4: Quad-core Cortex-A72 (ARM v8) 64-bit SoC @ 1.5GHz
• RAM: 2GB, 4GB, or 8GB options (8GB recommended for AI tasks)
• Wireless: Dual-band 802.11ac wireless (2.4 GHz and 5 GHz), Bluetooth 5.0
• Ports: 2 × micro HDMI ports, 2 × USB 3.0 ports, 2 × USB 2.0 ports, CSI camera port,
DSI display port
• Power: 5V/5A, 5V/3A limits peripherals to 600mA, via USB-C connector 27W USB-C
power supply
In the labs, we will use different names to address the Raspberry Pi: Raspi,
Raspi-5, Raspi-Zero, etc. Usually, “Raspi” or “Raspberry Pi” is used when the
instructions or comments apply to all models.
35
Installing the Operating System
An operating system (OS) is essential software that manages computer hardware and software
resources, providing standard services for computer programs. It is the core software that runs
on a computer, serving as an intermediary between hardware and application software. The
OS oversees the computer’s memory, processes, device drivers, files, and security protocols.
1. Key functions:
• Process management: Allocating CPU time to different programs
• Memory management: Allocating and freeing up memory as needed
• File system management: Organizing and keeping track of files and directories
• Device management: Communicating with connected hardware devices
• User interface: Providing a way for users to interact with the computer
2. Components:
• Kernel: The core of the OS that manages hardware resources
• Shell: The user interface for interacting with the OS
• File system: Organizes and manages data storage
• Device drivers: Software that allows the OS to communicate with hardware
The Raspberry Pi runs a specialized version of Linux designed for embedded systems. This
operating system, typically a variant of Debian called Raspberry Pi OS (formerly Raspbian),
is optimized for the Pi’s ARM-based architecture and limited resources.
Key features:
36
Installation
To use the Raspberry Pi, we will need an operating system. By default, Raspberry Pis check
for an operating system on any SD card inserted in the slot, so we should install an operating
system using Raspberry Pi Imager.
In November 2025, the Raspberry Pi Imager 2.0 was launched. It brings a new
wizard interface, the opportunity to pre-configure Raspberry Pi Connect, and im-
proved accessibility for screen readers and other assistive technologies.
Raspberry Pi Imager is a tool for downloading and writing images on macOS, Windows, and
Linux. It includes many popular operating system images for Raspberry Pi. We will also use
the Imager to preconfigure credentials and remote access settings.
Follow the steps to install the OS on your Raspberry Pi.
37
4. Choose the appropriate operating system:
• For Raspi-Zero: For example, you can select under Raspberry Pi OS (Other),
Raspberry Pi OS Lite (64-bit).
38
Due to the Raspberry Pi Zero’s limited SDRAM (512 MB), the recommended
OS is the 32-bit version. However, to run some machine learning models, such
as the YOLO from Ultralitics, we should use the 64-bit version. Although the
Raspi-Zero can run a desktop, we will choose the LITE version (no Desktop)
to reduce the RAM needed for regular operation.
• For Raspi-5: We can select the full 64-bit version, which includes a desktop:
Raspberry Pi OS (64-bit)
39
5. Select your microSD card as the storage device.
6. Click Next, then go to the Customization tab. The imager will guide you through
setting the hostname, the Raspberry Pi username and password, configuring WiFi, and
enabling SSH (Very important!).
7. Write the image to the microSD card.
In the examples here, we will use different hostnames depending on the device used:
raspi, raspi-5, raspi-Zero, etc. Please replace it with the one you’re currently using.
Initial Configuration
40
3. Please wait for the initial boot process to complete (it may take a few minutes).
You can find the most common Linux commands for the Raspberry Pi here or here.
Remote Access
SSH Access
The easiest way to interact with the Raspi-Zero is via SSH (“Headless”). You can use a
Terminal (MAC/Linux), PuTTy (Windows), or any other.
The Raspberry Pi and the notebook should be on the same WiFi network. Note
that the Raspberry Pi 5 supports dual-band 802.11ac Wi-Fi (2.4 GHz and 5 GHz),
whereas the Raspberry Pi Zero 2W supports only 2.4 GHz.
1. Find your Raspberry Pi’s IP address (for example, check your router).
2. On your computer, open a terminal and connect via SSH:
ssh username@[raspberry_pi_ip_address]
Alternatively, if you do not have the IP address, you can try, for example, ssh
mjrovai@[Link] , ssh mjrovai@[Link] , etc.: bash ssh username@[Link]
When you see the prompt:
mjrovai@rpi-5:~ $
41
sudo apt-get update
sudo apt upgrade
sudo reboot # Reboot to ensure all updates take effect
You should confirm the Raspberry Pi IP address. On the terminal, you can use:
hostname -I
python3 --version
Once we use the latest Raspberry Pi OS (based on Debian Trixie), it should be: 3.13:
As of today (January 2026), some packages, such as ExecuTorch, officially support only Python
versions 3.10-3.12. Python 3.13.5 is too new and will likely cause compatibility issues. Since
42
Debian Trixie ships with Python 3.13 by default, we’ll need to install a compatible
Python version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python. We will do it in the appropriate Lab.
For now, we will keep the system Python.
If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.
When you want to turn off your Raspberry Pi, there are better ideas than just pulling the
power cord. This is because the Raspi may still be writing data to the SD card, in which case
merely powering down may result in data loss or, even worse, a corrupted SD card.
For a safety shutdown, use the command line:
To avoid possible data loss and SD card corruption, before removing power, wait
a few seconds after shutdown for the Raspberry Pi’s LED to stop blinking and go
dark. Once the LED goes out, it’s safe to power down.
Transferring files between the Raspberry Pi and our main computer can be done using a USB
drive, via the terminal (scp), or an FTP program over the network.
43
You can use any text editor. In the same terminal, an option is nano.
To copy the file named [Link] from your personal computer to a user’s home folder on your
Raspberry Pi, run the following command from the directory containing [Link], replacing
the <username> placeholder with the username you use to log in to your Raspberry Pi and
the <pi_ip_address> placeholder with your Raspberry Pi’s IP address:
Note that ~/ means we will move the file to the ROOT of our Raspberry Pi. You
can choose any folder in your Raspberry Pi. But you should create the folder before
you run scp, since scp won’t create folders automatically.
For example, let’s transfer the file [Link] to the ROOT of my Raspberry Pi Zero, which
44
has an IP of [Link]:
I use a different profile to differentiate the terminals. The action above occurs on your
computer. Now, let’s go to our Raspi (using SSH) and check if the file is there:
$ scp <username>@<pi_ip_address>:[Link] .
For example:
On the Raspi, let’s create a copy of the file with another name:
cp [Link] test_2.txt
45
scp mjrovai@[Link]:test_2.txt .
Transferring files via FTP, such as FileZilla FTP Client, is also possible and much easier to
use. Follow the instructions to install the program on your Desktop, then use the Raspberry
Pi’s IP address as the Host. For example:
s[Link]
Enter your Raspberry Pi username and password. Pressing Quickconnect opens two win-
dows, one for your host computer desktop (right) and another for the Raspberry Pi (left).
46
Increasing SWAP Memory
Using htop, a cross-platform interactive process viewer, we can easily monitor resources on our
Raspberry Pi in real time, including the list of processes, the running CPUs, and the memory
usage. To lunch hop, enter with the command on the terminal:
htop
47
Regarding memory, among the Raspberry Pi family, the Raspberry Pi Zero has the least
SRAM (500 MB), compared to 2GB to 16GB on the Raspberry Pi 4 or 5.
On any Raspberry Pi, it is possible to increase/modify the system’s available memory using
“Swap.” Swap memory, also known as swap space, is a technique used in computer operating
systems to temporarily store data from RAM (Random Access Memory) on the SD card when
the physical RAM is fully utilized. This allows the operating system (OS) to continue running
even when RAM is full, which can prevent system crashes or slowdowns.
Swap memory benefits devices with limited RAM, such as the Raspberry Pi Zero. Increasing
swap can help run more demanding applications or processes, but it’s essential to balance this
with the potential performance impact of frequent disk access.
We can check the swap memory using htopor by the command:
48
On the Debian Trixie, 2 MB of swap memory is configured by default on the Rasp-5
and 512MB on the Raspi-Zero
By default, the Rapi-Zero’s SWAP (zram) memory is 512MB, which may be insufficient for
running more complex and demanding Machine Learning applications (for example, YOLO).
So, let’s increase it to 2MB:
On Zero 2 W with Raspberry Pi OS Trixie, you configure rpi-swap by dropping a small
config file into /etc/rpi/[Link].d/. The same mechanism works on Pi 5 and Zero‑2.
Below are two common patterns: fixed swap file (on SD/SSD) and fixed zram size (in RAM).
### Fixed‑size swap file (e.g., 2 GB on SD)
[Main]
Mechanism=swapfile
[File]
FixedSizeMiB=2048
49
4. Apply:
If we want to keep everything in compressed RAM (no SD wear), we can set a fixed zram size
instead:
[Main]
Mechanism=zram
[Zram]
FixedSizeMiB=2048
• Default Trixie on Zero‑2 uses rpi-swap with zram sized roughly to RAM×1, capped
by MaxSizeMiB (usually 2048), and no swap file unless you add it.
• If we create multiple drop‑ins in /etc/rpi/[Link].d/, they are merged; keep it
simple with one clear override file per board/use‑case.
To keep htop running, open another terminal window to continuously interact with
the Raspberry Pi.
50
Installing a Camera
The Raspberry Pi is an excellent device for computer vision applications that require a camera.
We can install a camera module connected to the Raspberry Pi CSI (Camera Serial Interface)
port, or a standard USB webcam on the micro-USB port using a USB OTG adapter (Raspi-
Zero) or directly on the USB port on the Raspi-5.
USB Webcams generally have inferior image quality compared to camera modules
that connect to the CSI port. They can not be controlled using the raspistill and
rasivid commands in the terminal, or the picamera recording package in Python.
Nevertheless, there may be reasons you want to connect a USB camera to your
Raspberry Pi, such as setting up multiple cameras with a single Raspberry Pi,
avoiding long cables, or simply because you have such a camera on hand.
There are now several Raspberry Pi camera modules. The original 5-megapixel model was
releasedin 2013, followed by the 8-megapixel Camera Module 2, released in 2016. The latest
camera model is the 12-megapixel Camera Module 3, released in 2023.
The original 5MP camera (Arducam OV5647) is no longer available from Raspberry Pi but
can be found from several alternative suppliers. Below is an example of such a camera on a
Raspberry Pi Zero.
51
Here is another example of a v2 Camera Module, which has a Sony IMX219 8-megapixel
sensor:
rpicam-hello --list-cameras
Try to list the installed camera again. If Ok, you should see something like:
52
Any camera module will work on the Raspberry Pi, but for that, it is possible that the
[Link] file must be updated:
At the bottom of the file, for example, to use the 5MP Arducam OV5647 camera, add the
line:
dtoverlay=ov5647,cam0
Or for the v2 module, wich has the 8MP Sony IMX219 camera:
dtoverlay=imx219,cam0
Save the file (CTRL+O [ENTER] CRTL+X) and reboot the Raspi:
Sudo reboot
rpicam-hello --list-cameras
53
libcamerais an open-source software library that supports camera systems directly
on Linux for Arm processors. It minimizes the amount of proprietary code running
on the Broadcom GPU.
Let’s capture a JPEG image with a resolution of 640 x 480 for testing and save it to a file
named test_cli_camera.jpg
To view the saved file, we should use ls, which lists the contents of the current directory:
54
As before, we can use scp to view the image:
sudo shutdown -h no
2. Connect the USB Webcam (USB Camera Module 30 fps, 1280x720) to your Raspberry
Pi (in this example, I am using a Raspberry Pi Zero, but the instructions work for all
Raspberry Pis).
55
3. Power on again and run the SSH
4. To check if your USB camera is recognized, run:
lsusb
56
5. To take a test picture with your USB camera, use:
fswebcam test_image.jpg
6. Since we are using SSH to connect to our Rapsi, we must transfer the image to our main
computer so we can view it. We can use FileZilla or SCP for this:
scp mjrovai@[Link]:~/test_image.jpg .
Replace “mjrovai” with your username and “raspi-zero” with Pi’s hostname.
7. If the image quality isn’t satisfactory, you can adjust various settings; for example, define
a resolution that is suitable for YOLO (640x640):
57
fswebcam -r 640x640 --no-banner test_image_yolo.jpg
58
And verified using lsusb
Video Streaming
For stream video (which is more resource-intensive), we can install and use mjpg-streamer:
59
First, install Git:
Now, we should install the necessary dependencies for mjpg-streamer, clone the repository,
and proceed with the installation:
We can then access the stream by opening a web browser and navigating to:
[Link] In my case: [Link]
We should see a webpage with options to view the stream. Click on the link that says “Stream”
or try accessing:
[Link]
60
Running the Raspi Desktop remotely
While we’ve primarily interacted with the Raspberry Pi via SSH terminal commands, we
can access the full graphical desktop environment remotely if we have installed the complete
Raspberry Pi OS (for example, Raspberry Pi OS (64-bit). This can be particularly useful
for tasks that benefit from a visual interface. To enable this functionality, we must set up a
VNC (Virtual Network Computing) server on the Raspberry Pi. Here’s how to do it:
61
• Connect to your Raspberry Pi via SSH.
• Run the Raspberry Pi configuration tool by entering:
sudo raspi-config
62
• Exit the configuration tool (use [Tab]), saving changes when prompted.
• Download and install a VNC viewer application on your main computer. Popular
options include RealVNC Viewer, TightVNC, or VNC Viewer by RealVNC. We
63
will install VNC Viewer by RealVNC.
3. Once installed, confirm the Raspberry Pi’s IP address. For example, on the terminal,
you can use:
hostname -I
64
• When prompted, enter your Raspberry Pi’s username and password.
65
6. Adjust Display Settings (if needed):
• Once connected, adjust the display resolution for optimal viewing. This can be
done through the Raspberry Pi’s desktop settings or by modifying the [Link]
file.
• Let’s do it using the desktop settings. Reach the menu (the Raspberry Icon at the
left upper corner) and select the best screen definition for your monitor:
66
Updating and Installing Software
sudo rm /usr/lib/python3.11/EXTERNALLY-MANAGED
67
• Anything that needs to interface directly with hardware
Rule of thumb: Use sudo apt install only for system dependencies and hardware inter-
faces. Use pip install (without sudo) inside an activated virtual environment for everything
else. Inside the vent, PIP or PIP3 are the same.
Model-Specific Considerations
Remember to adjust your project requirements based on the specific Raspberry Pi model you’re
using. The Raspi-Zero is great for low-power, space-constrained projects, while the Raspi-4 or
5 models are better suited for more computationally intensive tasks.
Here’s a concise, copy‑pasteable tutorial you can share or publish.
In this section, we will explore how to measure (and monitor) CPU temperature and power
consumption on a Raspberry Pi 5 using only built‑in tools and a small shell script.
68
Check CPU Temperature
vcgencmd measure_temp
Example:
temp=47.2'C
cat /sys/class/thermal/thermal_zone0/temp
vcgencmd pmic_read_adc
Example (truncated):
69
3V3_SYS_A current(1)=0.06245952A
1V8_SYS_A current(2)=0.16102850A
...
3V3_SYS_V volt(13)=3.29687500V
1V8_SYS_V volt(14)=1.79687500V
...
• Applies a linear correction to estimate real board power, based on the RPi5‑power cali-
bration.
nano avg_temp_power.sh
Paste:
#!/bin/bash
# avg_temp_power.sh
# Measure average CPU temperature and power on Raspberry Pi 5
TEMP_FILE=$(mktemp)
70
PMIC_FILE=$(mktemp)
sum_pmic=0
for idx in "${!volts[@]}"; do
v=${volts[$idx]}
i_amp=${currents[$idx]}
sum_pmic=$(awk -v a="$sum_pmic" -v b="$v" -v c="$i_amp" \
'BEGIN{printf "%.8f", a + b*c}')
done
71
}
return mean "|" std
}
BEGIN{
split(stats("'"${TEMP_FILE}"'"), t, "|")
t_mean = t[1]
t_std = t[2]
split(stats("'"${PMIC_FILE}"'"), p, "|")
p_pmic_mean = p[1]
p_pmic_std = p[2]
rm -f "${TEMP_FILE}" "${PMIC_FILE}"
Make it executable
chmod +x avg_temp_power.sh
Run a measurement
Idle baseline:
./avg_temp_power.sh
72
Sampling for 50 seconds...
Average_temperature = 47.25 +/- 0.40 °C
Average_pmic_power = 2.312 +/- 0.298 W
Estimated_real_power = 3.235 +/- 0.341 W
Then start a heavy workload (e.g., Llama 3.2:3B) in another terminal and run the script again
during inference. You should see higher temperature and power, often around 70–75 °C and
10–11 W with active cooling.
73
• Throttling usually starts around 80 °C, with a hard limit at 85 °C.[5]
• If we are well below that under full load, your cooling solution is working properly.
74
75
Image Classification Fundamentals
Figure 2: DALL·E prompt - “Create a Cartoon with style from the 50’s doing Image Classi-
fication on a Raspberrry Pi - based on the image uploaded.”
76
Introduction
Image classification has found its way into numerous real-world applications, revolutionizing
various sectors:
Implementing image classification on edge devices such as the Raspberry Pi offers several
compelling advantages:
1. Low Latency: Processing images locally eliminates the need to send data to cloud servers,
significantly reducing response times.
2. Offline Functionality: Classification can be performed without an internet connection,
making it suitable for remote or connectivity-challenged environments.
3. Privacy and Security: Sensitive image data remains on the local device, addressing data
privacy concerns and compliance requirements.
4. Cost-Effectiveness: Eliminates the need for expensive cloud computing resources, espe-
cially for continuous or high-volume classification tasks.
77
5. Scalability: Enables distributed computing architectures in which multiple devices can
operate independently or in a network.
6. Energy Efficiency: Optimized models on dedicated hardware can be more energy-efficient
than cloud-based solutions, which is crucial for battery-powered or remote applications.
7. Customization: Deploying specialized or frequently updated models tailored to specific
use cases is more manageable.
We can create more responsive, secure, and efficient computer vision solutions by leveraging
the power of edge devices such as the Raspberry Pi for image classification. This approach
opens new possibilities for integrating intelligent visual processing across diverse applications
and environments.
In the following sections, we’ll explore how to implement and optimize image classification
on the Raspberry Pi, leveraging these advantages to build powerful, efficient computer vision
systems.
78
Setting up a Virtual Environment
source ~/tflite_env/bin/activate
deactivate
Verify installation
79
System vs pip Package Installation Rule
Rule of thumb: Use sudo apt install only for system dependencies and hard-
ware interfaces. Use pip install (without sudo) inside an activated virtual envi-
ronment for everything else. Inside the vent, PIP or PIP3 are the same.
The virtual environment will automatically include both system packages and pip-
installed packages thanks to the --system-site-packages flag.
Let’s set up Jupyter Notebook optimized for headless Raspberry Pi camera work and develop-
ment:
To run Jupyter Notebook, run the command (change the IP address for yours):
80
On the terminal, you can see the local URL address to open the notebook:
You can access it from another device by entering the Raspberry Pi’s IP address and the
provided token in a web browser (you can copy the token from the terminal).
Define the working directory in the Raspi and create a new Python 3 notebook. For example:
81
cd Documents
mkdir Python
import time
import numpy as np
from PIL import Image
import [Link] as plt
from picamera2 import Picamera2
Load an image from the internet, for example (note that it is possible to run a command line
from the Notebook, using ! before the command:
!wget [Link]
img_path = "[Link]"
img = [Link](img_path)
82
Now, let’s use the camera to capture a local image:
# Initialize camera
picam2 = Picamera2()
[Link]()
83
# Wait for camera to warm up
[Link](2)
# Capture image
picam2.capture_file("class3_test.jpg")
print("Image captured: class3_test.jpg")
# Stop camera
[Link]()
[Link]()
And use a similar code as before to show it (adapting the img_path and title):
84
Installing LiteRT
We are interested in inference, which involves running trained models on a device to make pre-
dictions from input data. To perform an inference with a model, we must run it through an
interpreter. For that, we will use LiteRT, Google’s on-device framework for high-performance
ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimiza-
tion.
LiteRT features advanced GPU/NPU acceleration, delivers superior ML & GenAI performance,
making on-device ML inference easier than ever.
For installation on the Raspi, let’s use the command:
If you are working on the Raspi-Zero with the minimum OS (No Desktop), you do not have a
user-pre-defined directory tree (you can check it with ls. So, let’s create one:
mkdir Documents
cd Documents/
mkdir TFLITE
85
cd TFLITE/
mkdir IMG_CLASS
cd IMG_CLASS
mkdir models
cd models
wget [Link]
wget [Link]
Let’s test our setup by running a simple Python script on TFLITE/IMG_CLASS folder:
86
from ai_edge_litert.interpreter import Interpreter
import numpy as np
from PIL import Image
print("NumPy:", np.__version__)
print("Pillow:", Image.__version__)
We can create the Python script using nano on the terminal, saving it with CTRL+0 + ENTER
+ CTRL+X
python setup_test.py
87
Or you can run it directly on the Notebook:
88
Making inferences with Mobilenet V2
In the last section, we set up the environment, including downloading a popular pre-trained
model, Mobilenet V2, trained on ImageNet’s 224x224 images (1.2 million) for 1,001 classes
(1,000 object categories plus 1 background). The model was converted to a compact 3.5MB
tflite format, making it suitable for the limited storage and memory of a Raspberry Pi.
In the IMG_CLASS working directory, let’s start a new notebook to follow all the steps to classify
one image:
Import the needed libraries:
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
The message means LiteRTsuccessfully enabled an optimized CPU backend (XNNPACK) for
our model, which is good and expected.
What XNNPACK is
89
• XNNPACK is a library of highly optimized operators (conv, FC, etc.) for running neural
networks on CPUs, especially ARM and x86.
• LiteRT can “delegate” supported ops to XNNPACK so they run using these faster kernels
instead of the default reference CPU implementation.
So, it means the interpreter has attached the XNNPACK delegate and will run all compatible
parts of the graph on the CPU using it.
• On devices like the Raspberry Pi, this usually results in lower inference latency at the
cost of slightly longer delivery time and a bit more RAM for packed weights.
• We are currently using CPU acceleration, not GPU, which is the standard/optimal path
for many TFLite/LiteRT models on Pi-class hardware.
input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()
Input details provide information on how the model should be fed an image. The shape of
(1, 224, 224, 3) informs us that an image with dimensions (224x224x3) should be input one by
one (Batch Dimension: 1).
The output details indicate that the inference will produce an array of 1,001 integer values.
Those values result from image classification, where each value is the probability that the
corresponding label is associated with the image.
90
Let’s also inspect the dtype of the input details of the model
input_dtype = input_details[0]['dtype']
input_dtype
dtype('uint8')
This shows that the input image should be represented as raw pixels (0-255).
Let’s get a test image. We can either transfer it from our computer or download one for testing,
as we did before. Let’s first create a folder under our working directory:
mkdir images
cd images
wget [Link]
91
We can see the image size by running the command:
That shows that the image is an RGB image with a width and height of 1600 pixels each. To
use our model, we should reshape it to (224, 224, 3) and add a batch dimension of 1, as defined
in the input details: (1, 224, 224, 3). The inference result, as shown in the output details, will
be an array of size 1001, as shown below:
92
So, let’s reshape the image, add the batch dimension, and see the result:
input_data.dtype
dtype('uint8')
The input data dtype is ‘uint8’, which is compatible with the dtype expected for the model.
Using the input_data, let’s run the interpreter and get the predictions (output):
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
predictions = interpreter.get_tensor(output_details[0]['index'])[0]
The prediction is an array with 1001 elements. Let’s get the Top-5 indices where their elements
have high values:
top_k_results = 5
top_k_indices = [Link](predictions)[::-1][:top_k_results]
top_k_indices
93
The top_k_indices is an array with 5 elements: array([283, 286, 282])
So, 283, 286, 282, 288, and 479 are the image’s most probable classes. Having the index, we
must find to which class it belongs (such as car, cat, or dog). The text file downloaded with
the model includes a label for each index from 0 to 1,000. Let’s use a function to load the .txt
file as a list:
def load_labels(filename):
with open(filename, 'r') as f:
return [[Link]() for line in [Link]()]
And get the list, printing the labels associated with the indexes:
labels_path = "./models/[Link]"
labels = load_labels(labels_path)
print(labels[286])
print(labels[283])
print(labels[282])
print(labels[288])
print(labels[479])
As a result, we have:
Egyptian cat
tiger cat
tabby
lynx
carton
At least four of the top indices are related to felines. The prediction content is the probability
associated with each one of the labels. As we saw in the output details, those values are
quantized and should be dequantized:
The output (positive and negative numbers) shows that the output probably does not have
a Softmax. Checking the model documentation ([Link] Mo-
94
bileNet V2 typically doesn’t include a softmax layer at the output. It usually ends with a
1x1 convolution followed by average pooling and a fully connected layer. So, for getting the
probabilities (0 to 1), we should apply Softmax:
print (probabilities[286])
print (probabilities[283])
print (probabilities[282])
print (probabilities[288])
print (probabilities[479])
0.265947
0.39499295
0.17906114
0.08961108
0.022443123
For clarity, let’s create a function to relate the labels to the probabilities:
for i in range(top_k_results):
print("\t{:20}: {}%".format(
labels[top_k_indices[i]],
(int(probabilities[top_k_indices[i]]*100))))
Let’s create a general function to give an image as input, and we get the Top-5 possible
classes:
95
def image_classification(img_path, model_path, labels, top_k_results=5):
# load the image
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
# Preprocess
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims(img, axis=0)
# Inference on Raspi
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {}%".format(
96
labels[top_k_indices[i]],
(int(probabilities[top_k_indices[i]]*100))))
Let’s modify the Python script used before to capture an image from the camera (size:
224x224), saving it in the images folder:
def capture_image(image_path):
# Initialize camera
picam2 = Picamera2() # default is index 0
97
[Link]()
# Capture image
picam2.capture_file(image_path)
print("Image captured: "+"image_path")
# Stop camera
[Link]()
[Link]()
img_path = './images/cam_img_test.jpg'
model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
labels = load_labels("./models/[Link]")
capture_image(img_path)
image_classification(img_path, model_path, labels, top_k_results=5)
98
Exploring a Model Trained from Zero
Let’s get a TFLite model trained from scratch. For that, we can follow the Notebook:
99
CNN to classify Cifar-10 dataset
In the notebook, we trained a model using the CIFAR10 dataset, which contains 60,000 images
from 10 classes of CIFAR (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and
truck). CIFAR has 32x32 color images (3 color channels) where the objects are not centered
and can have the object with a background, such as airplanes that might have a cloudy sky
behind them! In short, small but real images.
The CNN trained model (cifar10_model.keras) had a size of 2.0MB. Using the TFLite Con-
verter, the model [Link] became with 674MB (around 1/3 of the original size).
Conclusion:
This chapter has established a solid foundation for understanding and implementing image
classification on Raspberry Pi devices using Python and LiteRT. Throughout this journey, we
100
have explored the essential components that make edge-based computer vision both practical
and powerful.
We began by understanding the theoretical foundations of image classification and its real-
world applications across diverse sectors, from healthcare to environmental monitoring. The
advantages of running classification on edge devices like the Raspberry Pi—including low
latency, offline functionality, enhanced privacy, and cost-effectiveness—make it an attractive
solution for many practical applications.
The hands-on experience of setting up the development environment provided crucial insights
into the requirements and constraints of embedded systems. We successfully configured LiteRT,
installed essential Python libraries, and established a working directory structure that serves
as the foundation for computer vision projects.
Working with the pre-trained MobileNet V2 model demonstrated several key concepts:
This foundational knowledge prepares us for more advanced topics, including custom model
training and deployment. The skills developed here—understanding model architectures, im-
plementing inference pipelines, and working with embedded Python environments—are trans-
ferable to a wide range of computer vision applications.
101
The chapter serves as a stepping stone toward building more sophisticated AI systems on
edge devices, demonstrating that powerful computer vision capabilities are accessible even on
modest hardware platforms when properly optimized and implemented.
Resources
• Dataset Example
• Setup Test Notebook on a Raspi
• Image Classification Notebook on a Raspi
• CNN to classify Cifar-10 dataset at CoLab
• Cifar 10 - Image Classification on a Raspi
• Python Scripts
102
103
Custom Image Classification Project
Figure 3: DALL·E prompt - A cover image for an ‘Image Classification’ chapter in a Raspberry
Pi tutorial, designed in the same vintage 1950s electronics lab style as previous
covers. The scene should feature a Raspberry Pi connected to a camera module, with
the camera capturing a photo of the small blue robot provided by the user. The
robot should be placed on a workbench, surrounded by classic lab tools like soldering
irons, resistors, and wires. The lab background should include vintage equipment like
oscilloscopes and tube radios, maintaining the detailed and nostalgic feel of the era.
No text or logos should be included.
104
Image Classification Project
In this chapter, we will develop a complete Image Classification project using the Edge Impulse
Studio. As we did with the MobiliNet V2, the trained and converted TFLite model will be
used for inference using a Python script.
Here is a typical ML workflow that we will use in our project:
The Goal
The first step in any ML project is to define its goal. In this case, it is to detect and classify
two specific objects present in one image. For this project, we will use two small toys: a robot
and a small Brazilian parrot (named Periquito). We will also collect images of a background
where those two objects are absent.
Data Collection
Once we have defined our Machine Learning project goal, the next and most crucial step is
collecting the dataset. We can use a phone for the image capture, but we will use the Raspi
here. Let’s set up a simple web server on our Raspberry Pi to view the QVGA (320 x 240)
captured images in a browser.
105
1. First, let’s install Flask, a lightweight web framework for Python:
2. Go to the working folder (IMG_CLASS) and create a new Python script combining image
capture with a web server. We’ll call it get_img_data.py:
app = Flask(__name__)
# Global variables
base_dir = "dataset"
picam2 = None
frame = None
frame_lock = [Link]()
capture_counts = {}
current_label = None
shutdown_event = [Link]()
def initialize_camera():
global picam2
picam2 = Picamera2()
config = picam2.create_preview_configuration(
main={"size": (320, 240)}
)
[Link](config)
[Link]()
[Link](2) # Wait for camera to warm up
def get_frame():
global frame
while not shutdown_event.is_set():
stream = [Link]()
picam2.capture_file(stream, format='jpeg')
106
with frame_lock:
frame = [Link]()
[Link](0.1) # Adjust as needed for smooth preview
def generate_frames():
while not shutdown_event.is_set():
with frame_lock:
if frame is not None:
yield (b'--frame\r\n'
b'Content-Type: image/jpeg\r\n\r\n' +
frame + b'\r\n')
[Link](0.1) # Adjust as needed for smooth streaming
def shutdown_server():
shutdown_event.set()
if picam2:
[Link]()
# Give some time for other threads to finish
[Link](2)
# Send SIGINT to the main process
[Link]([Link](), [Link])
107
<input type="text" name="label" required>
<input type="submit" value="Start Capture">
</form>
</body>
</html>
''')
@[Link]('/capture')
def capture_page():
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture</title>
<script>
var shutdownInitiated = false;
function checkShutdown() {
if (!shutdownInitiated) {
fetch('/check_shutdown')
.then(response => [Link]())
.then(data => {
if ([Link]) {
shutdownInitiated = true;
[Link](
'video-feed').src = '';
[Link](
'shutdown-message')
.[Link] = 'block';
}
});
}
}
setInterval(checkShutdown, 1000); // Check
every second
</script>
</head>
<body>
<h1>Dataset Capture</h1>
<p>Current Label: {{ label }}</p>
<p>Images captured for this label: {{ capture_count
}}</p>
108
<img id="video-feed" src="{{ url_for('video_feed')
}}" width="640"
height="480" />
<div id="shutdown-message" style="display: none;
color: red;">
Capture process has been stopped.
You can close this window.
</div>
<form action="/capture_image" method="post">
<input type="submit" value="Capture Image">
</form>
<form action="/stop" method="post">
<input type="submit" value="Stop Capture"
style="background-color: #ff6666;">
</form>
<form action="/" method="get">
<input type="submit" value="Change Label"
style="background-color: #ffff66;">
</form>
</body>
</html>
''', label=current_label, capture_count=capture_counts.get(
current_label, 0))
@[Link]('/video_feed')
def video_feed():
return Response(generate_frames(),
mimetype='multipart/x-mixed-replace;
boundary=frame')
@[Link]('/capture_image', methods=['POST'])
def capture_image():
global capture_counts
if current_label and not shutdown_event.is_set():
capture_counts[current_label] += 1
timestamp = [Link]("%Y%m%d-%H%M%S")
filename = f"image_{timestamp}.jpg"
full_path = [Link](base_dir, current_label,
filename)
picam2.capture_file(full_path)
109
return redirect(url_for('capture_page'))
@[Link]('/stop', methods=['POST'])
def stop():
summary = render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture - Stopped</title>
</head>
<body>
<h1>Dataset Capture Stopped</h1>
<p>The capture process has been stopped.
You can close this window.</p>
<p>Summary of captures:</p>
<ul>
{% for label, count in capture_counts.items() %}
<li>{{ label }}: {{ count }} images</li>
{% endfor %}
</ul>
</body>
</html>
''', capture_counts=capture_counts)
return summary
@[Link]('/check_shutdown')
def check_shutdown():
return {'shutdown': shutdown_event.is_set()}
if __name__ == '__main__':
initialize_camera()
[Link](target=get_frame, daemon=True).start()
[Link](host='[Link]', port=5000, threaded=True)
python get_img_data.py
110
4. Access the web interface:
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your
Raspberry Pi’s IP address). For example: [Link]
This Python script creates a web-based interface for capturing and organizing image datasets
using a Raspberry Pi and its camera. It’s handy for machine learning projects that require
labeled image data.
Key Features:
1. Web Interface: Accessible from any device on the same network as the Raspberry Pi.
2. Live Camera Preview: This shows a real-time feed from the camera.
3. Labeling System: Allows users to input labels for different categories of images.
4. Organized Storage: Automatically saves images in label-specific subdirectories.
5. Per-Label Counters: Keeps track of how many images are captured for each label.
6. Summary Statistics: Provides a summary of captured images when stopping the
capture process.
Main Components:
1. Flask Web Application: Handles routing and serves the web interface.
2. Picamera2 Integration: Controls the Raspberry Pi camera.
3. Threaded Frame Capture: Ensures smooth live preview.
4. File Management: Organizes captured images into labeled directories.
Key Functions:
111
Usage Flow:
112
Technical Notes:
• The script uses threading to handle concurrent frame capture and web serving.
• Images are saved with timestamps in their filenames for uniqueness.
• The web interface is responsive and can be accessed from mobile devices.
Customization Possibilities:
Get around 60 images from each category (periquito, robot and background). Try to capture
different angles, backgrounds, and light conditions.
On the Raspi, we will end with a folder named dataset, which contains three sub-folders:
periquito, robot, and background, one for each class of images.
You can use Filezilla to transfer the created dataset to your main computer.
We will use the Edge Impulse Studio to train our model. Go to the Edge Impulse Page, enter
your account credentials, and create a new project:
113
Here, you can clone a similar project: Raspi - Img Class.
Dataset
We will walk through four main steps using the EI Studio (or Studio). These steps are crucial
in preparing our model for use on the Raspi: Dataset, Impulse, Tests, and Deploy (on the
Edge Device, in this case, the Raspi).
Regarding the Dataset, it is essential to point out that our Original Dataset, cap-
tured with the Raspi, will be split into Training, Validation, and Test. The Test
Set will be separated from the beginning and reserved for use only in the Test
phase after training. The Validation Set will be used during training.
1. Go to the Data acquisition tab, and in the UPLOAD DATA section, upload the files from
your computer in the chosen categories.
2. Leave to the Studio the splitting of the original dataset into train and test and choose
the label about
3. Repeat the procedure for all three classes. At the end, you should see your “raw data”
in the Studio:
114
The Studio allows you to explore your data, showing a complete view of all the data in your
project. You can clear, inspect, or change labels by clicking on individual data items. In our
case, a straightforward project, the data seems OK.
• Pre-process our data, which consists of resizing the individual images and determining
the color depth to use (be it RGB or Grayscale) and
115
• Specify a Model. In this case, it will be the Transfer Learning (Images) to fine-tune a
pre-trained MobileNet V2 image classification model on our data. This method performs
well even with relatively small image datasets (around 180 images in our case).
Transfer Learning with MobileNet offers a streamlined approach to model training, which is
especially beneficial for resource-constrained environments and projects with limited labeled
data. MobileNet, known for its lightweight architecture, is a pre-trained model that has already
learned valuable features from a large dataset (ImageNet).
By leveraging these learned features, we can train a new model for your specific task with
fewer data and computational resources and achieve competitive accuracy.
This approach significantly reduces training time and computational cost, making it ideal for
quick prototyping and deployment on embedded devices where efficiency is paramount.
Go to the Impulse Design Tab and create the impulse, defining an image size of 160 × 160 and
squashing them (squared form, without cropping). Select Image and Transfer Learning blocks.
Save the Impulse.
116
Image Pre-Processing
All the input QVGA/RGB565 images will be converted to 76,800 features (160 × 160 × 3).
117
Press Save parameters and select Generate features in the next tab.
Model Design
MobileNet is a family of efficient convolutional neural networks designed for mobile and em-
bedded vision applications. The key features of MobileNet are:
1. Lightweight: Optimized for mobile devices and embedded systems with limited compu-
tational resources.
2. Speed: Fast inference times, suitable for real-time applications.
118
3. Accuracy: Maintains good accuracy despite its compact size.
MobileNetV2, introduced in 2018, improves the original MobileNet architecture. Key features
include:
1. Inverted Residuals: Inverted residual structures are used where shortcut connections are
made between thin bottleneck layers.
2. Linear Bottlenecks: Removes non-linearities in the narrow layers to prevent the destruc-
tion of information.
3. Depth-wise Separable Convolutions: Continues to use this efficient operation from Mo-
bileNetV1.
In our project, we will do a Transfer Learning with the MobileNetV2 160x160 1.0, which
means that the images used for training (and future inference) should have an input Size of
160 × 160 pixels and a Width Multiplier of 1.0 (full width, not reduced). This configuration
balances between model size, speed, and accuracy.
Model Training
Another valuable deep learning technique is Data Augmentation. Data augmentation im-
proves the accuracy of machine learning models by creating additional artificial data. A data
augmentation system makes small, random changes to the training data during the training
process (such as flipping, cropping, or rotating the images).
Looking under the hood, here you can see how Edge Impulse implements a data Augmentation
policy on your data:
119
return image, label
Exposure to these variations during training can help prevent your model from taking shortcuts
by “memorizing” superficial clues in your training data, meaning it may better reflect the deep
underlying patterns in your dataset.
The final dense layer of our model will have 0 neurons with a 10% dropout for overfitting
prevention. Here is the Training result:
The result is excellent, with a reasonable 35 ms of latency (for a Raspi-4), which should result
in around 30 fps (frames per second) during inference. A Raspi-Zero should be slower, and
the Raspi-5, faster.
If faster inference is needed, we should train the model using smaller alphas (0.35, 0.5, and
0.75) or even reduce the image input size, trading with accuracy. However, reducing the input
image size and decreasing the alpha (width multiplier) can speed up inference for MobileNet
V2, but they have different trade-offs. Let’s compare:
Pros:
120
• Significantly reduces the computational cost across all layers.
• Decreases memory usage.
• It often provides a substantial speed boost.
Cons:
• It may reduce the model’s ability to detect small features or fine details.
• It can significantly impact accuracy, especially for tasks requiring fine-grained recogni-
tion.
Pros:
Cons:
Comparison:
1. Speed Impact:
• Reducing input size often provides a more substantial speed boost because it reduces
computations quadratically (halving both width and height reduces computations
by about 75%).
• Reducing alpha provides a more linear reduction in computations.
2. Accuracy Impact:
• Reducing input size can severely impact accuracy, especially when detecting small
objects or fine details.
• Reducing alpha tends to have a more gradual impact on accuracy.
3. Model Architecture:
• Changing input size doesn’t alter the model’s architecture.
• Changing alpha modifies the model’s structure by reducing the number of channels
in each layer.
Recommendation:
1. If our application doesn’t require detecting tiny details and can tolerate some loss in
accuracy, reducing the input size is often the most effective way to speed up inference.
121
2. Reducing alpha might be preferable if maintaining the ability to detect fine details is
crucial or if you need a more balanced trade-off between speed and accuracy.
3. For best results, you might want to experiment with both:
• Try MobileNet V2 with input sizes like 160 × 160 or 92 × 92
• Experiment with alpha values like 1.0, 0.75, 0.5 or 0.35.
4. Always benchmark the different configurations on your specific hardware and with your
particular dataset to find the optimal balance for your use case.
Remember, the best choice depends on your specific requirements for accuracy,
speed, and the nature of the images you’re working with. It’s often worth exper-
imenting with combinations to find the optimal configuration for your particular
use case.
Model Testing
Now, you should take the data set aside at the start of the project and run the trained model
using it as input. Again, the result is excellent (92.22%).
As we did in the previous section, we can deploy the trained model as .tflite and use Raspi to
run it using Python.
On the Dashboard tab, go to Transfer learning model (int8 quantized) and click on the down-
load icon:
122
Let’s also download the float32 version for comparison
Transfer the models from your computer to the Raspi (./models), for example, using FileZilla.
Also, capture some images for inference and save them in (./images), or use the images in
the ./dataset folder.
Let’s remember what we did in the last chapter:
Activate the environment:
source ~/tflite_env/bin/activate
123
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
img_path = "./images/[Link]"
model_path = "./models/ei-raspi-img-class-int8-quantized-\
[Link]"
labels = ['background', 'periquito', 'robot']
Note that the models trained on the Edge Impulse Studio will output values with
index 0, 1, 2, etc., where the actual labels will follow an alphabetic order.
Load the model, allocate the tensors, and get the input and output tensor details:
One important difference to note is that the dtype of the input details of the model is now
int8, which means that the input values go from –128 to +127, while each pixel of our image
goes from 0 to 255. This means that we should pre-process the image to match it. We can
check here:
input_dtype = input_details[0]['dtype']
input_dtype
numpy.int8
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
124
[Link]('off')
[Link]()
Checking the input data, we can verify that the input tensor is compatible with what is
expected by the model:
input_data.shape, input_data.dtype
Now, it is time to perform the inference. Let’s also calculate the latency of the model:
125
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert
# to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))
The model will take around 125ms to perform the inference in the Raspi-Zero, which is 3 to 4
times longer than a Raspi-5.
Now, we can get the output labels and probabilities. It is also important to note that the
model trained on the Edge Impulse Studio has a softmax activation function in its output
(different from the original Movilenet V2), and we can use the model’s raw output as the
“probabilities.”
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {:.2f}%".format(
labels[top_k_indices[i]],
probabilities[top_k_indices[i]] * 100))
126
Let’s modify the function created before so that we can handle different type of models:
# Preprocess
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_dtype = input_details[0]['dtype']
if input_dtype == np.uint8:
input_data = np.expand_dims([Link](img), axis=0)
elif input_dtype == np.int8:
scale, zero_point = input_details[0]['quantization']
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (
img_array / scale
+ zero_point
).clip(-128, 127).astype(np.int8)
127
input_data = np.expand_dims(img_array, axis=0)
else: # float32
input_data = np.expand_dims(
[Link](img, dtype=np.float32),
axis=0
) / 255.0
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time -
start_time
) * 1000 # Convert to milliseconds
# Obtain results
predictions = interpreter.get_tensor(output_details[0]
['index'])[0]
if apply_softmax:
# Apply softmax
exp_preds = [Link](predictions - [Link](predictions))
probabilities = exp_preds / [Link](exp_preds)
else:
probabilities = predictions
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {:.1f}%".format(
128
labels[top_k_indices[i]],
probabilities[top_k_indices[i]] * 100))
print ("\n\tInference time: {:.1f}ms".format(inference_time))
And test it with different images and the int8 quantized model (160x160 alpha =1.0).
Let’s download a smaller model, such as the one trained for the Nicla Vision Lab (int8 quantized
model, 96x96, alpha = 0.1), as a test. We can use the same function:
The model lost some accuracy, but it is still OK once our model does not look for many details.
Regarding latency, we are aboutt ten times faster on the Raspi-Zero.
129
Live Image Classification
Let’s develop an app that captures images with the camera in real-time and displays their
classification.
Using the nano on the terminal, save the code below, such as img_class_live_infer.py.
app = Flask(__name__)
# Global variables
picam2 = None
frame = None
frame_lock = [Link]()
is_classifying = False
confidence_threshold = 0.8
model_path = "./models/ei-raspi-img-class-int8-quantized-\
[Link]"
labels = ['background', 'periquito', 'robot']
interpreter = None
classification_queue = Queue(maxsize=1)
def initialize_camera():
global picam2
picam2 = Picamera2()
config = picam2.create_preview_configuration(
main={"size": (320, 240)}
)
[Link](config)
[Link]()
[Link](2) # Wait for camera to warm up
130
def get_frame():
global frame
while True:
stream = [Link]()
picam2.capture_file(stream, format='jpeg')
with frame_lock:
frame = [Link]()
[Link](0.1) # Capture frames more frequently
def generate_frames():
while True:
with frame_lock:
if frame is not None:
yield (
b'--frame\r\n'
b'Content-Type: image/jpeg\r\n\r\n'
+ frame + b'\r\n'
)
[Link](0.1)
def load_model():
global interpreter
if interpreter is None:
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
return interpreter
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims([Link](img), axis=0)\
.astype(input_details[0]['dtype'])
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
predictions = interpreter.get_tensor(output_details[0]
['index'])[0]
131
# Handle output based on type
output_dtype = output_details[0]['dtype']
if output_dtype in [np.int8, np.uint8]:
# Dequantize the output
scale, zero_point = output_details[0]['quantization']
predictions = ([Link](np.float32) -
zero_point) * scale
return predictions
def classification_worker():
interpreter = load_model()
while True:
if is_classifying:
with frame_lock:
if frame is not None:
img = [Link]([Link](frame))
predictions = classify_image(img, interpreter)
max_prob = [Link](predictions)
if max_prob >= confidence_threshold:
label = labels[[Link](predictions)]
else:
label = 'Uncertain'
classification_queue.put({
'label': label,
'probability': float(max_prob)
})
[Link](0.1) # Adjust based on your needs
@[Link]('/')
def index():
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Image Classification</title>
<script
src="[Link]
</script>
<script>
function startClassification() {
$.post('/start');
132
$('#startBtn').prop('disabled', true);
$('#stopBtn').prop('disabled', false);
}
function stopClassification() {
$.post('/stop');
$('#startBtn').prop('disabled', false);
$('#stopBtn').prop('disabled', true);
}
function updateConfidence() {
var confidence = $('#confidence').val();
$.post('/update_confidence',
{confidence: confidence}
);
}
function updateClassification() {
$.get('/get_classification', function(data) {
$('#classification').text([Link] + ': '
+ [Link](2));
});
}
$(document).ready(function() {
setInterval(updateClassification, 100);
// Update every 100ms
});
</script>
</head>
<body>
<h1>Image Classification</h1>
<img src="{{ url_for('video_feed') }}"
width="640"
height="480" />
<br>
<button id="startBtn"
onclick="startClassification()">
Start Classification
</button>
<button id="stopBtn"
onclick="stopClassification()"
disabled>
133
Stop Classification
</button>
<br>
<label for="confidence">Confidence Threshold:</label>
<input type="number"
id="confidence"
name="confidence"
min="0" max="1"
step="0.1"
value="0.8"
onchange="updateConfidence()" />
<br>
<div id="classification">
Waiting for classification...
</div>
</body>
</html>
''')
@[Link]('/video_feed')
def video_feed():
return Response(
generate_frames(),
mimetype='multipart/x-mixed-replace; boundary=frame'
)
@[Link]('/start', methods=['POST'])
def start_classification():
global is_classifying
is_classifying = True
return '', 204
@[Link]('/stop', methods=['POST'])
def stop_classification():
global is_classifying
is_classifying = False
return '', 204
134
@[Link]('/update_confidence', methods=['POST'])
def update_confidence():
global confidence_threshold
confidence_threshold = float([Link]['confidence'])
return '', 204
@[Link]('/get_classification')
def get_classification():
if not is_classifying:
return jsonify({'label': 'Not classifying',
'probability': 0})
try:
result = classification_queue.get_nowait()
except [Link]:
result = {'label': 'Processing', 'probability': 0}
return jsonify(result)
if __name__ == '__main__':
initialize_camera()
[Link](target=get_frame, daemon=True).start()
[Link](target=classification_worker,
daemon=True).start()
[Link](host='[Link]', port=5000, threaded=True)
python img_class_live_infer.py
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Rasp-
berry Pi’s IP address). For example: [Link]
135
Here, you can see the app running on YouTube.
The code creates a web application for real-time image classification using a Raspberry Pi,
its camera module, and a TensorFlow Lite model. The application uses Flask to serve a web
interface where is possible to view the camera feed and see live classification results.
Key Components:
1. Flask Web Application: Serves the user interface and handles requests.
2. PiCamera2: Captures images from the Raspberry Pi camera module.
3. LiteRT: Runs the image classification model.
4. Threading: Manages concurrent operations for smooth performance.
Main Features:
Code Structure:
136
• Camera and frame management
• Classification control
• Model and label information
3. Camera Functions:
• initialize_camera(): Sets up the PiCamera2
• get_frame(): Continuously captures frames
• generate_frames(): Yields frames for the web feed
4. Model Functions:
• load_model(): Loads the TFLite model
• classify_image(): Performs inference on a single image
5. Classification Worker:
• Runs in a separate thread
• Continuously classifies frames when active
• Updates a queue with the latest results
6. Flask Routes:
• /: Serves the main HTML page
• /video_feed: Streams the camera feed
• /start and /stop: Controls classification
• /update_confidence: Adjusts the confidence threshold
• /get_classification: Returns the latest classification result
7. HTML Template:
• Displays camera feed and classification results
• Provides controls for starting/stopping and adjusting settings
8. Main Execution:
• Initializes camera and starts necessary threads
• Runs the Flask application
Key Concepts:
137
Usage:
Summary:
Image classification has emerged as a powerful and versatile application of machine learning,
with significant implications for various fields, from healthcare to environmental monitoring.
This chapter has demonstrated how to implement a robust image classification system on
edge devices like the Raspi-Zero and Raspi-5, showcasing the potential for real-time, on-device
intelligence.
We’ve explored the entire pipeline of an image classification project, from data collection and
model training using Edge Impulse Studio to deploying and running inferences on a Raspi.
The process highlighted several key points:
1. The importance of proper data collection and preprocessing for training effective models.
2. The power of transfer learning, allowing us to leverage pre-trained models like MobileNet
V2 for efficient training with limited data.
3. The trade-offs between model accuracy and inference speed, especially crucial for edge
devices.
4. The implementation of real-time classification using a web-based interface, demonstrating
practical applications.
The ability to run these models on edge devices like the Raspi opens up numerous possibilities
for IoT applications, autonomous systems, and real-time monitoring solutions. It allows for
reduced latency, improved privacy, and operation in environments with limited connectivity.
As we’ve seen, even with the computational constraints of edge devices, it’s possible to achieve
impressive results in terms of both accuracy and speed. The flexibility to adjust model pa-
rameters, such as input size and alpha values, allows for fine-tuning to meet specific project
requirements.
Looking forward, the field of edge AI and image classification continues to evolve rapidly.
Advances in model compression techniques, hardware acceleration, and more efficient neural
network architectures promise to further expand the capabilities of edge devices in computer
vision tasks.
This project serves as a foundation for more complex computer vision applications and encour-
ages further exploration into the exciting world of edge AI and IoT. Whether it’s for industrial
138
automation, smart home applications, or environmental monitoring, the skills and concepts
covered here provide a solid starting point for a wide range of innovative projects.
Resources
• Dataset Example
• Python Scripts
• Edge Impulse Project
• Image Classification Project - Edge Impulse Notebook
139
140
Object Detection: Fundamentals
Figure 4: DALL·E prompt - A cover image for an ‘Object Detection’ chapter in a Raspberry Pi
tutorial, designed in the same vintage 1950s electronics lab style as previous covers.
The scene should prominently feature wheels and cubes, similar to those provided by
the user, placed on a workbench in the foreground. A Raspberry Pi with a connected
camera module should be capturing an image of these objects. Surround the scene
with classic lab tools like soldering irons, resistors, and wires. The lab background
should include vintage equipment like oscilloscopes and tube radios, maintaining the
detailed and nostalgic feel of the era. No text or logos should be included.
141
Introduction
Building on our exploration of image classification, we now turn to a more advanced computer
vision task: object detection. While image classification assigns a single label to an entire
image, object detection goes further by identifying and locating multiple objects within a
single image. This capability opens up many new applications and challenges, particularly in
edge computing and IoT devices like the Raspberry Pi.
Object detection combines classification and localization. It not only determines which objects
are present in an image but also pinpoints their locations, for example, by drawing bounding
boxes around them. This added complexity makes object detection a more powerful tool
for understanding visual scenes, but it also requires more sophisticated models and training
techniques.
In edge AI, where computational resources are constrained, implementing efficient object de-
tection models is crucial. The challenges we faced with image classification—balancing model
size, inference speed, and accuracy—are even more pronounced in object detection. However,
the rewards are also more significant, as object detection enables more nuanced and detailed
analysis of visual data.
Some applications of object detection on edge devices include:
As we put our hands into object detection, we’ll build on the concepts and techniques we
explored in image classification. We’ll examine popular object detection architectures designed
for efficiency, such as:
To learn more about object detection models, follow the tutorial A Gentle Intro-
duction to Object Recognition With Deep Learning.
142
Throughout this lab, we’ll cover the fundamentals of object detection and how it differs from
image classification. We’ll also learn how to train, fine-tune, test, optimize, and deploy popular
object detection architectures using a dataset created from scratch.
Object detection builds upon the foundations of image classification but extends its capabilities
significantly. To understand object detection, it’s crucial first to recognize its key differences
from image classification:
Image Classification:
Object Detection:
143
To visualize this difference, let’s consider an example:
This diagram illustrates the critical difference: image classification provides a single label for
the entire image, while object detection identifies multiple objects, their classes, and their
locations within the image.
1. Object Localization: This component identifies the location of objects within the image.
It typically outputs bounding boxes, rectangular regions encompassing each detected
object.
2. Object Classification: This component determines the class or category of each detected
object, similar to image classification but applied to each localized region.
• Multiple objects: An image may contain multiple objects of various classes, sizes, and
positions.
• Varying scales: Objects can appear at different sizes within the image.
144
• Occlusion: Objects may be partially hidden or overlapping.
• Background clutter: Distinguishing objects from complex backgrounds can be challeng-
ing.
• Real-time performance: Many applications require fast inference times, especially on
edge devices.
1. Two-stage detectors: These first propose regions of interest and then classify each region.
Examples include R-CNN and its variants (Fast R-CNN, Faster R-CNN).
2. Single-stage detectors: These predict bounding boxes (or centroids) and class probabili-
ties in a single forward pass through the network. Examples include YOLO (You Only
Look Once), EfficientDet, SSD (Single Shot Detector), and FOMO (Faster Objects, More
Objects). These are often faster and better suited to edge devices, such as the Raspberry
Pi.
Evaluation Metrics
• Intersection over Union (IoU) is a metric used to evaluate the accuracy of an object
detector. It measures the overlap between two bounding boxes: the Ground Truth
box (the manually labeled correct box) and the Predicted box (the box generated by
the object detection model). The IoU value is calculated by dividing the area of the
Intersection (the overlapping area) by the area of the Union (the total area covered
by both boxes). A higher IoU value indicates a better prediction.
145
• Mean Average Precision (mAP) is a widely used metric for evaluating the perfor-
mance of object detection models. It provides a single number that reflects a model’s
ability to accurately both classify and localize objects. The “mean” in mAP refers to
the average taken over all object classes in the dataset. The “average precision” (AP) is
calculated for each class, and then these AP values are averaged to get the final mAP
score. A high mAP score indicates that the model is excellent at identifying all objects
and placing a tight-fitting, accurate bounding box around them.
146
• Frames Per Second (FPS): Measures detection speed, crucial for real-time applica-
tions on edge devices.
As we saw in the introduction, given an image or a video stream, an object detection model
can identify which of a known set of objects might be present and provide information about
their positions within the image.
You can test some common models online by visiting Object Detection - MediaPipe
Studio
On Kaggle, we can find the most common pre-trained TFLite models to use with the Raspberry
Pi, ssd_mobilenet_v1, and efficiendet. Those models were trained on the COCO (Common
Objects in Context) dataset, which contains over 200,000 labeled images across 91 categories.
Download the models and upload them to the ./models folder on the Raspberry Pi.
Alternatively, you can find the models and the COCO labels on GitHub.
147
For the first part of this lab, we will focus on a pre-trained 300x300 SSD-Mobilenet V1 model
and compare it with the 320x320 EfficientDet-lite0, also trained using the COCO 2017 dataset.
Both models were converted to a TensorFlow Lite format (4.2MB for the SSD Mobilenet and
4.6MB for the EfficientDet).
The model outputs up to ten detections per image, including bounding boxes,
class IDs, and confidence scores.
We should confirm the steps done on the last Hands-On Lab, Image Classification, as follows:
source ~/tflite/bin/activate
Considering that we have created the Documents/TFLITE folder in the last Lab, let’s now
create the specific folders for this object detection lab:
148
cd Documents/TFLITE/
mkdir OBJ_DETECT
cd OBJ_DETECT
mkdir images
mkdir models
cd models
Let’s start a new notebook to follow all the steps to detect objects in an image:
Import the needed libraries:
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
Download the model and labels from the folder models and save them in the models folder
under OBJ_DETECT.
Load the model and allocate tensors:
model_path = "./models/[Link]"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()
149
Input details will inform us how the model should be fed with an image. The shape of (1,
300, 300, 3) with a dtype of uint8 tells us that a non-normalized (pixel value range from
0 to 255) image with dimensions (300x300x3) should be input one by one (Batch Dimension:
1).
The output details include not only the labels (“classes”) and probabilities (“scores”), but
also the relative window positions of the bounding boxes (“boxes”), indicating where the object
is located in the image, and the number of detected objects (“num_detections”). The output
details also indicate that the model can detect up to 10 objects in the image.
So, for the above example, using the same cat image used with the Image Classification
Lab, looking for the output, we have a 76% probability of having found an object with a
150
class ID of 16 on an area delimited by a bounding box of [0.028011084, 0.020121813,
0.9886069, 0.802299]. Those four numbers are related to ymin, xmin, ymax, and xmax, the
box coordinates.
Considering that y ranges from the top (ymin) to the bottom (ymax) and x ranges from left
(xmin) to right (xmax), we have, in fact, the coordinates of the top-left corner and the bottom-
right one. With both edges and knowing the shape of the picture, it is possible to draw a
rectangle around the object, as shown in the figure below:
151
Next, we should find what class ID 16 means. Opening the file coco_labels.txt, we see that
each element has an associated index; inspecting index 16, we get, as expected, cat. The
probability is the value returned from the score.
Let’s now upload some images with multiple objects on them for testing.
img_path = "./images/cat_dog.jpeg"
orig_img = [Link](img_path)
152
[Link](figsize=(8, 8))
[Link](orig_img)
[Link]("Original Image")
[Link]()
Based on the input details, let’s pre-process the image, changing its shape and expanding its
dimensions:
img = orig_img.resize((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims(img, axis=0)
input_data.shape, input_data.dtype
The new input_data shape is(1, 300, 300, 3) with a dtype of uint8, which is compatible
153
with what the model expects.
Using the input_data, let’s run the interpreter, measure the latency, and get the output:
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))
With a latency of around 800 ms on the Raspi-Zero and 100 ms n the Raspi-5 , we can get
four distinct outputs:
boxes = interpreter.get_tensor(output_details[0]['index'])[0]
classes = interpreter.get_tensor(output_details[1]['index'])[0]
scores = interpreter.get_tensor(output_details[2]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[3]['index'])[0])
On a quick inspection, we can see that the model detected two objects with a score over 0.5:
for i in range(num_detections):
if scores[i] > 0.5: # Confidence threshold
print(f"Object {i}:")
print(f" Bounding Box: {boxes[i]}")
print(f" Confidence: {scores[i]}")
print(f" Class: {classes[i]}")
154
[Link](figsize=(12, 8))
[Link](orig_img)
for i in range(num_detections):
if scores[i] > 0.5: # Adjust threshold as needed
ymin, xmin, ymax, xmax = boxes[i]
(left, right, top, bottom) = (xmin * orig_img.width,
xmax * orig_img.width,
ymin * orig_img.height,
ymax * orig_img.height)
rect = [Link]((left, top), right-left, bottom-top,
fill=False, color='red', linewidth=2)
[Link]().add_patch(rect)
class_id = int(classes[i])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {scores[i]:.2f}',
color='red', fontsize=12, backgroundcolor='white')
155
The choice of the confidence threshold is crucial. For example, setting it to 0.2
will show false positives. A proper code should handle it.
EfficientDet
EfficientDet is not technically an SSD (Single Shot Detector) model, but it shares some simi-
larities and builds upon ideas from SSD and other object detection architectures:
1. EfficientDet:
• Developed by Google researchers in 2019
• Uses EfficientNet as the backbone network
• Employs a novel bi-directional feature pyramid network (BiFPN)
• It uses compound scaling to efficiently scale the backbone network and object de-
tection components.
156
2. Similarities to SSD:
• Both are single-stage detectors, meaning they perform object localization and clas-
sification in a single forward pass.
• Both use multi-scale feature maps to detect objects at different scales.
3. Key differences:
• Backbone: SSD typically uses VGG or MobileNet, while EfficientDet uses Efficient-
Net.
• Feature fusion: SSD uses a simple feature pyramid, while EfficientDet uses the more
advanced BiFPN.
• Scaling method: EfficientDet introduces compound scaling for all components of
the network
4. Advantages of EfficientDet:
• Generally achieves better trade-offs between accuracy and efficiency than SSD and
many other object detection models.
• More flexible scaling enables a family of models with varying size-performance trade-
offs.
While EfficientDet is not an SSD model, it can be seen as an evolution of single-stage detec-
tion architectures, incorporating more advanced techniques to improve efficiency and accuracy.
When using EfficientDet, we can expect outputs similar to those of SSD (e.g., bounding boxes
and class scores).
On GitHub, you can find another notebook exploring the EfficientDet model that
we did with SSD MobileNet.
Object detection models can also detect objects in real-time using a camera. The captured
image should be the input for the trained and converted model. On the Raspberry Pi 4 or 5,
OpenCV can capture frames and display inference results on a desktop.
However, even without a desktop, it’s possible to create a live stream with a webcam to detect
objects in real time. For example, let’s start with the script developed for the Image Classifi-
cation app and adapt it for a Real-Time Object Detection Web Application Using TensorFlow
Lite and Flask.
Download the Python script object_detection_app.py from GitHub.
This app version should work for any TFLite/LiteRT models.
157
model_path = "./models/[Link]"
python object_detection_app.py
After starting, you should receive the message on the terminal (the IP is from my Raspberry):
* Running on [Link]
Press CTRL+C to quit
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Rasp-
berry Pi’s IP address). For example: [Link]
158
159
Let’s see a technical description of the key modules used in the object detection application:
1. LiteRT:
• Purpose: Efficient inference of machine learning models on edge devices.
• Why: LiteRT offers a smaller model size and optimized performance compared to
full TensorFlow, which is crucial for resource-constrained devices like the Raspberry
Pi. It supports hardware acceleration and quantization, further improving efficiency.
• Key functions: Interpreter for loading and running the model, get_input_details(),
and get_output_details() for interfacing with the model.
2. Flask:
• Purpose: Lightweight web framework for building backend servers.
• Why: Flask’s simplicity and flexibility make it ideal for rapidly developing and
deploying web applications. It’s less resource-intensive than larger frameworks suit-
able for edge devices.
• Key components: route decorators for defining API endpoints, Response objects
for streaming video, render_template_string for serving dynamic HTML.
3. Picamera2:
• Purpose: Interface with the Raspberry Pi camera module.
• Why: Picamera2 is the latest library for controlling Raspberry Pi cameras, offering
improved performance and features over the original Picamera library.
• Key functions: create_preview_configuration() for setting up the camera,
capture_file() for capturing frames.
4. PIL (Python Imaging Library):
• Purpose: Image processing and manipulation.
• Why: PIL provides a wide range of image processing capabilities. It’s used here to
resize images, draw bounding boxes, and convert between image formats.
• Key classes: Image for loading and manipulating images, ImageDraw for drawing
shapes and text on images.
5. NumPy:
• Purpose: Efficient array operations and numerical computing.
• Why: NumPy’s array operations are much faster than pure Python lists, which is
crucial for efficiently processing image data and model inputs/outputs.
• Key functions: array() for creating arrays, expand_dims() for adding dimensions
to arrays.
6. Threading:
• Purpose: Concurrent execution of tasks.
160
• Why: Threading enables simultaneous frame capture, object detection, and web
server operation, which is crucial for maintaining real-time performance.
• Key components: Thread class creates separate execution threads, and Lock is used
for thread synchronization.
7. [Link]:
• Purpose: In-memory binary streams.
• Why: Allows efficient handling of image data in memory without needing temporary
files, improving speed and reducing I/O operations.
8. time:
• Purpose: Time-related functions.
• Why: Used for adding delays ([Link]()) to control frame rate and for perfor-
mance measurements.
9. jQuery (client-side):
• Purpose: Simplified DOM manipulation and AJAX requests.
• Why: It makes it easy to update the web interface dynamically and communicate
with the server without page reloads.
• Key functions: .get() and .post() for AJAX requests, DOM manipulation meth-
ods for updating the UI.
1. Main Thread: Runs the Flask server, handling HTTP requests and serving the web
interface.
2. Camera Thread: Continuously captures frames from the camera.
3. Detection Thread: Processes frames using the LiteRT object detection model.
4. Frame Buffer: Shared memory space (protected by locks) storing the latest frame and
detection results.
This architecture enables efficient, real-time object detection while maintaining a responsive
web interface on a resource-constrained edge device, such as a Raspberry Pi. Threading and
efficient libraries, such as LiteRT and PIL, enable the system to process video frames in real-
time, while Flask and jQuery provide a user-friendly way to interact with them.
161
You can test the app with another pre-processed model, such as the EfficientDet, by changing
the app line:
model_path = "./models/lite-model_efficientdet_lite0_detection_metadata_1.tflite"
If we want to use the app with the SSD-MobileNetV2 model, trained in Edge
Impulse Studio with the “Box versus Wheel” dataset, the code should also be
adapted to the input details, as we explored in its notebook.
Conclusion
This lab has explored implementing object detection on edge devices such as the Raspberry
Pi, demonstrating the power and potential of running advanced computer vision tasks on
resource-constrained hardware. We examined the object detection models SSD-MobileNet
and EfficientDet, comparing their performance and trade-offs on edge devices.
The lab demonstrated a real-time object-detection web application, showing how these models
can be integrated into practical, interactive systems.
The ability to perform object detection on edge devices opens up numerous possibilities across
domains such as precision agriculture, industrial automation, quality control, smart home
applications, and environmental monitoring. By processing data locally, these systems can offer
reduced latency, improved privacy, and operation in environments with limited connectivity.
Looking ahead, potential areas for further exploration include: - Using a custom dataset
(labeled on Roboflow), walking through the process of training models using Edge Impulse
Studio and Ultralytics, and deploying them on Raspberry Pi. - To improve inference speed
on edge devices, explore various optimization methods, such as model quantization (TFLite
int8) and format conversion (e.g., to NCNN). - Implementing multi-model pipelines for more
complex tasks - Exploring hardware acceleration options for Raspberry Pi - Integrating object
detection with other sensors for more comprehensive edge AI systems - Developing edge-to-
cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the
power of AI directly into the physical world, opening up new frontiers in how we interact with
and understand our environment.
Resources
162
• Python Scripts
• Models
163
Custom Object Detection Project
164
Object Detection Project
In this chapter, we will develop a complete Object Detection project from data collection,
labelling, training, and deployment. As we did with the Image Classification project, the
trained and converted model will be used for inference.
We will use the same dataset to train 3 models: SSD-MobileNet V2, FOMO, and YOLO.
The Goal
All Machine Learning projects need to start with a goal. Let’s assume we are in an industrial
facility and must sort and count wheels and special boxes.
In other words, we should perform a multi-label classification, where each image can have
three classes:
165
• Background (no objects)
• Box
• Wheel
Once we have defined our Machine Learning project goal, the next and most crucial step is
collecting the dataset. We can use a phone, the Raspi, or a mix to create the raw dataset
(with no labels). Let’s use the simple web app on our Raspberry Pi to view the QVGA (320 x
240) captured images in a browser.
From GitHub, get the Python script get_img_data.py and open it in the terminal:
python3 get_img_data.py
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Rasp-
berry Pi’s IP address). For example: [Link]
The
Python script creates a web-based interface for capturing and organizing image datasets using
a Raspberry Pi and its camera. It’s handy for machine learning projects that require labeled
image data, or not, as in our case here.
166
Access the web interface from a browser, enter a generic label for the images you want to
capture, and press Start Capture.
Note that the images to be captured will have multiple labels that should be defined
later.
Use the live preview to position the camera, then click Capture Image to save the images under
the current label (in this case, box-wheel).
167
When we have enough images, we can press Stop Capture. The captured images are saved in
the folder dataset/box-wheel:
168
Get around 60 images. Try to capture different angles, backgrounds, and light con-
ditions. FileZilla can transfer the raw dataset you created to your main computer.
Labeling Data
The next step in an Object Detect project is to create a labeled dataset. We should label the
raw dataset images, creating bounding boxes around each picture’s objects (box and wheel).
We can use labeling tools such as LabelImg, CVAT, Roboflow, or even the Edge Impulse Studio.
Once we have explored the Edge Impulse tool in other labs, let’s use Roboflow here.
We are using Roboflow (free version) here for two main reasons. 1) We can have an
auto-labeler, and 2) The annotated dataset is available in several formats and can
be used both on Edge Impulse Studio (we will use it for MobileNet V2 and FOMO
train) and on CoLab (YOLOv8 or YOLOv11 train), for example. An annotated
dataset created on Edge Impulse (Free account) cannot be used for training on
other platforms.
We should upload the raw dataset to Roboflow. Create a free account there and start a new
project, for example, (“box-versus-wheel”).
169
We will not go into great detail about the Roboflow process, as many tutorials are
already available.
Annotate
Once the project is created and the dataset is uploaded, you can use the “Auto-Label” tool to
generate annotations, or do it manually.
170
The Label Assist tool can be handy for the labeling process.
171
Note that you should also upload images with only a background, which should be saved w/o
any annotations using the Null Tool option.
172
Once all images are annotated, split them into training, validation, and test sets.
173
Data Pre-Processing
The last step in the dataset is preprocessing to generate a final training version. Let’s resize all
images to 320x320 and generate augmented versions of each image (augmentation) to create
new training examples from which our model can learn.
For augmentation, we will rotate the images (+/-15o ), crop, and vary the brightness and
exposure.
174
At the end of the process, we will have 153 images.
175
Now, you should export the annotated dataset in a format that Edge Impulse, Ultralitics, and
other frameworks/tools understand, for example, YOLOv8 (or v11). Let’s download a zipped
version of the dataset to our desktop.
176
Here, it is possible to review how the dataset was structured
177
There are 3 separate folders, one for each split (train/test/valid). For each of them, there
are 2 subfolders, images, and labels. The pictures are stored as image_id.jpg and im-
ages_id.txt, where “image_id” is unique for every picture.
The labels file format will be class_id bounding box coordinates, where in our case,
class_id will be 0 for box and 1 for wheel. The numerical id (o, 1, 2…) will follow the
alphabetical order of the class name.
The [Link] file contains information about the dataset, such as the classes’ names (names:
['box', 'wheel']) following the YOLO format.
And that’s it! We are ready to start training using Edge Impulse Studio (as we will in the
next step), Ultralytics (as we will when discussing YOLO), or even training from scratch on
CoLab (as we did with the Cifar-10 dataset in the Image Classification lab).
178
Training an SSD MobileNet Model on Edge Impulse Studio
Go to Edge Impulse Studio, enter your credentials at Login (or create an account), and start
a new project.
Here, you can clone the project developed for this hands-on lab: Raspi - Object
Detection.
On the Project Dashboard tab, go down to Project info, and for Labeling method select
Bounding boxes (object detection)
In Studio, go to the Data acquisition tab, and in the UPLOAD DATA section, upload the raw
dataset from your computer.
We can use the Select a folder option, choosing, for example, the train folder on your com-
puter, which contains two sub-folders: images and labels. Select the Image label format,
“YOLO TXT”, upload it into the category Training, and press Upload data.
179
Repeat the process for the test data (upload both folders, test, and validation). At the end
of the upload process, you should end with the annotated dataset of 153 images split in the
train/test (84%/16%).
Note that labels will be stored at the labels files 0 and 1 , which are equivalent to
box and wheel.
The first thing to define when we enter the Create impulse step is to describe the target device
for deployment. A pop-up window will appear. We will select Raspberry 4, an intermediary
device between the Raspi-Zero and the Raspi-5.
This choice will not interfere with the training; it will only give us an idea about
the latency of the model on that specific target.
180
In this phase, you should define how to:
• Pre-processing consists of resizing the individual images. In our case, the images were
pre-processed on Roboflow, to 320x320 , so let’s keep it. The resize will not matter
here because the images are already squared. If you upload a rectangular image, squash
it (squared form, without cropping). Afterward, you could define if the images are
converted from RGB to Grayscale or not.
• Design a Model, in this case, “Object Detection.”
181
Preprocessing all dataset
In the section Image, select Color depth as RGB, and press Save parameters.
182
The Studio automatically moves to the next section, Generate features, where all samples will
be preprocessed, resulting in 480 objects: 207 boxes and 273 wheels.
183
The feature explorer shows that all samples exhibit a good separation after the feature gener-
ation.
For training, we should select a pre-trained model. Let’s use the MobileNetV2 SSD FPN-
Lite (320x320 only).
184
It is a pre-trained object detection model that locates up to 10 objects in an image and outputs
a bounding box for each. The model is approximately 3.7 MB. It supports an RGB input at
320x320px.
• Epochs: 25
• Batch size: 32
• Learning Rate: 0.15.
For validation during training, 20% of the dataset (validation_dataset) will be spared.
185
As a result, the model achieves an overall precision score (based on COCO mAP) of 88.8%,
higher than the score on the test data (83.3%).
• TFLite model, which lets deploy the trained model as .tflite for the Raspberry Pi
to run it using Python.
• Linux (AARCH64), a binary for Linux (AARCH64), implements the Edge Impulse
Linux protocol, which lets us run our models on any Linux-based development board,
with SDKs such as Python. See the documentation for more information and setup
instructions.
Let’s deploy the TFLite model. On the Dashboard tab, go to Transfer learning model (int8
quantized) and click on the download icon:
186
Transfer the model from your computer to the Raspi folder./models and capture or get some
images for inference and save them in the folder ./images.
187
Inference and Post-Processing
The inference can be made as discussed in the Pre-Trained Object Detection Models Overview.
Let’s start a new notebook to follow all the steps to detect cubes and wheels in an image.
Import the needed libraries:
import time
import numpy as np
import [Link] as plt
import [Link] as patches
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
model_path = "./models/ei-raspi-object-detection-SSD-MobileNetv2-320x0320-\
[Link]"
labels = ['box', 'wheel']
Remember that the model will output the class ID as values (0 and 1), following
an alphabetic order regarding the class names.
Load the model, allocate the tensors, and get the input and output tensor details:
One crucial difference to note is that the dtype of the input details of the model is now int8,
which means that the input values go from -128 to +127, while each pixel of our raw image
goes from 0 to 256. This means that we should pre-process the image to match it. We can
check here:
input_dtype = input_details[0]['dtype']
input_dtype
numpy.int8
188
So, let’s open the image and show it:
189
And perform the pre-processing:
190
Checking the input data, we can verify that the input tensor is compatible with what is
expected by the model:
input_data.shape, input_data.dtype
Now, it is time to perform the inference. Let’s also calculate the latency of the model:
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))
The model will take around 600ms to perform the inference in the Raspi-Zero, which is around
5 times longer than a Raspi-5.
Now, we can get the output classes of objects detected, its bounding boxes coordinates, and
probabilities.
boxes = interpreter.get_tensor(output_details[1]['index'])[0]
classes = interpreter.get_tensor(output_details[3]['index'])[0]
scores = interpreter.get_tensor(output_details[0]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[2]['index'])[0])
for i in range(num_detections):
if scores[i] > 0.5: # Confidence threshold
print(f"Object {i}:")
print(f" Bounding Box: {boxes[i]}")
print(f" Confidence: {scores[i]}")
print(f" Class: {classes[i]}")
191
From the results, we can see that 4 objects were detected: two with class ID 0 (box)and two
with class ID 1 (wheel), what is correct!
Let’s visualize the result for a threshold of 0.5
threshold = 0.5
[Link](figsize=(6,6))
[Link](orig_img)
for i in range(num_detections):
if scores[i] > threshold:
ymin, xmin, ymax, xmax = boxes[i]
(left, right, top, bottom) = (xmin * orig_img.width,
xmax * orig_img.width,
ymin * orig_img.height,
ymax * orig_img.height)
rect = [Link]((left, top), right-left, bottom-top,
fill=False, color='red', linewidth=2)
[Link]().add_patch(rect)
class_id = int(classes[i])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {scores[i]:.2f}',
color='red', fontsize=12, backgroundcolor='white')
192
But what happens if we reduce the threshold to 0.3, for example?
193
We start to see false positives and multiple detections, where the model detects the same
object multiple times with different confidence levels and slightly different bounding boxes.
Commonly, sometimes, we need to adjust the threshold to smaller values to capture all objects,
avoiding false negatives, which would lead to multiple detections.
To improve the detection results, we should implement Non-Maximum Suppression
(NMS), which helps eliminate overlapping bounding boxes and keeps only the most confident
detection.
194
For that, let’s create a general function named non_max_suppression(), with the role of
refining object detection results by eliminating redundant and overlapping bounding boxes.
It achieves this by iteratively selecting the detection with the highest confidence score and
removing other significantly overlapping detections based on an Intersection over Union (IoU)
threshold.
keep = []
while [Link] > 0:
i = order[0]
[Link](i)
xx1 = [Link](x1[i], x1[order[1:]])
yy1 = [Link](y1[i], y1[order[1:]])
xx2 = [Link](x2[i], x2[order[1:]])
yy2 = [Link](y2[i], y2[order[1:]])
return keep
How it works:
1. Sorting: It starts by sorting all detections by their confidence scores, highest to lowest.
2. Selection: It selects the highest-scoring box and adds it to the final list of detections.
3. Comparison: This selected box is compared with all remaining lower-scoring boxes.
195
4. Elimination: Any box that overlaps significantly (above the IoU threshold) with the
selected box is eliminated.
5. Iteration: This process repeats with the next highest-scoring box until all boxes are
processed.
Now, we can define a more precise visualization function that will take into consideration an
IoU threshold, detecting only the objects that were selected by the non_max_suppression
function:
# Apply NMS
keep = non_max_suppression(boxes_pixel, scores, iou_threshold)
[Link](image_np)
for i in keep:
if scores[i] > threshold:
ymin, xmin, ymax, xmax = boxes[i]
rect = [Link]((xmin * width, ymin * height),
(xmax - xmin) * width,
(ymax - ymin) * height,
linewidth=2, edgecolor='r', facecolor='none')
ax.add_patch(rect)
class_name = labels[int(classes[i])]
[Link](xmin * width, ymin * height - 10,
f'{class_name}: {scores[i]:.2f}', color='red',
fontsize=12, backgroundcolor='white')
196
[Link]()
Now we can create a function that will call the others, performing inference on any image:
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to ms
print ("Inference time: {:.1f}ms".format(inference_time))
Now, running the code, having the same image again with a confidence threshold of 0.3, but
with a small IoU:
img_path = "./images/box_2_wheel_2.jpg"
detect_objects(img_path, conf=0.3,iou=0.05)
197
Training a FOMO Model at Edge Impulse Studio
The inference with the SSD MobileNet model worked well, but the latency was significantly
high. The inference varied from 0.5 to 1.3 seconds on a Raspi-Zero, which means around or
less than 1 FPS (1 frame per second). One alternative to speed up the process is to use FOMO
(Faster Objects, More Objects).
This novel machine learning algorithm lets us count multiple objects and find their location in
198
an image in real-time using up to 30x less processing power and memory than MobileNet SSD
or YOLO. The main reason this is possible is that while other models calculate the object’s
size by drawing a square around it (bounding box), FOMO ignores the size of the image,
providing only the information about where the object is located in the image through its
centroid coordinates.
In a typical object detection pipeline, the first stage is extracting features from the input image.
FOMO leverages MobileNetV2 to perform this task. MobileNetV2 processes the input
image to produce a feature map that captures essential characteristics, such as textures, shapes,
and object edges, in a computationally efficient way.
Once these features are extracted, FOMO’s simpler architecture, focused on center-point de-
199
tection, interprets the feature map to determine where objects are located in the image. The
output is a grid of cells, where each cell represents whether or not an object center is detected.
The model outputs one or more confidence scores for each cell, indicating the likelihood of an
object being present.
Let’s see how it works on an image.
FOMO divides the image into blocks of pixels using a factor of 8. For the input of 96x96,
the grid would be 12x12 (96/8=12). For a 160x160, the grid will be 20x20, and so on. Next,
FOMO will run a classifier through each pixel block to calculate the probability that there
is a box or a wheel in each of them and, subsequently, determine the regions that have the
highest probability of containing the object (If a pixel block has no objects, it will be classified
as background). From the overlap of the final region, the FOMO provides the coordinates
(related to the image dimensions) of the centroid of this region.
200
• Grid Resolution: FOMO uses a grid of fixed resolution, meaning each cell can detect if
an object is present in that part of the image. While it doesn’t provide high localization
accuracy, it makes a trade-off by being fast and computationally light, which is crucial
for edge devices.
• Multi-Object Detection: Since each cell is independent, FOMO can detect multiple
objects simultaneously in an image by identifying multiple centers.
Return to Edge Impulse Studio, and in the Experiments tab, create another impulse. Now,
the input images should be 160x160 (this is the expected input size for MobilenetV2).
On the Image tab, generate the features and go to the Object detection tab.
We should select a pre-trained model for training. Let’s use the FOMO (Faster Objects,
More Objects) MobileNetV2 0.35.
201
Regarding the training hyper-parameters, the model will be trained with:
• Epochs: 30
• Batch size: 32
• Learning Rate: 0.001.
For validation during training, 20% of the dataset (validation_dataset) will be spared. We
will not apply Data Augmentation for the remaining 80% (train_dataset) because our dataset
was already augmented during the labeling phase at Roboflow.
As a result, the model ends with an overall F1 score of 93.3% with an impressive latency of
8ms (Raspi-4), around 60X less than we got with the SSD MovileNetV2.
202
Note that FOMO automatically added a third label background to the two previ-
ously defined boxes (0) and wheels (1).
On the Model testing tab, we can see that the accuracy was 94%. Here is one of the test
sample results:
203
In object detection tasks, accuracy is generally not the primary evaluation metric.
Object detection involves classifying objects and providing bounding boxes around
them, making it a more complex problem than simple classification. The issue is
that we do not have the bounding box, only the centroids. In short, using accuracy
as a metric could be misleading and may not provide a complete understanding of
how well the model is performing.
As we did in the previous section, we can deploy the trained model as TFLite or Linux
(AARCH64). Let’s do it now as Linux (AARCH64), a binary that implements the Edge
Impulse Linux protocol.
Edge Impulse for Linux models is delivered in .eim format. This executable contains our
“full impulse” created in Edge Impulse Studio. The impulse consists of the signal processing
block(s) and any learning and anomaly block(s) we added and trained. It is compiled with
optimizations for our processor or GPU (e.g., NEON instructions on ARM cores), plus a
straightforward IPC layer (over a Unix socket).
At the Deploy tab, select the option Linux (AARCH64), the int8model and press Build.
204
The model will be automatically downloaded to your computer.
205
On our Raspi, let’s create a new working area:
cd ~
cd Documents
mkdir EI_Linux
cd EI_Linux
mkdir models
mkdir images
The inference will be made using the Linux Python SDK. This library lets us run machine
learning models and collect sensor data on Linux machines using Python. The SDK is open
source and available on GitHub at edgeimpulse/linux-sdk-python.
Let’s set up a Virtual Environment for working with the Linux Python SDK
chmod +x [Link]
206
Install the Jupiter Notebook on the new environment
jupyter notebook
Let’s start a new notebook by following all the steps to detect cubes and wheels on an image
using the FOMO model and the Edge Impulse Linux Python SDK.
Import the needed libraries:
model_file = "[Link]"
model_path = "models/"+ model_file # Trained ML model from Edge Impulse
labels = ['box', 'wheel']
Remember that the model will output the class ID as values (0 and 1), following
an alphabetic order regarding the class names.
# Initialize model
model_info = [Link]()
The model_info will contain critical information about our model. However, unlike the TFLite
interpreter, the EI Linux Python SDK library will now prepare the model for inference.
207
So, let’s open the image and show it (Now, for compatibility, we will use OpenCV, the CV
Library used internally by EI. OpenCV reads the image as BGR, so we will need to convert it
to RGB :
Now we will get the features and the preprocessed image (cropped) using the runner:
208
features, cropped = runner.get_features_from_image_auto_studio_setings(img_rgb)
And perform the inference. Let’s also calculate the latency of the model:
res = [Link](features)
Let’s get the output classes of objects detected, their bounding boxes centroids, and probabil-
ities.
The results show that two objects were detected: one with class ID 0 (box) and one with class
ID 1 (wheel), which is correct!
Let’s visualize the result (The threshold is 0.5, the default value set during the model testing
on the Edge Impulse Studio).
209
height = bbox['height']
210
Conclusion
This chapter has explored the implementation of a custom object detector on edge devices, such
as the Raspberry Pi, demonstrating the power and potential of running advanced computer
vision tasks on resource-constrained hardware. We’ve covered several vital aspects:
211
and FOMO, and compared their performance and trade-offs on edge devices.
2. Training and Deployment: Using a custom dataset of boxes and wheels (labeled on
Roboflow), we walked through the process of training models with Edge Impulse Studio
and Ultralytics and deploying them on a Raspberry Pi.
3. Optimization Techniques: To improve inference speed on edge devices, we explored
various optimization methods, such as model quantization (int8).
4. Performance Considerations: Throughout the lab, we discussed the balance between
model accuracy and inference speed, a critical consideration for edge AI applications.
As discussed earlier, the ability to perform object detection on edge devices opens up numer-
ous possibilities across domains, such as precision agriculture, industrial automation, quality
control, smart home applications, and environmental monitoring. By processing data locally,
these systems can offer reduced latency, improved privacy, and operation in environments with
limited connectivity.
Looking ahead, potential areas for further exploration include: - Implementing multi-model
pipelines for more complex tasks - Exploring hardware acceleration options for Raspberry Pi
- Integrating object detection with other sensors for more comprehensive edge AI systems -
Developing edge-to-cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the
power of AI directly into the physical world, opening up new frontiers in how we interact with
and understand our environment.
Resources
212
Computer Vision Applications with YOLO
In this chapter, we will explore YOLOv8 and v11. Ultralytics YOLO (v8 and v11) are versions
of the acclaimed real-time object detection and image segmentation model, YOLO. YOLOv8
and v11 are built on cutting-edge advances in deep learning and computer vision, offering
unparalleled speed and accuracy. Its streamlined design makes it suitable for a wide range
of applications and easily adaptable across hardware platforms, from edge devices to cloud
APIs.
213
Talking about the YOLO Model
The YOLO (You Only Look Once) model is a highly efficient, widely used object detection
algorithm known for its real-time performance. Unlike traditional object detection systems
that repurpose classifiers or localizers to perform detection, YOLO frames the detection prob-
lem as a single regression task. This innovative approach enables YOLO to simultaneously
predict multiple bounding boxes and their class probabilities from full images during a single
evaluation, significantly boosting its speed.
Key Features:
• YOLO employs a single neural network to process the entire image. This network
divides the image into a grid and, for each grid cell, directly predicts bounding
boxes and associated class probabilities. This end-to-end training improves speed
and simplifies the model architecture.
2. Real-Time Processing:
• One of YOLO’s standout features is its ability to perform object detection in real-
time. Depending on the version and hardware, YOLO can process images at high
frames per second (FPS). This makes it ideal for applications requiring quick and
accurate object detection, such as video surveillance, autonomous driving, and live
sports analysis.
3. Evolution of Versions:
• Over the years, YOLO has undergone significant improvements, from YOLOv1 to
the latest YOLOv12. Each iteration has introduced enhancements in accuracy,
speed, and efficiency. YOLOv8, for instance, incorporates advancements in net-
work architecture, improved training methodologies, and better support for various
hardware, ensuring a more robust performance.
• YOLOv11 offers substantial improvements in accuracy, speed, and parameter effi-
ciency compared to prior versions such as YOLOv8 and YOLOv10, making it one
of the most versatile and powerful real-time object detection models available as of
2025
214
4. Accuracy and Efficiency:
• While early versions of YOLO traded off some accuracy for speed, recent versions
have made substantial strides in balancing both. The newer models are faster
and more accurate, detecting small objects (such as bees) and performing well on
complex datasets.
• YOLO’s versatility has led to its adoption in numerous fields. It is used in traffic
monitoring systems to detect and count vehicles, security applications to identify
potential threats and agricultural technology to monitor crops and livestock. Its
application extends to any domain requiring efficient and accurate object detection.
7. Model Capabilities
YOLO models support multiple computer vision tasks:
215
• Classification: Image classification tasks
Ultralitics YOLO Detect, Segment, and Pose models pre-trained on the COCO
dataset, and Classify on the ImageNet dataset.
Track mode is available for all Detect, Segment, and Pose models. The latest versions of
YOLO can also perform OBB, which stands for Oriented Bounding Box, a rectangular
box in computer vision that can rotate to match the orientation of an object within
an image, providing a much tighter and more precise fit than traditional axis-aligned
bounding boxes.
YOLO offers several model variants optimized for different use cases, for example. The
YOLOv8:
Installation
python --version
216
If we use the latest Raspberry Pi OS (based on Debian Trixie), it should be:
3.13.5
As of today (January 2026), Ultralytics officially supports only Python 3.9-3.12; Python
3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships with
Python 3.13 by default, we’ll need to install a compatible Python version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python.
If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.
Install pyenv
Configure Shell
# pyenv configuration
export PYENV_ROOT="$HOME/.pyenv"
[[ -d $PYENV_ROOT/bin ]] && export PATH="$PYENV_ROOT/bin:$PATH"
eval "$(pyenv init -)"
EOF
217
source ~/.bashrc
pyenv --version
cd Documents
mkdir YOLO
cd YOLO
# Verify
python --version # Should show Python 3.11.14
218
which python
python --version
deactivate
Installing Ultralytics/Yolo
cd Documents/
cd YOLO
mkdir models
mkdir images
And install the Ultralytics packages for local inference on the Raspberry Pi (inside the env)
1. Update the packages list, install and/or upgrade PIP to the latest:
sudo reboot
After the Raspi booting, let’s activate the yolo env, go to the working directory,
cd /Documents/YOLO
source ~/yolo_env/bin/activate
And run inference on an image that will be downloaded from the Ultralytics website, using,
for example, the YOLOV8n model (the smallest in the family) at the Terminal (CLI):
219
yolo predict model='yolov8n' source='[Link]
Note that the first time we invoke a model, it will automatically be downloaded to
the current directory.
The inference result will appear in the terminal. In the image ([Link]), 4 persons, 1 bus,
and 1 stop signal were detected:
Also, we got a message that Results saved to runs/detect/predict. Inspecting that di-
rectory, we can see a new image saved ([Link]). Let’s download it from the Raspi to our
desktop for inspection:
220
221
So, the Ultrayitics YOLO is correctly installed on our Raspberry Pi. Note that on the Rasp-
berry Pi Zero, an issue is the high latency for this inference, which takes several seconds, even
with the most compact model in the family (YOLOv8n).
The procedure is the same as we did with version v8. As a comparison, we can see that the
YOLOv11 is faster than the v8, but seems a little less precise, as it does not detect the “stop
sign” as the v8.
Deploying computer vision models on edge devices with limited computational power, such as
the Raspberry Pi Zero, can cause latency issues. One alternative is to use a format optimized
for optimal performance. This ensures that even devices with limited processing power can
handle advanced computer vision tasks well.
Of all the model export formats supported by Ultralytics, the NCNN is a high-performance
neural network inference computing framework optimized for mobile platforms. From the
beginning of the design, NCNN was deeply considerate of deployment and use on mobile
phones, and it did not have third-party dependencies. It is cross-platform and runs faster than
all known open-source frameworks (such as TFLite).
NCNN delivers the best inference performance when working with Raspberry Pi devices.
NCNN is highly optimized for mobile embedded platforms (such as ARM architecture).
Let’s move the downloaded YOLO models to the ./models folder and [Link] to
./images.
222
And convert our models and rerun the inferences:
The first inference, when the model is loaded, typically has a high latency; however,
from the second inference, it is possible to note that the inference time decreases.
We can now realize that neither model detects the “Stop Signal”, with YOLOv11 being the
fastest. The optimized models are more rapid but also less accurate.
To start, let’s call the Python Interpreter so we can explore how the YOLO model works, line
by line:
python
Now, we should call the YOLO library from Ultralitics and load the model:
223
from ultralytics import YOLO
model = YOLO('./models/yolov8n_ncnn_model')
img = './images/[Link]'
result = [Link](img, save=True, imgsz=640, conf=0.5, iou=0.3)
We can verify that the result is almost identical to the one we get running the inference at
the terminal level (CLI), except that the bus stop was not detected with the reduced NCNN
model. Note that the latency was reduced.
Let’s analyze the “result” content.
For example, we can see result[0].[Link], showing us the main inference result, which
is a tensor with a shape of (4, 6). Each line is one of the objects detected, being the first four
columns, the bounding boxes coordinates, the 5th, the confidence, and the 6th, the class (in
this case, 0: person and 5: bus):
224
We can access several inference results separately, as the inference time, and have it printed
in a better format:
inference_time = int(result[0].speed['inference'])
print(f"Inference Time: {inference_time} ms")
With Python, we can create a detailed output that meets our needs (See Model Prediction with
Ultralytics YOLO for more details). Let’s run a Python script instead of manually entering
it line by line in the interpreter, as shown below. Let’s use nano as our text editor. First, we
should create an empty Python script named, for example, yolov8_tests.py:
nano yolov8_tests.py
# Run inference
img = './images/[Link]'
result = [Link](img, save=False, imgsz=640, conf=0.5, iou=0.3)
225
inference_time = int(result[0].speed['inference'])
print(f"Inference Time: {inference_time} ms")
print(f'Number of objects: {len (result[0].[Link])}')
And enter with the commands: [CTRL+O] + [ENTER] +[CTRL+X] to save the Python script.
Run the script:
python yolov8_tests.py
The result is the same as running the inference at the terminal level (CLI) and with the built-in
Python interpreter.
Calling the YOLO library and loading the model for inference for the first time
takes a long time, but the inferences after that will be much faster. For example,
the first single inference can take several seconds, but after that, the inference time
should be reduced to less than 1 second.
226
Inference Arguments
[Link]() accepts multiple arguments that can be passed at inference time to override
defaults:
Inference arguments:
227
Argument Type Default Description
batch int 1 Specifies the batch size for inference (only
works when the source is a directory, video
file or .txt file). A larger batch size can
provide higher throughput, shortening the
total amount of time required for inference.
max_det int 300 Maximum number of detections allowed per
image. Limits the total number of objects
the model can detect in a single inference,
preventing excessive outputs in dense scenes.
vid_stride int 1 Frame stride for video inputs. Allows
skipping frames in videos to speed up
processing at the cost of temporal resolution.
A value of 1 processes every frame, higher
values skip frames.
stream_buffer
bool False Determines whether to queue incoming
frames for video streams. If False, old
frames get dropped to accommodate new
frames (optimized for real-time applications).
If True, queues new frames in a buffer,
ensuring no frames get skipped, but will
cause latency if inference FPS is lower than
stream FPS.
visualize bool False Activates visualization of model features
during inference, providing insights into what
the model is “seeing”. Useful for debugging
and model interpretation.
augment bool False Enables test-time augmentation (TTA) for
predictions, potentially improving detection
robustness at the cost of inference speed.
agnostic_nmsbool False Enables class-agnostic Non-Maximum
Suppression (NMS), which merges
overlapping boxes of different classes. Useful
in multi-class detection scenarios where class
overlap is common.
classes list[int] None Filters predictions to a set of class IDs. Only
detections belonging to the specified classes
will be returned. Useful for focusing on
relevant objects in multi-class detection
tasks.
228
Argument Type Default Description
retina_masksbool False Returns high-resolution segmentation masks.
The returned masks ([Link]) will match
the original image size if enabled. If disabled,
they have the image size used during
inference.
embed list[int] None Specifies the layers from which to extract
feature vectors or embeddings. Useful for
downstream tasks like clustering or similarity
search.
project str None Name of the project directory where
prediction outputs are saved if save is
enabled.
name str None Name of the prediction run. Used for
creating a subdirectory within the project
folder, where prediction outputs are stored if
save is enabled.
stream bool False Enables memory-efficient processing for long
videos or numerous images by returning a
generator of Results objects instead of
loading all frames into memory at once.
verbose bool True Controls whether to display detailed
inference logs in the terminal, providing
real-time feedback on the prediction process.
Visualization arguments:
229
Argument Type Default Description
save_txt bool False Saves detection results in a text file, following the
format [class] [x_center] [y_center]
[width] [height] [confidence]. Useful for
integration with other analysis tools.
save_conf bool False Includes confidence scores in the saved text files.
Enhances the detail available for post-processing
and analysis.
save_crop bool False Saves cropped images of detections. Useful for
dataset augmentation, analysis, or creating
focused datasets for specific objects.
show_labels bool True Displays labels for each detection in the visual
output. Provides immediate understanding of
detected objects.
show_conf bool True Displays the confidence score for each detection
alongside the label. Gives insight into the model’s
certainty for each detection.
show_boxes bool True Draws bounding boxes around detected objects.
Essential for visual identification and location of
objects in images or video frames.
line_width None or None Specifies the line width of bounding boxes. If
int None, the line width is automatically adjusted
based on the image size. Provides visual
customization for clarity.
Let’s set up Jupyter Notebook optimized for headless Raspberry Pi camera work and develop-
ment:
To run Jupyter Notebook, run the command (change the IP address for yours):
On the terminal, you can see the local URL address and its Token to open the notebook. Copy
and paste it into the Browser.
230
Environment Setup and Dependencies
import time
import numpy as np
from PIL import Image
from ultralytics import YOLO
import [Link] as plt
Here we have all the necessary libraries, which we installed automatically when we installed
Ultralytics.
model_path= "./models/[Link]"
task = "detect"
verbose = False
• Model Selection: YOLOv11n (nano) is chosen for its balance of speed and accuracy
• Task Specification: We will select detect, which in fact is the default for the model.
But remember that YOLO supports multiple computer vision tasks, which will be ex-
plored later.
• Verbose Control: output model information during model initialization
Performance Characteristics
source = [Link]("./images/[Link]")
231
From the inference results info, we can see that the first time an inference is run, the latency
is greater.
# First inference
0: 640x480 4 persons, 1 bus, 7528.3ms
# Second inference
0: 640x480 4 persons, 1 bus, 2822.1ms
The dramatic difference between the first inference (7.5s) and subsequent inferences (2.8s)
illustrates:
result = results[0]
# - boxes, keypoints, masks, names
# - orig_img, orig_shape, path
# - speed metrics
232
8. Visualization and Customization
The Ultralytics plot() can be customized to show as the detection result, for example, only
the bounding boxes:
[Link](figsize=(6, 6))
[Link](img)
#[Link]('off') # This turns off the axis numbers
[Link]("YOLO Result")
[Link]()
233
234
Customization Options:
The plot() method in Ultralytics YOLO Results object accepts several arguments to control
what is visualized on the image, including boxes, masks, keypoints, confidences, labels, and
more. Common Arguments for plot()
Image Classification
As explored in previous chapters, the output of an image classifier is a single class label and
a confidence score. Image classification is useful when we need to know only what class an
image belongs to and don’t need to know where objects of that class are located or what their
exact shape is.
model_path= "./models/[Link]"
task = "clasification"
Note that a specific variation of the model, for image classification, will be downloaded. Now,
let’s do an inference, using the same bus image:
0: 224x224 minibus 0.57, police_van 0.34, trolleybus 0.04, recreational_vehicle 0.01, stre
Speed: 5233.9ms preprocess, 3355.1ms inference, 28.2ms postprocess per image at shape (1,
235
We can check the top5 inference results using Python:
classes = [Link].top5
classes
for id in classes:
print([Link][id])
minibus
police_van
trolleybus
recreational_vehicle
streetcar
probs = [Link]()
probs
[0.5710113048553467,
0.33745330572128296,
0.04209813103079796,
0.014150412753224373,
0.005880324635654688]
print([Link][[Link].top1],
round([Link](), 2))
minibus 0.57
236
Instance Segmentation
model_path= "./models/[Link]"
task = "segment"
Note that a specific variation of the model, for instance segmentation, will be downloaded.
Now, lt’s use another image for testing:
source = [Link]("./images/[Link]")
[Link](figsize=(6, 6))
[Link](source)
#[Link]('off') # This turns off the axis numbers
[Link]("Original Image")
[Link]()
237
And run the inference:
Pose Estimation
238
model_path= "./models/[Link]"
task = "pose"
source = [Link]("./images/[Link]")
results = [Link](source, save=False)
result = results[0]
239
Training YOLO on a Customized Dataset
We will now develop a customized object detection project from the data collected and labelled
with Roboflow. The training and deployment will be done in Python using a CoLab and
Ultralytics functions.
We will use with YOLO, the same dataset previously used to train the SSD-
MobileNet V2 and FOMO models.
As a reminder, we are assuming we are in an industrial facility that must sort and count
wheels and special boxes.
240
Each image can have three classes:
The Dataset
Return to our “Boxe versus Wheel” dataset, labeled on Roboflow. On the Download Dataset,
instead of Download a zip to computer option done for training on Edge Impulse Studio,
we will opt for Show download code. This option will open a pop-up window with a code
snippet that should be pasted into our training notebook.
241
For training, let’s choose one model (let’s say YOLOv8) and adapt one of the publicly available
examples from Ultralytics, then run it on Google Colab. Below, you can find my adaptation:
242
2. Install Ultralytics using PIP.
3. Now, you can import the YOLO and upload your dataset to the CoLab, pasting the
Download code that we get from Roboflow. Note that our dataset will be mounted
under /content/datasets/:
4. It is essential to verify and change the file [Link] with the correct path for the images
(copy the path on each images folder).
names:
- box
- wheel
nc: 2
243
roboflow:
license: CC BY 4.0
project: box-versus-wheel-auto-dataset
url: [Link]
version: 5
workspace: marcelo-rovai-riila
test: /content/datasets/Box-versus-Wheel-auto-dataset-5/test/images
train: /content/datasets/Box-versus-Wheel-auto-dataset-5/train/images
val: /content/datasets/Box-versus-Wheel-auto-dataset-5/valid/images
5. Define the main hyperparameters that you want to change from default, for example:
MODEL = '[Link]'
IMG_SIZE = 640
EPOCHS = 25 # For a final project, you should consider at least 100 epochs
The model took a few minutes to be trained and has an excellent result (mAP50 of
0.995). At the end of the training, all results are saved in the folder listed, for example:
/runs/detect/train/. There, you can find, for example, the confusion matrix.
244
7. Note that the trained model ([Link]) is saved in the folder /runs/detect/train/weights/.
Now, you should validate the trained model with the valid/images.
8. Now, we should perform inference on the images left aside for testing
245
!yolo task=detect mode=predict model={HOME}/runs/detect/train/weights/[Link] conf=0.25 so
The inference results are saved in the folder runs/detect/predict. Let’s see some of them:
9. It is advised to export the train, validation, and test results for a Drive at Google. To
do so, we should mount the drive.
and copy the content of /runs folder to a folder that you should create in your Drive,
for example:
246
Let’s return to the YOLO folder and use the Python Interpreter:
cd ..
python
We will import the YOLO library and define the model to use::
Now, let’s define an image and call the inference (we will save the image result this time to
external verification):
Let’s repeat for several images. The inference result is saved on the variable result, and the
processed image on runs/detect/predict8
Using FileZilla FTP, we can send the inference result to our Desktop for verification:
247
We can see that the inference result is excellent! The model was trained based on the smaller
base model of the YOLOv8 family (YOLOv8n). The issue is the latency, around 1 second (or
1 FPS on the Raspi-Zero). We can reduce this latency and convert the model to TFLite or
NCNN.
The model trained with YOLO11 has a latency of around 800 ms, similar to the
result of v8 with ncnn.
In the last section of the notebook, we can find inferences made with the trained YOLO11n
model on a Raspberry 5, which took around 400ms:
248
The same model, when exported to NCNN, took around 80 ms.
Conclusion
This chapter has explored the YOLO model and the implementation of a custom object detec-
tor on a Raspberry Pi, demonstrating the power and potential of running advanced computer
vision tasks on resource-constrained hardware. We’ve covered several vital aspects:
249
4. Performance Considerations: Throughout the lab, we discussed the balance between
model accuracy and inference speed, a critical consideration for edge AI applications.
AS discussed before, the ability to perform object detection on edge devices opens up numer-
ous possibilities across various domains, including precision agriculture, industrial automation,
quality control, smart home applications, and environmental monitoring. By processing data
locally, these systems can offer reduced latency, improved privacy, and operation in environ-
ments with limited connectivity.
Looking ahead, potential areas for further exploration include: - Implementing multi-model
pipelines for more complex tasks - Exploring hardware acceleration options for Raspberry Pi
- Integrating object detection with other sensors for more comprehensive edge AI systems -
Developing edge-to-cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the
power of AI directly into the physical world, opening up new frontiers in how we interact with
and understand our environment.
Resources
250
Counting objects with YOLO
251
Introduction
At the Federal University of Itajuba in Brazil, with the master’s student José Anderson Reis
and Professor José Alberto Ferreira Filho, we are exploring a project that delves into the
intersection of technology and nature. This tutorial will review our first steps and share our
observations on deploying YOLOv8, a cutting-edge machine learning model, on the compact
252
and efficient Raspberry Pi Zero 2W (Raspi-Zero). We aim to estimate the number of bees
entering and exiting their hive—a task crucial for beekeeping and ecological studies.
Why is this important? Bee populations are vital indicators of environmental health, and their
monitoring can provide essential data for ecological research and conservation efforts. However,
manual counting is labor-intensive and prone to errors. By leveraging the power of embedded
machine learning, or tinyML, we automate this process, enhancing accuracy and efficiency.
Figure 5: img
This tutorial will cover setting up the Raspberry Pi, integrating a camera module, optimizing
and deploying YOLOv8 for real-time image processing, and analyzing the data gathered.
For our project at the university, we are preparing to collect a dataset of bees at the entrance
of a beehive using the same camera connected to the Raspberry Pi. The images should be
collected every 10 seconds. With the Arducam OV5647, the horizontal Field of View (FoV)
is 53.5o , which means that a camera positioned at the top of a standard Hive (46 cm) will
capture all of its entrance (about 47 cm).
253
Dataset
The dataset collection is the most critical phase of the project and should take several weeks or
months. For this tutorial, we will use a public dataset: “Sledevic, Tomyslav (2023), “[Labeled
dataset for bee detection and direction estimation on beehive landing boards,” Mendeley Data,
V5, doi: 10.17632/8gb9r2yhfc.5”
The original dataset contains 6,762 images (1920 x 1080), and around 8% (518) of them have
no bees (only background). This is very important in Object Detection, where we should keep
around 10% of the dataset with only background (no objects to be detected).
254
The images contain from zero to up to 61 bees:
We downloaded the dataset (images and annotations) and uploaded it to Roboflow. There, you
should create a free account and start a new project, for example, (“Bees_on_Hive_landing_boards”):
255
We will not enter details about the Roboflow process once many tutorials are
available.
Once the project is created and the dataset is uploaded, you should review the annotations
using the “Auto-Label” Tool. Note that all images with only a background should be saved
w/o any annotations. At this step, you can also add additional images.
256
Once all images are annotated, you should split them into training, validation, and testing.
257
Pre-Processing
The last step with the dataset is preprocessing to generate a final version for training. The
Yolov8 model can be trained with 640 x 640 pixels (RGB) images. Let’s resize all images and
generate augmented versions of each image (augmentation) to create new training examples
from which our model can learn.
For augmentation, we will rotate the images (+/-15o ) and vary the brightness and exposure.
258
This will create a final dataset of 16,228 images.
259
Now, you should export the annotateddataset in a YOLOv8 format. You can download a
zipped version of the dataset to your desktop or get a downloaded code to be used with a
Jupyter Notebook:
260
And that is it! We are prepared to start our training using Google Colab.
For training, let’s adapt one of the public examples available from Ultralitytics and run it on
Google Colab:
261
Critical points on the Notebook:
3. Now, you can import the YOLO and upload your dataset to the CoLab, pasting the
Download code that you get from Roboflow. Note that your dataset will be mounted
under /content/datasets/:
262
4. It is important to verify and change, if needed, the file [Link] with the correct path
for the images:
names:
- bee
nc: 1
roboflow:
license: CC BY 4.0
project: bees_on_hive_landing_boards
url: [Link]
version: 1
workspace: marcelo-rovai-riila
test: /content/datasets/Bees_on_Hive_landing_boards-1test/images
train: /content/datasets/Bees_on_Hive_landing_boards-1/train/images
val: /content/datasets/Bees_on_Hive_landing_boards-1/valid/images
5. Define the main hyperparameters that you want to change from default, for example:
MODEL = '[Link]'
IMG_SIZE = 640
EPOCHS = 25 # For a final project, you should consider at least 100 epochs
The model took 2.7 hours to train and has an excellent result (mAP50 of 0.984). At the end
of the training, all results are saved in the folder listed, for example: /runs/detect/train3/.
There, you can find, for example, the confusion matrix and the metrics curves per epoch.
263
7. Note that the trained model ([Link]) is saved in the folder /runs/detect/train3/weights/.
Now, you should validade the trained model with the valid/images.
8. Now, we should perform inference on the images left aside for testing
The inference results are saved in the folder runs/detect/predict. Let’s see some of them:
We can also perform inference with a completely new and complex image from another beehive
with a different background (the beehive of Professor Maurilio of our University). The results
were great (but not perfect and with a lower confidence score). The model found 41 bees.
264
9. The last thing to do is export the train, validation, and test results for your Drive at
Google. To do so, you should mount your drive.
and copy the content of /runs folder to a folder that you should create in your Drive,
for example:
Using the FileZilla FTP, let’s transfer the [Link] to our Rasp-Zero (before the transfer, you
may change the model name, for example, bee_landing_640_best.pt).
The first thing to do is convert the model to an NCNN format:
265
yolo export model=bee_landing_640_best.pt format=ncnn
mkdir test_images
Using the FileZilla FTP, let’s transfer a few images from the test dataset to our Rasp-Zero:
python
As before, we will import the YOLO library and define our converted model to detect bees:
Now, let’s define an image and call the inference (we will save the image result this time to
external verification):
266
img = 'test_images/15_bees.jpg'
result = [Link](img, save=True, imgsz=640, conf=0.2, iou=0.3)
The inference result is saved on the variable result, and the processed image on
runs/detect/predict9
Using FileZilla FTP, we can send the inference result to our Desktop for verification:
267
let’s go over the other images, analyzing the number of objects (bees) found:
268
Depending on the confidence level, we may see some false positives or negatives. But in
general, with a model trained on a smaller base model in the YOLOv8 family (YOLOv8n)
and converted to NCNN, the results are pretty good, running on an Edge device such as the
Rasp-Zero. Also, note that the inference latency is around 730ms.
For example, by running the inference on [Link], we can find 40 bees. During
the test phase on Colab, 41 bees were found (we only missed one here).
Our final project should be very simple in terms of code. We will use the camera to capture
an image every 10 seconds. As we did in the previous section, the captured image should serve
as the input to the trained and converted model. We should count the bees in each image and
store the counts in a database (e.g., timestamp: number of bees).
We can do it with a single Python script, or use a Linux system timer, such as cron, to
periodically capture images every 10 seconds, and have a separate Python script process them
269
as they are saved. This method can be particularly efficient at managing system resources and
is more robust against potential delays in image processing.
First, we should set up a cron job to use the rpicam-jpeg command to capture an image
every 10 seconds.
• Open the terminal and type crontab -e to edit the cron jobs.
• cron normally doesn’t support sub-minute intervals directly, so we should use a
workaround, such as a loop or a file watcher.
• Image Capture: This bash script captures images every 10 seconds using
rpicam-jpeg, a command in the raspijpeg tool. This command lets us control
the camera and capture JPEG images directly from the command line. This is
especially useful because we are looking for a lightweight, straightforward method
to capture images without requiring additional libraries like Picamera or external
software. The script also saves the captured image with a timestamp.
#!/bin/bash
# Script to capture an image every 10 seconds
while true
do
DATE=$(date +"%Y-%m-%d_%H%M%S")
rpicam-jpeg --output test_images/$[Link] --width 640 --height 640
sleep 10
done
Image Processing: The Python script continuously monitors the designated directory for
new images, processes each new image using the YOLOv8 model, updates the database with
the count of detected bees, and optionally deletes the image to conserve disk space.
270
Database Updates: The results, along with the timestamps, are saved in an SQLite database.
For that, a simple option is to use sqlite3.
In short, we need to write a script that continuously monitors the directory for new images,
processes them using a YOLO model, and then saves the results to a SQLite database. Here’s
how we can create and make the script executable:
#!/usr/bin/env python3
import os
import time
import sqlite3
from datetime import datetime
from ultralytics import YOLO
def setup_database():
"""
Establishes a database connection and creates the table
if it doesn't exist.
"""
conn = [Link](DB_PATH)
cursor = [Link]()
[Link]('''
CREATE TABLE IF NOT EXISTS bee_counts
(timestamp TEXT, count INTEGER)
''')
[Link]()
return conn
271
[Link]("INSERT INTO bee_counts (timestamp, count) VALUES (?, ?)",
(timestamp, num_bees)
)
[Link]()
print(f'Processed {image_path}: Number of bees detected = {num_bees}')
def main():
conn = setup_database()
model = YOLO(MODEL_PATH)
monitor_directory(model, conn)
[Link]()
if __name__ == "__main__":
main()
chmod +x process_images.py
272
3. Run the script directly from the command line:
./process_images.py
We should consider keeping the script running even after closing the terminal; for that, we can
use nohup or screen:
or
screen -S bee_monitor
./process_images.py
Note that we capture images with their own timestamps and log a separate timestamp when the
inference results are saved to the database. This approach can be beneficial for the following
reasons:
273
#!/usr/bin/env python3
import sqlite3
def main():
db_path = 'bee_count.db'
conn = [Link](db_path)
cursor = [Link]()
query = "SELECT * FROM bee_counts"
[Link](query)
data = [Link]()
for row in data:
print(f"Timestamp: {row[0]}, Number of bees: {row[1]}")
[Link]()
if __name__ == "__main__":
main()
Besides bee counting, environmental data, such as temperature and humidity, are essential
for monitoring the bee-have health. Using a Rasp-Zero, it is straightforward to add a digital
sensor such as the DHT-22 to get this data.
274
Environmental data will be part of our final project. If you want to know more about connect-
ing sensors to a Raspberry Pi and, even more, how to save the data to a local database and
send it to the web, follow this tutorial: From Data to Graph: A Web Journey With Flask and
SQLite.
275
Conclusion
In this tutorial, we have thoroughly explored integrating the YOLOv8 model with a Rasp-
berry Pi Zero 2W to address the practical, pressing task of counting (or, better, “estimating”)
bees at a beehive entrance. Our project underscores the robust capability of embedding ad-
vanced machine learning technologies within compact edge computing devices, highlighting
their potential impact on environmental monitoring and ecological studies.
This tutorial provides a step-by-step guide to deploying the YOLOv8 model in practice. We
demonstrate a tangible real-world application by optimizing it for edge computing, improving
efficiency and processing speed (using the NCNN format). This not only serves as a functional
solution but also as an instructional tool for similar projects.
The technical insights and methodologies shared in this tutorial are the basis for the complete
work to be developed at our university in the future. We envision further development, such
as integrating additional environmental sensing capabilities and refining the model’s accuracy
276
and processing efficiency. Implementing alternative energy solutions, such as the proposed
solar power setup, will enhance the project’s sustainability and applicability in remote or
underserved locations.
Resources
The Dataset paper, Notebooks, and PDF version are in the Project repository.
277
Image Classification with EXECUTORCH
Introduction
Image classification is a fundamental computer vision task that powers countless real-world
applications—from quality control in manufacturing to wildlife monitoring, medical diagnos-
tics, and smart home devices. In the edge AI landscape, the ability to run these models effi-
278
ciently on resource-constrained devices has become increasingly critical for privacy-preserving,
low-latency applications.
In the chapter Image Classification Fundamentals, we explored image classification with Ten-
sorFlow Lite and demonstrated how to deploy efficient neural networks on the Raspberry
Pi. That tutorial covered the complete workflow from model conversion to real-time camera
inference, achieving excellent results with the MobileNet V2 architecture and a real dataset
(CIFAR-10).
This chapter takes a parallel approach using PyTorch EXECUTORCH—Meta’s modern
solution for edge deployment. Rather than replacing our TFLite knowledge, this chapter
expands your edge AI toolkit, giving us the flexibility to choose the right framework for our
specific needs.
What is EXECUTORCH?
EXECUTORCH is PyTorch’s official solution for deploying machine learning models on edge
devices, from smartphones and embedded systems to microcontrollers and IoT devices. Re-
leased in 2023, it represents Meta’s commitment to bringing the entire PyTorch ecosystem to
edge computing.
Core Capabilities:
• Native PyTorch Integration: Seamless workflow from model training to edge deploy-
ment without switching frameworks
• Efficient Execution: Optimized runtime designed specifically for resource-constrained
devices
• Broad Portability: Runs on diverse hardware platforms (ARM, x86, specialized accel-
erators)
• Flexible Backend System: Extensible delegate architecture for hardware-specific op-
timizations
• Quantization Support: Built-in integration with PyTorch’s quantization tools for
model compression
279
2. Modern Architecture Built from the ground up for edge computing with contemporary
best practices, EXECUTORCH incorporates lessons learned from previous mobile deployment
frameworks.
3. Comprehensive Quantization Native support for various quantization techniques (dy-
namic, static, quantization-aware training) enables significant model size reduction with min-
imal accuracy loss.
4. Extensible Backend System The delegate system allows seamless integration with
hardware accelerators (XNNPACK for CPU optimization, QNN for Qualcomm chips, CoreML
for Apple devices, and more).
5. Active Development Backed by Meta with rapid iteration and strong community support,
ensuring the framework evolves with edge AI needs.
6. Growing Model Zoo Access to pretrained models specifically optimized for edge deploy-
ment, with consistent performance across devices.
Understanding when to choose each framework is crucial for effective edge deployment:
280
2. Team expertise and existing infrastructure
3. Specific hardware requirements
4. Project timeline and maturity needs
Install Python tools, camera libraries, and build dependencies for PyTorch:
rpicam-hello --list-cameras
281
We should see that the OV5647 cam is installed.
import numpy as np
from picamera2 import Picamera2
import time
# Initialize camera
picam2 = Picamera2()
config = picam2.create_preview_configuration(main={"size":(640,480)})
[Link](config)
[Link]()
# Capture image
picam2.capture_file("camera_capture.jpg")
print("Image captured: cam_test.jpg")
# Stop camera
282
[Link]()
[Link]()
python --version
If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.
Install pyenv
283
Configure Shell
# pyenv configuration
export PYENV_ROOT="$HOME/.pyenv"
[[ -d $PYENV_ROOT/bin ]] && export PATH="$PYENV_ROOT/bin:$PATH"
eval "$(pyenv init -)"
EOF
source ~/.bashrc
pyenv --version
cd Documents
mkdir EXECUTORCH
cd EXECUTORCH
284
# Set Python 3.11.14 for this directory
pyenv local 3.11.14
# Verify
python --version # Should show Python 3.11.14
deactivate
Verify installation:
285
PyTorch and EXECUTORCH Installation
For the Raspberry Pi Zero 2 W (32-bit ARM), we may need to build from
source or use lighter alternatives, which are not covered here.
# Install dependencies
./install_requirements.sh
286
Verifying the Setup
Let’s verify our setup with a test script. Create setup_test.py (for example, using nano):
import torch
import numpy as np
from PIL import Image
import executorch
print("=" * 50)
print("SETUP VERIFICATION")
print("=" * 50)
# Check versions
print(f"PyTorch version: {torch.__version__}")
print(f"NumPy version: {np.__version__}")
print(f"PIL version: {Image.__version__}")
print(f"EXECUTORCH available: {executorch is not None}")
# Test PIL
test_img = [Link]('RGB', (224, 224), color='red')
print(f"Created test PIL image: {test_img.size}")
Run it:
python setup_test.py
==================================================
SETUP VERIFICATION
==================================================
PyTorch version: 2.9.1+cpu
NumPy version: 2.2.6
PIL version: 12.1.0
287
EXECUTORCH available: True
Working directory:
cd Documents
cd EXECUTORCH
mkdir IMG_CLASS
cd IMG_CLASS
mkdir MOBILENET
cd MOBILENET
mkdir models images notebooks
wget "[Link] \
-O ./images/[Link]
Now, let’s create a test program where we should take into consideration:
288
and save it as img_class_test_torch.py:
import torch
import [Link] as transforms
from torchvision import models
from PIL import Image
import time
import json
import [Link]
import os
# Paths
MODEL_PATH = "models/mobilenet_v2.pth"
LABELS_PATH = "models/imagenet_labels.json"
IMAGE_PATH = "images/[Link]"
289
model = models.mobilenet_v2()
model.load_state_dict([Link](MODEL_PATH, map_location='cpu'))
[Link]()
with torch.no_grad():
output = model(batch)
# Get predictions
probabilities = [Link](output[0], dim=0)
top5_prob, top5_idx = [Link](probabilities, 5)
# Display results
print("\n" + "="*50)
print("CLASSIFICATION RESULTS")
print("="*50)
print(f"Inference Time: {inference_time:.2f} ms\n")
print("Top 5 Predictions:")
print("-"*50)
for i in range(5):
290
idx = top5_idx[i].item()
prob = top5_prob[i].item()
print(f"{i+1}. {labels[idx]:20s} - {prob*100:.2f}%")
print("="*50)
The result:
==================================================
CLASSIFICATION RESULTS
==================================================
Inference Time: 86.12 ms
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================
The inference was OK, taking 86ms (first time). We can also verify the size of the saved Torch
model
ls -lh ./models/mobilenet_v2.pth
Unlike TensorFlow Lite, where we downloaded pre-converted .tflite models, with EXECU-
TORCH, we typically export PyTorch models to the .pte (PyTorch EXECUTORCH) format
ourselves. This gives us full control over the export process.
291
Understanding the Export Process
Let’s export a MobileNet V2 model to EXECUTORCH basic format. Creating a Python script
as convert_mobv2_executorch.py
import torch
from torchvision import models
from [Link] import to_edge
from [Link] import export
# Paths
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
292
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2.pte"
print("\n" + "="*50)
print("MODEL SIZE COMPARISON")
print("="*50)
print(f"PyTorch model: {pytorch_size:.2f} MB")
293
print(f"ExecuTorch model: {executorch_size:.2f} MB")
print(f"Reduction: {((pytorch_size - executorch_size) \
/pytorch_size * 100):.1f}%")
print("="*50)
python export_mobv2_executorch.py
We will get:
==================================================
MODEL SIZE COMPARISON
==================================================
PyTorch model: 13.60 MB
ExecuTorch model: 13.58 MB
Reduction: 0.2%
==================================================
The basic ExecuTorch conversion doesn’t compress the model much - it’s mainly for runtime
efficiency. To get real size reduction, we need quantization, which we will explore later.
But first, let’s do an inference test using the converted model.
Runing the script mobv2_executorch.py:
import torch
import [Link] as transforms
from PIL import Image
import time
import json
from [Link].portable_lib import _load_for_executorch
# Paths
294
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2.pte"
LABELS_PATH = "models/imagenet_labels.json"
IMAGE_PATH = "images/[Link]"
# Load labels
print("Loading labels...")
with open(LABELS_PATH, 'r') as f:
labels = [Link](f)
# Get predictions
output_tensor = output[0] # ExecuTorch returns a list
probabilities = [Link](output_tensor[0], dim=0)
top5_prob, top5_idx = [Link](probabilities, 5)
295
# Display results
print("\n" + "="*50)
print("EXECUTORCH CLASSIFICATION RESULTS")
print("="*50)
print(f"Inference Time: {inference_time:.2f} ms\n")
print("Top 5 Predictions:")
print("-"*50)
for i in range(5):
idx = top5_idx[i].item()
prob = top5_prob[i].item()
print(f"{i+1}. {labels[idx]:20s} - {prob*100:.2f}%")
print("="*50)
As a result, we got a similar inference result, but a much higher latency (almost 2.5 seconds),
which was unexpected.
Loading labels...
Loading ExecuTorch model from models/mobilenet_v2.pte...
Loading image from images/[Link]...
Running ExecuTorch inference...
==================================================
EXECUTORCH CLASSIFICATION RESULTS
==================================================
Inference Time: 2445.78 ms
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================
That export path produces a generic ExecuTorch CPU graph with reference kernels and no
backend optimizations or fusions, so significantly higher latency than PyTorch is expected for
MobileNet_v2 on a Pi 5.
ExecuTorch is designed to shine when delegated to a backend (XNNPACK, OpenVINO, etc.),
296
where large subgraphs are lowered into highly optimized kernels. Without a delegate, most of
the graph runs on the generic portable path, which is known to be significantly slower than
PyTorch for many models.
So, let’s export the .pth model again with a CPU‑optimized backend (e.g., XNNPACK) and
run with that backend enabled; this alone should reduce latency when compared with the
naïve interpreter path.
Here’s the corrected conversion script with XNNPACK delegation (convert_mobv2_xnnpack.py):
import torch
from torchvision import models
from [Link] import to_edge
from [Link] import export
from [Link].xnnpack_partitioner \
import XnnpackPartitioner
# Paths
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2_xnnpack.pte"
297
# Step 4: Convert to ExecuTorch program
print(" 4. Lowering to ExecuTorch...")
executorch_program = edge_program.to_executorch()
print("\n" + "="*50)
print("MODEL SIZE COMPARISON")
print("="*50)
print(f"PyTorch model: {pytorch_size:.2f} MB")
print(f"ExecuTorch+XNNPACK: {executorch_size:.2f} MB")
print("="*50)
Runing it we get:
==================================================
MODEL SIZE COMPARISON
==================================================
PyTorch model: 13.60 MB
ExecuTorch+XNNPACK: 13.35 MB
==================================================
We did not gain in terms of size, but let’s run the same inference script as before, with this
298
new converted model, to inspect the latency:
the result:
Loading labels...
Loading ExecuTorch model from models/mobilenet_v2_xnnpack.pte...
Loading image from images/[Link]...
Running ExecuTorch inference...
==================================================
EXECUTORCH CLASSIFICATION RESULTS
==================================================
Inference Time: 19.95 ms
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================
Now, the ExecuTorch runtime detects the backend automatically from the .pte file metadata.
We have achieved much faster inference: 20ms instead of 2445ms. This latency is, in fact,
several times faster than PyTorch.
Why XNNPACK is so fast:
This demonstrates:
Now we can add quantization to get an even smaller model size while maintaining (or even
increasing) this speed!
299
Model Quantization
Quantization reduces model size and can further improve inference speed. EXECUTORCH
supports PyTorch’s native quantization.
Quantization Overview
Quantization is a technique that reduces the precision of numbers used in a model’s compu-
tations and stored weights—typically from 32-bit floats to 8-bit integers. This reduces the
model’s memory footprint, speeds up inference, and lowers power consumption, often with
minimal loss in accuracy.
Quantization is especially important for deploying models on edge devices such as wearables,
embedded systems, and microcontrollers, which often have limited compute, memory, and
battery capacity. By quantizing models, we can make them significantly more efficient and
better suited to these resource-constrained environments.
Quantization in ExecuTorch
ExecuTorch uses torchao as its quantization library. This integration allows ExecuTorch to
leverage PyTorch-native tools for preparing, calibrating, and converting quantized models.
Quantization in ExecuTorch is backend-specific. Each backend defines how models should
be quantized based on its hardware capabilities. Most ExecuTorch backends use the torchao
PT2E quantization flow, which works with models exported with [Link] and enables
tailored quantization for each backend.
For a quantized XNNPACK .pte we need a different pipeline: PT2E quantization (with
XNNPACKQuantizer), then lowering with XnnpackPartitioner before to_executorch(). Oth-
erwise, we will hit errors or get an undelegated model.
For the conversion, we need: (1) calibrate with real, preprocessed images, and (2) compute
the quantized .pte size after you actually write the file.
First, let us create a small calib_images/ folder (e.g., 50–100 natural images across a few
classes). A simple way is to reuse an existing dataset (e.g., CIFAR‑10) and save 50–100 images
into calib_images/ with an ImageNet‑style folder layout.
The script gen_calibr_images.py will: • Download CIFAR‑10. • Pick 10 classes × 10 images
each = 100 images. • Save them under calib_images/<class_name>/img_XXX.jpg.
import os
from pathlib import Path
import torch
from torchvision import datasets, transforms
from [Link] import save_image
300
# Where to store calibration images
OUT_ROOT = Path("calib_images")
OUT_ROOT.mkdir(parents=True, exist_ok=True)
idx = counts[cls_name]
out_path = class_dir / f"img_{idx:04d}.jpg"
save_image(img, out_path)
counts[cls_name] += 1
301
if all(counts[c] >= images_per_class for c in counts):
break
Let’s use the inference script convert_mobv2_xnnpack_int8.py, which is the same inference
script as before, with this new int8 converted model to inspect the latency:
import os
import torch
import [Link] as models
import [Link] as transforms
import [Link] as datasets
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
EXECUTORCH_QUANTIZED_PATH = "models/mobilenet_v2_quantized_xnnpack.pte"
CALIB_IMAGES_DIR = "calib_images" # <-- put some natural images here
302
example_inputs = ([Link](1, 3, 224, 224),)
calib_dataset = [Link](CALIB_IMAGES_DIR,
transform=calib_transform)
calib_loader = [Link](
calib_dataset, batch_size=1, shuffle=True
)
303
et_program = to_edge_transform_and_lower(
exported_quant,
partitioner=[XnnpackPartitioner()],
).to_executorch()
pytorch_size = [Link](PYTORCH_MODEL_PATH)/(1024*1024)
quantized_size = [Link](EXECUTORCH_QUANTIZED_PATH)/(1024*1024)
print("\n" + "="*60)
print("MODEL SIZE COMPARISON")
print("="*60)
print(f"PyTorch (FP32): {pytorch_size:6.2f} MB")
print(f"ExecuTorch Quantized (INT8): {quantized_size:6.2f} MB")
print(f"Size reduction: {((pytorch_size - quantized_size) \
/ pytorch_size * 100):5.1f}%")
print(f"Savings: {pytorch_size - quantized_size:6.2f} MB")
print("="*60)
============================================================
MODEL SIZE COMPARISON
============================================================
PyTorch (FP32): 13.60 MB
ExecuTorch Quantized (INT8): 3.59 MB
Size reduction: 73.6%
Savings: 10.01 MB
============================================================
The quantized (int8) model achieved 74% size reduction: ~3.5 MB (similar to TFLite). Let’s
see about the inference latency, runing mobv2_xnnpack_int8.py.
Loading labels...
Loading ExecuTorch model from models/mobilenet_v2_quantized_xnnpack.pte...
304
Loading image from images/[Link]...
Running ExecuTorch inference (Quantized INT8)...
==================================================
EXECUTORCH QUANTIZED INT8 RESULTS
==================================================
Inference Time: 13.56 ms
Output dtype: torch.float32
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 51.01%
2. Egyptian Mau - 34.11%
3. lynx - 7.54%
4. tabby cat - 6.17%
5. plastic bag - 0.37%
==================================================
Slightly higher top‑1 probabilities in the INT8 model are normal and do not indicate
a problem by themselves. Quantization slightly changes the logits, and softmax
can become a bit “sharper” or “flatter” even when top‑1 remains correct.
NOTE
• Looking at Htop, we can see that only one of the Pi’s cores is at 100%. This indicates
that the shipped Python runtime currently runs our ExecuTorch/XNNPACK model
effectively single‑threaded on Pi.
• To exploit all four cores, the next step would be to move inference into a small C++
wrapper that sets the ExecuTorch threadpool size before executing the graph. With the
pure‑Python path, there is no clean public knob to change it yet. We will not explore it
here.
305
Making Inferences with EXECUTORCH
Now that we have our EXECUTORCH models, let’s explore them in more detail for image
classification using a Jupyter Notebook!
jupyter notebook
Access it from another device using the provided token in your web browser.
EXECUTORCH/MOBILENET/
��� convert_mobv2_executorch.py
��� convert_mobv2_xnnpack.py
��� convert_mobv2_xnnpack_int8.py
��� mobv2_executorch.py
��� mobv2_xnnpack.py
��� mobv2_xnnpack_int8.py
��� calib_images/
��� data/
��� models/
� ��� mobilenet_v2.pth # Float32 pytorch model
306
� ��� mobilenet_v2.pte # Float32 conv model
� ��� mobilenet_v2_xnnpack.pte # Float32 conv model
� ��� mobilenet_v2_quantized_xnnpack.pte # Quantized conv model
� ��� imagenet_labels.json # Labels
��� images/ # Test images
� ��� [Link]
� ��� camera_capture.jpg
��� notebooks/
��� image_classification_executorch.ipynb
Inside the folder ‘notebooks’, on the project space IMAGE_CLASS/MOBILENET, create a new
notebook: image_classification_executorch.ipynb.
print("=" * 50)
print("SETUP VERIFICATION")
print("=" * 50)
# Check versions
print(f"PyTorch version: {torch.__version__}")
print(f"NumPy version: {np.__version__}")
print(f"PIL version: {Image.__version__}")
print(f"EXECUTORCH available: {executorch is not None}")
307
# Test basic PyTorch functionality
x = [Link](3, 224, 224)
print(f"\nCreated test tensor with shape: {[Link]}")
# Test PIL
test_img = [Link]('RGB', (224, 224), color='red')
print(f"Created test PIL image: {test_img.size}")
We get:
==================================================
SETUP VERIFICATION
==================================================
PyTorch version: 2.9.1+cpu
NumPy version: 2.2.6
PIL version: 12.1.0
EXECUTORCH available: True
img_path = "../images/[Link]"
308
#[Link]('off')
[Link]()
• python export_mobv2_executorch.py
309
imagenet_labels.json mobilenet_v2_quantized_xnnpack.pte
mobilenet_v2.pte mobilenet_v2_xnnpack.pte
mobilenet_v2.pth
The conversions were performed using the Python scripts in the previous sections.
try:
model = _load_for_executorch(model_path)
print(f"Model loaded successfully from: {model_path}")
#print(f" Available methods: {model.method_names}")
except FileNotFoundError:
print(f"� Model not found: {model_path}")
print("\nPlease run the export script first:")
print(" python export_mobilenet.py")
# Download and save ImageNet labels (if you do not have it)
LABELS_PATH = "../models/imagenet_labels.json"
if not [Link](LABELS_PATH):
print("Downloading ImageNet labels...")
LABELS_URL = "[Link]
imagenet-simple-labels/master/[Link]"
with [Link](LABELS_URL) as url:
labels = [Link](url)
310
[Link](labels, f)
print(f"Labels saved to {LABELS_PATH}")
else:
print("Loading labels from disk...")
with open(LABELS_PATH, 'r') as f:
labels = [Link](f)
Image Preprocessing
A preprocessing pipeline is needed because ExecuTorch only runs the exported core network;
it does not include the input normalization logic that MobileNet v2 expects, and the model
will give incorrect predictions if the input tensor is not in the exact format it was trained on.
What MobileNet v2 expects For typical PyTorch MobileNet v2 models (ImageNet‑pretrained):
• Input shape: 3‑channel RGB tensor of size. • Value range: floating-point values, usually in
float32 after dividing by 255. • Normalization: per‑channel mean/std (ImageNet) normaliza-
tion, e.g., mean=0.485, 0.456, 0.406, std=0.229, 0.224, 0.225.
These steps (resize, convert to tensor, normalize) are not “optional decorations”; they are part
of the functional definition of the model’s expected input distribution.
preprocess = [Link]([
[Link](256), # Resize to 256
[Link](224), # Center crop to 224x224
[Link](), # Convert to tensor [0, 1]
[Link]( # Normalize with ImageNet stats
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
311
Apply preprocessing
input_tensor = preprocess(img)
print(f" Input shape: {input_tensor.shape}")
print(f" Input dtype: {input_tensor.dtype}")
input_batch = input_tensor.unsqueeze(0)
Run Inference
For inference, we should run a forward pass of the model in inference mode (torch.no_grad()),
measure the time, and print basic information about the outputs.
torch.no_grad() is a context manager that disables gradient calculation inside its block.
During inference, we do not need gradients, so disabling them:
# Run inference
with torch.no_grad():
312
start_time = [Link]()
outputs = [Link]((input_batch,))
inference_time = [Link]() - start_time
type(outputs) tells us what container the model returned. Often this is a tuple or list when
working with exported/ExecuTorch‑style models, e.g., <class 'tuple'>.
That container may hold one or more tensors (e.g., logits, auxiliary outputs).
• outputs[0] accesses the first element of that container (usually the main output tensor),
and .shape prints its dimensions (For image classification, this is often batch_size,
num_classes).
Now we should take the model’s raw scores (logits) for a single image, convert them into prob-
abilities with softmax, select the top‑5 most likely classes, and print them nicely formatted.
• outputs[0][0] selects the first element in the batch, giving a 1D tensor of logits of
length num_classes.
• [Link](..., dim=0) applies the softmax function along that
1D dimension, turning logits into probabilities that sum to 1.
# Display results
print("\n" + "="*60)
print("TOP 5 PREDICTIONS")
print("="*60)
313
print(f"{'Class':<35} {'Probability':>10}")
print("-"*60)
for i in range(5):
label = labels[top5_indices[i]]
prob = top5_prob[i].item() * 100
print(f"{label:<35} {prob:>9.2f}%")
print("="*60)
============================================================
TOP 5 PREDICTIONS
============================================================
Class Probability
------------------------------------------------------------
tiger cat 12.85%
Egyptian cat 9.75%
tabby 6.09%
lynx 1.70%
carton 0.84%
============================================================
For simplicity and reuse across other tests, let’s create a reusable function that builds on what
was done so far.
Args:
img_path: Path to input image
model_path: Path to .pte model file
labels_path: Path to labels text file
top_k: Number of top predictions to return
show_image: Whether to display the image
Returns:
314
inference_time: Inference time in ms
top_indices: Indices of top k predictions
top_probs: Probabilities of top k predictions
"""
# Load image
img = [Link](img_path).convert('RGB')
# Display image
if show_image:
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
[Link]('Input Image')
[Link]()
# Load model
print(f"Model Path {model_path}")
model_size = [Link](model_path) / (1024 * 1024)
print(f"Model size: {model_size:6.2f} MB")
model = _load_for_executorch(model_path)
# Preprocess
preprocess = [Link]([
[Link](256),
[Link](224),
[Link](),
[Link](
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
input_tensor = preprocess(img)
input_batch = input_tensor.unsqueeze(0)
# Inference
with torch.no_grad():
start_time = [Link]()
315
outputs = [Link]((input_batch,))
inference_time = ([Link]() - start_time)*1000
# Process results
probabilities = [Link](outputs[0][0], dim=0)
top_prob, top_indices = [Link](probabilities, top_k)
# Load labels
with open(labels_path, 'r') as f:
labels = [Link](f)
# Display results
print(f"\nInference time: {inference_time:.2f} ms")
print("\n" + "="*60)
print(f"{'[PREDICTION]':<35} {'[Probability]':>15}")
print("-"*60)
for i in range(top_k):
label = labels[top_indices[i]]
prob = top_prob[i].item() * 100
print(f"{label:<35} {prob:>14.2f}%")
print("="*60)
316
# Test with the cat image
inf_time, indices, probs = classify_image_executorch(
img_path="../images/[Link]",
model_path="../models/mobilenet_v2.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)
317
(2445.200204849243,
tensor([282, 285, 287, 281, 728]),
tensor([0.4744, 0.3761, 0.0691, 0.0622, 0.0047]))
318
The inference time was reduced from +2.5s to around -20ms
319
==> Even faster inference with a lower model in size
Slightly higher probabilities in the INT8 model are normal and do not indicate a
problem by themselves. Quantization slightly changes the logits, and softmax can
become a bit “sharper” or “flatter” even when top‑1 remains correct.
Camera Integration
We essentially have two different Python worlds: system Python 3.13 (where the camera stack
is wired up) and our 3.11 virtual env (where ExecuTorch is installed). To run ExecuTorch on
live frames from the Pi camera, we need to bridge those worlds.
320
• Recent Raspberry Pi OS uses Picamera2 on top of libcamera as the recom-
mended interface.
• The Picamera2/libcamera Python bindings are usually installed into the sys-
tem Python and are not trivially pip‑installable into arbitrary venvs or other
Python versions.
• Once we create a separate 3.11 environment, it will not automatically see
the Picamera2/libcamera bindings under 3.13, so imports fail or the camera
device is not accessible from that environment.
We will use a two‑process solution: capture in 3.13, infer in 3.11. For that, we should run
a small capture service under Python 3.13 that:
The 3.11 process (under venev) receives the frame, decodes it, runs the preprocessing pipeline
(resize, normalize), then calls ExecuTorch for inference..
Image Capture
Outside of the ExecuTorch env and folder, we will create a folder (CAMERA).
Documents/
��� EXECUTORCH/MOBILENET/ # Python 3.11
��� CAMERA/ # Python 3.13
��� camera_capture.py
��� camera_capture.jpg
import numpy as np
from picamera2 import Picamera2
import time
# Initialize camera
picam2 = Picamera2()
321
config = picam2.create_preview_configuration(main={"size":(640,480)})
[Link](config)
[Link]()
[Link](2)
# Capture image
picam2.capture_file("camera_capture.jpg")
print("Image captured: camera_capture.jpg")
# Stop camera
[Link]()
[Link]()
• /Documents/CAMERA/camera_capture.jpg
Looking from the notebook folder, the image path will be:
../../../../CAMERA/camera_capture.jpg
Let’s run the same function used with the test image:
322
Performance Benchmarking
Let’s now define a function to run inference several times for each model and compare their
performance.
323
def benchmark_inference(model_path, num_runs=50):
"""
Benchmark model inference speed
"""
print(f"Benchmarking model: {model_path}")
print(f"Number of runs: {num_runs}\n")
# Load model
model = _load_for_executorch(model_path)
# Benchmark
print(f"Running benchmark...")
times = []
for i in range(num_runs):
start = [Link]()
with torch.no_grad():
_ = [Link]((dummy_input,))
[Link]([Link]() - start)
# Print statistics
print("\n" + "="*50)
print("BENCHMARK RESULTS")
print("="*50)
print(f" Mean: {[Link]():.2f} ms")
print(f" Median: {[Link](times):.2f} ms")
print(f" Std: {[Link]():.2f} ms")
print(f" Min: {[Link]():.2f} ms")
print(f" Max: {[Link]():.2f} ms")
print("="*50)
324
# Plot distribution
[Link](figsize=(12, 4))
# Histogram
[Link](1, 2, 1)
[Link](times, bins=20, edgecolor='black', alpha=0.7)
[Link]([Link](), color='red', linestyle='--',
label=f'Mean: {[Link]():.2f} ms')
[Link]('Inference Time (ms)')
[Link]('Frequency')
[Link]('Inference Time Distribution')
[Link]()
[Link](alpha=0.3)
# Time series
[Link](1, 2, 2)
[Link](times, marker='o', markersize=3, alpha=0.6)
[Link]([Link](), color='red', linestyle='--',
label=f'Mean: {[Link]():.2f} ms')
[Link]('Run Number')
[Link]('Inference Time (ms)')
[Link]('Inference Time Over Runs')
[Link]()
[Link](alpha=0.3)
plt.tight_layout()
[Link]()
return times
mobilenet_v2.pte
mobilenet_v2_xnnpack.pte
mobilenet_v2_quantized_xnnpack.pte
325
Basic (Float32): mobilenet_v2.pte
# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2.pte",
num_runs=50
)
# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2_xnnpack.pte",
num_runs=50
)
326
Quantization (INT8): mobilenet_v2_quantized_xnnpack.pte
# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
num_runs=50
)
327
Performance Comparison Table
Mean Median
Model Configuration (ms) (ms) Std Dev (ms) File Size (MB) Latency
Float32 (basic) 2440 2440 2.17 13.58 +600×
Float32 + 11.24 10.84 1.67 13.35 ~3×
XNNPACK
INT8 + XNNPACK 3.91 3.69 0.55 3.59 1×
Key Observations:
328
2. Quantization Benefit: INT8 quantization, besides size reduction, adds additional
speedup beyond XNNPACK
3. Variability: Quantized model shows lower standard deviation, indicating more stable
performance
4. Size-Speed Tradeoff: 75% size reduction (14MB → 3.5MB) with 3× speed improve-
ment
CIFAR-10 Dataset:
• 10 classes: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck
Figure 6: cifar10
329
• The images in CIFAR-10 are of size 3x32x32 (3-channel color images of 32x32 pixels in
size).
Let’s create a Project folder structure as below (some files are shown as they will appear
later)
EXECUTORCH/CIFAR-10/
��� export_cifar10_xnnpack.py
��� inference_cifar10_xnnpack.py
��� models/
� ��� cifar10_model_jit.pt # Float32 pytorch model
� ��� cifar10_xnnpack.pte # Float32 conv model
��� images/ # Test images
� ��� [Link]
��� notebooks/
��� CIFAR-10_Inference_RPI.ipynb
Let’s train a model from scratch on CIFAR-10. For that, we can run the Notebook below on
Google Colab:
cifar10_colab_training.ipynb
From the training, we will have the trained model:
cifar10_model_jit.pt, which should be saved on /models folder
Next, as we did before, we should export the PyTorch model to ExecuTorch, and let’s use
XNNPACK. Run the script: export_cifar10_xnnpack.py, as a result, we have:
330
Runing it, a converted model cifar10_xnnpack.pte will be saved in ./models/ folder.
Runing the script inference_cifar10_xnnpack.py, over the “cat” image, we can see that the
converted model is working fine:
python inference_cifar10_xnnpack.py ./images/[Link]
331
And runing 20 times….
332
Despite the exported model being OK, when we make an inference with the original PyTorch
model, in this case (a small model), we will find even lower latencies.
333
In short, our export script is conceptually the right pattern for ExecuTorch+XNNPACK on
Arm, but for this specific small CIFAR‑10 CNN, the overhead of ExecuTorch and partial
XNNPACK delegation on a Pi‑class device can easily make it slower than a well‑optimized
plain PyTorch JIT model.
Optionally, it is possible to explore those models with the notebook:
CIFAR-10_Inference_RPI_Updated.ipynb
Conclusion
This chapter adapted our image classification workflow from TensorFlow Lite to PyTorch
EXECUTORCH, demonstrating that the PyTorch ecosystem provides a powerful and modern
alternative for edge AI deployment on Raspberry Pi devices.
EXECUTORCH represents a significant evolution in edge AI deployment, bringing PyTorch’s
research-friendly ecosystem to production edge devices. While TensorFlow Lite remains excel-
lent and mature, having EXECUTORCH in your toolkit makes you a more versatile edge AI
practitioner.
The future of edge AI is multi-framework, multi-platform, and rapidly evolving. By mastering
both EXECUTORCH and TensorFlow Lite, you’re positioned to make informed technical
decisions and adapt as the landscape changes.
Remember: The best framework is the one that serves your specific needs. This
tutorial empowers you to make that choice confidently.
334
Key Takeaways
Technical Achievements:
EXECUTORCH Advantages:
Comparison with TFLite: Both frameworks achieve similar goals with different philoso-
phies:
The choice between them often comes down to your training framework and specific require-
ments.
Performance Considerations
On Raspberry Pi 4/5, you can expect: - Float32 models: 10-20ms per inference (MobileNet
V2)
335
Resources
Code Repository
Official Documentation
Quantization:
• PyTorch Quantization
• Quantization API Tutorial
Models:
• Torchvision Models
• Pretrained Model Deployment Guide
Hardware Resources
Books
336
Beyond CPU - Hardware Acceleration for Edge
AI
337
Introduction
Throughout this course, we’ve explored various approaches to deploying AI models at the edge.
We started with TensorFlow Lite running on the Raspberry Pi’s CPU, then moved to YOLO
and ExecuTorch with optimized backends like XNNPACK. While these software optimizations
significantly improve performance, they still rely on the general-purpose CPU to execute neural
network operations.
In this chapter, we’ll take the next step: dedicated hardware acceleration. We’ll use the
MemryX MX3 M.2 AI Accelerator Module—a specialized processor designed specifically for
neural network inference. The MX3 module contains four AI accelerator chips that can run
deep learning models with dramatically lower latency and power consumption compared to
CPU execution.
Consider the requirements for real-time edge AI applications: - Latency: Autonomous sys-
tems need predictions in milliseconds - Power efficiency: Battery-powered devices must
conserve energy - Throughput: Multi-camera systems may need to process several streams
simultaneously - Cost: System designs often cannot afford high-end GPUs
The MX3 addresses these challenges with a unique architecture: - At-memory computing:
All memory is integrated on the accelerator, eliminating bandwidth bottlenecks - Pipelined
dataflow: Optimized for streaming inputs with a batch size of 1 - Floating-point accu-
racy: No quantization required (though supported) - Low power: Maximum 10W for four
accelerator chips
338
For learning more about AI Acceleration, please refer to MLSys book and how the
MemryX module works, read the Architecture Overview.
Our goal
By the end of this lab, we will have installed and configured the MX3 hardware on a Raspberry
Pi 5, set up the MemryX SDK and development environment, and gained a clear understanding
of the MX3 compilation and deployment workflow. We will also compile neural network models
for execution on the MX3 accelerator, compare their performance against CPU-based inference
while analyzing the trade-offs, and finally build a complete end-to-end inference pipeline using
the MemryX Python API.
Prerequisites
IMPORTANT NOTE: MemryX recommends the GeeekPi N04 M.2 2280 HAT
as an excellent choice for the Raspberry Pi 5. It delivers solid power and fits the
2280 MX3 M.2 form factor. Some hats can lead to instabilities, mainly due to PCIe
speed (Gen3). The Raspberry Pi 5 can have stability issues on Gen3.
339
The Raspberry Pi M.2 HAT+ is a good option. It works very well, despite the fact that we
should adapt the MX3 board to it (The MX3 is longer than the hat).
MemryX MX3 M.2 module with the heatsink installed
For heatsink installation, follow the video instructions: [Link]
It is essential to ensure we have sufficient cooling for the MemryX MX3 M.2 module, or we
may experience thermal throttling and reduced performance. The chips will throttle their
performance if they hit 100 °C.
During normal operation, the current MemryX MX3 temperature and throttle status can be
viewed at any time with:
cat /sys/memx0/temperature
Verification
After installing the hardware, turn on the Raspberry Pi and verify the system setup.
340
ls /dev/memx*
The lab temperature at the time of the above measurement was 25 °C.
Software Installation
cd Documents
mkdir MEMRYX
cd MEMRYX
python --version
341
Important: As of January 2026, MemryX officially supports only Python 3.09 to 3.12.
Python 3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships
with Python 3.13 by default, we’ll need to install a compatible Python version alongside
it.
One solution is to install Pyenv, which allows us to easily manage multiple Python versions
for different projects without affecting the system Python.
If the Raspberry Pi OS is the legacy version, the Python version should be 3.11,
and it is not necessary to install Pyenv.
# Install dependencies
sudo apt install -y make build-essential libssl-dev zlib1g-dev \
libbz2-dev libreadline-dev libsqlite3-dev wget curl llvm \
libncursesw5-dev xz-utils tk-dev libxml2-dev libxmlsec1-dev \
libffi-dev liblzma-dev
# Install pyenv
curl [Link] | bash
# Add to ~/.bashrc
echo 'export PYENV_ROOT="$HOME/.pyenv"' >> ~/.bashrc
echo 'command -v pyenv >/dev/null || export PATH="$PYENV_ROOT/bin:$PATH"' \
>> ~/.bashrc
echo 'eval "$(pyenv init -)"' >> ~/.bashrc
# Reload shell
source ~/.bashrc
Once Pyenv and the selected Python version are installed, define it for the project direc-
tory:
342
pyenv local 3.11.14
• Drivers (memx-drivers): Kernel-level drivers for PCIe communication with the accel-
erator hardware
• SDK (memx-accl): Python libraries, neural compiler, runtime, and benchmarking tools
This command downloads the repository’s GPG key for package verification and adds the
MemryX package repository:
343
Configure Platform Settings
Run the ARM setup utility to configure platform-specific settings. This opens a menu to
select the platform and apply the necessary configurations (e.g., enabling PCIe Gen 3.0 on the
Raspberry Pi 5):
sudo mx_arm_setup
Select the appropriate option for your hardware, and press <OK> in the next page:
sudo reboot
344
Verify Driver Installation
After rebooting, verify that the MemryX driver is installed by checking its version:
Install Utilities
It’s best practice to use a virtual environment to avoid conflicts with system packages.
Create and activate a virtual environment:
345
pip3 install --upgrade pip wheel
pip3 install --extra-index-url [Link] memryx
mx_nc --version
Verification
Verify the complete installation by running the built-in “hello world” benchmark:
mx_bench --hello
With the benchmark results, our MemryX MX3 is properly installed and ready to use.
346
Our First Accelerated Model
Working with the MemryX MX3 follows a straightforward four-step workflow that differs from
traditional CPU-based inference:
347
Step 1: Select or Train a Model
Start with a pre-trained model or train your own. MemryX supports models from major
frameworks:
The model remains in its original format—no framework-specific conversions needed yet. For
this lab, we’re using MobileNetV2 from Keras Applications, but we could equally use a custom
model we have trained for a specific task, as we have seen before.
Supported Operations: The MX3 supports most common deep learning operators (convolu-
tions, pooling, activations, etc.). Check the supported operators if using custom architectures.
Unsupported operations will fall back to CPU, though this is rare for standard vision models.
The MemryX Neural Compiler (mx_nc) transforms the model into a DFP (Dataflow Pack-
age):
The MemryX Neural Compiler, mx_nc, is a command‑line tool that takes one or more neu-
ral‑network models (Keras, TensorFlow, TFLite, ONNX, etc.) and compiles them into a
MemryX Dataflow Program (DFP) that can run on MemryX accelerators (MXA). Internally,
it does framework import, graph optimization (fusion/splitting, operator expansion, activation
approximation), resource mapping on the MXA cores, and finally emits the DFP used by the
runtime or simulator.
What mx_nc does:
• Compiles models into a single DFP file per compilation, then loads it onto one or more
MXA chips to run inference.
• Supports multi‑model, multi‑stream, and multi‑chip mapping, automatically distributing
models and layers across available MX3 devices for higher throughput.
• Handles mixed‑precision weights (per‑channel 4/8/16‑bit) while keeping activations in
floating point on the accelerator. By default, MemryX quantizes weights to INT8 preci-
sion and activations to BFloat16.
348
• Can crop pre/post‑processing parts of the graph so the MXA focuses on the core
CNN/ML operators while the host CPU runs the cropped sections.
1. Model parsing: Loads the model and extracts the computational graph
2. Graph optimization: Fuses operations, eliminates redundancies
3. Operator mapping: Maps each layer to MX3 hardware instructions
4. Dataflow scheduling: Determines optimal execution order for pipelined processing
5. Memory allocation: Assigns on-chip memory for all intermediate activations
6. Multi-chip distribution: If using multiple chips, partitions the workload
• Model specification:
– -m / --models – input model file(s) (e.g. .h5, .pb, .onnx, TFLite).
349
– -v, -vv, etc. – increase verbosity, useful to inspect graph transformations and
cropping decisions.
• Extensions / unsupported patterns:
– --extensions – load Neural Compiler Extensions (.nce files or builtin names) to
add or patch graph handling (e.g., complex transformer subgraphs or unsupported
ops) without a new SDK release.
For the complete option list (including less common flags), run mx_nc -h or consult
the Neural Compiler page in the MemryX Developer Hub, which documents all
arguments and includes usage examples for single‑model, multi‑model, cropping,
and mixed‑precision flows.
Once compiled, the DFP file is portable across all MX3 hardware.
Before integrating into our application, we can verify performance with the benchmarking
tool:
The benchmarker: - Generates synthetic input data matching the model’s input shape - Runs
warm-up inferences to stabilize performance - Measures throughput (FPS), latency, and chip
utilization - Reports first-inference latency (includes loading overhead)
Why benchmark separately? Real-world applications involve preprocessing (image loading
and resizing) and postprocessing (parsing outputs). Benchmarking isolates pure inference
performance, letting to identify bottlenecks in our full pipeline.
Finally, integrate the accelerator into our Python application using the MemryX API:
350
from memryx import SyncAccl # or AsyncAccl for concurrent processing
# Run inference
output = [Link](input_data)
# Process results
# ...
# Clean up
[Link]()
The API handles all hardware communication, memory transfers, and scheduling. Our code
just provides input tensors and receives output tensors—the complexity is abstracted away.
# Step 3: Benchmark
mx_bench -d mobilenet_v2.dfp -f 1000
This workflow is remarkably consistent across models and use cases. Once we’ve done it for
one model, adapting to others is straightforward.
351
Key Takeaway: The MX3 workflow separates compilation (done once) from in-
ference (done repeatedly). This “compile-once, run-many” approach means the
optimization overhead is amortized over thousands or millions of inferences in pro-
duction.
In Keras Applications, we can find deep learning models that are provided with pre-trained
weights. These models can be used for prediction, feature extraction, and fine-tuning.
Let’s download MobileNetV2, which was used in previous labs:
mx_nc -v -m mobilenet_v2.h5
352
The compiled model, mobilenet_v2.dfp, is saved in the current folder.
The .dfp (Dataflow Package) file is MemryX’s proprietary compiled format. Unlike standard
model formats (H5, ONNX, etc.) that describe the network architecture, a DFP file contains:
The neural compiler (mx_nc) performs this transformation automatically, with no manual
tuning required. The compilation process: 1. Parses the input model (H5, ONNX, TFLite,
etc.) 2. Maps operators to MX3-supported operations 3. Optimizes the dataflow graph 4.
Allocates memory on-chip 5. Generates the DFP binary
This is why compilation takes a few minutes, but inference is blazingly fast—all the optimiza-
tion work happens once, upfront.
Benchmarking Performance
Now that the model is compiled, it’s time to deploy it and run a benchmark to test its
performance on the MXA hardware. We will run 1000 frames of random data through the
accelerator to measure performance metrics:
353
Let’s understand what these metrics mean:
• FPS (Frames Per Second): How many images the accelerator can process per second
(~1,200 FPS for MobileNetV2)
• Latency: Time for a single inference (shown as “Avg” in the output)
– Subsequent inferences: True steady-state performance (~2ms)
• Throughput: Total data processed per second
The benchmark runs with random input data, which is why we see consistent performance.
Real-world performance with actual images should be similar once the preprocessing pipeline
is optimized, but we have found bigger latency.
In true dataflow architecture, latency and FPS are not coupled in the traditional
sense; latency does not equal 1/FPS. Note that even though latency is ~2ms in
the above benchmarking results, FPS is not measured to be 1000 ms / 2 ms = 500
FPS; rather, the FPS from the benchmarking results is ~1160.
In MemryX’s dataflow architecture, the “usual” rule (latency = 1/FPS) only applies to
frame‑to‑frame latency, not to end‑to‑end in‑to‑out latency for a single frame. That is why we
see ~2 ms latency per frame, yet still measure around 1160 FPS in MX3 benchmarks.
Two different latencies MemryX explicitly distinguishes two metrics.
354
• Latency 1 (frame‑to‑frame latency): Time between consecutive outputs once the pipeline
is full. Its reciprocal is FPS.
• Latency 2 (full in‑to‑out latency): Time from when the first input frame enters the
system (host + MX3 pipeline) until its output appears. This can be larger, but it does
not set the FPS.
In a streaming, pipelined accelerator like MX3, multiple frames are in‑flight simultaneously,
so the pipeline “fills” once and then produces results at a steady cadence.
Now let’s build a complete inference application that processes real images and compares CPU
vs. MX3 performance.
mkdir models
mkdir images
Load an image from the internet, for example, a cat (for comparison, it is the same as used
on previous chapters):
wget "[Link] \
-O ./images/[Link]
355
Understanding Input Requirements
All neural networks expect input data in a specific format, determined during training. For
MobileNetV2 trained on ImageNet:
The preprocessing must match exactly what was used during training, or accuracy will suffer.
356
Getting the Labels
For inference, we will need the ImageNet labels. The following function checks if the file exists,
and if not, downloads it:
MODELS_DIR = Path("./models")
IMAGENET_JSON = MODELS_DIR / "imagenet_class_index.json"
IMAGENET_JSON_URL = (
"[Link]
imagenet_class_index.json"
)
def load_idx2label():
with open(IMAGENET_JSON, "r") as f:
class_idx = [Link](f)
idx2label = [class_idx[str(k)][1] for k in range(len(class_idx))]
return idx2label
Image Preprocessing
The image used for inference should be preprocessed in the same way as during model train-
ing. [Link].mobilenet_v2.preprocess_input() takes an image of shape (224,
224) and converts it to (1, 224, 224, 3):
357
import numpy as np
from PIL import Image
import tensorflow as tf
from tensorflow import keras
def load_and_preprocess_image(image_path):
img = [Link](image_path).convert("RGB").resize((224, 224))
arr = [Link](img).astype(np.float32)
arr = [Link].mobilenet_v2.preprocess_input(arr)
arr = np.expand_dims(arr, 0) # Add batch dimension
return arr
The processed image will serve as the model’s input tensor (x):
ensure_imagenet_labels()
idx2label_full = load_idx2label() # length 1000 for ImageNet
IMAGE_PATH = Path("./images/[Link]")
x = load_and_preprocess_image(IMAGE_PATH)
Move the models (the original and compiled) to the models folder and set up the paths:
MODELS_DIR = Path("./models")
DFP_PATH = MODELS_DIR / "mobilenet_v2.dfp"
KERAS_PATH = MODELS_DIR / "mobilenet_v2.h5"
accl = SyncAccl(dfp=str(DFP_PATH))
mxa_outputs = [Link](x)
We get a list/array of outputs. In this case, with a shape of (1, 1000) and a dtype of float32.
This output should be normalized to a NumPy array:
358
mxa_outputs = [Link](mxa_outputs)
if mxa_outputs.ndim == 3:
mxa_outputs = mxa_outputs[0]
Expected output:
359
MXA top-5:
# 282: tiger_cat (38.6%)
# 281: tabby (18.3%)
# 285: Egyptian_cat (15.2%)
# 287: lynx (3.9%)
# 478: carton (1.7%)
We can also run the unconverted model (mobilenet_v2.h5) on the CPU, applying the code
to the same input tensor:
cpu_model = [Link].load_model(KERAS_PATH)
cpu_outputs = cpu_model.predict(x)
num_classes = cpu_outputs.shape[-1]
idx2label = idx2label_full if num_classes == len(idx2label_full) else None
Expected output:
CPU top-5:
# 282: tiger_cat (58.4%)
# 285: Egyptian_cat (12.9%)
# 281: tabby (11.6%)
# 287: lynx (3.4%)
# 588: hamper (1.3%)
Despite the probabilities not being identical, both models reach the same top prediction. The
slight differences are due to numerical precision variations between CPU and accelerator im-
plementations.
Measuring Latency
360
Note: The following sections break down the complete inference script into logical
components. The full working script is available separately and integrates all these
pieces together.
import time
# Warm-up run
_ = [Link](x)
# Timed inference
start = [Link]()
mxa_outputs = [Link](x)
mxa_latency = [Link]() - start
python run_inference_comp_mobilenetv2.py
Expected results:
361
The Accelerator runs 11 times faster than the CPU!
# Download ResNet50
python3 -c "import tensorflow as tf; \
[Link].ResNet50().save('resnet50.h5');"
# Compile
mx_nc -v -m resnet50.h5
362
The script can be found in the lab repo: run_inference_comp_resnet50.py in the
terminal:
The performance improvements are even more dramatic with larger models!
Clean Shutdown
[Link]()
363
Folders Structure
Documents/MEMRYX/
��� run_inference_comp_mobilenetv2.py # MobileNetV2 script
��� run_inference_comp_resnet50.py # ResNet50 script
��� images/
� ��� [Link] # Test image
��� models/
� ��� mobilenet_v2.h5 # Original model
� ��� mobilenet_v2.dfp # Compiled model
� ��� resnet50.h5 # Original model
� ��� [Link] # Compiled model
� ��� imagenet_class_index.json # Labels (auto-downloaded)
��� mx-env/
Here’s how the MX3 compares across different deployment approaches we’ve covered in this
course:
Key Observations
364
When to Use the MX3?
In this part of the lab, we’ll deploy YOLOv8n (nano) for real-time object detection on the
Raspberry Pi 5 using the MemryX MX3 AI accelerator. We’ll cover the complete workflow
from model export to inference optimization.
365
pip3 install --extra-index-url [Link] memryx
The model and the image [Link] will be download and tested with the YOLOV8n:
4 persons, 1 bus, and one stop signal were detected in 522 ms.
366
Model Export and Compilation
YOLOv8 must be converted to ONNX format before compilation for the MX3:
367
#### Export to ONNX format
[Link](format='onnx', simplify=True)
We can use the MemryX Neural Compiler to generate the DFP file:
Key flags:
Output files:
368
2. Detection head (yolov8n_post.onnx): Bounding box decoding on CPU
369
Understanding YOLOv8 Output Format
Decoding Process
We should now create a script to run an object detector (YOLOv8 with a pre/post-processing
pipeline), print each detection (label, confidence, bounding box), and save a copy of the image
with the boxes drawn.
Configuration section
DFP_PATH = "./models/[Link]"
POST_MODEL_PATH = "./models/yolov8n_post.onnx"
IMAGE_PATH = "./images/[Link]"
CONF_THRESHOLD = 0.25
370
• DFP_PATH: path to the compiled model used for inference.
• CONF_THRESHOLD: minimum confidence score; detections below this are filtered out.
Running detection
Here we call a helper function detect_objects that encapsulates the heavy lifting:
• Runs inference.
• Returns:
– detections: list/array where each element is [x1, y1, x2, y2, conf,
class_id].
371
Printing results
print(f"\n{'='*60}")
print("Detection Results:")
print(f"{'='*60}")
for i, det in enumerate(detections):
x1, y1, x2, y2, conf, class_id = det
print(f" {i+1}. {COCO_CLASSES[int(class_id)]}: {conf:.3f}")
print(f" Box: [{int(x1)}, {int(y1)}, {int(x2)}, {int(y2)}]")
• The loop goes over each detection, unpacks the bounding box coordinates, confidence,
and class ID.
if len(detections) > 0:
output_path = IMAGE_PATH.rsplit('.', 1)[0] + '_detected.jpg'
annotated_image.save(output_path)
print(f"\nSaved: {output_path}")
• The output filename is built by taking the original name and appending _detected
before the extension (e.g., bus_detected.jpg).
• annotated_image.save(...) writes the image with drawn boxes and labels to disk.
print(f"\n{'='*60}")
print(f"Total: {len(detections)} objects")
print(f"Time: {inference_time:.2f} ms")
print(f"{'='*60}")
372
• Prints how many objects were found in total.
• Prints the inference time, which is useful to talk about performance (e.g., model size
vs. speed, hardware differences).
python yolov8_m3_detect.py
As a result, we can see that the models found 4 persons and 1 bus, missing only the stop signal.
Regarding latency, the MX3 runs inference about 11 times faster than a CPU-only
system.
Basically, the same accuracy result that we got on the YOLO chapter running
yolov11
373
• Load and normalize:
– Open the image with PIL, convert to RGB, and get its original size.
– Compute a scale ratio so the image fits into 640×640 without distortion (preserving
aspect ratio).
• Letterboxing:
– Resize the image to (new_w, new_h) = (int(w * ratio), int(h * ratio)).
– Paste it onto a 640×640 canvas filled with color (114, 114, 114) (same as
Ultralytics).
– Compute the padding offsets (pad_w, pad_h) so we can undo this later.
• Tensor conversion:
– Convert to numpy, normalize to [0,1], permute from HWC to CHW, and add a
batch dimension to get shape [1, 3, 640, 640], which matches YOLOv8’s ex-
pected input.[1][2]
• For COCO YOLOv8n ONNX, the detection head outputs a tensor of shape (1, 84,
8400).
• 84 = 4 (bbox) + 80 (class scores). Each of the 8400 positions corresponds to one
candidate box.
Function Walkthrough:
• Transpose:
– From (1, 84, 8400) to (8400, 84) so each row is: [x_center, y_center,
width, height, class_0_score, ..., class_79_score].
• Best class per box:
– Take max_scores = [Link](class_scores, axis=1) and class_ids =
[Link](class_scores, axis=1) to select the most likely class and its
score for each of the 8400 candidates.
• Confidence filtering:
– Drop boxes whose max class score is below conf_threshold.
• Coordinate conversion:
374
– Convert from YOLO’s center-format (x, y, w, h) to corner-format (x1, y1, x2,
y2) to make drawing and IoU calculation simpler.
• NMS:
– Call apply_nms to remove overlapping boxes and keep only the best ones.
IoU:
area of intersection
IoU =
area of union
• compute_iou_batch does this between one box and many boxes at once using vectorized
Numpy operations.
NMS:
• apply_nms:
– Sort boxes by score descending.
– Repeatedly pick the highest-score box, compute its IoU with the remaining boxes,
and discard those whose IoU is above iou_threshold.
• The result is a list of indices for boxes that don’t overlap too much and represent unique
objects.
375
5. Drawing results (draw_detections)
• Preprocess once: call preprocess_image to get the model-ready tensor and the info
needed for rescaling.
• Create the accelerator:
– accl = AsyncAccl(dfp_path) loads the compiled Memryx DFP model.
– accl.set_postprocessing_model(post_model_path, model_idx=0) attaches
the ONNX post-processing graph.
• Streaming-style design:
– frame_queue is a queue of inputs; you put your tensor in it.
– generate_frame is a generator feeding frames into the accelerator.
– process_output is a callback that collects outputs into results.
– The code wires them with connect_input and connect_output, then waits for
completion with [Link]().
• Post-processing:
– Grab the first output, call decode_predictions, rescale boxes, and draw.
Making Inferences
Let’s change the script to easily handle different images and confidence threshold
(yolov8_m3_detect_v2.py). We should replace the hardcoded IMAGE_PATH with a command-
line argument:
376
import argparse
if __name__ == "__main__":
parser = [Link]()
parser.add_argument(
"-i", "--image",
type=str,
required=True,
help="Path to input image"
)
parser.add_argument(
"-c", "--conf",
type=float,
default=0.25,
help="Confidence threshold"
)
args = parser.parse_args()
# Configuration
DFP_PATH = "./models/[Link]"
POST_MODEL_PATH = "./models/yolov8n_post.onnx"
IMAGE_PATH = [Link]
CONF_THRESHOLD = [Link]
# Run detection
detections, annotated_image, inference_time = detect_objects(
DFP_PATH,
POST_MODEL_PATH,
IMAGE_PATH,
CONF_THRESHOLD
)
# Print results
print(f"\n{'='*60}")
print("Detection Results:")
print(f"{'='*60}")
for i, det in enumerate(detections):
x1, y1, x2, y2, conf, class_id = det
print(f" {i+1}. {COCO_CLASSES[int(class_id)]}: {conf:.3f}")
print(f" Box: [{int(x1)}, {int(y1)}, {int(x2)}, {int(y2)}]")
377
# Save annotated image
if len(detections) > 0:
output_path = IMAGE_PATH.rsplit('.', 1)[0] + '_detected.jpg'
annotated_image.save(output_path)
print(f"\nSaved: {output_path}")
print(f"\n{'='*60}")
print(f"Total: {len(detections)} objects")
print(f"Time: {inference_time:.2f} ms")
print(f"{'='*60}")
As we saw in the YOLO chapter, we are assuming we are in an industrial facility that must
sort and count wheels and special boxes.
378
Each image can have three classes:
We have captured a raw dataset using the Raspberry Pi Camera and labeled it with the
ROBOFLOW. The Yolo model was trained on a Google Colab using Ultralytics.
Using the FileZilla FTP, transfer a few images from the test dataset to .\images:
379
Let’s return to the ./MEMRYX/YOLO folder and using the Python Interpreter, to quickly do
some inferences:
python
We will import the YOLO library and define the model to use:
Now, let’s define an image and call the inference (we will save the image result this time to
external verification):
380
We can see that the model is working and that the latency was 168 ms.
Let’s now export the model first to ONNX and after to FFPls
, to run it in the MX3 device:
cd ./models
yolo export model=box_wheel_320_yolo.pt format=onnx
mx_nc -v --autocrop -m box_wheel_320_yolo.onnx
cd ..
Naturally we should enter with the new models ’names and instead of
COCO_LABELS, the script was changed to:
381
]
Thant’s all!
Run it with:
The Result was great! And the latency (~38 ms) was 4 times lower than with the
CPU-only approach (even smaller than the model exported to NCNN, runing 100% at CPU
- 80 ms).
382
Advanced Topics
accl = AsyncAccl(dfp_path)
accl.set_postprocessing_model(post_model_path)
Thermal Management
Model Selection
383
Clone the MemryX eXamples Repository
After cloning the repository, you’ll find several subdirectories with different categories of ap-
plications:
• Preprocessing pipelines
• Multi-threaded inference
• Output visualization
• Performance optimization
• Multi-model orchestration
Exploring these examples is an excellent way to learn production-ready patterns for deploying
MemryX applications.
384
sudo raspi-config
# Navigate to: Advanced Options → PCIe Speed → Enable PCIe Gen 3
sudo reboot
• Ensure sufficient power: Use the official Raspberry Pi 27W power supplyCheck
HAT installation: Ensure the M.2 HAT is properly seated.
• Lower Frequency: Try running sudo mx_set_powermode with a lower frequency, such
as 200 or 300 MHz. Then restart mxa-manager for good measure with sudo service
mxa-manager restart
sudo mx_set_powermode
385
sudo service mxa-manager restart
cat /etc/memryx/[Link]
We can check it by the first field of the file, FREQ4. If the Raspberry Pi is set to the module’s
default operating frequency (500 MHz), we should see FREQ4C=500, indicating that the module
is set to a 500 MHz clock speed for 4-chip DFPs.
If decreasing the frequency solves the issue, then you can either keep the default
frequency for all DFPs at 300 MHz (or 400, 450, etc.), or you can raise it back to
500 MHz and use the C++ API’s set_operating_frequency function to change the
clock speed on a per-DFP basis.
Compilation Errors
import tensorflow as tf
import tf2onnx
model = [Link].load_model('model.h5')
onnx_model, _ = [Link].from_keras(model)
386
with open("[Link]", "wb") as f:
[Link](onnx_model.SerializeToString())
• Check the compilation log (-v flag) to identify which specific layer is causing issues
Thermal Throttling
• Verify heatsink installation: Ensure thermal paste is properly applied and heatsink
is firmly attached
• Improve airflow: Position the Raspberry Pi for better air circulation
• Check ambient temperature: Ensure the room temperature is reasonable (<30°C)
• Monitor continuously:
deactivate
rm -rf mx-env
python -m venv mx-env
source mx-env/bin/activate
387
pip install --upgrade pip wheel
pip install --extra-index-url [Link] memryx
sudo raspi-config
# Advanced Options → PCIe Speed
htop
• Verify Frequency
By default, the frequency should be at 500 MHz. Smaller frequencies will reduce the FPS
(increase the latency)
cat /etc/memryx/[Link]
388
Import Errors
• Reinstall if necessary:
import sys
print([Link])
• Check input normalization: Confirm the value range matches training (e.g., [0, 1] vs
[-1, 1])
• Test with known inputs: Use the validation dataset to verify accuracy
• Compare outputs numerically: Print raw logits/probabilities to identify differences
• Check for quantization effects: If using -q flag, try without quantization first
389
Next Steps and Extensions
Project Ideas
• Quantization: Experiment with 8-bit and 4-bit quantization for even better perfor-
mance
• Model Zoo: Explore pre-optimized models in the MemryX Model Explorer
• Async API: Use AsyncAccl for non-blocking, concurrent processing
• Custom Operators: Learn to handle models with custom layers
• Multi-chip Scaling: Understand how workload distributes across the four accelerators
390
Conclusion
In this lab, we’ve explored hardware acceleration for edge AI using the MemryX MX3 acceler-
ator. We’ve learned:
The MX3 demonstrates that dedicated AI accelerators can deliver significant performance
improvements for edge applications, achieving FPS several times higher (up to 25x for ResNet-
50) than CPU inference while maintaining accuracy and providing deterministic latency.
As edge AI continues to evolve, hardware acceleration will become increasingly important for
real-time, power-efficient deployments. The skills we’ve developed in this lab—understanding
the compilation workflow, benchmarking methodologies, and performance optimization—will
transfer to other accelerator platforms as well.
Official Documentation
Background Reading
391
Community and Support
#
Generative AI (Proactive)
392
Text Generation with RNNs
Introduction
In this chapter, we will explore how to build a character-level text generation model using
Recurrent Neural Networks (RNNs), specifically inspired by the works of Jules Verne.
The “Jules Verne Bot” project will help to show us the fundamental concepts of sequence
modeling and text generation using deep learning techniques, as a preview of how the modern
LLMs work.
Project Overview:
393
• Goal: Create an AI model that generates text in the style of Jules Verne
• Architecture: RNN with GRU (Gated Recurrent Unit) layers
• Approach: Character-level text prediction
• Framework: TensorFlow/Keras
• Platform: Google Colab with Tesla T4 GPU
Imagine we could teach a computer to write like Jules Verne, the famous author of “Twenty
Thousand Leagues Under the Sea” and “Around the World in Eighty Days.” That’s precisely
what we’re doing with the Jules Verne Bot. This project creates an artificial intelligence system
that learns the patterns, style, and vocabulary from Jules Verne’s novels, then generates new
text that sounds like it could have come from his pen.
Think of it like this: if we read enough of someone’s writing, we start to recognize their style.
We notice they use certain phrases, prefer specific sentence structures, or have favorite topics.
Our neural network does something similar, but with mathematical precision. It analyzes
millions of characters from Verne’s works and learns to predict what character should come
next in any given sequence.
Before we dive into the technical details, let’s understand why we use neural networks for this
task and why we chose the specific type we did.
When you read a sentence like “The submarine descended into the dark…” your brain auto-
matically starts predicting what might come next. Maybe “depths” or “ocean” or “waters.”
Your brain does this because it has learned patterns from all the text you’ve ever read. Neu-
ral networks work similarly, but they learn these patterns through mathematical calculations
rather than biological processes.
Before diving into our RNN implementation, let’s understand where RNNs fit in the neural
network ecosystem:
Key Neural Network Architectures:
394
• MLP (Multi-Layer Perceptron): Basic feedforward networks for general tasks, for
example, vibration analysis
• CNN (Convolutional Neural Networks): Specialized for image processing as Image
Classification tasks and spatial data
• RNN (Recurrent Neural Networks): Designed for sequential data like text and
time series
• GAN (Generative Adversarial Networks): Two networks competing for realistic
data generation, as images
• Transformers (Attention Networks): Modern architecture using attention mecha-
nisms, as in LLMs (Large Language Models, such as GPT)
We chose a Recurrent Neural Network (RNN) for this project because text has a crucial prop-
erty: order matters tremendously. The sequence “The cat sat on the mat” means something
completely different from “Mat the on sat cat the.” Regular neural networks process all inputs
simultaneously, like looking at a photograph. But for text, we need a network that processes
information sequentially, remembering what came before to understand what should go next.
In text generation, we aim to predict the most probable word to follow a sentence.
395
Think of reading a book. You don’t just look at all the words on a page simultaneously. You
read word by word, sentence by sentence, and your understanding builds as you progress. Each
new word is interpreted in the context of everything you’ve read before in that chapter. RNNs
work the same way.
Early RNNs had a significant problem: they couldn’t remember information for very long.
Imagine trying to understand a story where you could only remember the last few words you
read. You’d lose track of characters, plot points, and context very quickly.
This is where the Gated Recurrent Unit (GRU) comes in. Think of GRU as an improved
memory system with two special abilities:
Reset Gate: This decides when to “forget” old information. If the story switches to a new
scene or character, the reset gate helps the network forget irrelevant details from the previous
context.
Update Gate: This decides how much new information to incorporate. When encountering
important plot points or character names, the update gate helps the network remember these
crucial details for longer.
It’s like having a smart note-taking system that automatically decides what’s worth remem-
bering and what can be forgotten.
Why RNNs for Text Generation?
Recurrent Neural Networks are designed explicitly for sequential data processing. Key
characteristics:
396
• Memory: RNNs maintain an internal state (memory) to remember previous inputs
• Sequential Processing: Process data one element at a time, making them ideal for
text
• Variable Length Input: Can handle sequences of different lengths
• Parameter Sharing: Same weights applied across different time steps
397
Dataset Preparation
Our model is trained on a curated collection of 10 classic Jules Verne novels, downloaded from
public domain texts of the Gutenberg Project:
“A Journey to the Centre of the Earth” teaches the model about geological descriptions and
underground adventures. “Twenty Thousand Leagues Under the Sea” provides vocabulary
about marine life and submarine technology. “Around the World in Eighty Days” offers ge-
ographical references and travel descriptions. Each book contributes unique vocabulary and
stylistic elements while maintaining Verne’s consistent voice.
The complete dataset contains 5,768,791 characters, of which 123 are unique. To put this
in perspective, that’s roughly equivalent to 3,000 pages of text. This provides our neural
398
network with ample material to learn from, enabling it to capture both common patterns and
unique expressions in Verne’s writing.
return text
Character-Level Tokenization
Here’s where our approach differs from how humans typically think about text. While we
naturally think in words and sentences, our model processes text character by character. This
means it learns that certain letters frequently follow others, that spaces separate words, and
that punctuation marks signal sentence boundaries.
Why choose character-level processing? Consider the word “extraordinary,” which appears
frequently in Verne’s work. A word-level model would need to have seen this exact word
during training to use it. But a character-level model can generate this word by learning that
‘e’ often starts words, ‘x’ can follow ‘e’, ‘t’ often follows ‘x’, and so on. This allows our model
to create new words or handle misspellings gracefully.
The downside is that character-level processing requires more computational steps to generate
the same amount of text. Generating “Hello world” requires 11 prediction steps instead of just
2. However, for our educational purposes, this trade-off provides valuable insights into how
language generation works at its most fundamental level.
399
Unlike word-level tokenization, character-level tokenization treats each character
as a token.
Please see the following site for a great general visual explanation, from Andrej
Karpathy, The Unreasonable Effectiveness of Recurrent Neural Networks.
Computers work with numbers, not letters, so we need to convert our text into a numerical
representation. We start by finding every unique character in our dataset. This includes
not just letters A-Z and a-z, but also numbers, punctuation marks, spaces, and even special
characters that might appear in the original texts.
Our Jules Verne collection contains 123 unique characters. These include obvious ones like
letters and common punctuation, but also less common characters like accented letters from
French names or special typography marks from the original publications.
We create two dictionaries: one that converts characters to numbers (encoding) and another
that converts numbers back to characters (decoding). For example:
‘a’ might become 47, ‘b’ becomes 48, ‘c’ becomes 49, and so on. The space character might
be 1, and the period might be 72. These assignments are arbitrary but consistent throughout
our project.
When we want to process the phrase “The sea”, we convert it to something like [84, 72, 69,
1, 83, 69, 47]. When the model generates numbers like [84, 72, 69, 1, 87, 47, 83], we convert
them back to “The was” (as an example).
400
print(f"Vocabulary size: {len(vocab)}")
print(f"Unique characters: {vocab}")
Training Sequences
Our model learns by playing a sophisticated prediction game. We show it sequences of 120
characters and ask it to predict what the 121st character should be. Think of it like a fill-in-
401
the-blank exercise, but instead of missing words, we’re missing the next character.
For example, if our text contains “The submarine descended into the dark depths of the
ocean”, we might show the model “The submarine descended into the dark depths of the ocea”
and ask it to predict “n”. Then we slide our window forward by one character and show it
“he submarine descended into the dark depths of the ocean” and ask it to predict the next
character.
Training Configuration
We chose 120 characters as our context window because it roughly corresponds to one
paragraph of text in English. This gives the model enough context to understand local patterns
(like completing words and phrases) while remaining computationally manageable. In practical
terms, 120 characters might look like:
“The Nautilus had been cruising in these waters for some time. Captain Nemo stood on the
bridge, observing the vast exp”
From this context, the model might predict “a” to complete “expanse” or “l” to form “ex-
plore”.
The longer the context window, the better the model can maintain coherence, but
the more computer memory and processing time it requires.
Training Example
This means our dataset of 5.8 million characters becomes millions of individual training exam-
ples, each teaching the model about character sequence patterns.
402
Creating Training Data
Character Embeddings
Initially, each character is represented as a one-hot vector, which is mostly zeros with a single
one indicating which character it is. For 123 characters, this means each character is repre-
sented by a vector with 123 elements, where 122 are zero and 1 is one. This is wasteful and
doesn’t capture any relationships between characters.
Character embeddings solve this problem by representing each character as a dense vector of
real numbers. Instead of 123 mostly-zero values, each character becomes 256 meaningful num-
bers. These numbers are learned during training and end up encoding relationships between
characters.
Something fascinating happens during training: characters that behave similarly end up with
similar embedding vectors. Vowels tend to cluster together because they can often substitute
for each other in similar contexts. Consonants that frequently appear together (like ‘th’ or
‘ch’) develop related embeddings.
The model learns that uppercase and lowercase versions of the same letter are related but
distinct. It discovers that digits form their own cluster since they appear in similar contexts
403
(dates, measurements, chapter numbers). Punctuation marks develop embeddings based on
their grammatical functions.
404
You can play with Word2Vec - Embedding Projector
405
Model Architecture
Our Jules Verne Bot consists of three main components, each serving a specific purpose in the
text generation pipeline.
Embedding Layer: This is our translation layer. It takes character indices (numbers like
47, 83, 72) and converts them into dense 256-dimensional vectors that capture character rela-
tionships. Think of this as converting raw symbols into a format that captures meaning and
relationships.
GRU Layer: This is the brain of our operation. With 1024 hidden units, this layer processes
sequences and maintains memory about what it has seen. When processing the sequence “The
submarine descended”, the GRU maintains a hidden state that encodes information about the
submarine, the action of descending, and the overall maritime context.
Dense Output Layer: This is our decision-making layer. It takes the GRU’s 1024-
dimensional hidden state and converts it into 123 probabilities, one for each character in our
vocabulary. These probabilities represent the model’s confidence about what character should
come next.
Model Summary
Model: "sequential_4"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
embedding_4 (Embedding) (1, 120, 256) 31,488
406
_________________________________________________________________
gru_3 (GRU) (1, 120, 1024) 3,938,304
_________________________________________________________________
dense_3 (Dense) (1, 120, 123) 126,075
=================================================================
Total params: 4,095,867 (15.62 MB)
Trainable params: 4,095,867 (15.62 MB)
Non-trainable params: 0 (0.00 B)
Our model has 4,095,867 parameters. These are the individual numbers that the model adjusts
during training to improve its predictions. To put this in perspective, each parameter is like
a tiny dial that affects how the model processes information.
The GRU layer contains most of these parameters (about 3.9 million) because it needs to learn
complex patterns about how characters relate to each other across different time steps. The
embedding layer has about 31,000 parameters (123 characters × 256 dimensions), and the
output layer has about 126,000 parameters.
When processing text, information flows through the model like this:
A character index enters the embedding layer and becomes a 256-dimensional vector. This
vector enters the GRU, which combines it with its current memory state to produce a new
1024-dimensional hidden state. This hidden state captures everything the model “knows” at
this point in the sequence.
The hidden state goes to the dense layer, which produces probability scores for each of the
123 possible next characters. The character with the highest probability becomes the model’s
prediction.
Crucially, the GRU’s hidden state becomes its memory for the next character prediction.
This creates a chain of memory that allows the model to maintain context across the entire
sequence.
GRU Advantages:
407
• Computational Efficiency: Fewer parameters than LSTM
• Better Performance: More stable training than basic RNNs
Training a neural network means adjusting its millions of parameters so it makes better pre-
dictions. We use a loss function called sparse categorical crossentropy, which measures how
far off the model’s predictions are from the correct answers.
Think of it like teaching someone to play darts. Each throw (prediction) has a target (the
correct next character). The loss function measures how far each dart lands from the bullseye.
Training adjusts the player’s technique (the model’s parameters) to improve accuracy over
time.
We trained our model on a Tesla T4 GPU, which can perform thousands of calculations
simultaneously. This parallelization is crucial because each training step involves matrix mul-
tiplications with millions of numbers. The training took 33 minutes for 30 complete passes
through the entire dataset.
To understand why we need a GPU, consider that training involves calculating gradients for
all 4 million parameters, potentially thousands of times per second. A regular CPU would
take many hours to complete the same training that a GPU accomplishes in minutes.
Monitoring Progress
During training, we watch the loss decrease from about 1.9 to 0.9. This represents the model’s
improving ability to predict the next character. Early in training, the model makes essentially
random predictions. By the end, it has learned sophisticated patterns about English spelling,
grammar, and Jules Verne’s writing style.
The learning curve typically shows rapid improvement in the first few epochs as the model
learns basic patterns like common letter combinations. Later epochs show slower but steady
improvement as the model refines its understanding of more complex patterns like narrative
structure and thematic elements.
408
Preventing Overfitting
One challenge in training is overfitting, in which the model memorizes the training data rather
than learning generalizable patterns. We use techniques like monitoring validation loss and
potentially stopping training early if the model stops improving on unseen text.
Training Configuration
Hardware Setup:
Training Parameters:
409
Training Implementation
# Model compilation
[Link](
optimizer='adam',
loss='sparse_categorical_crossentropy',
metrics=['accuracy']
)
Text Generation
Once trained, our model becomes a text generation engine. We start with a seed phrase like:
“THE FLYING SUBMARINE”
and ask the model to continue the story. The process works character by character:
The model receives “THE FLYING SUBMARINE” and predicts the most likely next character
based on everything it learned from Jules Verne’s works. Maybe it predicts a space, starting a
new word. Then we feed “THE FLYING SUBMARINE” (with the space) back to the model
and ask for the next character.
This process continues indefinitely, with each new character becoming part of the context
for predicting the next one. The model might generate “THE FLYING SUBMARINE de-
scended into the mysterious depths…” as it draws upon patterns learned from Verne’s nautical
adventures.
410
Temperature Control
Here’s where we can control the model’s creativity through a parameter called temperature.
Temperature affects how the model chooses between different possible next characters.
With temperature set to 0.1, the model almost always picks the most probable next character.
This produces very predictable, conservative text that closely mimics the training data but
might be repetitive or boring.
With temperature set to 1.0, the model considers all possible next characters according to
their learned probabilities. This produces more varied and creative text, but sometimes makes
unusual choices that lead to interesting narrative directions.
With temperature above 1.5, the model becomes quite random, often producing text that
starts coherently but gradually becomes nonsensical as unlikely character combinations accu-
mulate.
In short:
Implementation
text_generated = []
model.reset_states()
for i in range(num_generate):
predictions = model(input_eval)
predictions = [Link](predictions, 0)
# Apply temperature
predictions = predictions / temperature
predicted_id = [Link](predictions, num_samples=1)[-1,0].numpy()
411
text_generated.append(idx_to_char[predicted_id])
This eBook is for the use of anyone anywhere in the United States and most
other parts of the earth and miserable eruptions. The solar rays should be
entirely under the shock of the intensity of the sea. We were all sorts. Are
we to prepare for our feelings?"
"Well, then, John, for I get to the Pampas, that we ought to obey the same
time. In the country of this latitude changed my brother, and the
_Nautilus_ floated in a sea which contained the rudder and
lower colour visibly. The loiter was a fatalint region the two
scientific discoverers. Several times turning toward the river, the cry
of doors and over an inclined plains of the Angara, with a threatening
water and disappeared in the midst of the solar rays.
The weather was spread and strewn with closed bottoms which soon appeared
that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives, and the prisoners
were successively assuming the sides of this agreement for fifteen days
with a threatening voice.
...
Let’s examine some generated text: “The weather was spread and strewn with closed bottoms
which soon appeared that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives…”
412
This excerpt shows both the model’s strengths and limitations. It successfully captures Verne’s
descriptive style and maritime vocabulary (“seas,” “wind,” “natives”). The sentence structure
feels appropriately Victorian and elaborate. However, the meaning becomes confused with
phrases like “closed bottoms” and “sheets of wind was soon and linen.”
This illustrates the fundamental challenge of character-level generation: the model learns local
patterns (how words are spelled, common phrases) much better than global coherence (logical
narrative flow, consistent meaning).
Our 120-character context window creates a fundamental limitation. The model can only “see”
about one paragraph of previous text when making predictions. This means it might introduce
a character named Captain Smith, then 200 characters later introduce another character with
the same name, having “forgotten” the first introduction.
Humans writing stories maintain mental models of characters, plot lines, and world-building
details across entire novels. Our model’s memory effectively resets every 120 characters, making
long-term narrative consistency nearly impossible.
Character-level generation requires many more prediction steps than word-level generation.
Generating the phrase “extraordinary adventure” requires 22 character predictions instead
of just 2 word predictions. This makes character-level generation much slower and more
computationally expensive.
However, character-level generation offers unique advantages. The model can generate new
words it has never seen before by combining character patterns. It can handle misspellings,
made-up words, or technical terms more gracefully than word-level models that have fixed
vocabularies.
Coherence Challenges
Perhaps the biggest limitation is maintaining semantic coherence. The model might generate
grammatically correct text that makes no logical sense. It can describe “The submarine floating
in the air above the mountain peaks” because it has learned that submarines float and that
Verne often described mountains, but it hasn’t learned the physical constraint that submarines
float in water, not air.
413
This happens because the model learns statistical patterns without understanding meaning.
It knows that certain word combinations are common without understanding why they make
sense.
Summary
Potential Improvements
Scale Comparison
To appreciate how far language modeling has advanced, consider the scale differences between
our Jules Verne Bot and modern language models:
414
Our model has 4 million parameters and was trained on about 5.8 million characters (10 books).
GPT-3 has 175 billion parameters and was trained on 45 terabytes of text (roughly equivalent
to millions of books). That’s a difference of over 40,000 times more parameters and millions
of times more training data.
Modern small language models (SLMs) like Phi-3-mini still dwarf our model with 3.8 billion
parameters, but they represent more efficient designs that achieve impressive performance with
“only” 1,000 times more parameters than our model.
Architectural Evolution
The biggest advancement since RNNs is the Transformer architecture, which uses atten-
tion mechanisms instead of recurrent processing. While RNNs process text sequentially (like
reading word by word), Transformers can examine all parts of a text simultaneously and learn
relationships between any two words, regardless of how far apart they are.
This solves the long-term memory problem that limits our RNN model. A Transformer can
maintain awareness of a character introduced in a hypothetical “chapter 1” while writing
“chapter 10”, something our 120-character context window makes impossible.
Training Efficiency
Modern models also benefit from more sophisticated training techniques. They’re pre-trained
on massive, diverse datasets to learn general language patterns, then fine-tuned on specific
tasks. They use techniques like instruction tuning, where they learn to follow human com-
mands, and reinforcement learning from human feedback, where they learn to generate text
that humans find helpful and appropriate.
1. Transformer Architecture
415
• Attention Mechanism: Can look at any part of the input sequence
• Parallel Processing: Much faster training and inference
• Better Long-range Dependencies: Maintains context over thousands of tokens
2. Scale
• More Data: Trained on vastly more diverse text
• More Parameters: Can memorize and generalize better
• More Compute: Allows for more sophisticated training techniques
3. Advanced Techniques
• Pre-training + Fine-tuning: Learn general language then specialize
• Instruction Tuning: Trained to follow human instructions
• RLHF: Reinforcement Learning from Human Feedback
Conclusion:
Building the Jules Verne Bot teaches us that creating artificial intelligence systems capable
of generating human-like text requires careful consideration of multiple components working
together. The embedding layer learns to represent characters meaningfully, the RNN layer
processes sequences and maintains memory, and the output layer makes predictions based on
learned patterns.
The project also illustrates the fundamental trade-offs in machine learning: between model
complexity and training speed, between creativity and coherence, between local accuracy and
global consistency. These trade-offs appear in every AI system, from simple character-level
generators to the most sophisticated language models.
Most importantly, this project demonstrates that impressive AI capabilities emerge from rel-
atively simple components combined thoughtfully. Our 4-million parameter model, while
limited compared to modern systems, genuinely learns to write in Jules Verne’s style through
nothing more than statistical pattern recognition and mathematical optimization.
The techniques we’ve explored, sequence processing, embedding learning, and generation strate-
gies, form the foundation for understanding any language model. Whether you encounter
RNNs, Transformers, or future architectures yet to be invented, the core concepts remain
consistent: learn patterns from data, encode meaning in mathematical representations, and
generate new content by predicting what should come next.
Understanding these fundamentals provides the foundation for working with, improving, or
creating the next generation of language models that will shape how humans and computers
communicate in the future.
416
Resourses
• Gutenberg Project
• The Unreasonable Effectiveness of Recurrent Neural Networks
• Word2Vec - Embedding Projector
• OpenAI’s tokenizer tool
• Generating Text with RNNs: The Jules Verne Bot - CoLab
417
Knowledge Distillation in Practice
418
Figure 7: Image created by DALLE-3
Knowledge distillation is a powerful technique in machine learning that enables the transfer
of knowledge from a large, complex model (the “teacher”) to a smaller, more efficient model
(the “student”). This process allows us to create compact models that maintain much of the
419
performance of their larger counterparts while being significantly faster and requiring fewer
computational resources.
In today’s AI landscape, models are becoming increasingly large and complex. While these
models achieve remarkable performance, they often require substantial computational re-
sources, making deployment challenging in resource-constrained environments such as mobile
devices, edge computing systems, or real-time applications. Knowledge distillation addresses
this challenge by:
The core concept of knowledge distillation revolves around the teacher-student relationship:
• Teacher Model: A large, well-trained model with high capacity and performance
• Student Model: A smaller, more efficient model trained to mimic the teacher’s behavior
• Knowledge Transfer: The process of transferring the teacher’s “dark knowledge” to
the student
420
Figure 8: Figure from “Knowledge Distillation: A Survey, Jianping Gou Baosheng Yu Stephen
J. Maybank Dacheng Tao, 2021
Theoretical Foundations
Traditional supervised learning uses “hard targets” - one-hot encoded labels that provide
limited information. For example, in MNIST digit classification, the label for the digit “5”
would be represented as [0, 0, 0, 0, 0, 1, 0, 0, 0, 0].
Knowledge distillation leverages “soft targets” - the probability distributions produced by the
teacher model. These soft targets contain richer information about the relationships between
classes. For instance, the teacher might output [0.01, 0.02, 0.01, 0.05, 0.1, 0.78,
0.02, 0.01, 0.0, 0.0] for a “5”, indicating that it’s most confident about “5” but also
considers “4” somewhat similar.
The temperature parameter (�) is crucial in knowledge distillation. It controls the “softness”
of the probability distribution by modifying the softmax function:
421
Where:
Effects of Temperature:
1. Distillation Loss (L_KD): Measures how well the student mimics the teacher’s soft
targets
2. Student Loss (L_CE): Traditional cross-entropy loss with hard targets
Where � is a weighting parameter that balances the two objectives, in our implementation, we
use � = 0.3, giving more weight to the soft targets from the teacher.
The distillation loss is typically computed using the KL divergence:
The �² factor compensates for the gradient scaling effect of temperature. This is crucial for
stable training.
Dataset Overview
422
• 28x28 grayscale images
• 10 classes (digits 0-9)
• Well-established baseline performances
Our teacher model has substantial capacity with multiple convolutional layers, batch normal-
ization, and dropout for regularization:
def build_teacher_model():
model = [Link]([
layers.Conv2D(64, 3,
activation='relu',
padding='same',
input_shape=(28, 28, 1)),
[Link](),
layers.Conv2D(64, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),
[Link](256, activation='relu'),
[Link](),
[Link](0.5),
[Link](128, activation='relu'),
[Link](),
[Link](0.5),
[Link](10, activation='softmax')
])
return model
423
424
This architecture achieved 99.44 % accuracy on MNIST.
Our student model is intentionally much simpler, with fewer layers and significantly fewer
parameters:
def build_student_model():
model = [Link]([
layers.Conv2D(16, 3,
activation='relu',
padding='same',
input_shape=(28, 28, 1)),
layers.MaxPooling2D(2, 2),
layers.Conv2D(32, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),
[Link](64, activation='relu'),
[Link](10, activation='softmax')
])
return model
425
The student model has fewer parameters than the teacher (105K versus 658K), or
6.2 times smaller.
In our implementation, we use a custom training loop that explicitly calculates both hard and
soft losses:
# Forward pass
with [Link]() as tape:
predictions = kd_student(x_batch, training=True)
Training Process
1. Train the Teacher: First, we train the complex teacher model using early stopping
and a reduced learning rate.
2. Train a Vanilla Student: We train a student model normally on the hard labels for
comparison.
3. Generate Soft Targets: We use the teacher to create softened probability distributions.
4. Train the Distilled Student: We train another student using our custom distillation
training loop.
5. Evaluate and Compare: We compare the performance, size, and speed of all three
models.
426
Results and Analysis
• Teacher Model: 99.44% accuracy, largest size, slowest inference (0.8632 seconds)
• Vanilla Student: 98.77% accuracy, smaller size, faster inference (0.3579 seconds)
• Distilled Student: 99.32% accuracy, same size as vanilla student, but better perfor-
mance (+0.55%) and faster inference than the Teacher (0.5467 seconds)
These results demonstrate the key benefit of knowledge distillation: the distilled student
achieves performance closer to the teacher while maintaining the efficiency benefits of the
smaller architecture.
[Link].0.1 * We analyze the models on challenging examples where the teacher succeeds but
the vanilla student fails. This reveals how knowledge distillation enables students to handle
complex cases by learning the teacher’s “dark knowledge.”
427
Advanced Techniques
Feature-Based Distillation
428
s_feat_aligned = align_feature_dimensions(s_feat, t_feat.shape)
loss += [Link](t_feat, s_feat_aligned)
return loss
Attention-Based Distillation
1. Layer Alignment: When using feature distillation, carefully align the feature dimen-
sions
2. Feature Selection: Not all features are equally important; focus on the most informa-
tive ones
3. Multi-Teacher Distillation: Combine knowledge from multiple teachers for better
results
4. Online Distillation: Train teacher and student simultaneously for mutual improvement
429
LLM Distillation Techniques
1. Sequence-Level Distillation
1. Teacher Model: The larger Llama 3.2 models (70B/8B parameters) serve as the teach-
ers
2. Student Model: Llama 3.2 3B and 1B are the distilled student models. They were cre-
ated using pruning techniques, which systematically remove less meaningful connections
(weights) in the neural network.
430
Key Points
1. Scale Difference: The parameter reduction from 70B to 3B (~23x) or 1B (~70x) demon-
strates industrial-scale distillation
2. Performance Preservation: Despite massive size reduction, the smaller models main-
tain impressive capabilities:
• The 3B model preserves most of the reasoning abilities of larger models.
• The 1B model remains highly functional for many tasks.
3. Practical Benefits:
• The 1B model can run on consumer laptops and even some mobile devices.
• The 3B model offers a balance of performance and accessibility.
4. Distillation Techniques Used:
• Meta likely used a combination of response-based and feature-based distillation.
• They may have employed temperature scaling and specialized loss functions similar
to what we demonstrated in our MNIST example.
• The principles we covered (soft targets, temperature, loss weighting) all apply at
this larger scale.
431
While our MNIST example demonstrates the application of knowledge distillation principles us-
ing a simple dataset, these same principles can be applied directly to state-of-the-art language
models.
Meta’s Llama 3.2 family provides a perfect real-world example:
Size
Model Parameters Reduction Use Case
Llama 3.2 70 billion (Teacher) Data centers, high-performance applications
70B
Llama 3.2 8 billion ~9x Server deployment, high-end workstations
8B
Llama 3.2 3 billion ~23x Consumer laptops, desktop applications
3B
Llama 3.2 1 billion ~70x Edge devices, mobile applications, embedded
1B systems
The 3B and 1B models represent successful distillations of the larger models, preserving core
capabilities while dramatically reducing computational requirements. This demonstrates the
industrial importance of knowledge distillation techniques we’ve explored.
It is possible to note that while the underlying principles remain the same as our MNIST
example, industrial LLM distillation includes additional techniques:
Despite these additional complexities, the core concept remains the same: using a larger, more
capable model to guide the training of a smaller, more efficient one.
Large models like ChatGPT (175B+ parameters) can be distilled to create mobile-friendly
assistants (1-2B parameters) that maintain core capabilities while running locally on smart-
phones.
432
2. Domain-Specific Distillation
• Medical LLMs: Distill medical knowledge from large models to smaller, specialized
ones
• Legal Assistants: Create compact models focused on legal reasoning and terminology
• Educational Tools: Develop small models optimized for teaching specific subjects
The journey from research models to production deployment often involves distillation:
1. Teacher-Student Size Ratio: Aim for a 10-20x parameter reduction for significant
efficiency gains
2. Architectural Similarity: Maintain similar architectural patterns between teacher
and student
3. Bottleneck Identification: Ensure the student has adequate capacity at critical layers
Hyperparameter Selection
1. Temperature (�):
• For MNIST: 3-5 works well
• For complex tasks: 5-10 may be better
• If outputs are already soft: Lower temperatures (2-3) may suffice
2. Alpha Weighting (�):
433
• For simpler tasks: 0.3-0.5 (balanced approach)
• For complex reasoning: 0.1-0.3 (more emphasis on teacher’s knowledge)
• When teacher is extremely accurate: Lower � values work better
3. Training Duration:
• Distilled students often benefit from longer training (1.5-2x the epochs)
• Use early stopping with patience to avoid overfitting
1. Teacher Performance: Ensure the teacher actually outperforms the student (we fixed
this in our implementation)
2. Temperature Selection: If knowledge transfer is poor, experiment with different tem-
peratures
3. Loss Weighting: If the student ignores soft targets, reduce � to emphasize distillation
loss
4. Gradient Scaling: Always apply the �² correction factor to the soft loss
Performance Evaluation
1. Accuracy: Primary performance metric (should be closer to teacher than vanilla stu-
dent)
2. Model Size: Parameter count and memory footprint (should match vanilla student)
3. Inference Speed: Time per prediction (should be significantly faster than teacher)
4. Challenging Cases: Performance on difficult examples (should be better than vanilla
student)
Conclusion
434
Key Takeaways
The principles we’ve demonstrated with MNIST directly scale to Large Language Models:
Final Thoughts
Knowledge distillation isn’t just an academic technique—it’s essential for practical AI deploy-
ment. The same principles that helped us compress our MNIST classifier can be scaled to
compress models with hundreds of billions of parameters. This universality makes knowledge
distillation an indispensable skill for engineering students entering the field of AI.
Resources
Notebook
435
436
Small Language Models (SLM)
Setup
We could use any Raspberry Pi model in the previous labs, but here, the choice must be the
Raspberry Pi 5 (Raspi-5). It is a robust platform that substantially upgrades the last version
4, equipped with the Broadcom BCM2712, a 2.4GHz quad-core 64-bit Arm Cortex-A76 CPU
featuring Cryptographic Extension and enhanced caching capabilities. It boasts a VideoCore
VII GPU, dual 4Kp60 HDMI® outputs with HDR, and a 4Kp60 HEVC decoder. Memory
options include 4GB, 8GB, and 16GB of high-speed LPDDR4X SDRAM, with 8GB as our
choice for running SLMs. It also features expandable storage via a microSD card slot and a
PCIe 2.0 interface for fast peripherals such as M.2 SSDs (Solid State Drives).
For real SSL applications, SSDs are a better option than SD cards.
By the way, as Alasdair Allan discussed, inferencing directly on the Raspberry Pi 5 CPU—with
no GPU acceleration—is now on par with the performance of the Coral TPU.
438
For more info, please see the complete article: Benchmarking TensorFlow and TensorFlow Lite
on Raspberry Pi 5.
We suggest installing an Active Cooler, a dedicated clip-on cooling solution for Raspberry Pi
5 (Raspi-5), for this lab. It combines an aluminum heatsink with a temperature-controlled
blower fan to keep the Raspi-5 operating comfortably under heavy loads, such as running
SLMs.
439
The Active Cooler has pre-applied thermal pads for heat transfer and is mounted directly to
the Raspberry Pi 5 board using spring-loaded push pins. The Raspberry Pi firmware actively
manages it: at 60°C, the blower’s fan is turned on; at 67.5°C, the fan speed is increased; and
finally, at 75°C, the fan increases to full speed. The blower’s fan will spin down automatically
when the temperature drops below these limits.
Generative AI (GenAI)
Generative AI is an artificial intelligence system capable of creating new, original content across
various media such as text, images, audio, and video. These systems learn patterns from
existing data and use that knowledge to generate novel outputs that didn’t previously exist.
Large Language Models (LLMs), Small Language Models (SLMs), and multimodal
models can all be considered types of GenAI when used for generative tasks.
GenAI provides the conceptual framework for AI-driven content creation, with LLMs serving
as powerful general-purpose text generators. SLMs adapt this technology for edge computing,
while multimodal models extend GenAI capabilities across different data types. Together, they
440
represent a spectrum of generative AI technologies, each with its strengths and applications,
collectively driving AI-powered content creation and understanding.
Large Language Models (LLMs) are advanced artificial intelligence systems that understand,
process, and generate human-like text. These models are characterized by their massive scale
in terms of the amount of data they are trained on and the number of parameters they contain.
Critical aspects of LLMs include:
1. Size: LLMs typically contain billions of parameters. For example, GPT-3 has 175 billion
parameters, while some newer models exceed a trillion parameters.
2. Training Data: They are trained on vast amounts of text data, often including books,
websites, and other diverse sources, amounting to hundreds of gigabytes or even terabytes
of text.
3. Architecture: Most LLMs use transformer-based architectures, which allow them to
process and generate text by paying attention to different parts of the input simultane-
ously.
4. Capabilities: LLMs can perform a wide range of language tasks without specific fine-
tuning, including:
• Text generation
• Translation
• Summarization
• Question answering
• Code generation
• Logical reasoning
5. Few-shot Learning: They can often understand and perform new tasks with minimal
examples or instructions.
6. Resource-Intensive: Due to their size, LLMs typically require significant computa-
tional resources to run, often needing powerful GPUs or TPUs.
7. Continual Development: The field of LLMs is rapidly evolving, with new models and
techniques constantly emerging.
8. Ethical Considerations: The use of LLMs raises important questions about bias,
misinformation, and the environmental impact of training such large models.
9. Applications: LLMs are used in various fields, including content creation, customer
service, research assistance, and software development.
441
10. Limitations: Despite their power, LLMs can produce incorrect or biased information
and lack true understanding or reasoning capabilities.
We must note that we use large models beyond text, which we call multi-modal models. These
models integrate and process information from multiple types of input simultaneously. They
are designed to understand and generate content across various data types, such as text, images,
audio, and video.
Certainly. Let’s define open and closed models in the context of AI and language models:
Closed models, also called proprietary models, are AI models whose internal workings, code,
and training data are not publicly disclosed. Examples: GPT-3 and beyond (by OpenAI),
Claude (by Anthropic), Gemini (by Google).
Open models, also known as open-source models, are AI models whose code, architecture,
and, often, training data are publicly available. Examples: Gemma (by Google), LLaMA (by
Meta), and Phi (by Microsoft).
Open models are particularly relevant for running models on edge devices like Raspberry Pi
as they can be more easily adapted, optimized, and deployed in resource-constrained environ-
ments. Still, it is crucial to verify their Licenses. Open models come with various open-source
licenses that may affect their use in commercial applications, while closed models have clear,
albeit restrictive, terms of service.
442
Small Language Models (SLMs)
In the context of edge computing on devices like Raspberry Pi, full-scale LLMs are typically
too large and resource-intensive to run directly. This limitation has driven the development
of smaller, more efficient models, such as the Small Language Models (SLMs).
SLMs are compact versions of LLMs designed to run efficiently on resource-constrained devices
such as smartphones, IoT devices, and single-board computers like the Raspberry Pi. These
models are significantly smaller in size and computational requirements than their larger coun-
terparts while still retaining impressive language understanding and generation capabilities.
Key characteristics of SLMs include:
1. Reduced parameter count: Typically ranging from a few hundred million to a few
billion parameters, compared to two-digit billions in larger models.
2. Lower memory footprint: Requiring, at most, a few gigabytes of memory rather than
tens or hundreds of gigabytes.
3. Faster inference time: Can generate responses in milliseconds to seconds on edge
devices.
4. Energy efficiency: Consuming less power, making them suitable for battery-powered
devices.
5. Privacy-preserving: Enabling on-device processing without sending data to cloud
servers.
6. Offline functionality: Operating without an internet connection.
SLMs achieve their compact size through various techniques such as knowledge distillation,
model pruning, and quantization. While they may not match the broad capabilities of larger
models, SLMs excel in specific tasks and domains, making them ideal for targeted applications
on edge devices.
We will generally consider SLMs —language models with fewer than 5 to 8 billion
parameters — quantized to 4 bits.
443
Examples of SLMs include compressed versions of models like Meta Llama, Microsoft PHI,
and Google Gemma. These models enable a wide range of natural language processing tasks
directly on edge devices, from text classification and sentiment analysis to question answering
and limited text generation.
For more information on SLMs, the paper, LLM Pruning and Distillation in Practice: The
Minitron Approach, provides an approach applying pruning and distillation to obtain SLMs
from LLMs. And, SMALL LANGUAGE MODELS: SURVEY, MEASUREMENTS, AND
INSIGHTS, presents a comprehensive survey and analysis of Small Language Models (SLMs),
which are language models with 100 million to 5 billion parameters designed for resource-
constrained devices.
444
Ollama
The primary and most user-friendly tool for running Small Language Models (SLMs) directly
on a Raspberry Pi, especially the Pi 5, is Ollama. It is an open-source framework that allows us
to install, manage, and run various SLMs (such as TinyLlama, smollm, Microsoft Phi, Google
Gemma, Meta Llama, MoonDream, LLaVa, among others) locally on our Raspberry Pi for
tasks such as text generation, image captioning, and translation.
Alternatives options:
[Link] the Hugging Face Transformers library are well-supported for run-
ning SLMs on Raspberry Pi. [Link] is particularly efficient for running quan-
tized models natively. At the same time, Hugging Face Transformers offers a
broader range of models and tasks, which are best suited to smaller architectures
due to hardware limitations. Note that Ollama run [Link] under the hood.
1. Local Model Execution: Ollama enables running LMs on personal computers or edge
devices such as the Raspi-5, eliminating the need for cloud-based API calls.
445
2. Ease of Use: It provides a simple command-line interface for downloading, running,
and managing different language models.
3. Model Variety: Ollama supports various LLMs, including Phi, Gemma, Llama, Mistral,
and other open-source models.
4. Customization: Users can create and share custom models tailored to specific needs
or domains.
5. Lightweight: Designed to be efficient and run on consumer-grade hardware.
6. API Integration: Offers an API that allows integration with other applications and
services.
7. Privacy-Focused: By running models locally, it addresses privacy concerns associated
with sending data to external servers.
8. Cross-Platform: Available for macOS, Windows, and Linux systems (our case, here).
9. Active Development: Regularly updated with new features and model support.
10. Community-Driven: Benefits from community contributions and model sharing.
To learn more about what Ollama is and how it works under the hood, you should see this
short video from Matt Williams, one of the founders of Ollama:
[Link]
Matt has an entirely free course about Ollama that we recommend: [Link]
[Link]/9KEUFe4KQAI?si=D_-q3CMbHiT-twuy
Installing Ollama
Let’s set up and activate a Virtual Environment for working with Ollama:
As a result, an API will run in the background on [Link]:11434. From now on, we can
run Ollama via the terminal. For starting, let’s verify the Ollama version, which will also tell
us that it is correctly installed:
ollama -v
446
On the Ollama Library page, we can find the models Ollama supports. For example, by
filtering by Most popular, we can see Meta Llama, Google Gemma, Microsoft Phi, LLaVa,
etc.
447
Meta Llama 3.2 1B/3B
Let’s install and run our first small language model, Llama 3.2 1B (and 3B). The Meta Llama,
3.2 collections of multilingual large language models (LLMs), is a collection of pre-trained
and instruction-tuned generative models in 1B and 3B sizes (text in/text out). The Llama 3.2
instruction-tuned text-only models are optimized for multilingual dialogue use cases, including
agentic retrieval and summarization tasks.
The 1B and 3B models were pruned from the Llama 8B, and then logits from the 8B and 70B
models were used as token-level targets (token-level distillation). Knowledge distillation was
used to recover performance (they were trained with 9 trillion tokens). The 1B model has
1,24B, quantized to integer (Q8_0), and the 3B, 3.12B parameters, with a Q4_0 quantization,
which ends with a size of 1.3 GB and 2GB, respectively. Its context window is 131,072 tokens.
448
Install and run the Model
Running the model with the command before, we should have the Ollama prompt available
for us to input a question and start chatting with the LLM model; for example,
>>> What is the capital of France?
Almost immediately, we get the correct answer:
The capital of France is Paris.
Using the option --verbose when calling the model will generate several statistics about its
performance (The model will be polling only the first time we run the command).
449
Each metric gives insights into how the model processes inputs and generates outputs. Here’s
a breakdown of what each metric means:
• Total Duration (2.620170326s): This is the complete time taken from the start of
the command to the completion of the response. It encompasses loading the model,
processing the input prompt, and generating the response.
• Load Duration (39.947908ms): This duration indicates the time to load the model
or necessary components into memory. If this value is minimal, it can suggest that the
model was preloaded or that only a minimal setup was required.
• Prompt Eval Count (32 tokens): The number of tokens in the input prompt. In
NLP, tokens are typically words or subwords, so this count includes all the tokens that
the model evaluated to understand and respond to the query.
• Prompt Eval Duration (1.644773s): This measures the model’s time to evaluate or
process the input prompt. It accounts for the bulk of the total duration, implying that
understanding the query and preparing a response is the most time-consuming part of
the process.
• Prompt Eval Rate (19.46 tokens/s): This rate indicates how quickly the model
processes tokens from the input prompt. It reflects the model’s speed in terms of natural
450
language comprehension.
• Eval Count (8 token(s)): This is the number of tokens in the model’s response, which
in this case was, “The capital of France is Paris.”
• Eval Duration (889.941ms): This is the time taken to generate the output based
on the evaluated input. It’s much shorter than the prompt evaluation, suggesting that
generating the response is less complex or computationally intensive than understanding
the prompt.
• Eval Rate (8.99 tokens/s): Similar to the prompt eval rate, this indicates the speed
at which the model generates output tokens. It’s a crucial metric for understanding the
model’s efficiency in output generation.
This detailed breakdown can help understand the computational demands and performance
characteristics of running SLMs like Llama on edge devices like the Raspberry Pi 5. It shows
that while prompt evaluation is more time-consuming, the actual generation of responses is
relatively quicker. This analysis is crucial for optimizing performance and diagnosing potential
bottlenecks in real-time applications.
Loading and running the 3B model, we can see the difference in performance for the same
prompt;
The eval rate is lower, 5.3 tokens/s versus 9 tokens/s with the smaller model.
When question about
>>> What is the distance between Paris and Santiago, Chile?
The 1B model answered 9,841 kilometers (6,093 miles), which is inaccurate, and the 3B
model answered 7,300 miles (11,700 km), which is close to the correct (11,642 km).
Let’s ask for the Paris’s coordinates:
451
>>> what is the latitude and longitude of Paris?
Google Gemma
Google Gemma, is a collection of lightweight, state-of-the-art open models built from the same
technology that powers our Gemini models. Today, the Gemma family has the Gemma3 and
Gemma3n models.
452
Gemma3
We can, for example, install gemma3:latest, using ollama run gemma3:latest. This model
has 4.3B parameters, with a context length of 8,192 and an embedding length of 2,560. A
typical quantization schema is the Q4_K_M. This model has vision capabilities. Besides the
4B, we can also install the 1B parameter model.
Install and run the Model
Running the model with the command before, we should have the Ollama prompt available
for us to input a question and start chatting with the LLM model; for example,
>>> What is the capital of France?
Almost immediately, we get the correct answer:
The capital of France is **Paris**. It's a global center for art, fashion,
gastronomy, and culture. � Do you want to know anything more about Paris?
And its statistics.
We can see that Gemma 3:4B has roughly the same performance as Lama 3.2:3B,
despite having more parameters.
453
Other examples:
A good and accurate answer (a little more verbose than the Llama answers).
An advantage of the Gemma3 models are their vision capability, for example, we can ask it to
caption an image:
454
This is a very accurate model, but it still has high latency, as we can see in the example above
(more than 3 minutes to caption the image).
The Gemma 3 1B size models are text only and don’t support image input.
Gemma 3n
Gemma3 is a powerful, efficient open-source model that runs locally on phones, tablets, and
laptops. The models are listed with parameter counts, such as E2B and E4B, that are lower than
the total number of parameters contained in the models. The E prefix indicates these models
can operate with a reduced set of Effective parameters. This reduced-parameter operation can
be achieved using the flexible parameter technology built into Gemma 3n models, which helps
them run efficiently on lower-resource devices.
The parameters in Gemma 3n models are divided into four main groups: text, visual, audio,
and per-layer embedding (PLE) parameters. In the standard execution of the E2B model,
over 5 billion parameters are loaded. However, by using parameter skipping and PLE caching,
this model can be operated with an effective memory load of just under 2 billion (1.91B)
parameters, as illustrated below:
455
Figure 12: Immage from Google: Gemma 3n diagram of parameter usage
Once installed, (using ollama run gemma3n:e2b), inspecting the model we get:
• Architecture: gemma3n
• Parameters: 4.5B
• Quantization Q4_K_M
• Capabilities: completion only
When we run it, we can see that, despite having 4.5B parameters, it is faster than
Llama3.3:3B.
456
Microsoft Phi3.5 3.8B
Let’s pull now the PHI3.5, a 3.8B lightweight state-of-the-art open model by Microsoft. The
model belongs to the Phi-3 model family and supports 128K token context length and the
languages: Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, He-
brew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish,
Swedish, Thai, Turkish, and Ukrainian.
The model size, in terms of bytes, will depend on the specific quantization format used. The
size can go from 2-bit quantization (q2_k) of 1.4 GB (higher performance/lower quality) to
16-bit quantization (fp-16) of 7.6 GB (lower performance/higher quality).
Let’s run the 4-bit quantization (Q4_0), which will need 2.2 GB of RAM, with an intermediary
trade-off regarding output quality and performance.
457
ollama run phi3.5:3.8b --verbose
You can use run or pull to download the model. What happens is that Ollama
keeps note of the pulled models, and once the PHI3 does not exist, before running
it, Ollama pulls it.
...
458
In this case, the answer was still longer than we expected, with an eval rate of 2.25 tokens/s,
more than double that of Gemma and Llama.
Choosing the most appropriate prompt is one of the most important skills to be
used with LLMs, no matter its size.
When we asked the same questions about distance and Latitude/Longitude, we did not get a
good answer for a distance of 13,507 kilometers (8,429 miles), but it was OK for coordi-
nates. Again, it could have been less verbose (more than 200 tokens for each answer).
We can use any model as an assistant since their speed is relatively decent, but in October 2024,
the Llama2:3B or Gemma 3 are better choices. Try other models, depending on your needs.
� Open LLM Leaderboard can give you an idea about the best models in size, benchmark,
license, etc.
The best model to use is the one fit for your specific necessity. Also, take into
consideration that this field evolves with new models everyday.
459
MoonDream
Moondream is an open-source visual language model that understands images using simple
text prompts. It’s fast and wildly capable. It is 3 to 4 times faster than the Gemma3, for
example.
460
Moondream can interpret images much like a human would, enabling tasks like: • Captioning
and describing images in detail • Visual question answering (VQA) such as “What is the
color of the shirt?” • Object detection and pointing out locations • Counting and contextual
reasoning about visual scenes
461
Lava-Phi-3
Another multimodal model is the LLaVA-Phi-3, a fine-tuned LLaVA model from Phi 3 Mini
4k. It has strong performance benchmarks that are on par with the original LLaVA (Large
Language and Vision Assistant) model.
In terms of latency, it is a pair with Gemma3 and much slower than MoonDream
The response took around 30s, with an eval rate of 3.93 tokens/s! Not bad!
But let us know to enter with an image as input. For that, let’s create a directory for working:
cd Documents/
mkdir OLLAMA
cd OLLAMA
Let’s download a 640x320 image from the internet, for example (Wikipedia: Paris, France):
462
Using FileZilla, for example, let’s upload the image to the OLLAMA folder on the Raspi-5
and name it image_test_1.jpg. We should have the whole image path (we can use pwd to
get it).
/home/mjrovai/Documents/OLLAMA/image_test_1.jpg
If you use a desktop, you can copy the image path by right-clicking the image.
463
Let’s enter with this prompt:
The result was great, but the overall latency was significant; almost 4 minutes to perform the
inference.
464
Inspecting Model parameters
It is possible to know more about each model downloaded by Ollama, using the command
ollama show <model>. For example:
465
We’re looking at a 3.2B‑parameter Llama‑family instruction model, quantized for efficient local
use. Let’s walk through each field and what it implies in practice.
High‑level identity
466
Capabilities
– tools: the model has been trained/fine‑tuned to call tools/functions (e.g., JSON
function calls) when the host runtime exposes them, so it can decide “I should call
a function to get data” vs just answering from its own weights.
This doesn’t mean the tools are built into the model; it means it knows how to
format tool calls when the surrounding system supports them.
Quantization: Q4_K_M
• llama.context_length: 131072
Context length is 131,072 tokens (~131k):
– This is a very large context window (comparable to “128k context” models).
– It can ingest long documents, multiple files, or extended conversations without
immediately forgetting earlier content.
– Under the hood, this usually implies some form of advanced attention scaling (e.g.,
RoPE scaling, attention optimizations, maybe sparse/flash attention), but the key
user effect is: we can feed a lot of text at once.
467
• llama.embedding_length: 3072
The hidden/embedding dimension is 3072:
– Each token is represented as a 3072‑dimensional vector internally.
– This roughly sets the “width” of the model: wider models can represent richer
patterns but cost more per token.
– Often, this is also the size of embeddings we’d extract if using the model as an
encoder (assuming our runtime supports that mode).
– We should read the exact license text bundled with the model before using it in
commercial products, especially at scale.
Parameters
The only parameters explicitly set in this Modelfile are three stop strings:
• <|start_header_id|>
• <|end_header_id|>
• <|eot_id|>
No temperature, top_k, top_p, etc., are defined in the Model file. Those, therefore, use
Ollama’s defaults for this model at runtime.
• temperature 0.3,
• top_p 0.9and
• top_k 40
468
Inspecting local resources
htop
During the time that the model is running, we can inspect the resources:
All four CPUs run at almost 100% of their capacity, and the memory used with the model
loaded is 3.24GB. Exiting Ollama, the memory goes down to around 377MB (with no desk-
top).
It is also essential to monitor the temperature. When running the Raspberry with a desktop,
you can have the temperature shown on the taskbar:
If you are “headless”, the temperature can be monitored continuously (for example, every 2
seconds) with the command:
469
watch -n 2 vcgencmd measure_temp
If you are doing nothing, the temperature is around 50°C for CPUs running at 1%. During
inference, with the CPUs at 100%, the temperature can rise to almost 70°C. This is OK and
means the active cooler is working, keeping the temperature below 80°C / 85°C (its limit).
So far, we have explored SLMs’ chat capability using the command line on a terminal. However,
we want to integrate those models into our projects, so Python seems to be the right path.
The good news is that Ollama has such a library.
The Ollama Python library simplifies interaction with advanced LLM models, enabling more
sophisticated responses and capabilities, besides providing the easiest way to integrate Python
3.8+ projects with Ollama.
For a better understanding of how to create apps using Ollama with Python, we can follow
Matt Williams’s videos, as the one below:
[Link]
Installation:
In the terminal (and inside the virtual environment), run the command:
470
For better using, we will need a text editor or an IDE to create a Python script. If we
run Raspberry Pi OS on a desktop, several options, such as Thonny and Geany, are already
installed by default (accessed via [Menu][Programming]). You can download other IDEs, such
as Visual Studio Code, from [Menu][Recommended Software]. When the window pops up,
go to [Programming], select your option, and press [Apply].
A good option to run Python scripts during development is to use the Jupyter Notebook:
We can use the Jupyter Notebook using SSH on our host computer. Run the command below,
changing the IP address for the Raspi:
On the terminal, you can see the local URL address to open the notebook:
471
We can access it from another computer by entering the Raspberry Pi’s IP address and the
provided token in a web browser (we should copy it from the terminal).
In our working directory (Documents/OLLAMA/) on the Raspberry Pi, we will create a new
Python 3 notebook.
472
Let’s enter with a very simple script to verify the installed models:
import ollama
models = [Link]()
for model in models['models']:
print(model['model'])
gemma3:4b
riven/smolvlm:latest
gemma3n:e2b
moondream:latest
llama3.2:3b
Same as on the terminal using ollama show llama3.2: 3b, we can get model information
with the Python command [Link]('llama3.2:3b'):
info = [Link]('llama3.2:3b')
473
print("Quantization Level :", getattr([Link], 'quantization_level', None))
print("Family :", getattr([Link], 'family', None))
print("Supported Capabilities:", getattr(info, 'capabilities', None))
print("Date Modified (Local) :", getattr(info, 'modified_at', None))
print("License Short :", (str(getattr(info, 'license', None)).split('\n')[0]))
for key in [
'[Link]', '[Link]',
'llama.context_length', 'llama.embedding_length', 'llama.block_count'
]:
print(f" - {key}: {[Link](key)}")
As a result, we have:
Model/Format : None
Parameter Size : 3.2B
Quantization Level : Q4_K_M
Family : llama
Supported Capabilities : ['completion', 'tools']
Date Modified (Local) : 2025-09-15 15:39:48.153510-03:00
License Short : LLAMA 3.2 COMMUNITY LICENSE AGREEMENT
Key Architecture Details :
- [Link]: llama
- [Link]: Instruct
- llama.context_length: 131072
- llama.embedding_length: 3072
- llama.block_count: 28
• More layers typically improve the model’s ability to perform multi‑step reasoning and to
build hierarchical representations.
474
• For 3.2B parameters, 28 layers with width 3072 is a sensible trade‑off (moderately deep,
not extremely wide).
• Compared with a 7B model (~32–40 layers, 4k+ width), this will be less capable on very
complex reasoning or coding, but noticeably more capable than 1B‑class models.
Model/Format: None : This usually means the frontend (e.g., our UI or manager) didn’t
detect a specific “format template” (like “chat”, “instruct”, “plain”) beyond knowing it’s
Llama:
• The underlying file is probably something like a GGUF or similar binary weight file.
• Prompt formatting will depend on our runtime’s default “llama‑instruct” template; if
nothing is configured, we may need to wrap instructions manually (system/user style or
clear “You are an assistant…” preamble).
Ollama Generate
Let’s repeat one of the questions that we did before, but now using [Link]() from
the Ollama Python library. This API generates a response for the given prompt using the
provided model. This is a streaming endpoint, so there will be a series of responses. The final
response object will include statistics and additional data from the request.
response = [Link](
model="llama3.2:3b",
prompt="What is the capital of Brazil"
)
print(response['response'])
If you are running the code as a Python script, you should save it as, for example,
test_ollama.py. You can run it in the IDE or run it directly in the terminal. Also,
remember always to call the model and define it when running a stand-alone script.
python test_ollama.py
475
return response
MODEL="llama3.2:3b"
response = simple_query ("What is the capital of Peru?")
response
Let’s print the full response now. As a result, we will have the model response in a JSON
format:
import json
print([Link](response.__dict__, indent=2))
{
"model": "llama3.2:3b",
"created_at": "2026-02-26T18:58:30.730414614Z",
"done": true,
"done_reason": "stop",
"total_duration": 2307368739,
"load_duration": 343767495,
"prompt_eval_count": 32,
"prompt_eval_duration": 722764406,
"eval_count": 8,
"eval_duration": 1229567370,
"response": "The capital of Peru is Lima.",
"thinking": null,
"context": [
128006,
9125,
...
13
],
"logprobs": null
}
476
• response: the main output text generated by the model in response to our prompt.
• context: the token IDs representing the input and context used by the model. Tokens
are numerical representations of text that the language model uses to process it.
But what we want is the plain ‘response’ and, perhaps for analysis, the total duration of the
inference, so let’s change the code to extract it from the dictionary:
print(f"\n{response['response']}")
print(f"\nTotal Duration: {(response['total_duration']/1e9):.2f} seconds")
Now, we got:
print(f"eval_count: {response['eval_count']}")
print(f"eval_duration: {(response['eval_duration']/1e9):.2f} s")
print(f"eval_rate: {response['eval_count']/(response['eval_duration']/1e9):.2f} tokens/s")
eval_count: 8
eval_duration: 1.23 s
eval_rate: 6.51 tokens/s
477
Streaming with [Link]()
To stream the output from [Link]() in Python, set stream=True and iterate over
the generator to print each chunk as it’s produced. This enables real-time response streaming,
similar to chat models.
stream = [Link](
model='llama3.2:3b',
prompt='Tell me an interesting fact about Brazil',
stream=True
)
This approach is ideal for long or complex generations, making the user experience feel faster
and more interactive.
System Prompt
Add a system parameter to set overall instructions or behavior for the model (useful for role
assignment and tone control):
response = [Link](model='llama3.2:3b',
prompt='Tell about industry',
system='You are an expert on Brazil.',
stream=False)
478
print(response['response'])
Control creativity/randomness via temperature, and customize output style with extra set-
tings like top_p and num_ctx (context window size)
response = [Link](
model='llama3.2:3b',
prompt='Why the sky is blue?',
options={'temperature':0.1},
stream=True)
for chunk in response:
print(chunk['response'], end='', flush=True)
479
response = [Link](
messages=[
{"role": "user",
"content": "Poetically describe Paris in one short sentence"},
],
model='llama3.2:3b',
options={"temperature": 1.0} # Set. temp. to 1.0 for more creativity
)
print(response['message']['content'])
More Options
Besides temperature, we can control a lot via the options={…} parameter in [Link].
Key ones:
Sampling/style
Length/context
480
Putting it together in your code:
response = [Link](
model='llama3.2:3b',
prompt='Why is the sky blue?',
options={
'temperature': 0.1,
'top_p': 0.9,
'top_k': 40,
'repeat_penalty': 1.1,
'num_predict': 50, # Limit the answer
'num_ctx': 4096,
'seed': 42,
},
stream=True,
)
for chunk in response:
print(chunk['response'], end='', flush=True)
The sky appears blue because of a phenomenon called Rayleigh scattering, named after the
British physicist Lord Rayleigh, who first described it in the late 19th century.
top_k and top_p both control how random and diverse the model’s next tokens can be, but
they do it in different ways.
top_k: fixed shortlist size
• At each step, the model ranks all possible next tokens by probability.
• top_k = K means: keep only the K most likely tokens, set all others to probability
0, then sample from those K.
• Effects:
481
– Small k (1–20) → very focused, deterministic, “on‑rails”; less creative, fewer weird
tokens.
– Large k (50–200+) → more variety and creativity; higher chance of unusual or
off‑topic words.
• Extreme:
– top_k = 1 → greedy decoding: always pick the single most likely token, almost no
randomness.
• Starting from the top, you add tokens until their cumulative probability � p. Then
you sample only from that set.
• top_p = 0.9 means: “consider just enough top tokens to cover 90% of the probability
mass”.
• Effects:
– In confident situations (one token is clearly best), the shortlist may be very small
→ behavior similar to low k.
– In uncertain situations (probabilities spread out), more tokens enter the shortlist
→ more exploration.
• Typical:
– top_p � 0.9–0.95 is a common sweet spot for natural but not too wild text.
• Higher top_k / top_p → more diverse, more creative, but also more risk of nonsense.
We can use them together—for example, top_k=40, top_p=0.9—where top_k gives a hard
cap and top_p then trims that set by probability mass.
[Link]()
Another way to get our response is to use [Link](), which generates the next message
in a chat with a provided model. This is a streaming endpoint, so a series of responses will
occur. Streaming can be disabled using "stream": false. The final response object will also
include statistics and additional data from the request.
482
PROMPT_1 = 'What is the capital of France?'
In the above code, we are running two queries, and the second prompt considers the result of
the first one.
Here is how the model responded:
483
The above code works with two prompts. Let’s include a conversation variable to really provide
the chat with a memory:
# Question
prompt = "What is the capital of Brazil"
print(chat_with_memory(prompt))
Image Description:
As we did with the visual models and the command line to analyze an image, the same
can be done here with Python. Let’s use the same image of Paris, but now with the
[Link]():
484
MODEL = 'llava-phi3:3.8b'
PROMPT = "Describe this picture"
response = [Link](
model=MODEL,
prompt=PROMPT,
images= [img]
)
print(f"\n{response['response']}")
print(f"\n [INFO] Total Duration: {(res['total_duration']/1e9):.2f} seconds")
This image captures the iconic cityscape of Paris, France. The vantage point
is high, providing a panoramic view of the Seine River that meanders through
the heart of the city. Several bridges arch gracefully over the river,
connecting different parts of the city. The Eiffel Tower, an iron lattice
structure with a pointed top and two antennas on its summit, stands tall in the
background, piercing the sky. It is painted in a light gray color, contrasting
against the blue sky speckled with white clouds.
The buildings that line the river are predominantly white or beige, their uniform
color palette broken occasionally by red roofs peeking through. The Seine River
itself appears calm and wide, reflecting the city's architectural beauty in its
surface. On either side of the river, trees add a touch of green to the urban
landscape.
The image is taken from an elevated perspective, looking down on the city. This
viewpoint allows for a comprehensive view of Paris's beautiful architecture and
layout. The relative positions of the buildings, bridges, and other structures
create a harmonious composition that showcases the city's charm.
In summary, this image presents a serene day in Paris, with its architectural
marvels - from the Eiffel Tower to the river-side buildings - all bathed in soft
colors under a clear sky.
The model took about 4 minutes (256.45 s) to return with a detailed image description.
485
Let’s capture an image from the Raspberry Pi camera and get the description, now using the
MoonDream model:
import time
import numpy as np
import [Link] as plt
from picamera2 import Picamera2
from PIL import Image
def capture_image(image_path):
# Initialize camera
picam2 = Picamera2() # default is index 0
# Capture image
picam2.capture_file(image_path)
print("Image captured: "+image_path)
# Stop camera
[Link]()
[Link]()
Using the above code, we can capture an image, which can be displayed with:
def show_image(image_path):
img = [Link](image_path)
486
def image_description(img_path, model):
with open(img_path, 'rb') as file:
response = [Link](
model=model,
messages=[
{
'role': 'user',
'content': '''return the description of the image''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,
}
)
return response
Now, let’s put all togheter and capture an image from the camera:
IMG_PATH = "/home/mjrovai/Documents/OLLAMA/SST/capt_image.jpg"
MODEL = "moondream:latest"
apture_image(IMG_PATH)
show_image(IMG_PATH)
response = image_description(IMG_PATH, MODEL)
caption = response['message']['content']
print ("\n==> AI Response:", caption)
print(f"\n[INFO] ==> Total Duration: {
(response['total_duration']/1e9):.2f} seconds")
487
We got the description:
The image features a green table with various items on it. A white mug adorned with
black faces is prominently displayed, and there are several other mugs scattered around
the table as well. In addition to the mugs, there's also a microphone placed near them,
suggesting that this might be an office or workspace setting where someone could enjoy
their coffee while recording podcasts or audio content.
488
A computer keyboard can be seen in the background, indicating that it is likely
connected to a computer for work purposes. A mouse and a cell phone are also present on
the table, further emphasizing the technology-oriented nature of this scene.
We can now change the image_description function to ask “Who are the faces in the mug?”.
The answer:
==> AI Response:
The mug has a picture of the Beatles on it.
One alternative to running an SLM in Python using Ollama is to call the API directly. Let’s
explore some advantages and disadvantages of both methods.
Python Library:
response = [Link](
model=MODEL,
prompt=QUERY)
result = response['response']
import requests
import json
# Configuration
OLLAMA_URL = "[Link]
MODEL = MODEL
response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": QUERY,
489
"stream": False
}
)
response = [Link](
f"{OLLAMA_URL}/generate",
json={"model": MODEL,
"prompt": query,
"stream": False}
)
result = [Link]().get("response", "")
One clear advantage of the Python library is that it handles URL construction,
request formatting, and response parsing.
Error Handling
Python Library:
Connection Management
Python Library:
490
Features & Functionality
Python Library:
• Clean access to all Ollama features (generate, chat, embeddings, list models, pull, etc.)
• Streaming is simple: for chunk in [Link](..., stream=True)
• Type hints and better IDE support
Dependencies
Python Library:
Advanced Features
Python Library:
# Streaming
for chunk in [Link](model=MODEL, prompt=query, stream=True):
print(chunk['response'], end='')
# Chat history
response = [Link](
model=MODEL,
messages=[
491
{'role': 'user', 'content': 'Hello!'},
{'role': 'assistant', 'content': 'Hi there!'},
{'role': 'user', 'content': 'How are you?'}
]
)
# List models
models = [Link]()
# Pull models
[Link]('llama3.2:3b')
Performance Difference
Minimal difference in practice! The Python library uses httpx which is comparable to
requests. Both make the same underlying HTTP calls to Ollama.
492
Bottom line: For most use cases, the Python library is the better choice due
to its simplicity and built-in features. Use direct API calls only when you need
specific control or have constraints that prevent adding the dependency.
Going Further
The small LLM models tested worked well at the edge, both with text and with images, but,
of course, the last one had high latency. A combination of specific, dedicated models can lead
to better results; for example, in real cases, an Object Detection model (such as YOLO) can
provide a general description and count of objects in an image, which, once passed to an LLM,
can help extract essential insights and actions.
According to Avi Baum, CTO at Hailo,
In the vast landscape of artificial intelligence (AI), one of the most intriguing
journeys has been the evolution of AI on the edge. This journey has taken us from
classic machine vision to the realms of discriminative AI, enhancive AI, and now,
the groundbreaking frontier of generative AI. Each step has brought us closer to a
future where intelligent systems seamlessly integrate with our daily lives, offering
an immersive experience of not just perception but also creation at the palm of our
hand.
493
Conclusion
This chapter has demonstrated how a Raspberry Pi 5 can be transformed into a potent AI
hub capable of running large language models (LLMs) for real-time, on-site data analysis and
insights using Ollama and Python. The Raspberry Pi’s versatility and power, coupled with
the capabilities of lightweight LLMs like Llama 3.2 and MoonDream, make it an excellent
platform for edge computing applications.
The potential of running LLMs on the edge extends far beyond simple data processing, as in
this lab’s examples. Here are some innovative suggestions for using this project:
1. Smart Home Automation:
• Integrate SLMs to interpret voice commands or analyze sensor data for intelligent home
automation. This could include real-time monitoring and control of home devices, se-
curity systems, and energy management, all processed locally without relying on cloud
services.
• Deploy SLMs on Raspberry Pi in remote or mobile setups for real-time data collection
and analysis. This can be used in agriculture to monitor crop health, in environmental
studies for wildlife tracking, or in disaster response for situational awareness and resource
management.
3. Educational Tools:
• Create interactive educational tools that leverage SLMs to provide instant feedback,
language translation, and tutoring. This can be particularly useful in developing regions
with limited access to advanced technology and internet connectivity.
4. Healthcare Applications:
• Use SLMs for medical diagnostics and patient monitoring. They can provide real-time
analysis of symptoms and suggest potential treatments. This can be integrated into
telemedicine platforms or portable health devices.
6. Industrial IoT:
494
• Integrate SLMs into industrial IoT systems for predictive maintenance, quality control,
and process optimization. The Raspberry Pi can serve as a localized data processing
unit, reducing latency and improving the reliability of automated systems.
7. Autonomous Vehicles:
• Use SLMs to process sensory data from autonomous vehicles, enabling real-time decision-
making and navigation. This can be applied to drones, robots, and self-driving cars for
enhanced autonomy and safety.
• Implement SLMs to provide interactive and informative cultural heritage sites and mu-
seum guides. Visitors can use these systems to get real-time information and insights,
enhancing their experience without internet connectivity.
• Use SLMs to analyze and generate creative content, such as music, art, and literature.
This can foster innovative projects in the creative industries and allow for unique inter-
active experiences in exhibitions and performances.
Resources
495
SLM: Basic Optimization Techniques
Figure 13: DALL·E prompt - I am writing a tutorial using Raspberry Pi. I am talking about
Optimization Techniques, such as Function Calling and RAG, using Ollama and
SLMs. I want a landscape-format image for the tutorial cover (without title). Should
be a cartoon styled in 5Os
Introduction
Large Language Models (LLMs) have revolutionized natural language processing, but their de-
ployment and optimization come with unique challenges. One significant issue is the tendency
496
for LLMs (and more, the SLMs) to generate plausible-sounding but factually incorrect infor-
mation, a phenomenon known as hallucination. This occurs when models produce content
that appears coherent but lacks grounding in truth or real-world facts.
Other challenges include the immense computational resources required for training and run-
ning these models, the difficulty in maintaining up-to-date knowledge within the model, and
the need for domain-specific adaptations. Privacy concerns also arise when handling sensitive
data during training or inference. Additionally, ensuring consistent performance across diverse
tasks and maintaining ethical use of these powerful tools present ongoing challenges. Address-
ing these issues is crucial for the effective and responsible deployment of LLMs in real-world
applications.
The fundamental and more common techniques for enhancing LLM (and SLM) performance
and efficiency are Function (or Tool) Calling, Prompt engineering, Fine-tuning, and Retrieval-
Augmented Generation (RAG).
• Function (Tool) calling allows models to perform actions beyond generating text. By
integrating with external functions or APIs, SLMs can access real-time data, automate
tasks, and perform precise calculations—addressing the reliability issues that arise from
the model’s limitations in mathematical operations.
• Prompt engineering is at the forefront of LLM optimization. By carefully crafting
input prompts, we can guide models to produce more accurate and relevant outputs. This
technique involves structuring queries that leverage the model’s pre-trained knowledge
and capabilities, often incorporating examples or specific instructions to shape the desired
response.
• Retrieval-Augmented Generation (RAG) represents a powerful approach that’s
ideal for resource-constrained edge devices. This method combines the knowledge em-
bedded in pre-trained models with the ability to access external, up-to-date information
without requiring fine-tuning. By retrieving relevant data from a local knowledge base,
RAG significantly enhances accuracy and reduces hallucinations—all without the com-
putational overhead of model retraining.
• Fine-tuning, while more resource-intensive, offers a way to specialize LLMs for specific
domains or tasks. This process involves further training the model on carefully curated
datasets, allowing it to adapt its vast general knowledge to particular applications. Fine-
tuning can lead to substantial performance improvements, especially in specialized fields
or for unique use cases.
In this chapter, we’ll start focusing on two techniques that are particularly well-suited for edge
devices like the Raspberry Pi: Function Calling and RAG.
We will learn more in detail about optimization techniques for SLMs, in the chapter:
Advancing EdgeAI: Beyond Basic SLMs
497
Function Calling Introduction
So far, we can see that, with the model’s (“response”) answer to a variable, we can efficiently
work with it and integrate it into real-world projects. However, a big problem is that the
model can respond differently to the same prompt. Let’s say, as in the last examples, that we
want the model’s response to be only the name of a given country’s capital and its coordinates,
nothing more, even with very verbose models such as the Microsoft Phi. We can use the
Ollama function's calling to guarantee the same answers, which is perfectly compatible
with the OpenAI API.
In modern artificial intelligence, function calling with Large Language Models (LLMs) allows
these models to perform actions beyond generating text. By integrating with external functions
or APIs, LLMs can access real-time data, automate tasks, and interact with various systems.
For instance, instead of merely responding to a weather query, an LLM can call a weather API
to fetch the current conditions and provide accurate, up-to-date information. This capability
enhances the relevance and accuracy of the model’s responses, making it a powerful tool for
driving workflows and automating processes, thereby transforming it into an active participant
in real-world applications.
For more details about Function Calling, please see this video made by Marvin Prison:
[Link]
And on this link: HuggingFace Function Calling
123456*123456
The result would be: 15,241,383,936. No issues on it, but let’s ask a SLM to do the same
simple task:
import ollama
response = [Link](
model='llama3.2:3B',
messages=[{
"role": "user",
498
"content": "What is 123456 multiplied by 123456? Only give me the answer"
}],
options={"temperature": 0}
)
• LLMs work by predicting the next most likely token based on patterns in training data
• They don’t perform actual arithmetic operations
• They’re essentially “guessing” what a plausible answer looks like
This makes it nearly impossible for the model to “see” the actual numbers properly for com-
putation.
The LLM will likely give something that “looks” like a big number but is mathematically
incorrect.
499
The Solution: Function Calling / Tool Use
Setting temperature=0 makes the output deterministic (same input → same output), but
it doesn’t make it correct. The model will confidently give the same wrong answer every
time.
Best Practices
500
result = multiply(123456, 123456) # Python does the math
Bottom line: Use LLMs for natural language understanding and intent classifica-
tion, but delegate actual computations to proper tools/functions. This is the core
principle behind tool use and function calling in modern LLM applications!
multiply_tool = {
"type": "function",
"function": {
"name": "multiply_numbers",
"description": "Multiply two numbers together",
"parameters": {
"type": "object",
"required": ["a", "b"],
"properties": {
"a": {"type": "number", "description": "First number"},
"b": {"type": "number", "description": "Second number"}
}
}
}
}
Now, let’s create a function to handle the user query and calling for the tool when needed.
def answer_query(QUERY):
response = [Link](
501
'llama3.2:3B',
messages=[{"role": "user", "content": QUERY}],
tools=[multiply_tool]
)
Result: 15,241,383,936.00
Great! And now, can I use the same code to answer general questions? Let’s test it:
Result: 1,000,000.00
The result is wrong. So, the above approach works fine for using the tool, but to answer it
correctly (even without a tool), we should implement an “agentic approach”, which is a subject
for later (See the Chapter: Advancing EdgeAI: Beyond Basic SLMs)
Suppose we want an SLM to return the distance in km from the capital city of the country
specified by the user to the user’s current location. We can see that the first is not so simple:
502
it is not always enough to enter only the country’s name; the SLM can also give us a different
(and incorrect) answer every time.
OK, for trying to mitigate it, let’s create an app where the user enters a country’s name and
gets, as an output, the distance in km from the capital city of such a country and the app’s
location (for simplicity, we will use Santiago, Chile, as the app location).
Once the user enters a country name, the model should return the capital city’s name (as
503
a string) and its latitude and longitude (as floats). Using those coordinates, we can use
a simple Python library (haversine) to compute the great‑circle distance between the two
latitude/longitude points.
The idea of this project is to demonstrate a combination of language model interaction (IA)
and geospatial calculations using the Haversine formula (traditional computing).
First, let us install the Haversine library:
Now, we should create a Python script designed to interact with our model (LLM) to determine
the coordinates of a country’s capital city and calculate the distance from Santiago de Chile
to that capital.
Let’s go over the code:
Importing Libraries
import time
from haversine import haversine
from ollama import chat
• MODEL: Specifies the model being used, which is, in this example, the Lhama3.2.
• mylat and mylon: Coordinates of Santiago de Chile, used as the starting point for the
distance calculation.
This is the real Python function that Ollama will be allowed to call. It performs the following
steps: • Takes latitude, longitude, and city name as input arguments. • Uses the haversine
504
library to calculate the distance from Santiago to the target city. • Returns a JSON‑like
dictionary containing the computed distance and a human‑readable text summary.
In Ollama’s terminology, this is a tool — a callable external function that the LLM
may invoke automatically
tools = [
{
"type": "function",
"function": {
"name": "calc_distance",
"description": "Calculates the distance from Santiago, Chile to a \
given city's coordinates.",
"parameters": {
"type": "object",
"properties": {
"lat": {"type": "number", "description": "Latitude of the city"},
"lon": {"type": "number", "description": "Longitude of the city"},
"city": {"type": "string", "description": "Name of the city"}
},
"required": ["lat", "lon", "city"]
}
}
}
]
This JSON object describes the metadata and input schema of the tool so that the LLM knows:
• Name: which function to call. • Description: what purpose it serves. • Parameters: input
argument types and their descriptions. This schema mirrors the OpenAI function‑calling
format and is fully supported in Ollama � 0.4 .
Defining tools this way allows Ollama to validate arguments before sending a call
request back.
response = chat(
model=MODEL,
messages=[{
"role": "user",
"content": f"Find the decimal latitude and longitude of the capital of \
505
I am running a few minutes late; my previous meeting is running over. running a few\
minutes late; my previous meeting is running over.
{country},"
" then use the calc_distance tool to determine how far it is from \
Santiago de Chile."
}],
tools=tools
)
• The chat() function is called with the chosen model, a message, and the tools list.
• The prompt instructs the model first to identify the capital and its coordinates, and then
invoke the tool (calc_distance) with those values.
• Ollama returns a structured response that may include a tool_calls section, indicating
which tool to execute.
city = raw_args['city']
lat = float(raw_args['lat'])
lon = float(raw_args['lon'])
506
distance = haversine((mylat, mylon), (lat, lon), unit='km')
print(f"Santiago de Chile is about {int(round(distance, -1)):,}
kilometers away from {city}.")
In this case:
NOTE: Sometimes the model returns parameter names that differ from what your function
expects. Specifically, Ollama occasionally returns argument objects like:
{"lat1": -33.33, "lon1": -70.51, "lat2": 48.8566, "lon2": 2.3522, "city":
"Paris"}
Instead of the schema-defined keys (lat, lon, city).
This happens because some LLMs (such as Llama 3.2 and Qwen 3) attempt to be “helpful”
by naming coordinates explicitly—lat1/lon1 for origin and lat2/lon2 for destination—even
when the schema only defines lat/lon.
To handle this, optionally a mapping-correction step can be added after decoding the tocall
arguments.
# Convert numbers
args["lat"] = float(args["lat"])
507
args["lon"] = float(args["lon"])
result = calc_distance(**args)
print(result["message"])
This records how long the operation took from prompt submission to tool execution, useful
for benchmarking response performance.
Example Usage
If we enter different countries, for example, France, Colombia, and the United States, We can
note that we always receive the same structured information:
ask_and_measure("France")
ask_and_measure("Colombia")
ask_and_measure("United States")
If you run the code as a script, the result will be printed on the terminal:
508
The complete script can be found at: func_call_dist_calc.py and on the 10-Ollama_Function_Calling
notebook.
The models that will run with the described approach are the ones that can handle tools. For
example, Gemma 3 and 3n will not work.
An alternative is to use the Pydantic library to serialize the schema using model_json_schema().
Using the Pydantic library, models as Gemma can also be used, as explored in the:
20-Ollama_Function_Calling_Pydantic notebook
Adding images
Now it is time to wrap up everything so far! Let’s modify the script using Pydantic so that
instead of entering the country name (as a text), the user enters an image, and the application
(based on SLM) returns the city in the image and its geographic location. With that data, we
can calculate the distance as before.
509
For simplicity, we will implement this new code in two steps. First, the LLM will analyze the
image and create a description (text). This text will be passed on to another instance, where
the model will extract the information needed to pass along.
import time
from haversine import haversine
from ollama import chat
from pydantic import BaseModel, Field
We can see the image if you run the code on the Jupyter Notebook. For that, we also need to
import:
MODEL = 'gemma3:4b'
mylat = -33.33
mylon = -70.51
We can download a new image, for example, Machu Picchu from Wikipedia. On the Notebook
we can see it:
510
[Link](figsize=(8, 8))
[Link](img)
[Link]('off')
#[Link]("Image")
[Link]()
Now, let’s define a function that will receive the image and will return the decimal
latitude and decimal longitude of the city in the image, its name, and what
country it is located
def image_description(img_path):
with open(img_path, 'rb') as file:
response = chat(
model=MODEL,
messages=[
{
'role': 'user',
'content': '''return the decimal latitude and decimal longitude
of the city in the image, its name, and
what country it is located''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,
511
}
)
#print(response['message']['content'])
return response['message']['content']
We can print the entire response for debug purposes. In this case, we can get
something as:
'{\n "city": "Machu Picchu",\n "country": "Peru",\n "lat":
-13.1631,\n "lon": -72.5450\n}\n'
Let’s define a Pydantic model (CityCoord) that describes the expected structure of the SLM’s
response. It expects four fields: country, city (city name), lat (latitude), and lon (longitude).
class CityCoord(BaseModel):
city: str = Field(..., description="Name of the city in the image")
country: str = Field(..., description="Name of the country where the city in the\ imag
lat: float = Field(..., description="Decimal Latitude of the city in the image")
lon: float = Field(..., description="Decimal Longitude of the city in the image")
The image description generated for the function will be passed as a prompt for the model
again.
response = chat(
model=MODEL,
messages=[{
"role": "user",
"content": image_description # image_description from previous model's run
}],
format=CityCoord.model_json_schema(), # Structured JSON format
options={"temperature": 0}
)
resp = CityCoord.model_validate_json([Link])
And so, we can calculate and print the distance, using haversine():
512
about {int(round(distance, -1)):,} kilometers away from Santiago, Chile.\n")
The image shows Machu Picchu, with lat:-13.16 and long: -72.55, located in Peru and about
Enter with the Machu Picchu image full patch as an argument. We will get the same previous
result.
513
The app is working fine with both models, with the Gemma being faster.
How about Paris?
Of course, there are many ways to optimize the code used here. Still, the idea is to explore
the considerable potential of function calling with SLMs at the edge, allowing those models
to integrate with external functions or APIs. Going beyond text generation, SLMs can access
real-time data, automate tasks, and interact with various systems.
In a basic interaction between a user and a language model, the user asks a question, which is
sent to the model as a prompt. The model generates a response based solely on its pre-trained
knowledge.
514
In a RAG process, there’s an additional step between the user’s question and the model’s
response. The user’s question triggers a retrieval process from a knowledge base.
Here are the steps to implement a basic Retrieval Augmented Generation (RAG):
• Determine the type of documents you’ll be using: The best types are documents
from which we can get clean and unobscured text. PDFs can be problematic because
they are designed for printing, not for extracting sensible text. To work with PDFs, we
should get the source document or use tools to handle it.
• Chunk the text: We can’t store the text as one long stream because of context size
limitations and the potential for confusion. Chunking involves splitting the text into
smaller pieces. Chunk text has many ways, such as character count, tokens, words,
paragraphs, or sections. It is also possible to overlap chunks.
515
• Create embeddings: Embeddings are numerical representations of text that capture
semantic meaning. We create embeddings by passing each chunk of text through a
particular embedding model. The model outputs a vector, the length of which depends
on the embedding model used. We should pull one (or more) embedding models from
Ollama, to perform this task. Here are some examples of embedding models available at
Ollama.
Generally, larger embedding sizes capture more nuanced information about the
input. Still, they also require more computational resources to process, and a
higher number of parameters should increase the latency (but also the quality
of the response).
• Store the chunks and embeddings in a vector database: We will need a way to
efficiently find the most relevant chunks of text for a given prompt, which is where a vector
database comes in. We will use Chromadb, an AI-native open-source vector database,
which simplifies building RAGs by creating knowledge, facts, and skills pluggable for
LLMs. Both the embedding and the source text for each chunk are stored.
• Build the prompt: When we have a question, we create an embedding and query the
vector database for the most similar chunks. Then, we select the top few results and
include their text in the prompt.
The goal of RAG is to provide the model with the most relevant information from our docu-
ments, allowing it to generate more accurate and informative responses. So, let’s implement a
simple example of an SLM incorporating a particular set of facts about bees (“Bee Facts”).
Inside the ollama env, enter the command in the terminal for Chromadb instalation:
cd Documents/OLLAMA/
mkdir RAG-simple-bee
516
cd RAG-simple-bee/
import ollama
import chromadb
import time
EMB_MODEL = "nomic-embed-text"
MODEL = 'llama3.2:3B'
Initially, a knowledge base about bee facts should be created. This involves collecting relevant
documents and converting them into vector embeddings. These embeddings are then stored in
a vector database, allowing for efficient similarity searches later. Enter with the “document,”
a base of “bee facts” as a list:
documents = [
"Bee-keeping, also known as apiculture, involves the maintenance of bee \
colonies, typically in hives, by humans.",
"The most commonly kept species of bees is the European honey bee (Apis \
mellifera).",
...
We do not need to “chunk” the document here because we will use each element
of the list as a chunk.
517
Now, we will create our vector embedding database bee_facts and store the document in
it:
client = [Link]()
collection = client.create_collection(name="bee_facts")
Now that we have our “Knowledge Base” created, we can start making queries, retrieving data
from it:
User Query: The process begins when a user asks a question, such as “How many bees are
in a colony? Who lays eggs, and how much? How about common pests and diseases?”
518
prompt = "How many bees are in a colony? Who lays eggs and how much? How about\
common pests and diseases?"
Query Embedding: The user’s question is converted into a vector embedding using the
same embedding model used for the knowledge base.
response = [Link](
prompt=prompt,
model=EMB_MODEL
)
Relevant Document Retrieval: The system searches the knowledge base using the query
embedding to find the most relevant documents (in this case, the 5 more probable). This is done
using a similarity search, which compares the query embedding to the document embeddings
in the database.
results = [Link](
query_embeddings=[response["embedding"]],
n_results=5
)
data = results['documents']
Prompt Augmentation: The retrieved relevant information is combined with the original
user query to create an augmented prompt. This prompt now contains the user’s question and
pertinent facts from the knowledge base.
Answer Generation: The augmented prompt is then fed into a language model, in this case,
the llama3.2:3b model. The model uses this enriched context to generate a comprehensive
answer. Parameters like temperature, top_k, and top_p are set to control the randomness
and quality of the generated response.
output = [Link](
model=MODEL,
prompt=f"Using this data: {data}. Respond to this prompt: {prompt}",
options={
"temperature": 0.0,
"top_k":10,
"top_p":0.5 }
)
519
Response Delivery: Finally, the system returns the generated answer to the user.
print(output['response'])
Based on the provided data, here are the answers to your questions:
results = [Link](
query_embeddings=[response["embedding"]],
n_results=n_results
)
data = results['documents']
520
print(output['response'])
Yes, bees are found in Brazil. According to the data, Brazil has more than 300
different bee species, and indigenous people in Brazil used bees for medicine and
food purposes. Additionally, reports from 1577 mention three native bees used by
indigenous people in Brazil.
[INFO] ==> The code for model: llama3.2:3b, took 22.7s to generate the answer.
By the way, if the model used supports multiple languages, we can use it (for example, Por-
tuguese), even if the dataset was created in English:
Sim, existem abelhas no Brasil! De acordo com o relato de Hans Staden, há três
espécies de abelhas nativas do Brasil que foram mencionadas: mandaçaia (Melipona
quadrifasciata), mandaguari (Scaptotrigona postica) e jataí-amarela (Tetragonisca
angustula). Além disso, o Brasil é conhecido por ter mais de 300 espécies
diferentes de abelhas, a maioria das quais não é agressiva e não põe veneno.
[INFO] ==> The code for model: llama3.2:3b, took 54.6s to generate the answer.
In the Chapter Advancing EdgeAI: Beyond Basic SLMs, we will learn how to
implement a Naive RAG System
521
Conclusion
Throughout this chapter, we’ve explored two fundamental optimization techniques that signifi-
cantly enhance the capabilities of Small Language Models (SLMs) running on edge devices like
the Raspberry Pi: Function Calling and Retrieval-Augmented Generation (RAG).
We began by addressing a critical limitation of language models—their inability to perform
accurate calculations. By implementing function calling, we demonstrated how to transform
SLMs from text generators into actionable agents that can interact with external tools and
APIs. Whether extracting structured data like geographic coordinates, fetching real-time
weather information, or performing precise mathematical operations, function calling bridges
the gap between natural language understanding and deterministic computation. The key
principle remains:
Use SLMs for intent classification and understanding, while delegating specific
tasks to specialized functions that guarantee accuracy.
Resources
• 10-Ollama_Function_Calling notebook
• 20-Ollama_Function_Calling_Pydantic
• 30-Function_Calling_with_images notebook
522
• 40-RAG-simple-bee notebook
• calc_distance_image python script
523
Vision-Language Models at the Edge
We will learn Vison-Language Models across tasks such as captioning, object detection,
grounding, and segmentation on a Raspberry Pi.
Figure 14: DALL·E prompt - A Raspberry Pi setup featuring vision tasks. The image shows
a Raspberry Pi connected to a camera, with various computer vision tasks displayed
visually around it, including object detection, image captioning, segmentation, and
visual grounding. The Raspberry Pi is placed on a desk, with a display showing
bounding boxes and annotations related to these tasks. The background should be a
home workspace, with tools and devices typically used by developers and hobbyists.
In this hands-on lab, we will continuously explore AI applications at the Edge, going from the
basic setup of the Florence-2, Microsoft’s state-of-the-art vision foundation model, to advanced
implementations on devices like the Raspberry Pi.
524
Why Florence-2 at the Edge?
Florence-2 is a vision-language model open-sourced by Microsoft under the MIT license, which
significantly advances vision-language models by combining a lightweight architecture with
robust capabilities. Thanks to its training on the massive FLD-5B dataset, which contains
126 million images and 5.4 billion visual annotations, it achieves performance comparable to
larger models. This makes Florence-2 ideal for deployment at the edge, where power and
computational resources are limited.
In this tutorial, we will explore how to use Florence-2 for real-time computer vision applications,
such as:
• Image captioning
• Object detection
• Segmentation
• Visual grounding
525
• Image Encoder: The image encoder is based on the DaViT (Dual Attention Vision
Transformers) architecture. It converts input images into a series of visual token embed-
dings. These embeddings serve as the foundational representations of the visual content,
capturing both spatial and contextual information about the image.
• Multi-Modal Transformer Encoder-Decoder: Florence-2’s core is the multi-modal
transformer encoder-decoder, which combines visual token embeddings from the image
encoder with textual embeddings generated by a BERT-like model. This combination
526
allows the model to simultaneously process visual and textual inputs, enabling a unified
approach to tasks such as image captioning, object detection, and segmentation.
The model’s training on the extensive FLD-5B dataset ensures it can effectively handle diverse
vision tasks without requiring task-specific modifications. Florence-2 uses textual prompts to
activate specific tasks, making it highly flexible and capable of zero-shot generalization. For
tasks like object detection or visual grounding, the model incorporates additional location
tokens to represent regions within the image, ensuring a precise understanding of spatial
relationships.
Technical Overview
Architecture
527
– Florence-2-Base: 232 million parameters
– Florence-2-Large: 771 million parameters
• Unified Representation: Handles multiple vision tasks through a single architecture
• DaViT Vision Encoder: Converts images into visual token embeddings
• Transformer-based Multi-modal Encoder-Decoder: Processes combined visual
and text embeddings
528
• Automated annotation pipeline using specialist models
• Iterative refinement process for high-quality labels
Key Capabilities
Zero-shot Performance
Fine-tuned Performance
Practical Applications
1. Content Understanding
• Automated image captioning for accessibility
• Visual content moderation
• Media asset management
2. E-commerce
• Product image analysis
• Visual search
• Automated product tagging
3. Healthcare
• Medical image analysis
• Diagnostic assistance
• Research data processing
4. Security & Surveillance
529
• Object detection and tracking
• Anomaly detection
• Scene understanding
Florence-2 stands out from other visual language models due to its impressive zero-shot capa-
bilities. Unlike models like Google PaliGemma, which rely on extensive fine-tuning to adapt
to various tasks, Florence-2 works right out of the box, as we will see in this lab. It can also
compete with larger models like GPT-4V and Flamingo, which often have many more param-
eters but only sometimes match Florence-2’s performance. For example, Florence-2 achieves
better zero-shot results than Kosmos-2 despite having over twice the parameters.
In benchmark tests, Florence-2 has shown remarkable performance in tasks like COCO cap-
tioning and referring expression comprehension. It outperformed models like PolyFormer and
UNINEXT in object detection and segmentation tasks on the COCO dataset. It is a highly
competitive choice for real-world applications where both performance and resource efficiency
are crucial.
Our choice of edge device is the Raspberry Pi 5 (Raspi-5). Its robust platform is equipped
with the Broadcom BCM2712, a 2.4GHz quad-core 64-bit Arm Cortex-A76 CPU featuring
Cryptographic Extension and enhanced caching capabilities. It boasts a VideoCore VII GPU,
dual 4Kp60 HDMI® outputs with HDR, and a 4Kp60 HEVC decoder. Memory options include
4GB and 8GB of high-speed LPDDR4X SDRAM, with 8GB being our choice to run Florence-2.
It also features expandable storage via a microSD card slot and a PCIe 2.0 interface for fast
peripherals such as M.2 SSDs (Solid State Drives).
We suggest installing an Active Cooler, a dedicated clip-on cooling solution for Raspberry Pi
5 (Raspi-5), for this lab. It combines an aluminum heatsink with a temperature-controlled
blower fan to keep the Raspi-5 operating comfortably under heavy loads, such as running
Florense-2.
530
Environment configuration
1. Transformers:
• Florence-2 uses the transformers library from Hugging Face for model loading and
inference. This library provides the architecture for working with pre-trained vision-
language models, making it easy to perform tasks like image captioning, object
detection, and more. Essentially, transformers helps in interacting with the model,
processing input prompts, and obtaining outputs.
2. PyTorch:
• PyTorch is a deep learning framework that provides the infrastructure needed to
run the Florence-2 model, which includes tensor operations, GPU acceleration (if
a GPU is available), and model training/inference functionalities. The Florence-2
model is trained in PyTorch, and we need it to leverage its functions, layers, and
computation capabilities to perform inferences on the Raspberry Pi.
3. Timm (PyTorch Image Models):
531
• Florence-2 uses timm to access efficient implementations of vision models and pre-
trained weights. Specifically, the timm library is utilized for the image encoder
part of Florence-2, particularly for managing the DaViT architecture. It provides
model definitions and optimized code for common vision tasks and allows the easy
integration of different backbones that are lightweight and suitable for edge devices.
4. Einops:
• Einops is a library for flexible and powerful tensor operations. It makes it easy to
reshape and manipulate tensor dimensions, which is especially important for the
multi-modal processing done in Florence-2. Vision-language models like Florence-
2 often need to rearrange image data, text embeddings, and visual embeddings
to align correctly for the transformer blocks, and einops simplifies these complex
operations, making the code more readable and concise.
• Transformers and PyTorch are needed to load the model and run the inference.
• Timm is used to access and efficiently implement the vision encoder.
• Einops helps reshape data, facilitating the integration of visual and text features.
All these components work together to help Florence-2 run seamlessly on our Raspberry Pi,
allowing it to perform complex vision-language tasks relatively quickly.
Considering that the Raspberry Pi already has its OS installed, let’s use SSH to reach it from
another computer:
ssh mjrovai@[Link]
hostname -I
[Link]
532
Updating the Raspberry Pi
First, ensure your Raspberry Pi is up to date:
Install Dependencies
Let’s set up and activate a Virtual Environment for working with Florence-2:
Install PyTorch
533
pip3 install setuptools numpy Cython
pip3 install requests
pip3 install torch torchvision --index-url [Link]
pip3 install torchaudio --index-url [Link]
534
Running the above command on the SSH terminal, we can see the local URL address to open
the notebook:
The notebook with the code used on this initial test can be found on the Lab GitHub:
• 10-florence2_test.ipynb
We can access it on the remote computer by entering the Raspberry Pi’s IP address and the
provided token in a web browser ( copy the entire URL from the terminal).
From the Home page, create a new notebook [Python 3 (ipykernel) ] and copy and paste
the example code from Hugging Face Hub.
The code is designed to run Florence-2 on a given image to perform object detection. It
loads the model, processes an image and a prompt, and then generates a response to identify
and describe the objects in the image.
535
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("microsoft/Florence-2-base",
torch_dtype=torch_dtype,
trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base",
trust_remote_code=True)
prompt = "<OD>"
url = "[Link]
images/resolve/main/transformers/tasks/[Link]?download=true"
image = [Link]([Link](url, stream=True).raw)
generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
do_sample=False,
num_beams=3,
)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=False)[0]
print(parsed_answer)
536
1. Importing Required Libraries
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForCausalLM
• requests: Used to make HTTP requests. In this case, it downloads an image from a
URL.
• PIL (Pillow): Provides tools for manipulating images. Here, it’s used to open the
downloaded image.
• torch: PyTorch is imported to handle tensor operations and determine the hardware
availability (CPU or GPU).
• transformers: This module provides easy access to Florence-2 by using AutoProcessor
and AutoModelForCausalLM to load pre-trained models and process inputs.
model = AutoModelForCausalLM.from_pretrained("microsoft/Florence-2-base",
torch_dtype=torch_dtype,
trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base",
trust_remote_code=True)
• Model Initialization:
537
– AutoModelForCausalLM.from_pretrained() loads the pre-trained Florence-2
model from Microsoft’s repository on Hugging Face. The torch_dtype is set
according to the available hardware (GPU/CPU), and trust_remote_code=True
allows the use of any custom code that might be provided with the model.
– .to(device) moves the model to the appropriate device (either CPU or GPU). In
our case, it will be set to CPU.
• Processor Initialization:
– AutoProcessor.from_pretrained() loads the processor for Florence-2. The pro-
cessor is responsible for transforming text and image inputs into a format the model
can work with (e.g., encoding text, normalizing images, etc.).
prompt = "<OD>"
• Prompt Definition: The string "<OD>" is used as a prompt. This refers to “Object
Detection”, instructing the model to detect objects on the image.
url = "[Link]
images/resolve/main/transformers/tasks/[Link]?download=true"
image = [Link]([Link](url, stream=True).raw)
• Downloading the Image: The [Link]() function fetches the image from the
specified URL. The stream=True parameter ensures the image is streamed rather than
downloaded completely at once.
• Opening the Image: [Link]() opens the image so the model can process it.
6. Processing Inputs
• Processing Input Data: The processor() function processes the text (prompt) and
the image (image). The return_tensors="pt" argument converts the processed data
into PyTorch tensors, which are necessary for inputting data into the model.
538
• Moving Inputs to Device: .to(device, torch_dtype) moves the inputs to the
correct device (CPU or GPU) and assigns the appropriate data type.
generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
do_sample=False,
num_beams=3,
)
539
• Post-Processing: processor.post_process_generation() is called to process the
generated text further, interpreting it based on the task ("<OD>" for object detection)
and the size of the image.
• This function extracts specific information from the generated text, such as bounding
boxes for detected objects, making the output more useful for visual tasks.
print(parsed_answer)
• Finally, print(parsed_answer) displays the output, which could include object detec-
tion results, such as bounding box coordinates and labels for the detected objects in the
image.
Result
540
By the Object Detection result, we can see that:
It seems that at least a few objects were detected. we can also implement a code to draw the
bounding boxes in the find objects:
541
# Unpack the bounding box coordinates
x1, y1, x2, y2 = bbox
# Create a Rectangle patch
rect = [Link]((x1, y1), x2-x1, y2-y1, linewidth=1,
edgecolor='r', facecolor='none')
# Add the rectangle to the Axes
ax.add_patch(rect)
# Annotate the label
[Link](x1, y1, label, color='white', fontsize=8,
bbox=dict(facecolor='red', alpha=0.5))
Box (x0, y0, x1, y1): Location tokens correspond to the top-left and bottom-
right corners of a box.
And running
plot_bbox(image, parsed_answer['<OD>'])
We get:
542
Florence-2 Tasks
543
1. Object Detection (OD)
• Prompt: "<OD>"
• Description: Identifies objects in an image and provides bounding boxes for each de-
tected object. This task is helpful for applications like visual inspection, surveillance,
and general object recognition.
2. Image Captioning
• Prompt: "<CAPTION>"
• Description: Generates a textual description for an input image. This task helps the
model describe what is happening in the image, providing a human-readable caption for
content understanding.
3. Detailed Captioning
• Prompt: "<DETAILED_CAPTION>"
• Description: Generates a more detailed caption with more nuanced information about
the scene, such as the objects present and their relationships.
4. Visual Grounding
• Prompt: "<CAPTION_TO_PHRASE_GROUNDING>"
• Description: Links a textual description to specific regions in an image. For example,
given a prompt like “a green car,” the model highlights where the red car is in the image.
This is useful for human-computer interaction, where you must find specific objects based
on text.
5. Segmentation
• Prompt: "<REFERRING_EXPRESSION_SEGMENTATION>"
• Description: Performs segmentation based on a referring expression, such as “the blue
cup.” The model identifies and segments the specific region containing the object men-
tioned in the prompt (all related pixels).
• Prompt: "<DENSE_REGION_CAPTION>"
• Description: Provides captions for multiple regions within an image, offering a detailed
breakdown of all visible areas, including different objects and their relationships.
544
7. OCR with Region
• Prompt: "<OCR_WITH_REGION>"
• Description: Performs Optical Character Recognition (OCR) on an image and provides
bounding boxes for the detected text. This is useful for extracting and locating textual
information in images, such as reading signs, labels, or other forms of text in images.
• Prompt: "<OPEN_VOCABULARY_OD>"
• Description: The model can detect objects without being restricted to a predefined list
of classes, making it helpful in recognizing a broader range of items based on general
visual understanding.
• 20-florence_2.ipynb
Let’s use a couple of images created by Dall-E and upload them to the Rasp-5 (FileZilla can
be used for that). The images will be saved on a sub-folder named images :
dogs_cats = [Link]('./images/[Link]')
table = [Link]('./images/[Link]')
545
Let’s create a function to facilitate our exploration and to keep track of the latency of the
model for different tasks:
546
image_size=([Link], [Link])
)
return parsed_answer
Caption
run_example(task_prompt='<CAPTION>',image=dogs_cats)
2. Table
run_example(task_prompt='<CAPTION>',image=table)
{'<CAPTION>': 'A wooden table topped with a plate of fruit and a glass of wine.'}
DETAILED_CAPTION
run_example(task_prompt='<DETAILED_CAPTION>',image=dogs_cats)
{'<DETAILED_CAPTION>': 'The image shows a group of cats and dogs sitting on top of a
lush green field, surrounded by plants with flowers, trees, and a house in the
547
background. The sky is visible above them, creating a peaceful atmosphere.'}
2. Table
run_example(task_prompt='<DETAILED_CAPTION>',image=table)
{'<DETAILED_CAPTION>': 'The image shows a wooden table with a bottle of wine and a
glass of wine on it, surrounded by a variety of fruits such as apples, oranges, and
grapes. In the background, there are chairs, plants, trees, and a house, all slightly
blurred.'}
MORE_DETAILED_CAPTION
run_example(task_prompt='<MORE_DETAILED_CAPTION>',image=dogs_cats)
{'<MORE_DETAILED_CAPTION>': 'The image shows a group of four cats and a dog in a garden.
The garden is filled with colorful flowers and plants, and there is a pathway leading up
to a house in the background. The main focus of the image is a large German Shepherd dog
standing on the left side of the garden, with its tongue hanging out and its mouth open,
as if it is panting or panting. On the right side, there are two smaller cats, one orange
and one gray, sitting on the grass. In the background, there is another golden retriever
dog sitting and looking at the camera. The sky is blue and the sun is shining, creating a
warm and inviting atmosphere.'}
2. Table
run_example(task_prompt='< MORE_DETAILED_CAPTION>',image=table)
{'<MORE_DETAILED_CAPTION>': 'The image shows a wooden table with a wooden tray on it. On
the tray, there are various fruits such as grapes, oranges, apples, and grapes. There is
also a bottle of red wine on the table. The background shows a garden with trees and a
548
house. The overall mood of the image is peaceful and serene.'}
We can note that the more detailed the caption task, the longer the latency and
the possibility of mistakes (like “The image shows a group of four cats and a dog
in a garden”, instead of two dogs and three cats).
OD - Object Detection
We can run the same previous function for object detection using the prompt <OD>.
task_prompt = '<OD>'
results = run_example(task_prompt,image=dogs_cats)
print(results)
Only by the labels ['cat,' 'cat,' 'cat,' 'dog,' 'dog'] is it possible to see that the main
objects in the image were captured. Let’s apply the function used before to draw the bounding
boxes:
plot_bbox(dogs_cats, results['<OD>'])
549
Let’s also do it with the Table image:
task_prompt = '<OD>'
results = run_example(task_prompt,image=table)
plot_bbox(table, results['<OD>'])
550
DENSE_REGION_CAPTION
It is possible to mix the classic Object Detection with the Caption task in specific sub-regions
of the image:
task_prompt = '<DENSE_REGION_CAPTION>'
results = run_example(task_prompt,image=dogs_cats)
plot_bbox(dogs_cats, results['<DENSE_REGION_CAPTION>'])
results = run_example(task_prompt,image=table)
plot_bbox(table, results['<DENSE_REGION_CAPTION>'])
551
CAPTION_TO_PHRASE_GROUNDING
With this task, we can enter with a caption, such as “a wine glass”, “a wine bottle,” or “a half
orange,” and Florence-2 will localize the object in the image:
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
552
[INFO] ==> Florence-2-base (<CAPTION_TO_PHRASE_GROUNDING>), took 15.7 seconds to execute
each task.
Cascade Tasks
We can also enter the image caption as the input text to push Florence-2 to find more objects:
task_prompt = '<CAPTION>'
results = run_example(task_prompt,image=dogs_cats)
text_input = results[task_prompt]
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(task_prompt, text_input,image=dogs_cats)
plot_bbox(dogs_cats, results['<CAPTION_TO_PHRASE_GROUNDING>'])
553
OPEN_VOCABULARY_DETECTION
task_prompt = '<OPEN_VOCABULARY_DETECTION>'
text = ["a house", "a tree", "a standing cat at the left",
"a sleeping cat on the ground", "a standing cat at the right",
"a yellow cat"]
for txt in text:
results = run_example(task_prompt, text_input=txt,image=dogs_cats)
bbox_results = convert_to_od_format(results['<OPEN_VOCABULARY_DETECTION>'])
plot_bbox(dogs_cats, bbox_results)
554
[INFO] ==> Florence-2-base (<OPEN_VOCABULARY_DETECTION>), took 15.1 seconds
to execute each task.
Note: Trying to use Florence-2 to find objects that were not found can leads to
mistakes (see exaamples on the Notebook).
We can also segment a specific object in the image and give its description (caption), such as
“a wine bottle” on the table image or “a German Sheppard” on the dogs_cats.
Referring expression segmentation results format: {'<REFERRING_EXPRESSION_SEGMENTATION>':
{'Polygons': [[[polygon]], ...], 'labels': ['', '', ...]}}, one object is repre-
sented by a list of polygons. each polygon is [x1, y1, x2, y2, ..., xn, yn].
Polygon (x1, y1, …, xn, yn): Location tokens represent the vertices of a polygon
in clockwise order.
555
So, let’s first create a function to plot the segmentation:
Parameters:
- image_path: Path to the image file.
- prediction: Dictionary containing 'polygons' and 'labels' keys.
'polygons' is a list of lists, each containing vertices
of a polygon.
'labels' is a list of labels corresponding to each polygon.
- fill_mask: Boolean indicating whether to fill the polygons with color.
"""
# Load the image
draw = [Link](image)
556
# Draw the polygon
if fill_mask:
[Link](_polygon, outline=color, fill=fill_color)
else:
[Link](_polygon, outline=color)
task_prompt = '<REFERRING_EXPRESSION_SEGMENTATION>'
557
[INFO] ==> Florence-2-base (<REFERRING_EXPRESSION_SEGMENTATION>),
took 207.0 seconds to execute each task.
Region to Segmentation
With this task, it is also possible to give the object coordinates in the image to segment it.
The input format is '<loc_x1><loc_y1><loc_x2><loc_y2>', [x1, y1, x2, y2] , which is
the quantized coordinates in [0, 999].
For example, when running the code:
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(task_prompt, text_input="a half orange",image=table)
results
558
task_prompt = '<REGION_TO_SEGMENTATION>'
results = run_example(task_prompt,
text_input="<loc_343><loc_690><loc_531><loc_874>",
image=table)
output_image = [Link](table)
draw_polygons(output_image, results['<REGION_TO_SEGMENTATION>'], fill_mask=True)
Region to Texts
We can also give the region (coordinates and ask for a caption):
559
task_prompt = '<REGION_TO_CATEGORY>'
results = run_example(task_prompt, text_input="<loc_343><loc_690><loc_531>
<loc_874>",image=table)
results
{'<REGION_TO_CATEGORY>': 'orange<loc_343><loc_690><loc_531><loc_874>'}
The model identified an orange in that region. Let’s ask for a description:
task_prompt = '<REGION_TO_DESCRIPTION>'
results = run_example(task_prompt, text_input="<loc_343><loc_690><loc_531>
<loc_874>",image=table)
results
{'<REGION_TO_CATEGORY>': 'orange<loc_343><loc_690><loc_531><loc_874>'}
In this case, the description did not provide more details, but it could. Try another example.
OCR
With Florence-2, we can perform Optical Character Recognition (OCR) on an image, getting
what is written on it (task_prompt = '<OCR>' and also get the bounding boxes (location) for
the detected text (ask_prompt = '<OCR_WITH_REGION>'). Those tasks can help extract and
locate textual information in images, such as reading signs, labels, or other forms of text in
images.
Let’s upload a flyer from a talk in Brazil to Raspi. Let’s test works in another language, here
Portuguese):
flayer = [Link]('./images/[Link]')
# Display the image
[Link](figsize=(8, 8))
[Link](flayer)
[Link]('off')
#[Link]("Image")
[Link]()
560
Let’s examine the image with '<MORE_DETAILED_CAPTION>' :
The description is very accurate. Let’s get to the more important words with the task OCR:
task_prompt = '<OCR>'
run_example(task_prompt,image=flayer)
561
[INFO] ==> Florence-2-base (<OCR>), took 37.7 seconds to execute.
task_prompt = '<OCR_WITH_REGION>'
results = run_example(task_prompt,image=flayer)
Let’s also create a function to draw bounding boxes around the detected words:
fill=color)
display(image)
output_image = [Link](flayer)
draw_ocr_bboxes(output_image, results['<OCR_WITH_REGION>'])
562
We can inspect the detected words:
results['<OCR_WITH_REGION>']['labels']
'</s>Machine Learning',
'Café',
'com',
'Embarcado',
'Embarcados',
'Democratizando a Inteligência',
'Artificial para Paises em',
'25 de Setembro ás 17h',
'Desenvolvimento',
'Toda quarta-feira',
'Marcelo Roval',
'Professor na UNIFIEI e',
'Transmissão via',
'in',
'Co-Director do TinyML4D']
563
Latency Summary
The latency observed for different tasks using Florence-2 on the Raspberry Pi (Raspi-5) varied
depending on the complexity of the task:
These latency times highlight the resource constraints of edge devices like the Raspberry Pi
and emphasize the need to optimize the model and the environment to achieve real-time
performance.
564
Running complex tasks can use all 8GB of the Raspi-5’s memory. For example,
the above screenshot during the Florence OD task shows 4 CPUs at full speed and
over 5GB of memory in use. Consider increasing the SWAP memory to 2 GB.
Checking the CPU temperature with vcgencmd measure_temp , showed that temperature can
go up to +80oC.
Fine-Tunning
As explored in this lab, Florence supports many tasks out of the box, including caption-
ing, object detection, OCR, and more. However, like other pre-trained foundational models,
Florence-2 may need domain-specific knowledge. For example, it may need to improve with
medical or satellite imagery. In such cases, fine-tuning with a custom dataset is necessary.
The Roboflow tutorial, How to Fine-tune Florence-2 for Object Detection Tasks, shows how
to fine-tune Florence-2 on object detection datasets to improve model performance for our
specific use case.
Based on the above tutorial, it is possible to fine-tune the Florence-2 model to detect boxes
and wheels used in previous labs:
It is important to note that after fine-tuning, the model can still detect classes that don’t
belong to our custom dataset, like cats, dogs, grapes, etc, as seen before).
The complete fine-tunning project using a previously annotated dataset in Roboflow and exe-
cuted on CoLab can be found in the notebook:
• 30-Finetune_florence_2_on_detection_dataset_box_vs_wheel.ipynb
565
report that Florence 2 can perform visual question answering (VQA), but the released models
don’t include VQA capability.
Conclusion
Florence-2 offers a versatile and powerful approach to vision-language tasks at the edge, pro-
viding performance that rivals larger, task-specific models, such as YOLO for object detection,
BERT/RoBERTa for text analysis, and specialized OCR models.
Thanks to its multi-modal transformer architecture, Florence-2 is more flexible than YOLO in
terms of the tasks it can handle. These include object detection, image captioning, and visual
grounding.
Unlike BERT, which focuses purely on language, Florence-2 integrates vision and language,
allowing it to excel in applications that require both modalities, such as image captioning and
visual grounding.
Moreover, while traditional OCR models such as Tesseract and EasyOCR are designed solely
for recognizing and extracting text from images, Florence-2’s OCR capabilities are part of a
broader framework that includes contextual understanding and visual-text alignment. This
makes it particularly useful for scenarios that require both reading text and interpreting its
context within images.
Overall, Florence-2 stands out for its ability to seamlessly integrate various vision-language
tasks into a unified model that is efficient enough to run on edge devices like the Raspberry
Pi. This makes it a compelling choice for developers and researchers exploring AI applications
at the edge.
1. Unified Architecture
• Single model handles multiple vision tasks vs. specialized models (YOLO, BERT,
Tesseract)
• Eliminates the need for multiple model deployments and integrations
• Consistent API and interface across tasks
2. Performance Comparison
• Object Detection: Comparable to YOLOv8 (~37.5 mAP on COCO vs. YOLOv8’s
~39.7 mAP) despite being general-purpose
• Text Recognition: Handles multiple languages effectively like specialized OCR mod-
els (Tesseract, EasyOCR)
566
• Language Understanding: Integrates BERT-like capabilities for text processing
while adding visual context
3. Resource Efficiency
• The Base model (232M parameters) achieves strong results despite smaller size
• Runs effectively on edge devices (Raspberry Pi)
• Single model deployment vs. multiple specialized models
Trade-offs
1. Resource-Constrained Environments
• Edge devices requiring multiple vision capabilities
• Systems with limited storage/deployment capacity
• Applications needing flexible vision processing
2. Multi-modal Applications
• Content moderation systems
• Accessibility tools
• Document analysis workflows
3. Rapid Prototyping
• Quick deployment of vision capabilities
• Testing multiple vision tasks without separate models
• Proof-of-concept development
567
Future Implications
Florence-2 represents a shift toward unified vision models that could eventually replace task-
specific architectures in many applications. While specialized models maintain advantages in
specific scenarios, the convenience and efficiency of unified models like Florence-2 make them
increasingly attractive for real-world deployments.
The lab demonstrates Florence-2’s viability on edge devices, suggesting future IoT, mobile
computing, and embedded systems applications where deploying multiple specialized models
would be impractical.
Resources
• 10-florence2_test.ipynb
• 20-florence_2.ipynb
• 30-Finetune_florence_2_on_detection_dataset_box_vs_wheel.ipynb
568
Audio and Vision AI Pipeline
Introduction
In this chapter, we extend our SLM and SVL capabilities by creating a complete audio pro-
cessing pipeline that transforms voice or image input into intelligent vocal responses. We will
learn to integrate Speech-to-Text (STT), Small Language (or Visual) Models, and Text-to-
Speech (TTS) technologies to build conversational AI systems that run entirely on Raspberry
Pi hardware.
This chapter bridges the gap between our existing computer vision knowledge and multimodal
AI applications, demonstrating how different AI components work together in real-world edge
deployments.
569
The Audio to Audio AI Pipeline Architecture
We will understand how to architect and implement multimodal AI systems by building a com-
plete voice interaction pipeline. The goal is to gain practical experience with audio processing
on edge devices while learning to efficiently integrate multiple AI models within resource con-
straints. Additionally, you will develop troubleshooting skills for complex AI pipelines and
understand the engineering trade-offs involved in edge audio processing.
When we built computer vision systems earlier in the course, we processed visual data to extract
meaningful information. Audio AI systems follow a similar principle but work with temporal
audio signals instead of static images. The key insight is that speech processing requires
multiple specialized models working together, rather than a single end-to-end system.
Modern small models, such as Gemma 3n, can process audio directly and its prompt.
Today (September 2025), Gemma 3n can transcribe text from audio files using
Hugging Face Transformers, but it is not available with Ollama
Consider how humans process spoken language. We simultaneously parse the acoustic signal,
understand the linguistic content, reason about the meaning, and formulate responses. Our
AI pipeline mimics this process by breaking it into distinct, manageable components.
570
[Microphone] → [STT Model] → [SLM] → [TTS Model] → [Speaker]
Audio Text Text Audio
Audio captured by the microphone is processed through a Speech-to-Text model, which con-
verts sound waves into text transcriptions. This text becomes input for our Small Language
Model, which generates intelligent responses. Finally, a Text-to-Speech system converts the
written response back into spoken audio.
Each component has specific requirements and limitations. The STT model must handle
various accents and noise conditions. The SLM needs sufficient context to generate coherent
responses. The TTS system must produce speech that sounds natural. Understanding these
individual requirements helps us optimize the overall system performance.
Edge AI Considerations
Begin by identifying the audio capabilities of our system. The Raspberry Pi can work with
various audio input and output devices, but proper configuration is essential for reliable oper-
ation.
Use the command arecord -l to list available recording devices. You should see output
showing your microphone’s card and device numbers. For USB microphones, this typically
appears as something like card 2: Microphone [USB Condenser Microphone], device 0:
USB Audio [USB Audio]. The critical information is the card number and device number,
which you’ll reference as hw:2,0 in the code.
571
Testing Basic Audio Functionality
Before writing Python code, we should verify that our audio setup works correctly at the
system level. Let’s record a short test file using:
arecord --device="plughw:2,0" --format=S16_LE --rate=16000 -c2 [Link]
572
Play back the recording with aplay [Link] to confirm that both capture and playback
work correctly (use [CTRL]+[C] to stop the recording or add a duration in seconds to the
command line).
The .WAV file can be played on another device (such as a computer) or on the
Raspberry Pi, as a speaker can be connected via USB or Bluetooth.
source ~/ollama/bin/activate
The PyAudio library requires system-level audio libraries; therefore, install them using sudo
apt-get.
Let’s create a working directory: Documents/OLLAMA/SST and verify the USB device index,
with the below script (verify_usb_index.py:
import pyaudio
p = [Link]()
for ii in range(p.get_device_count()):
print(ii, p.get_device_info_by_index(ii).get('name'))
573
As a result, we should get:
A lot of messages should appear. They are mostly ALSA and JACK warnings
about missing or undefined virtual/surround sound devices—they are common on
Raspberry Pi systems with minimal or headless sound configs and typically do
not impact basic USB microphone capture. If our USB Microphone appears as a
recording device (as it does: “hw:2,0”), we can safely ignore most of these unless
audio capture fails.
import pyaudio
import wave
FORMAT = pyaudio.paInt16
CHANNELS = 1
RATE = 16000 # 16 kHz
CHUNK = 1024
RECORD_SECONDS = 10
DEVICE_INDEX = 2 # replace this with your detected USB mic's index
WAVE_OUTPUT_FILENAME = "[Link]"
574
audio = [Link]()
print("Recording...")
frames = []
print("Finished recording.")
stream.stop_stream()
[Link]()
[Link]()
wf = [Link](WAVE_OUTPUT_FILENAME, 'wb')
[Link](CHANNELS)
[Link](audio.get_sample_size(FORMAT))
[Link](RATE)
[Link](b''.join(frames))
[Link]()
Understanding the audio configuration parameters helps prevent common problems. We use
16-bit PCM format (pyaudio.paInt16) because it provides good quality while remaining
computationally efficient. The 16kHz sampling rate balances audio quality with processing
requirements - most speech recognition models expect this rate.
The buffer size (CHUNK = 1024) affects latency and reliability. Smaller buffers reduce latency
but may cause audio dropouts on busy systems. Larger buffers increase latency but provide
more stable recording.
Let’s Playback to verify if we get it correctly:
575
Your browser does not support the audio element.
Traditional speech recognition systems, such as OpenAI’s Whisper, are highly accurate but re-
quire substantial computational resources. Moonshine is specifically designed for edge devices,
using optimized model architectures and quantization techniques to achieve good performance
on resource-constrained hardware.
The ONNX (Open Neural Network Exchange) version of Moonshine provides additional op-
timization benefits. ONNX Runtime includes hardware-specific optimizations that can sig-
nificantly improve inference speed on ARM processors, such as those found in Raspberry Pi
devices.
Moonshine offers different model sizes with clear trade-offs between accuracy and computa-
tional requirements. The “tiny” model processes audio quickly but may struggle with difficult
audio conditions. The “base” model provides better accuracy but requires more processing
time and memory.
For initial development, we should start with the tiny model to ensure that the pipeline works
correctly. Once the complete system is functional, we can experiment with larger models to
find the optimal balance for our specific use case and hardware capabilities.
576
Implementation and Preprocessing
This specific installation method ensures compatibility with the ONNX runtime optimiza-
tions.
Let’s run the test script below (transcription_test.py):
import moonshine_onnx
text = moonshine_onnx.transcribe('[Link]', 'moonshine/tiny')
print(text[0])
As a result, we will get the corresponding text, which was recorded before:
The text output from your speech recognition system becomes input for your Small Language
Model. However, the characteristics of spoken language differ significantly from written text,
especially when filtered through speech recognition systems.
Spoken language tends to be more informal, may contain false starts and repetitions, and might
include transcription errors. Our SLM integration should account for these characteristics. On
577
a final implementation, we should consider preprocessing the STT output to clean up obvious
transcription errors or providing context to the SLM about the voice interaction nature of the
input.
Voice-based interactions have different expectations than text-based chats. Responses should
be concise since users must listen to the entire output. Avoid complex formatting or long lists
that work well in text but become cumbersome when spoken aloud.
We should design our system prompts to encourage responses appropriate for voice interaction.
For example, “Provide a brief, conversational response suitable for speaking aloud” can help
guide the SLM toward more appropriate output formatting.
Unlike single-query text interactions, voice conversations often involve multiple exchanges. Im-
plementing conversation context memory significantly enhances the user experience. However,
context management on edge devices requires careful consideration of memory usage.
Consider implementing a sliding window approach, where you maintain the last few exchanges
in memory but discard older context to prevent memory exhaustion, balancing context length
with available system resources.
Let’s create a function to handle this. For test, run slm_test.py:
import ollama
578
"""
response = [Link](
model=model,
prompt=full_prompt
)
return response['response']
Text-to-speech systems face different challenges than speech recognition systems. While STT
must handle various input conditions, TTS must generate consistent, natural-sounding output
across diverse text inputs. The quality of TTS has a significant impact on the user experience
in voice interaction systems.
PIPER provides an excellent balance between voice quality and computational efficiency for
edge deployments. Unlike cloud-based TTS services, PIPER runs entirely locally, ensuring
privacy and eliminating network dependencies.
579
Voice Model Selection and Installation
PIPER offers various voice models with different characteristics. The “low”, “medium”, and
“high” quality designations primarily refer to model size and computational requirements rather
than dramatic quality differences. For most applications, the low-quality models provide ac-
ceptable voice output while running efficiently on Raspberry Pi hardware.
Install PIPER with pip install piper-tts, create a voices directory:
Download voice models from the Hugging Face repository. Each voice model requires both the
model file (.onnx) and a configuration file (.json). The configuration file contains model-specific
parameters essential for generating proper audio.
We should download both files for our chosen voice; for example, the English female “lessac”
voice provides clear, natural speech suitable for most applications.
import subprocess
import os
Args:
text (str): Text to convert to speech
output_file (str): Output WAV file path
"""
# Path to your voice model
model_path = "voices/en_US-[Link]"
580
print(f"Error: Model file not found at {model_path}")
return False
try:
# Run PIPER command
process = [Link](
['piper', '--model', model_path, '--output_file', output_file],
stdin=[Link],
stdout=[Link],
stderr=[Link],
text=True
)
if [Link] == 0:
print(f"\nSpeech generated successfully: {output_file}")
return True
else:
print(f"Error: {stderr}")
return False
except Exception as e:
print(f"Error running PIPER: {e}")
return False
if text_to_speech_piper(txt):
print("You can now play the file with: aplay piper_output.wav")
else:
print("Failed to generate speech")
Runing the script, a piper_output.wav file will be generated, which is the text converted into
speech.
Your browser does not support the audio element.
To listen to the sound, we can run: aplay piper_output.wav.
581
Handling Long Text and Special Cases
TTS systems may struggle with very long input text or special characters. Implement text
preprocessing to handle these cases gracefully. Break long responses into shorter segments,
handle abbreviations and numbers appropriately, and filter out problematic characters that
might cause TTS failures.
Consider implementing text chunking for responses longer than a reasonable speaking length.
This prevents both TTS processing issues and user fatigue from overly long audio responses.
Integrating all components requires careful attention to error handling and resource manage-
ment. Each stage of the pipeline can fail independently, and robust systems must handle these
failures gracefully rather than crashing.
Design our integration with modularity in mind. Test each component independently before
combining them. This approach simplifies debugging and allows you to optimize individual
components separately.
We should also implement proper logging throughout our pipeline. When complex systems
fail, detailed logs help identify whether the issue occurs in audio capture, speech recognition,
language model processing, text-to-speech conversion, or audio playback.
582
Performance Optimization Strategies
Measure the timing of each pipeline component to identify bottlenecks. Typically, the SLM
inference takes the longest time, followed by TTS generation. Understanding these timing
characteristics helps prioritize optimization efforts.
Consider implementing concurrent processing where possible. For example, you might begin
TTS processing for the first part of an extended response while the SLM is still generating the
remainder. However, be cautious about memory usage when implementing parallel processing
on resource-constrained devices.
Edge devices have limited RAM, and loading multiple large models simultaneously can cause
memory pressure. Implement strategies to manage memory efficiently, such as loading models
only when needed or using model swapping for infrequently used components.
Monitor system memory usage during operation and implement safeguards to prevent mem-
ory exhaustion. Consider implementing graceful degradation where your system switches to
smaller, more efficient models if memory becomes constrained.
Production-quality voice interaction systems must handle various failure modes gracefully.
Network interruptions, hardware disconnections, model loading failures, and unexpected input
conditions should not cause your system to crash.
Implement comprehensive error handling at each pipeline stage. When speech recognition
produces empty output, provide the user with meaningful feedback rather than processing
empty strings. When TTS fails, consider falling back to text display or simplified audio
feedback.
Design user feedback mechanisms that work within your voice interaction paradigm. Audio
beeps, LED indicators, or simple voice messages can communicate system status without
requiring visual displays.
Multi-stage systems present unique debugging challenges. When the overall system fails, iden-
tifying the specific failure point requires systematic testing approaches.
583
Implement test modes that allow you to inject known inputs at each pipeline stage. This
capability enables you to isolate problems to specific components rather than repeatedly testing
the entire system.
Create diagnostic outputs that help understand system behavior. For example, displaying
transcription confidence scores, SLM response times, or TTS processing status helps identify
performance issues or quality problems.
Considering the previous points, let’s assemble all the essential components that work together:
audio capture, transcription, language model processing, and text-to-speech. We should com-
bine these into a complete voice pipeline that flows naturally from one step to the next.
The key insight here is that each of our developed scripts represents a stage in what’s called
an “audio processing pipeline”. Let’s walk through how we can connect these pieces.
Understanding the Pipeline Architecture
The pipeline follows a logical sequence: the voice becomes audio data, that audio is converted
into text, the text is processed into an AI response, the response is converted into speech audio,
and finally, that speech audio is transformed into sound that we can hear.
We should have a run_voice_pipeline() function in addition to the previous ones that acts
as a coordinator, ensuring each step completes successfully before proceeding to the next. If
any step fails, the entire pipeline stops gracefully rather than trying to continue with missing
data.
Key Integration Points
We should connect the scripts by ensuring the output of each function becomes the input
for the following function. For example, record_audio() creates “user_input.wav”, which
transcribe_audio() reads to produce text, which generate_response() processes to de-
velop an AI response, and so on.
The error handling at each step ensures that if our microphone isn’t working, or if the AI
model is busy, or if the voice model files are missing, we get clear feedback about what went
wrong rather than mysterious crashes.
Optimizations for Raspberry Pi
The Raspberry Pi has limited resources compared to a desktop computer, so we should include
several optimizations. The cleanup_temp_files() function prevents our storage from filling
up with temporary audio files. The audio configuration uses 16kHz sampling (which matches
Moonshine’s expectations) rather than CD-quality 44kHz, reducing processing overhead.
584
The continuous assistant mode includes a manual trigger (pressing Enter) rather than voice
activation detection, which saves CPU cycles that would otherwise be spent constantly mon-
itoring audio input.
585
We can start by testing individual components using detect_audio_devices() first to confirm
our USB microphone is still at index 2. Then, we can run run_voice_pipeline() with a
simple question to verify the complete flow works.
Once we are confident in single interactions, we can use continuous_voice_assistant()
for extended conversations. This mode lets us have back-and-forth exchanges with your AI
assistant, making it feel more like a natural conversation partner.
If you want to experiment with different SLM models, we need to change the model
parameter in generate_response().
Voice interaction systems have significant potential for educational applications and accessibil-
ity improvements by designing interfaces that adapt to different user needs and capabilities. By
integrating small visual models, such as Moondream, with existing TTS pipelines, we can cre-
ate multimodal assistants that describe images for people with visual impairments, converting
visual content into detailed spoken descriptions of scenes, objects, and spatial relationships.
586
To use the camera, we must ensure that the NumPy version is compatible with the Raspberry
Pi system (1.24.2). If this version was changed (what it should be due to the Moonshine
installation), revert it to the 1.24.2 version.
Now, let’s apply what we have explored in previous chapters, along with what we learn in this
one, to create a simple code that converts an image captured by a camera into its corresponding
spoken caption by the speaker.
The code is straightforward and is intended solely to test the solution’s potential.
import os
import time
import subprocess
import ollama
from picamera2 import Picamera2
587
def capture_image(img_path):
# Initialize camera
picam2 = Picamera2()
[Link]()
# Capture image
picam2.capture_file(img_path)
print("\n==> Image captured: "+img_path)
# Stop camera
[Link]()
[Link]()
588
# Check if model exists
if not [Link](model_path):
print(f"Error: Model file not found at {model_path}")
return False
try:
# Run PIPER command
process = [Link](
['piper', '--model', model_path, '--output_file', output_file],
stdin=[Link],
stdout=[Link],
stderr=[Link],
text=True
)
if [Link] == 0:
print(f"\nSpeech generated successfully: {output_file}")
return True
else:
print(f"Error: {stderr}")
return False
except Exception as e:
print(f"Error running PIPER: {e}")
return False
def play_audio(filename="assistant_response.wav"):
try:
# Use aplay to play the audio file
result = [Link](['aplay', filename],
capture_output=True,
text=True)
if [Link] == 0:
print("\nAudio playback completed")
return True
else:
print(f"\nPlayback error: {[Link]}")
return False
589
except Exception as e:
print(f"\nError playing audio: {e}")
return False
IMG_PATH = "/home/mjrovai/Documents/OLLAMA/SST/capt_image.jpg"
MODEL = "moondream:latest"
capture_image(IMG_PATH)
caption = image_description(IMG_PATH, MODEL)
print ("\n==> AI Response:", caption)
text_to_speech_piper(caption)
play_audio()
590
Your browser does not support the audio element.
Troubleshooting
To work simultaneously with STT and the camera in the same environment, we should ensure
that all core scientific packages are version-aligned for the Raspberry Pi environment. To use
NumPy 1.24.2 on our Raspberry Pi, we must ensure that both our SciPy and Librosa versions
are compatible with that NumPy version.
So, to avoid an eventual [Link] import error, and other possible incompatibilities,
we should downgrade scipy (and librosa) to versions that support numpy 1.24.2. According to
official compatibility tables, scipy 1.11.x works with numpy 1.24.x. Librosa versions released
after 0.9.0 also provide better support for recent NumPy releases. However, some older versions
of Librosa are not compatible with NumPy 1.24.2 due to deprecated NumPy attributes. So,
to fix it, we should run the lines below:
Conclusion
This chapter introduced us to multimodal AI system development through an audio and vision
processing pipeline. We learned to integrate speech recognition, language models, and speech
synthesis into a cohesive system that runs efficiently on edge hardware.
We explored how to architect systems with multiple AI components, handle complex error
conditions, and optimize performance within resource constraints.
These skills are essential for the advanced topics in upcoming chapters, including RAG systems,
agent architectures, and the integration of physical computing. The system thinking approach
used here will be essential for future AI engineering work.
We should consider how the voice interaction capabilities we built might enhance other AI
systems. Many applications benefit from voice interfaces, and the foundation established here
can be adapted and extended for various use cases. We can, for example, transform our
audio pipeline into a smart home assistant by integrating physical computing elements. Voice
commands can trigger LED indicators, read sensor values, or control actuators connected to
our Raspberry Pi GPIO pins. Voice queries about environmental conditions can trigger sensor
readings, while voice commands can control connected devices.
591
This chapter extends our SLM and VLM work by adding input and/or output modalities
beyond text and images. The same language models we used previously now process voice-
derived input and generate responses for speech synthesis.
Consider how RAG systems from later chapters might integrate with voice interactions. Voice
queries could trigger document retrieval, with synthesized responses incorporating retrieved
information.
Resources
Python Scripts
592
593
Physical Computing with Raspberry Pi
594
Introduction
Physical computing creates interactive systems that sense and respond to the analog world.
While this field has traditionally focused on direct sensor readings and programmed responses,
we’re entering an exciting new era where Large Language Models (LLMs) can add sophisticated
decision-making and natural language interaction to physical computing projects.
In the Small Language Models (SLM) chapter, we learned how to run an LLM (or, more
precisely, an SLM) on a Single Board Computer (SBC) such as the Raspberry Pi. In this
chapter, we will go through the process of setting up a Raspberry Pi for physical computing,
with an eye toward future AI integration. We’ll cover:
We will also use a Jupyter notebook (programmed in Python) to interact with sensors and
actuators—an important and necessary first step toward the goal of integrating the Raspi with
an SLM.
The combination of Raspberry Pi’s versatility and the power of SLMs opens up exciting pos-
sibilities for creating more intelligent and responsive physical computing systems.
The diagram below gives us an overview of the project:
595
Prerequisites
• Raspberry Pi (model 4 or 5)
• DHT22 Temperature and Relative Humidity Sensor
• BMP280 Barometric Pressure, Temperature and Altitude Sensor
• Colored LEDs (3x)
• Push Button (1x)
• Resistor 4K7 ohm (2x)
• Resistor 220 or 330 ohm (3x)
The Raspberry Pi’s GPIO (General Purpose Input/Output) pins allow us to connect elec-
tronic components and control them with Python code. This opens up endless possibilities for
creating interactive projects, home automation systems, robotics, and more.
This chapter covers the modern GPIO Zero library for interactions with buttons and
LEDs.
� IMPORTANT: [Link] does NOT support Raspberry Pi 5!
596
With the Raspberry Pi 5, we must use GPIO Zero or the newer lgpio library.
[Link] only works on Pi models 1-4, the Pi Zero, Pi Zero 2, and the Pi Zero
2W.
The Raspberry Pi has 40 pins on its header, but not all of them are GPIO pins. Some provide
power (3.3V and 5V), others are ground pins, and the rest are programmable GPIO pins.
In this chapter, we’ll use BCM numbering as it’s more commonly used in Python
programming.
Safety First
Before connecting any components, please read these important safety guide-
lines:
597
• Never connect 5V directly to GPIO pins - GPIO pins are 3.3V tolerant only
• Always use current-limiting resistors with LEDs - Without them, you risk dam-
aging the LED or GPIO pin
• Double-check your connections before powering on
• Disconnect power when making circuit changes
• Respect polarity - LEDs for example, have positive (long leg) and negative (short leg)
sides
A modern, high-level library that makes GPIO programming much simpler and more intuitive.
It uses object-oriented programming and includes built-in features like automatic pin cleanup,
device abstraction, and event detection.
It is essential to note that the GPIO Zero Library uses Broadcom (BCM) pin num-
bering for GPIO pins, rather than physical (board) numbering. Any pin marked
“GPIO” in the previous diagram can be used as a PIN. For example, if an LED were
attached to GPIO13, we would specify the PIN as 13 rather than 33 (the physical
one).
It was created by Ben Nuttall of the Raspberry Pi Foundation, Dave Jones, and other contrib-
utors (GitHub).
Advantages of GPIO Zero:
Installation:
598
“Hello World”: Blinking an LED
Let’s start with the classic ‘Hello World’ of physical computing - making an LED blink!
To connect our RPi to the world, let’s first connect:
• Physical Pin 6 (GND) to GND Breadboard Power Grid (Blue -), using a black jumper
• Physical Pin 1 (3.3V) to +VCC Breadboard Power Grid (Red +), using a red jumper
Now, let’s connect an LED (red) using the physical pin 33 (GPIO13) connected to the LED
cathode (longer LED leg). Connect the LED anode to the breadboard GND using a 220 ohms
resistor to reduce the current drawn from the Raspberry Pi, as shown below:
An LED (Light-Emitting Diode) requires current to flow through it to produce light. However,
without a resistor, too much current can flow, damaging the LED or your Raspberry Pi. We
use a resistor to limit the current to a safe level.
599
Why 220Ω or 330Ω Resistors?
These values limit the current to approximately 10-15mA, which is safe for most standard
LEDs and GPIO pins. The exact value isn’t critical—anything from 220 Ω to 1 kΩ will work
fine.
We can use the built-in Python interpreter to test the LED. In the terminal, enter python,
and once in the interpreter, enter with the commands below:
python
>>> from gpiozero import LED
>>> led = LED(13)
>>> [Link]()
>>> [Link]
600
On the Raspberry, start at home and go to Documents.
cd Documents
Create a directory to save the scripts and install the libraries. Move to there:
mkdir GPIO
cd GPIO
We can use any text editor (such as Nano) to create and run the script. Save the file, for
example, as led_test.py, and then execute it using the terminal:
python led_test.py
Now, let’s blink the LED (the actual “Hello world”) when talking about physical computing.
To do that, we must also import another library: time. We need it to define how long the
LED will be ON and OFF. In the case below, the LED will blink every 1 second.
601
We can use any text editor (such as Nano) to create and run the script. Save the file, for
example, as [Link], and then execute it using the terminal:
python [Link]
The LEDs can be used as “actuators”; depending on the condition of a code running on our
Pi, we can command one of the LEDs to fire! We will install two more LEDs, in addition to
the red one already installed. Follow the diagram and install the yellow (on GPIO 19 ) and
the green (on GPIO 26).
602
For testing we can run a similar code as the used with the single red led, changing the pin
accordantly, for example.
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
603
[Link]()
sleep(5)
[Link]()
[Link]()
[Link]()
In this section, we will setup the Raspberry Pi to capture data from several different sensors:
Sensors and Communication type:
604
Button
Now, let’s learn how to read input from a button. This allows your Raspberry Pi to respond
to physical interactions!
When a button is not pressed, the GPIO pin is ‘floating’ - it’s not connected to anything and
can read random values. We use pull-up or pull-down resistors to set the pin’s default state.
• Pull-Down: Default state is LOW (0V), becomes HIGH when button pressed
• Pull-Up: Default state is HIGH (3.3V), becomes LOW when button pressed
The simplest way to run an external command is with a push button, and the GPIO Zero
Library makes it easy to include in the project. We do not need to think about Pull-up or
Pull-down resistors, etc. In terms of HW, the only thing to do is to connect one leg of our
push-button to any one of the Raspi GPIOs and the other one to GND, as shown in the
diagram:
605
• Push-Button leg1 to GPIO 20
• Push-Button leg2 to GND
606
On a Raspberry Pi Zero 2W, for example, we could use the [Link], whose code is:
[Link]([Link])
[Link](False)
BUTTON_PIN = 20
try:
while True:
if [Link](BUTTON_PIN) == [Link]:
print('Button pressed!')
else:
print('Button not pressed')
[Link](0.1)
except KeyboardInterrupt:
[Link]()
607
Installing Adafruit CircuitPython
The GPIO Zero library is an excellent hardware interfacing library for Raspberry Pi. It’s great
for digital in/out, analog inputs, servos, basic sensors, etc. However, it doesn’t cover SPI/I2C
sensors or drivers. By using CircuitPython via adafruit_blinka, we can take advantage of all
of the drivers and example code developed by Adafruit!
Note that we will keep using GPIO Zero for pins, buttons, and LEDs.
Enable Interfaces
Run these commands to enable the various interfaces such as I2C and SPI:
Install Blinka
Let’s enter the Ollama environment, alheady created to install Blinka:
source ~/ollama/bin/activate
ls /dev/i2c* /dev/spi*
608
Verify Blinka Version
import board
import digitalio
import busio
print("Hello, blinka!")
609
# Try to create a Digital input
pin = [Link](board.D4)
print("Digital IO ok!")
print("done!")
python blinka_test.py
The first sensor to be installed will be the DHT22 for capturing air temperature and relative
humidity data.
Overview
The low-cost DHT temperature and humidity sensors are elementary and slow, but great for
logging basic data. They consist of a capacitive humidity sensor and a thermistor. A bare
chip inside performs the analog-to-digital conversion and outputs a digital signal containing
610
the temperature and humidity. The digital signal is relatively easy to read using any micro-
controller.
DHT22 Main characteristics:
Once we use the sensor at distances less than 20m, a 4K7 ohm resistor should be connected
between the Data and VCC pins. The DHT22 output data pin will be connected to Raspberry
GPIO 16. Check the electrical diagram, connecting the sensor to RPi pins as below:
Do not forget to Install the 4K7 ohm resistor between the VCC and Data pins.
611
Once the sensor is connected, we must install its library on our Raspberry Pi. First, we
should install the Adafruit CircuitPython library, which we have already done, and the
Adafruit_CircuitPython_DHT.
Create a new Python script as below and name it, for example, dht_test.py:
import time
import board
import adafruit_dht
dhtDevice = adafruit_dht.DHT22(board.D16)
612
while True:
try:
# Print the values to the serial port
temperature_c = [Link]
temperature_f = temperature_c * (9 / 5) + 32
humidity = [Link]
print(
"Temp: {:.1f} F / {:.1f} C Humidity: {}% ".format(
temperature_f, temperature_c, humidity
)
)
[Link](2.0)
Placing a finger on the sensor, we can see that both temperature and humidity begin to rise.
613
Installing the BMP280: Barometric Pressure & Altitude Sensor
Sensor Overview:
Environmental sensing has become increasingly important in various industries, from weather
forecasting to indoor navigation and consumer electronics. At the forefront of this technological
advancement are sensors like the BMP280 and BMP180 (deprected), which excel in measuring
temperature and barometric pressure with exceptional precision and reliability.
614
As its predecessor, the BMP180, the BMP280 is an absolute barometric pressure sensor, which
is especially feasible for mobile applications. Its diminutive dimensions and low power con-
sumption allow for its implementation in battery-powered devices such as mobile phones, GPS
modules, or watches. The BMP280 is based on Bosch’s proven piezo-resistive pressure sensor
technology featuring high accuracy and linearity as well as long-term stability and high EMC
robustness. Numerous device operation options guarantee the highest flexibility. The device
is optimized for power consumption, resolution, and filter performance.
Technical data
615
Enabling I2C Interface
Go to RPi Configuration and confirm that the I2C interface is enabled. If not, enable it.
616
If everything has been installed and connected correctly, you can turn on your Rapspi and
start interpreting the BMP180’s information about the environment.
The first thing to do is to check if the Raspi sees your BMP280. Try the following in a
terminal:
sudo i2cdetect -y 1
Create a new Python script as below and name it, for example, bmp280_test.py:
import time
import board
import adafruit_bmp280
i2c = board.I2C()
bmp280 = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address = 0x76)
617
bmp280.sea_level_pressure = 1013.25
while True:
print("\nTemperature: %0.1f C" % [Link])
print("Pressure: %0.1f hPa" % [Link])
print("Altitude = %0.2f meters" % [Link])
[Link](2)
python [Link]
Note that pressure is presented in hPa. See the next section to better understand
this unit.
618
Measuring Weather and Altitude With BMP280
Let’s take some time to understand more about what we will get with the BMP readings.
619
You can skip this part of the tutorial, or return later.
The BMP280 (and its predecessor, the BMP180) was designed to measure atmospheric pressure
accurately. Atmospheric pressure varies with both weather and altitude.
What is Atmospheric Pressure?
Atmospheric pressure is a force that the air around you exerts on everything. The weight of
the gasses in the atmosphere creates atmospheric pressure. A standard unit of pressure is
“pounds per square inch” or psi. We will use the international notation, newtons per square
meter, called pascals (Pa).
This weight, pressing down on the footprint of that column, creates the atmospheric pressure
that we can measure with sensors like the BMP280. Because that cm-wide column of air
weighs about 1 kg, the average sea level pressure is about 101,325 pascals, or better, 1013.25
hPa (1 hPa is also known as milibar - mbar). This will drop about 4% for every 300 meters
you ascend. The higher you get, the less pressure you’ll see because the column to the top
of the atmosphere is much shorter and weighs less. This is useful because you can determine
your altitude by measuring the pressure and doing math.
The air pressure at 3,810 meters is only half that at sea level.
The BMP280 outputs absolute pressure in hPa (mbar). One pascal is a minimal amount of
pressure, approximately the amount that a sheet of paper will exert resting on a table. You
will often see measurements in hectopascals (1 hPa = 100 Pa). The library here provides
outputs of floating-point values in hPa, equaling one millibar (mbar).
Here are some conversions to other pressure units:
Temperature Effects
Because temperature affects the density of a gas, density affects the mass of a gas, and mass
affects the pressure (whew), atmospheric pressure will change dramatically with temperature.
Pilots know this as “density altitude”, which makes it easier to take off on a cold day than a hot
one because the air is denser and has a more significant aerodynamic effect. To compensate for
temperature, the BMP280 includes a rather good temperature sensor and a pressure sensor.
620
To perform a pressure reading, you first take a temperature reading, then combine that with a
raw pressure reading to come up with a final temperature-compensated pressure measurement.
(The library makes all of this very easy.)
Measuring Absolute Pressure
If your application requires measuring absolute pressure, all you have to do is get a temperature
reading, then perform a pressure reading (see the test script for details). The final pressure
reading will be in hPa = mbar. You can convert this to a different unit using the above
conversion factors.
Note that the absolute pressure of the atmosphere will vary with both your altitude
and the current weather patterns, both of which are useful things to measure.
Weather Observations
The atmospheric pressure at any given location on Earth (or anywhere with an atmosphere)
isn’t constant. The complex interaction between the earth’s spin, axis tilt, and many other
factors result in moving areas of higher and lower pressure, which in turn cause the variations
in weather we see every day. By watching for changes in pressure, you can predict short-
term changes in the weather. For example, dropping pressure usually means wet weather or a
storm is approaching (a low-pressure system is moving in). Rising pressure usually means clear
weather is coming (a high-pressure system is moving through). But remember that atmospheric
pressure also varies with altitude. The absolute pressure in my home, Lo Barnechea, in Chile
(altitude 960m), will always be lower than that in San Francisco (less than 2 meters, almost
sea level). If weather stations just reported their absolute pressure, it would be challenging
to compare pressure measurements from one location to another (and large-scale weather
predictions depend on measurements from as many stations as possible).
To solve this problem, weather stations continuously remove the effects of altitude from their
reported pressure readings by mathematically adding the equivalent fixed pressure to make it
appear that the reading was taken at sea level. When you do this, a higher reading in San
Francisco than in Lo Barnechea will always be because of weather patterns and not because
of altitude.
Sea Level Pressure Calculation
The See Level Pressure can be calculated with the formula:
Where,
po = SeaLevel Pressure
p = Atmospheric Pressure
621
L = Temperature Lapse Rate
h = Altitude
To = Sea Level Standard Temperature
g = Earth Surface Gravitational Acceleration
M = Molar Mass Of Dry Air
R = Universal Gas Constant
Having the absolute pressure in Pa, you check the sea level pressure using the
Calculator.
Or calculating in Python, where the altitude is the real altitude in meters where the sensor
is located.
Determining Altitude
Since pressure varies with altitude, you can use a pressure sensor to measure altitude (with a
few caveats). The average pressure of the atmosphere at sea level is 1013.25 hPa (or mbar).
This drops off to zero as you climb towards the vacuum of space. Because the curve of this
drop-off is well understood, you can compute the altitude difference between two pressure
measurements (p and p0) by using a specific equation. The BMP280 gives the measured
altitude using [Link].
The above explanation was based on the BMP 180 Sparkfun tutorial.
In this section, using the Jupyter Notebook, we will read sensors and act on actuators directly
on the Pi.
On the terminal, start the Jupyter notebook server with the command (change the IP address
with the one for your Raspi):
622
You will need the Token; you can copy it from the terminal as shown above.
http:localhost:8888
The first time you connect, you’ll need the token that appears in the Pi terminal
when you start the notebook server.
623
When you start your Pi and want to use Jupyter Notebook, type the “Jupyter
Notebook” command on your terminal and keep it running. This is very important!
If you need to use the terminal for another task, such as running a program, open
a new Terminal window.
To stop the server and close the “kernels” (the Jupyter notebooks), press [Ctrl] + [C].
Let’s create a new notebook (Kernel: Python 3). Open dht_test.py, copy the code, and paste
it into the notebook. That’s it. We can see the temperature and humidity values appearing
on the cell. To interrupt the execution, go to the [stop] button at the top menu.
624
OK, this means we can access the physical world from our notebook! Let’s create a more
structured code for dealing with sensors and actuators.
Initialization
# time library
import time
import datetime
625
import adafruit_bmp280
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address = 0x76)
bmp280Sensor.sea_level_pressure = 1013.25
# LEDs
from gpiozero import LED
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
[Link]()
[Link]()
[Link]()
# Push-Button
from gpiozero import Button
button = Button(20)
626
ledGrnSts = ledGrn.is_lit
[Link]()
[Link]()
[Link]()
627
If you press the push-button, its status will also be shown:
[Link]()
[Link]()
[Link]()
628
We can create a function to simplify turning LEDs on and off:
getGpioStatus()
PrintGpioStatus()
629
Getting and displaying Sensor Data
First, we should create a function to read the BMP280 and calculate the pressure value at sea
level, once the sensor only gives us the absolute pressure based on the actual altitude:
temp = [Link]
pres = [Link]
alt = [Link]
presSeaLevel = pres / pow(1.0 - real_altitude/44330.0, 5.255)
Entering the BMP280 real altitude where it is located, run the code:
bmp280GetData(960)
• Temperature of 26.9 oC
• Absolute Pressure of 906.73 hPa
• Measured Altitude (from Pressure) of 927 m
630
• Sea Level converted Pressure: 1,017.29 hPa
Now, we will generate a unique function to get the BMP280 and the DHT data, including a
timestamp:
tempDHT = [Link]
humDHT = [Link]
Runing them:
631
real_altitude = 960 # real altitude of where the BMP280 is installed
getSensorData(real_altitude)
printData()
Results:
Using Python, we can command the actuators (LEDs) and read the sensors and GIPOs status
at this stage. This is important, for example, to generate a data log to be read by an SLM in
the future.
� IMPORTANT: The Notebook Kernel should end after using to liberate GPIOs
The problem is that Jupyter Notebook is still holding onto the GPIO pins
even after our code finishes running. When we try to run it from the terminal,
those pins are already claimed by the Jupyter process. So, before running a code
that deals with GPIO in the terminal, In Jupyter:
Widgets
pywidgets, or jupyter-widgets orwidgets, are interactive HTML widgets for Jupyter note-
books and the IPython kernel. Notebooks come alive when interactive widgets are used. We
can gain control of our data and visualize changes in them.
632
Widgets are eventful Python objects that have a representation in the browser, often as a
control like a slider, text box, etc. We can use widgets to build interactive GUIs for our
project.
In this lab, for example, we will use a slide bar to control the state of actuators in real time,
such as by turning on or off the LEDs. Widgets are great for adding more dynamic behavior
to Jupyter Notebooks.
Installation
To use Widgets, we must install the Ipywidgets library using the commands:
# widget library
from ipywidgets import interactive
import ipywidgets as widgetsfrom
[Link] import display
And running the below line, we can control the LEDs in real-time:
This interactive widget is very easy to implement and very powerful. You can learn
more about Interactive on this link: Interactive Widget.
633
Advanced GPIO Zero Features
GPIO Zero includes many advanced features that make complex projects much easier to
build.
led = PWMLED(13)
Composite Devices
while True:
[Link]()
sleep(2)
[Link]()
sleep(1)
[Link]()
[Link]()
[Link]()
sleep(3)
[Link]()
[Link]()
sleep(1)
634
[Link]()
This section demonstrates in a simple way how to integrate a Small Language Model (SLM)
with the sensors and LEDs we have set up. The diagram below shows how data flows from
sensors through processing and AI analysis to control the actuators and ultimately provide
user feedback.
635
We will use the Transformers library from Hugging Face for model loading and inference.
This library provides the architecture for working with pre-trained language models, helping
interact with the model, processing input prompts, and obtaining outputs.
Installation
Let’s create a simple SLM test in the Jupyter Notebook that checks if the model loads and
measures inference time. The model used here is the TinyLLama 1.1B. We will ask a straight-
forward question:
As a result, besides the SLM answer, we will also measure the latency.
636
Run this script:
import time
from transformers import pipeline
import torch
model='TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T'
generator = pipeline('text-generation',
model=model,
device=device)
load_time = [Link]() - start_time
print(f"Model loading time: {load_time:.2f} seconds")
# Test prompt
test_prompt = "The weather today is"
As we can see, the SLM works, but the latency is very high (+3 minutes). It is OK because this
particular test is on a Raspberry Pi 4. With a Raspberry Pi 5, the result would be better.:
637
The Raspi uses around 1GB of memory (model + process) and all four cores to process the
answer. The model alone needs around 800MB.
Now, let us create a code showing a basic interaction pattern where the SLM can respond to
sensor data and interact with the LEDs.
Install the Libraries:
import time
import datetime
import board
import adafruit_dht
import adafruit_bmp280
from gpiozero import LED, Button
from transformers import pipeline
Initialize sensors
638
DHT22Sensor = adafruit_dht.DHT22(board.D16)
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)
model='TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T'
generator = pipeline('text-generation',
model=model,
device='cpu')
Support Functions
Now, let’s create support functions for readings from all sensors and control the LEDs:
def get_sensor_data():
"""Get current readings from all sensors"""
try:
temp_dht = [Link]
humidity = [Link]
temp_bmp = [Link]
pressure = [Link]
return {
'temperature_dht': round(temp_dht, 1) if temp_dht else None,
'humidity': round(humidity, 1) if humidity else None,
'temperature_bmp': round(temp_bmp, 1),
'pressure': round(pressure, 1)
}
except RuntimeError:
return None
639
def control_leds(red=False, yellow=False, green=False):
"""Control LED states"""
[Link] = red
[Link] = yellow
[Link] = green
def process_conditions(sensor_data):
"""Process sensor data and control LEDs based on conditions"""
if not sensor_data:
control_leds(red=True) # Error condition
return
temp = sensor_data['temperature_dht']
humidity = sensor_data['humidity']
def generate_response(sensor_data):
"""Generate response based on sensor data using SLM"""
if not sensor_data:
return "Unable to read sensor data"
640
# Generate response from SLM
response = generator(prompt,
max_length=100,
num_return_sequences=1,
temperature=0.7)[0]['generated_text']
return response
Main Function
And now, let’s create a main() function to wait for the user to, for example, press a button
and, capture the data generated by the sensors, delivering some observation or recommendation
from the SLM:
def main_loop():
"""Main program loop"""
print("Starting Physical Computing with SLM Integration...")
print("Press the button to get a reading and SLM response.")
try:
while True:
if button.is_pressed:
# Get sensor readings
sensor_data = get_sensor_data()
if sensor_data:
# Get SLM response
response = generate_response(sensor_data)
641
[Link](0.1) # Reduce CPU usage
except KeyboardInterrupt:
print("\nShutting down...")
control_leds(False, False, False) # Turn off all LEDs
Test Result
The sensors are read after the user presses the button to trigger a reading, and LEDs are
controlled based on conditions. Sensor data is formatted into a prompt for the SLM to generate
a response analyzing the current conditions. The results are displayed in the terminal, and
the LED indicators are shown.
This simple code integrates a Small Language Model (TinyLlama model (1.1B parameters)
with our physical computing setup, providing raw sensor data and intelligent responses from
the SLM about the environmental conditions.
642
We can extend this first test to more sophisticated and valuable uses of the SLM integration,
for example: adding:
Other Models
We can use other SLMs in a Raspberry Pi that have distinct ways of handling them. For
example, many modern models use GGUF formats, and to use them, we need to install
llama-cpp-python, which is designed to work with GGUF models.
Also, as we saw in a previous lab, Ollama is a great way to download and test SLMs on the
Raspberry Pi.
Conclusion
Key Achievements
Throughout this tutorial, we’ve successfully: - Set up a complete physical computing environ-
ment using Raspberry Pi - Integrated multiple environmental sensors (DHT22 and BMP280)
- Implemented visual feedback through LED actuators - Created interactive controls using
push buttons - Integrated a Small Language Model (TinyLLama 1.1B) for intelligent analysis
- Developed a foundation for AI-enhanced environmental monitoring
Technical Insights
Hardware Integration
The combination of digital (DHT22) and I2C (BMP280) sensors demonstrated different com-
munication protocols and their implementations. This multi-sensor approach provides redun-
dancy and comprehensive environmental monitoring capabilities. The LED actuators and
push-button interface created a responsive and interactive system that bridges the digital and
physical worlds.
643
Software Architecture
AI Integration Learnings
The integration of TinyLLama 1.1B revealed several important insights: - Small Language
Models can effectively run on edge devices like Raspberry Pi - Natural language processing
can enhance sensor data interpretation - Real-time analysis is possible, though with some
latency considerations - The system can provide human-readable insights from complex sensor
data
Practical Applications
This project serves as a foundation for numerous real-world applications: - Environmental mon-
itoring systems - Smart home automation - Industrial sensor networks - Educational platforms
for IoT and AI integration - Prototyping platforms for larger-scale deployments
2. Data Integration:
• Developed robust sensor data validation
• Created effective data preprocessing pipelines
• Implemented error handling for sensor failures
3. AI Integration:
• Designed effective prompting strategies
• Managed inference latency
• Balanced accuracy with response time
644
Future Enhancements
Final Thoughts
This chapter demonstrates that integrating physical computing with AI is feasible and practical
on readily accessible hardware such as the Raspberry Pi. Combining sensors, actuators, and
AI creates a powerful platform for developing intelligent environmental monitoring and control
systems.
While the current implementation focuses on environmental monitoring, the principles and
techniques can be adapted to various applications. The modular nature of hardware and
software components allows for customization and expansion based on specific needs.
Integrating small language models into physical computing opens new possibilities for creating
more intuitive and intelligent IoT devices. As edge AI capabilities evolve, projects like this
will become increasingly important in developing the next generation of smart devices and
systems.
Remember that this is just the beginning. Our foundation can be extended in countless ways
to create more sophisticated and capable systems. The key is to build on these basics while
balancing functionality, reliability, and resource usage.
Resources
• GPIOs - Scripts
• Sensors - Scripts
• Notebooks
645
Experimenting with SLMs for IoT Control
Introduction
This chapter explores the implementation of Small Language Models (SLMs) in IoT control
systems, demonstrating the possibility of creating a monitoring and control system using edge
AI. We’ll integrate these models with physical sensors and actuators, creating an intelligent
IoT system capable of natural language interaction. While this implementation shows
the potential of integrating AI with physical systems, it also highlights current limitations and
areas for improvement.
This chapter builds on the concepts introduced in “Small Language Models (SLMs)”
and “Physical Computing with Raspberry Pi.”
646
The Physical Computing chapter laid the groundwork for interfacing with hardware compo-
nents using the Raspberry Pi’s GPIO pins. We’ll revisit these concepts, focusing on connecting
and interacting with sensors (DHT22 for temperature and humidity, BMP280 for temperature
and pressure, and a push-button for digital inputs), as well as controlling actuators (LEDs) in
a more sophisticated setup.
We will progress from a simple IoT system to a more advanced platform that combines real-
time monitoring, historical data analysis, and natural language processing (NLP).
647
• Simple observation and reporting of system state
• Demonstration of SLM’s ability to interpret sensor data
3. Active Control Implementation
• Direct LED control based on SLM decisions
• Temperature threshold monitoring
• Emergency state detection via button input
• Real-time system state analysis
4. Natural Language Interaction
• Free-form command interpretation
• Context-aware responses
• Multiple SLM model support
• Flexible query handling
5. Data Logging and Analysis
• Continuous system state recording
• Trend analysis and pattern detection
• Historical data querying
• Performance monitoring
Let’s begin by setting up our hardware and software environment, building upon the foundation
established in our previous labs.
Setup
Hardware Setup
Connection Diagram
648
• Raspberry Pi 5 (with an OS installed, as detailed in previous labs)
• DHT22 temperature and humidity sensor
• BMP280 temperature and pressure sensor
• 3 LEDs (red, yellow, green)
• Push button
• 330Ω resistors (3)
• Jumper wires and breadboard
649
Software Prerequisites
Let’s create a Python script ([Link]) to handle the sensors and actuators. This script
will contain functions to be called from other scripts later:
import time
import board
import adafruit_dht
import adafruit_bmp280
from gpiozero import LED, Button
DHT22Sensor = adafruit_dht.DHT22(board.D16)
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)
def collect_data():
try:
temperature_dht = [Link]
humidity = [Link]
temperature_bmp = [Link]
pressure = [Link]
button_pressed = button.is_pressed
return temperature_dht, humidity, temperature_bmp, pressure, button_pressed
except RuntimeError:
return None, None, None, None, None
def led_status():
650
ledRedSts = ledRed.is_lit
ledYlwSts = ledYlw.is_lit
ledGrnSts = ledGrn.is_lit
return ledRedSts, ledYlwSts, ledGrnSts
while True:
ledRedSts, ledYlwSts, ledGrnSts = led_status()
temp_dht, hum, temp_bmp, press, button_state = collect_data()
[Link](2)
651
SLM Basic Analysis
Now, let’s create a new script, slm_basic_analysis.py, which will be responsible for
analysing the hardware components’ status, according to the following diagram:
The diagram shows the basic analysis system, which consists of:
1. Hardware Layer:
• Sensors: DHT22 (temperature/humidity), BMP280 (temperature/pressure)
652
• Input: Emergency button
• Output: Three LEDs (Red, Yellow, Green)
2. [Link]:
• Handles all hardware interactions
• Provides two main functions:
– collect_data(): Reads all sensor values
– led_status(): Checks current LED states
3. slm_basic_analysis.py:
• Creates a descriptive prompt using sensor data
• Sends prompt to SLM (for example, the Llama 3.2 1B)
• Displays analysis results
• In this step we will not control the LEDs (observation only)
Okay, let’s implement the code, starting for importing the Ollama library and the functions
to monitor the HW (from the previous script):
import ollama
from monitor import collect_data, led_status
Now, the heart of out code, we will generate the Prompt, using the data captured on the
previous variables:
prompt = f"""
You are an experienced environmental scientist.
Analyze the information received from an IoT system:
Where,
- The button, not pressed, shows a normal operation
- The button, when pressed, shows an emergency
653
- Red LED when is on, indicates a problem/emergency.
- Yellow LED when is on indicates a warning situation.
- Green LED when is on, indicates system is OK.
"""
Now, the Prompt will be passed to the SLM, which will generate a response:
MODEL = 'llama3.2:3b'
PROMPT = prompt
response = [Link](
model=MODEL,
prompt=PROMPT
)
The last stage will be show the real monitored data and the SLM’s response:
654
In this initial experiment, the system successfully collected sensor data (temperatures of 26.3°C
and 26.1°C from DHT22 and BMP280, respectively, 40.2% humidity, and 908.84hPa pressure)
and processed this information through the SLM, which produced a coherent response recom-
mending the activation of the yellow LED due to elevated temperature conditions.
The model’s ability to interpret sensor data and provide logical, rule-based decisions shows
promise. Still, the simplistic nature of the current implementation (using basic thresholds and
binary LED outputs) suggests significant room for improvement through more sophisticated
prompting strategies, historical data integration, and the implementation of safety mechanisms.
Also, the result is probabilistic, meaning it should change after execution.
Let’s use the output generated and use it to actuate on the LEDs, our “actuators”:
655
We can add a new function, parse_llm_response(), to return a command to the the LEDs
based on the SLM’s response:
def parse_llm_response(response_text):
"""Parse the LLM response to extract LED control instructions."""
response_lower = response_text.lower()
red_led = 'activate red led' in response_lower
yellow_led = 'activate yellow led' in response_lower
green_led = 'activate green led' in response_lower
return (red_led, yellow_led, green_led)
Would be:(False, True, False), which can be the input to the function control_leds(red,
yellow, green)'
656
Creating new functions for the Actuation:
657
print(f"\nSYSTEM ACTUATOR STATUS")
ledRedSts, ledYlwSts, ledGrnSts = led_status()
print(f" - Red LED {'is on' if ledRedSts else 'is off'}")
print(f" - Yellow LED {'is on' if ledYlwSts else 'is off'}")
print(f" - Green LED {'is on' if ledGrnSts else 'is off'}")
By updating the system and calling the two functions in sequence, we can have the full cycle
of analysis by the SLM and the actuation on the output LEDs:
Let’s press the button, call the functions, and see what happens:
658
We can see that, despite the button being pressed, the SLM did not consider it and misin-
terpreted the temperature value as ABOVE the threshold, not below. Also, despite the fact
that we asked for a straight answer about which LED to turn on, the model lost time with
analysis,
This issue relates to the prompt we wrote. Let’s cover it in the next section.
659
Prompting Engineering
Looking at the answers we got with the previous implementation, which are not always correct,
we can see that the main issue is unreliable text parsing. Using JSON to parse the answer
should be much better!
660
sensor data.
3. Changed to JSON format - The prompt now asks for a structured JSON response:
4. Updated parser - Now parses JSON instead of searching for text strings. Includes
error handling and fallback to safe state (all LEDs off) if parsing fails.
Let’s revise the previous code, with a better prompt and parcing funcion, now based on
JSON.
This project demonstrates how a Small Language Model (SLM) can make intelligent decisions
in an IoT system. The system monitors environmental conditions using sensors and controls
LED indicators based on the SLM’s analysis.
661
New System Architecture
662
– Yellow LED: Warning state
– Green LED: Normal operation
This layer provides the bridge between hardware and software with three key functions:
• collect_data(): Reads all sensor values (temperature, humidity, pressure) and button
state
• led_status(): Checks the current state of all LEDs
• control_leds(red, yellow, green): Controls which LED is turned on based on
boolean values
This is where the SLM makes decisions. The workflow follows these steps:
a) Data Collection & Preparation
b) Prompt Generation
c) SLM Inference
663
{
"red_led": false,
"yellow_led": false,
"green_led": true
}
d) Response Processing
e) Actuation
664
Configurable Threshold: The temperature threshold is a parameter that makes the system
adaptable to different environments without code changes. Can be updated by the user.
Clear Priority Rules: The system follows explicit priority rules that the SLM understands:
Now, this architecture showcases how Small Language Models can be integrated into IoT
systems to provide intelligent, context-aware decision-making while maintaining simplicity
and reliability.
import ollama
import json
from monitor import collect_data, led_status, control_leds
SENSOR DATA:
- DHT22 Temperature: {temp_dht:.1f}°C
- BMP280 Temperature: {temp_bmp:.1f}°C
- Humidity: {hum:.1f}%
- Pressure: {press:.2f}hPa
- Button: {"PRESSED" if button_state else "NOT PRESSED"}
665
CURRENT ANALYSIS:
- Button status: {"PRESSED" if button_state else "NOT PRESSED"}
- DHT22 temp ({temp_dht:.1f}°C) is {"OVER" if temp_dht > TEMP_THRESHOLD
else "AT OR BELOW"} threshold ({TEMP_THRESHOLD}°C)
- BMP280 temp ({temp_bmp:.1f}°C) is {"OVER" if temp_bmp > TEMP_THRESHOLD
else "AT OR BELOW"} threshold ({TEMP_THRESHOLD}°C)
Based on these rules, respond with ONLY a JSON object (no other text):
{{"red_led": true, "yellow_led": false, "green_led": false}}
Only ONE LED should be true, the other two must be false.
"""
def parse_llm_response(response_text):
"""Parse the LLM JSON response to extract LED control instructions."""
try:
# Clean the response - remove any markdown code blocks if present
response_text = response_text.strip()
if response_text.startswith('```'):
# Extract JSON from markdown code block
lines = response_text.split('\n')
response_text = '\n'.join(lines[1:-1])
if len(lines) > 2
else response_text
# Parse JSON
data = [Link](response_text)
red_led = [Link]('red_led', False)
yellow_led = [Link]('yellow_led', False)
green_led = [Link]('green_led', False)
return (red_led, yellow_led, green_led)
except ([Link], KeyError) as e:
print(f"Error parsing JSON response: {e}")
print(f"Response was: {response_text}")
# Fallback to safe state (all LEDs off)
return (False, False, False)
666
print(f"SYSTEM REAL DATA")
print(f" - DHT22 ==> Temp: {temp_dht:.1f}°C, Humidity: {hum:.1f}%")
print(f" - BMP280 => Temp: {temp_bmp:.1f}°C, Pressure: {press:.2f}hPa")
print(f" - Button {'pressed' if button_state else 'not pressed'}")
667
Code Flow Diagram
Definitions
# Model to be used
MODEL = 'llama3.2:3b'
668
Calling the program
slm_analyse_act(MODEL, TEMP_THRESHOLD)
Result
669
TEST2: Temp below the threshold
slm_analyse_act(MODEL, TEMP_THRESHOLD)
670
TEST3: Alarm Button pressed
slm_analyse_act(MODEL, TEMP_THRESHOLD)
Now, let’s transform our IoT monitoring setup into an interactive assistant that accepts
natural language commands and queries. Instead of autonomously monitoring conditions, the
671
system will now respond to our requests in real-time.
672
How It will Work
1. Interactive Loop
User Input → Sensor Reading → SLM Analysis → LED Control → Status Display → Wait for Next Inp
2. Dual-Purpose Response
{
"message": "Helpful text response to the user",
"leds": {
"red_led": false,
"yellow_led": true,
"green_led": false
}
}
3. Command Types
A. Information Queries
Ask questions about sensor readings - LEDs remain unchanged.
Examples:
Response format:
{
"message": "The current temperature is 21.5°C from DHT22 and 22.3°C from BMP280.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false} // keeps current sta
673
}
Response format:
{
"message": "Yellow LED turned on.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false}
}
C. Conditional Commands
Commands that depend on sensor readings or button state.
Examples:
Response format:
{
"message": "Temperature is 21.5°C, which is above 20°C. Yellow LED turned on.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false}
}
D. Toggle/Switch Commands
Commands that change LED states based on current conditions.
Examples:
Response format:
674
{
"message": "Button is pressed. LED states switched.",
"leds": {"red_led": true, "yellow_led": false, "green_led": false} // inverted from cur
}
E. Analysis Queries
Ask the SLM to analyze sensor data.
Examples:
Response format:
{
"message": "Based on pressure of 910.18hPa and humidity of 28.8%, conditions are dry. Ra
"leds": {"red_led": false, "yellow_led": false, "green_led": true} // keeps current sta
}
Key Functions
[Link].1 * create_interactive_prompt()
[Link].2 * parse_interactive_response()
675
[Link].3 * display_system_status()
[Link].4 * interactive_mode()
python slm_act_leds_interactive.py
Tests
============================================================
IoT Environmental Monitoring System - Interactive Mode
Using Model: llama3.2:3b
============================================================
676
- Type 'status' to see system status
- Type 'exit' or 'quit' to stop
============================================================
You: status
============================================================
SYSTEM STATUS
============================================================
DHT22 Sensor: Temp = 21.5°C, Humidity = 28.8%
BMP280 Sensor: Temp = 22.3°C, Pressure = 910.18hPa
Button: NOT PRESSED
LED Status:
Red LED: � OFF
Yellow LED: � ON
Green LED: � OFF
============================================================
You: exit
Exiting interactive mode. Goodbye!
677
Special Commands
Languages
One advantage of using SLMs is that, once multilingual models are used (as Llama 3,2 or
Gemma), the user can choose the better language for them, independent of the language used
during coding.
Error Handling
• If sensor data cannot be read, the system will notify you and wait for the next command
• If the SLM response cannot be parsed, the system keeps LEDs in their current state
678
Tips for Best Results
1. Be specific: “Turn on the yellow LED” works better than “turn on the light”
2. Use conditions clearly: “If temperature is above 20°C” is clearer than “when it’s hot”
3. Ask for status: Use the status command frequently to verify system state
4. One LED at a time: Unless you specifically say “all LEDs”, the system defaults to
one LED on
Flow Diagram
The development of our Interactive IoT-SLM system can also be followed using the
Jupyter Notebook: SLM_IoT.ipynb.
679
Adding Data Logging
Now, we will develop an enhanced version that adds data logging, analysis, and historical
query capabilities. The system automatically logs all sensor readings and commands, allowing
it to analyze trends and query historical data using natural language.
Key Features
680
3. Statistical Analysis
• Min/max/average calculations
• LED state change tracking
• Button press counting
• Trend analysis
python slm_act_leds_with_logging.py
Example Queries
Real-time:
Historical:
Built-in commands:
681
Data Files
sensor_readings.csv
timestamp,temp_dht,humidity,temp_bmp,pressure,
button_pressed,red_led,yellow_led,green_led
command_history.csv
timestamp,user_command,slm_response,red_led,yellow_led,green_led
Tips
The complete scripts for the datalogger version arehere: data_logger.py and
slm_act_leds_with_logging.py
682
Flow Diagram
683
Examples:
684
Prompt Optimization and Efficiency
685
tokens/s")
We will get:
Based on the prompt_eval_duration value, the PROMPT is our main bottleneck, so we must
make it as concise as possible. The SLM must process all of the context before it can generate
the first output token.
Quick Solution:
• Condense System Status: Remove unnecessary descriptive text and present the status
information in a compact, structured format.
– Example (Before):
– Example (After):
686
• Reduce Examples: While examples are crucial for instruction-following, eliminate
redundancy. Keep only the most diverse and representative examples. The current
prompt is very long due to verbose examples and instructions. Focus on the single-
shot example that shows the required JSON output format.
• Simplify Instructions: Make the instructions as direct and short as possible. Use
keywords instead of full sentences where clarity is maintained.
• System message: It defines the assistant’s behavior and should sent once at initializa-
tion, not at PROMPT
RULES:
Always respond with valid JSON containing both "message" and "leds" fields."""
687
Model Pre-loading
New Code
Original Functions:
New functions:
688
Using the optimized code: slm_act_leds_interactive_optimized.py, we get as a response:
Using Pydantic
Using Pydantic is a robust way to improve the reliability, efficiency, and maintainability of
our system, mainly since we rely on the LLM to output precise JSON.
Here’s how Pydantic can help and what you would need to do:
689
1. How Pydantic Reduces Latency and Improves Reliability
Pydantic doesn’t directly speed up the model’s token generation, but it can indirectly reduce
latency and eliminate error-handling overhead by enabling cleaner, faster parsing and
more robust communication.
• JSON Schema: Pydantic can generate a JSON Schema from your Python classes. We
can include this schema directly in our prompt, which acts as an unambiguous, machine-
readable instruction for the SLM. This often leads to fewer errors in the model’s output,
reducing the need for costly retries or complex string manipulation.
690
Step 1: Define the Pydantic Models
class LEDControl(BaseModel):
"""LED control configuration."""
red_led: bool = Field(description="Red LED state (on/off)")
yellow_led: bool = Field(description="Yellow LED state (on/off)")
green_led: bool = Field(description="Green LED state (on/off)")
class AssistantResponse(BaseModel):
"""Complete assistant response with message and LED control."""
message: str = Field(description="Helpful response to the user")
leds: LEDControl = Field(description="LED control configuration")
def parse_interactive_response(response_text):
"""Parse the interactive SLM response using Pydantic (guaranteed valid)."""
try:
# Parse directly into Pydantic model - guaranteed valid JSON structure
691
data = AssistantResponse.model_validate_json(response_text)
except Exception as e:
print(f"Error parsing response: {e}")
print(f"Response was: {response_text}")
return "Error: Could not parse SLM response.", (False, False, False)
With such modifications, the final latency was reduced from 90 to around 60 seconds
Note that when we sent two different commands to turn on the LEDs, the new one did not
turn off the previous one. This is due to the change in the rules. If we want, we can modify
it to match whatever we wish to.
692
The final code can be found on GitHub: slm_act_leds_interactive_pydantic.py
and in notebook SLM_IoT.ipynb
In short, using Pydantic over tradicional approuch with JSON, we have as benefits:
Next Steps
This chapter involved experimenting with simple applications and verifying the feasibility of
using an SLM to control IoT devices. The final result is far from usable in the real world, but
it can serve as a starting point for more interesting applications. Below are some observations
and suggestions for improvement:
693
• Some simple commands could be handled without SLM intervention. We can do it
programmatically.
• Consider implementing a proper state machine for LED control to ensure consistent
behavior.
• Implement more sophisticated trend analysis using statistical methods.
• Add support for more complex queries combining multiple data points.
694
Conclusion
This chapter has demonstrated the progressive evolution of an IoT system from basic sen-
sor integration to an intelligent, interactive platform powered by Small Language Models.
Through our journey, we’ve explored several key aspects of combining edge AI with physical
computing:
Key Achievements
1. SLM Reliability
• Probabilistic nature of responses
695
• Consistency issues in decision making
• Need for better validation and verification
2. System Performance
• Response time considerations
• Resource usage on edge devices
• Efficiency of data logging and analysis
3. Architectural Constraints
• Simple state management
• Basic error handling
• Limited data validation
Final Thoughts
While this implementation demonstrates the potential of combining SLMs with IoT systems,
it also highlights the exciting possibilities and challenges ahead. Though experimental, the
system we’ve built provides a solid foundation for understanding how edge AI can enhance IoT
applications. As SLMs evolve and improve, their integration with physical computing systems
will likely become more robust and practical for real-world applications.
This chapter has shown that, despite current limitations, SLMs can provide intelligent, natural-
language interfaces to IoT systems, opening new possibilities for human-machine interaction
in the physical world.
The future of IoT systems is shaped by intelligent, edge-based solutions that combine AI’s
power with the practicality of physical computing.
Resourses
696
Advancing EdgeAI: Beyond Basic SLMs
Figure 16: Image from author, with a Raspberry Pi from ImageFX - prompt - Create a cartoon
image with a single Raspberry Pi with a white background
697
Understanding SLM Limitations
Small Language Models, while impressive in their ability to run on edge devices, face several
key limitations:
1. Knowledge Constraints
SLMs have limited knowledge based on their training data, often outdated and incomplete. Un-
like their larger counterparts, they cannot store the vast information needed for comprehensive
expertise across all domains.
Let’s run the below example to verify this limitation.
import ollama
response = [Link](
model="llama3.2:1b",
prompt="Who won the 2024 Summer Olympics men's 100m sprint final?"
)
print(response['response'])
The output of the previous code will likely show hallucination or admission of not knowing, as
in the case below:
This constraint could be solved simply by having an Agent search the Internet for the answer
or using Retrieval-Augmented Generation (RAG), as we will see later.
698
2. Reasoning Limitations
Complex reasoning tasks often exceed the capabilities of SLMs, which struggle with multi-
step logical deductions, mathematical computations, and a nuanced understanding of context.
Agents can be used to mitigate such limitations.
For example, let’s reuse the previous code and ask to the SLM to multiply two numbers :
import ollama
response = [Link](
model="llama3.2:3b",
prompt="Multiply 123456 by 123456"
)
print(response['response'])
The response is wrong; once the multiplication result should be 15,241,383,936. This is
expected once the language models are not suitable for mathematical computations. Still, we
can use an “agent” to determine whether a user asks for multiplication or a general question.
We will learn how to create an agent later.
3. Inconsistent Outputs
SLMs may produce inconsistent responses to the same query, making them unreliable for
critical applications requiring deterministic outputs. Several enhancements, such as Function
Calling and Response Validation, can improve reliability.
4. Domain Specialization
SLMs perform worse than specialized models in domain-specific tasks like visual recognition or
time-series analysis. Fine-tuning can adapt models to specific domains or tasks, improving
performance for targeted applications.
699
Techniques for Enhancing SLM at the Edge
Small Language Models (SLMs) offer remarkable capabilities for edge devices, but various
techniques can significantly enhance their effectiveness. Here, we present a comprehensive
framework for optimizing SLMs on resource-constrained devices like the Raspberry Pi, orga-
nized from fundamental to advanced approaches.
We will divide those technics into 3 segments:
– Chain-of-Thought Prompting
– Few-Shot Learning
– Task Decomposition
The true power of these techniques emerges when they’re strategically combined:
1. Agent Architecture with RAG: Create agents that can access both tools and knowl-
edge bases
2. Validation-Enhanced RAG: Apply response validation to ensure RAG outputs are
accurate
3. Fine-Tuned Routers: Use specialized fine-tuned models to handle routing decisions
4. Chain-of-Thought with Function Calling: Combine reasoning traces with struc-
tured outputs
700
• Function calling to structure sensor data analysis
• Response validation to verify recommendations
• Task decomposition to handle complex multi-part weather analysis
Chain-of-Thought Prompting
Chain-of-thought prompting encourages SLMs to break down complex problems into step-by-
step reasoning, leading to more accurate results:
def solve_math_problem(problem):
prompt = f"""
Problem: {problem}
Let's think about this step by step:
1. First, I'll identify what we're looking for
2. Then, I'll identify the relevant information
3. Next, I'll set up the appropriate equations
4. Finally, I'll solve the problem carefully
Solving:
"""
response = [Link](model="llama3.2:3b", prompt=prompt)
return response['response']
Few-Shot Learning
Few-shot learning provides examples within the prompt, helping SLMs understand the ex-
pected response format and reasoning pattern:
def classify_sentiment(text):
prompt = f"""
Task: Classify the sentiment of the text as positive, negative, or neutral.
Examples:
Text: "I love this product, it works perfectly!"
Sentiment: positive
Text: "This is the worst experience I've ever had."
Sentiment: negative
701
Text: "The package arrived on time."
Sentiment: neutral
Text: "{text}"
Sentiment:
"""
response = [Link](model="llama3.2:1b", prompt=prompt)
return response['response'].strip()
This approach is particularly effective for classification tasks and standardized outputs.
Task Decomposition
For complex tasks, breaking them into smaller subtasks helps SLMs manage complexity:
def analyze_product_review(review):
# Step 1: Extract main points
points_prompt = f"Extract the main points from this product review: {review}"
points_response = [Link](model="llama3.2:1b", prompt=points_prompt)
main_points = points_response['response']
# Final synthesis
final_prompt = f"""
Create a concise analysis of this product review based on:
Main points: {main_points}
Overall sentiment: {sentiment}
Improvement suggestions: {improvements}
"""
final_response = [Link](model="llama3.2:1b",
702
prompt=final_prompt)
return final_response['response']
This technique distributes cognitive load across multiple simpler prompts, enabling SLMs to
handle tasks that might otherwise exceed their capabilities.
To address some of these limitations, we can develop agents that leverage SLMs as part of a
more extensive system with additional capabilities.
Let’s think about the multiplication problem that we faced before. An Agent can be used for
that.
An agent is a system that uses an AI Model as its core reasoning engine to:
703
For example, if it is a multiplication, we can use a Python function as a “tool” to calculate it,
as shown in the diagram:
704
Our code works through the following steps:
1. User Input: The user types a query like “What is 7 times 8?” or “What is the capital
of France?”
2. Process Query: The process_query() function handles the input and decides what
to do with it.
3. Classification: The ask_ollama_for_classification() function sends the user’s
query to the SLM (using Ollama) with a prompt asking it to classify whether the query
705
is requesting multiplication or asking a general question.
4. Decision: Based on the SLM’s classification:
• If it’s a multiplication request, the SLM also extracts the numbers, and we use our
multiply() function.
• If it’s a general question, we send the original query to the SLM for a direct answer.
5. Response: The system returns either the multiplication result or the SLM’s answer to
the general question.
Here’s a Python script that creates a simple agent (or router) between multiplication operations
and general questions as described:
import requests
import json
# Configuration
OLLAMA_URL = "[Link]
MODEL = "llama3.2:3b" # You can change this to any model you have installed
VERBOSE = True
def ask_ollama_for_classification(user_input):
"""
Ask Ollama to classify whether the query is a multiplication request or a \
general question.
"""
classification_prompt = f"""
Analyze the following query and determine if it's asking for multiplication \
or if it's a general question.
Query: "{user_input}"
If it's asking for multiplication, respond with a JSON object in this format:
{{
"type": "multiplication",
"numbers": [number1, number2]
}}
706
Respond ONLY with the JSON object, nothing else.
"""
try:
if VERBOSE:
print(f"Sending classification request to Ollama")
response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": classification_prompt,
"stream": False
}
)
if response.status_code == 200:
response_text = [Link]().get("response", "").strip()
if VERBOSE:
print(f"Classification response: {response_text}")
except Exception as e:
if VERBOSE:
707
print(f"Error connecting to Ollama: {str(e)}")
return {"type": "general_question"}
def ask_ollama(query):
"""
Send a query to Ollama for general question answering.
"""
try:
if VERBOSE:
print(f"Sending query to Ollama")
response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": query,
"stream": False
}
)
if response.status_code == 200:
return [Link]().get("response", "")
else:
return f"Error: Received status code {response.status_code} \
from Ollama."
except Exception as e:
return f"Error connecting to Ollama: {str(e)}"
def process_query(user_input):
"""
Process the user input by first asking Ollama to classify it,
then either performing multiplication or sending it back as a
general question.
708
"""
# Let Ollama classify the query
classification = ask_ollama_for_classification(user_input)
if VERBOSE:
print("Ollama classification:", classification)
if [Link]("type") == "multiplication":
numbers = [Link]("numbers", [0, 0])
if len(numbers) >= 2:
return multiply(numbers[0], numbers[1])
else:
return "I understood you wanted multiplication, but couldn't \
extract the numbers properly."
else:
return ask_ollama(user_input)
def main():
"""
Main function to run the agent interactively.
"""
print("Ollama Agent (Type 'exit' to quit)")
print("-----------------------------------")
while True:
user_input = input("\nYou: ")
response = process_query(user_input)
print(f"\nAgent: {response}")
# Example usage
if __name__ == "__main__":
# Set to True to see detailed logging
VERBOSE = True
main()
When we run the script, we can see that, first, the SLM chooses multiplication, passing the
numbers entered by the user to the “tool,” which, in this case, is the multiply() function. As
709
a result, we got 15,241,383,936, which it is correct.
Let’s now enter with another question that has no relation with arithmetic, for example: What
is the capital of Brazil? In this case, the SLM will decide that the query is a general
question and pass it on to the SLM to answer it.
This simple agent (or router) demonstrates the fundamental concept of using an SLM to make
decisions about processing different types of user inputs. It shows both the power of SLMs for
natural language understanding and their limitations in structured tasks.
710
Limitations and Considerations
This agent seems to resolve our problem, but it has several limitations that are common when
working with SLMs:
1. JSON Parsing Issues: SLMs don’t always perfectly format JSON responses as re-
quested. The code includes error handling for this.
2. Classification Reliability: The SLM might not always correctly classify the query,
especially with ambiguous questions.
3. Number Extraction: The SLM might extract numbers incorrectly or miss them en-
tirely.
4. Error Handling: Robust error handling is essential when working with SLMs because
their outputs can be unpredictable.
5. Latency: Significant latency is involved in making multiple calls to the SLM. For exam-
ple, for the above simple agent, the latency was about 50s when using the llama3.2:3B
on a Raspberry Pi 5.
Here, you can see the SLM latency (simple query) per device (in tokens/s):
In my simple tests, the 1B models struggled to classify the tasks correctly. The
the 3B and 4B models worked fine
Improvements
1. Expand Capabilities: Add support for more operations (addition, subtraction, divi-
sion).
2. Better Error Handling: Improve fallback mechanisms when the SLM fails to extract
numbers or classify correctly.
3. Model Preloading: Initialize the model at startup to reduce latency.
4. Adding Regex Fallbacks: Use regular expressions as a fallback to extract numbers
when the SLM fails.
5. Context Preservation: Maintain conversation context for multi-turn interactions.
711
A more robust script can be used with the above improvements. The diagram shows how it
would work:
1. Initialization:
712
• The system starts by initializing both models in parallel threads
• This prevents cold starts and reduces latency
2. Query Processing Flow:
• User input is first sent to a classification step
• A model (llama3.2:3B) determines if it’s a calculation or a general question (we can
choose a different model here).
• If it’s a calculation:
– The system extracts the operation type and numbers
– Numbers are converted from strings to floats
– The appropriate calculation is performed
– Results are formatted with comma separators (e.g., 1,234,567.89)
• If it’s a general question:
– The query is sent to the main model (llama3.2:3b) for answering (we can choose
a different model here)
3. Optimizations (highlighted in the subgraph):
• Persistent HTTP session for connection reuse
• Keep-alive parameter to prevent model unloading
• Simplified classification prompt for faster processing
• Using a smaller model for the classification task
• Rule-based fallback logic if the model classification fails
This approach maintains the intelligent classification capability while significantly reducing
execution time compared to the original implementation.
Runing the script [Link], we get correct results with reduced la-
tency of about 60%.
713
General Knowledge Router
Remember when we asked our SLM: Who won the 2024 Summer Olympics men's 100m
sprint final? We could not receive an answer because the modes were trained with
information previously in late 2023.
To solve this issue, let’s build a more advanced agent to classify whether it should use its
knowledge to answer a question or fetch updated information from the Internet. This addresses
a key limitation of Small Language Models: their knowledge cutoff date.
714
The general architecture of our agent will be similar to the calculator, but now, we will use a
web search API as a tool.
This agent addresses a critical limitation of SLMs - their knowledge cutoff date - by determining
when to use the model’s built-in knowledge versus when to search for up-to-date information
from the web.
How it works:
Uses SLM for Classification: Relies entirely on the SLM to determine whether a query
needs web search or can be answered from the model’s knowledge.
Provides Date Context: This section supplies the current date to help the SLM make
informed decisions about whether information is outdated.
Integrates Tavily Search: Uses Tavily’s powerful search API to find relevant information
for queries that need external data.
Handles Timeouts: Includes fallback mechanisms when the model takes too long to re-
spond.
Maintains Source Attribution: Clearly indicates to the user whether the answer comes
from the model’s knowledge or web search.
Let’s run the script: [Link]
But first, we should install the required libraries:
715
Runing the script and entering with the same questions that could not be answered before,
we now have: Noah Lyles won the men's 100m sprint final at the 2024 Summer
Olympics. He set a new personal best time of 9.79 seconds. This victory
marked the United States' first win in the event since 2004.
When the user enters a common-knowledge question, the agent will send it directly to the
SLM. For example, if the user asks, "Who is Albert Einstein?", we get:
716
Improving Agent Reliability
There are several ways to enhance an agent’s reliability. One is to implement effective, ap-
proved, structured function calling, which makes agents’ responses more consistent and pre-
dictable.
In the SLM chapter, we explored function calling when we created an app where the user
enters a country’s name and gets, as an output, the distance in km from the capital city of
such a country and the app’s location.
717
Once the user enters a country name, the model will return the name of its capital city (as a
string) and the latitude and longitude of such city (in float). Using those coordinates, the app
used a simple Python library (haversine) to calculate the distance between those 2 points.
The critical library used was Pydantic (and instructor), a robust data validation and settings
management library engineered by Python to enhance the robustness and reliability of our
codebase. In short, Pydantic helps ensure that the model’s response will always be consis-
tent.
Function calling can improve an agent’s reliability by ensuring structured outputs and clear
tool selection logic. Here’s a generic template about how we can implement it :
import time
from haversine import haversine
from pydantic import BaseModel, Field
from ollama import chat
class CityCoord(BaseModel):
city: str = Field(..., description="Name of the city")
lat: float = Field(..., description="Decimal Latitude of the city")
lon: float = Field(..., description="Decimal Longitude of the city")
718
# Ask Ollama for structured data
response = chat(
model=model,
messages=[{
"role": "user",
"content": f"Return the capital city of {country}, \
with its decimal latitude and longitude."
}],
format=CityCoord.model_json_schema(), # Structured JSON format
options={"temperature": 0}
)
resp = CityCoord.model_validate_json([Link])
# Test
calc_dist('france')
calc_dist('colombia')
calc_dist('united states')
719
2. Response Validation
• Response Relevancy: Determines if the LLM output addresses the input informatively
and concisely.
• Prompt Alignment: Check if the LLM output follows instructions from the prompt
template.
• Correctness: Assesses factual accuracy based on ground truth.
• Hallucination Detection: Identifies fake or made-up information in LLM outputs.
Adding validation prevents incorrect or harmful responses, and here, we can test it with a
simple script:
import ollama
import json
try:
validation = [Link](
model="llama3.2:3b",
prompt=validation_prompt
)
720
result = [Link](validation['response'])
return result
except Exception as e:
print(f"Error during validation: {e}")
return {"valid": False, "reason": "Validation error", "score": 0}
# Test
query = "What is the Raspberry Pi 5?"
response = "It is a pie created with raspberry and cooked in an oven"
validation = validate_response(query, response)
print(validation)
RAG systems enhance Small Language Models (SLMs) by providing relevant information
from external sources before generation. This is particularly valuable for edge devices with
limited model sizes, as it allows them to access knowledge beyond their training data without
increasing the model size.
Understanding RAG
In a basic interaction between a user and a language model, the user asks a question, which is
sent as a prompt to the model. The model generates a response based solely on its pre-trained
knowledge. In a RAG process, there’s an additional step between the user’s question and the
model’s response. The user’s question triggers a retrieval process from a knowledge base.
721
The RAG process consists of these key steps:
1. Query Processing: When a user asks a question, the system converts it into an em-
bedding (a numerical representation).
2. Document Retrieval: The system searches a knowledge base for documents with
similar embeddings.
3. Context Enhancement: Relevant documents are retrieved and combined with the
original query.
4. Generation: The SLM generates a response using both the query and the retrieved
context.
722
• Loads the saved vector database
• Accepts user queries
• Retrieves relevant documents based on query similarity
• Combines documents with the query in a prompt
• Generates a response using the SLM
Instalation
Let’s examine how these components work together to implement a RAG system on edge
devices.
1. Document Processing
723
def create_vectorstore():
# Load documents from PDFs and URLs
docs_list = []
# [Document loading code]
# Persist to disk
[Link]()
This function processes our documents (chunk size of 300 with an overlap of 30), cre-
ating a searchable knowledge base. Notice we’re using OllamaEmbeddings with the
nomic-embed-text model, which can run efficiently on edge devices like the Raspberry Pi.
print(f"Question: {question}")
print("Retrieving documents...")
docs = [Link](question)
docs_content = "\n\n".join(doc.page_content for doc in docs)
print(f"Retrieved {len(docs)} document chunks")
print("Generating answer...")
724
client = Client()
rag_prompt = client.pull_prompt("rlm/rag-prompt")
end_time = [Link]()
latency = end_time - start_time
print(f"Response latency: {latency:.2f} seconds using model: {local_llm}")
return answer
This function retrieves relevant documents based on the query and combines them with a
specialized RAG prompt to generate a response. The RAG prompt is particularly important
as it tells the model how to use the context documents to answer the question.
1. SLM Integration
We’re using Ollama to run the SLM locally on our edge device, in this case using the 3B
parameter version of Llama 3.2.
1. Knowledge Extension: RAG allows small models to access knowledge beyond their
training data, effectively extending their capabilities without increasing model size.
2. Reduced Hallucination: By providing factual context, RAG significantly reduces the
likelihood of SLMs generating incorrect information.
3. Up-to-date Information: Unlike the fixed knowledge in a model’s weights, RAG
knowledge bases can be updated regularly with new information.
4. Domain Specialization: RAG can make general SLMs perform like domain specialists
by providing domain-specific knowledge bases.
725
5. Resource Efficiency: RAG allows smaller models (which require less memory and
computation) to achieve performance comparable to much larger models.
When implementing RAG on resource-constrained edge devices like the Raspberry Pi, consider
these optimizations:
1. Chunk Size: Smaller chunks (300-500 tokens) reduce memory usage during retrieval
and generation.
2. Retrieval Limits: Limit the number of retrieved documents (k=3 to 5) to reduce
context size.
3. Embedding Model Selection: Choose lightweight embedding models like
nomic-embed-text (137M parameters) or all-minilm (23M parameters).
4. Persistent Storage: As shown in our examples, using persistent storage prevents re-
computing embeddings every time that the RAG system is initiated.
5. Query Optimization: Implement query preprocessing to improve retrieval accuracy
while reducing computational load.
def optimize_query(query):
"""Optimize the query for better retrieval results"""
# Remove filler words, focus on key terms
stop_words = {"and", "or", "the", "a", "an", "in", "on", "at", "to",
"for", "with"}
terms = [term for term in [Link]().split() if term not in stop_words]
return " ".join(terms)
Building on our advanced weather station (see the chapter “Experimenting with SLMs for IoT
Control”), we can, for example, integrate RAG to provide more contextual responses about
weather conditions and historical patterns:
726
def weather_station_with_rag(retriever, model="llama3.2:3b"):
# Get current sensor readings
temp_dht, humidity, temp_bmp, pressure, button_state = collect_data()
727
# Create a prompt that combines current readings with retrieved context
prompt = f"""
Current Weather Station Data:
- Temperature (DHT22): {temp_dht:.1f}°C
- Humidity: {humidity:.1f}%
- Pressure: {pressure:.2f}hPa
Reference Information:
{context}
return [Link]
This function enhances our weather station by providing context-aware responses incorporating
current sensor readings and relevant information from our knowledge base. This is only an
example. To use it, we should have “Weather Reference Data,” which we do not currently
have. Instead, let’s create a general RAG system specializing in Edge AI Engineering.
For our RAG system, we will create a database with all chapters alheady written for the
EdgeAI Engineering book (chapters as URLs) and a PDF Wevolver 2025 Edge AI Technology
Report.
728
/image_classification.html",
"[Link]
"[Link]
/counting_objects_yolo.html",
"[Link]
"[Link]
"[Link]
/RPi_Physical_Computing.html",
"[Link]
]
Using the RAG system is straightforward. First, ensure you’ve created the vector database:
729
# Start the interactive query interface
python [Link]
Example interactions:
Generating answer...
ANSWER:
==================================================
EdgeAI refers to the application of artificial intelligence (AI) at the edge
of a network, typically in real-time applications such as IoT sensors, industrial
robots, and smart cameras. The Edge AI ecosystem includes edge devices, edge
servers, and cloud platforms that work together to enable low-latency AI
inferencing and processing of data on-site without relying on continuous cloud connectivit
in AI by leveraging energy-efficient, affordable, and scalable solutions for machine
learning and advanced edge computing.
==================================================
730
Those responses demonstrate how RAG enhances the SLM’s response with specific information
from our knowledge base about Edge AI applications on Raspberry Pi. One issue that should
be addressed is the latency.
To reduce latency, we can use for embedding, the all-minilm model which is much smaller
(23M parameters vs. 137M for nomic-embed-text) and creates 384-dimensional embeddings
instead of 768, significantly reducing computation time.
Also, smaller chunks can be helpful but have some disadvantages. For example, let’s say that
we can use a small chunk size (100 tokens with 50 overlap). Here are some considerations:
Advantages
1. Memory Efficiency: Smaller chunks require less memory during retrieval and process-
ing, which is beneficial for resource-constrained devices like the Raspberry Pi.
2. More Granular Retrieval: Smaller chunks can potentially provide more precise
matches to specific questions, especially for targeted queries about very specific details.
3. Reduced Context Window Usage: SLMs have limited context windows; smaller
chunks allow you to include more distinct pieces of information while staying within
these limits.
731
Disadvantages
1. Loss of Context: 100 tokens is approximately 75-80 words, which is often insufficient to
capture complete concepts or explanations. Many paragraphs and technical descriptions
require more space to convey their full meaning.
2. Increased Vector Store Size: More chunks mean more embeddings to store, poten-
tially increasing the overall size of your vector database.
3. Fragmented Information: With such small chunks, related information will be split
across multiple chunks, making it harder for the model to synthesize coherent answers.
A good practice would be to experiment with different chunk sizes and embedding models and
measure:
We can create a simple benchmarking function to have one embedding model defined test the
best chunk size:
def benchmark_chunk_sizes(document_list,
query_list,
sizes=[(100, 50), (300, 30), (500, 50), (1000, 100)]):
"""Test different chunk sizes and measure performance"""
results = {}
# Split documents
start_time = [Link]()
doc_splits = text_splitter.split_documents(document_list)
split_time = [Link]() - start_time
732
# Create embeddings and store
embedding_function = OllamaEmbeddings(model="nomic-embed-text")
temp_db_path = f"temp_db_{chunk_size}_{overlap}"
start_time = [Link]()
vectorstore = Chroma.from_documents(
documents=doc_splits,
collection_name="benchmark",
embedding=embedding_function,
persist_directory=temp_db_path
)
db_time = [Link]() - start_time
# Create retriever
retriever = vectorstore.as_retriever(k=3)
# Test queries
query_times = []
for query in query_list:
start_time = [Link]()
docs = [Link](query)
query_time = [Link]() - start_time
query_times.append(query_time)
# Store results
results[(chunk_size, overlap)] = {
"num_chunks": len(doc_splits),
"splitting_time": split_time,
"db_creation_time": db_time,
"avg_query_time": sum(query_times) / len(query_times),
"max_query_time": max(query_times),
"min_query_time": min(query_times)
}
# Clean up temporary DB
[Link](temp_db_path)
return results
Regarding the query side, some optimations can also reduce the latency at the edge. Let’s
modify the previous script, with:
1. Direct Ollama API Calls: Bypasses the LangChain abstraction layer for embedding
733
and LLM generation to reduce overhead.
2. Embedding Caching: Uses lru_cache to prevent recalculating embeddings for re-
peated queries.
3. Preloading Models: Initializes models at startup to avoid cold-start latency.
4. Optimized Retriever Settings: Uses minimal k-value (2) and adds a score threshold
to filter out irrelevant matches.
5. Reduced Dependency Usage: Removes unnecessary imports and simplifies the
pipeline.
6. Concurrent Processing: Uses ThreadPoolExecutor for batch document embedding
(when needed).
7. Early Termination: Checks for empty document results before running the LLM.
8. Simplified Prompt: Uses a more concise prompt template focused on getting direct
answers.
9. Fixed Seed: Uses a consistent seed for the LLM to reduce variability in response times.
The direct Ollama API approach removes several layers of abstraction in the
LangChain implementations.
We can see latency improvements from 2 minutes down to approximately 50-110 seconds,
depending on the complexity of the queries.
734
In the next section, we’ll explore how RAG can be combined with our agent architecture to
create even more powerful edge AI systems.
Note that any tool could be used here; the calculator is only a simple example to
demonstrate the concept.
735
When the user asks a question, the system first determines if it needs to use a tool or the RAG
approach. For knowledge queries, the RAG system enhances the response with information
from the database. The system then validates the answer quality, and if it’s not sufficient,
tries again with an improved prompt. In cases where questions fall outside the database’s
scope, the system will clearly inform the user rather than attempting to generate potentially
misleading answers.
System Architecture
736
The system functions through several key components:
1. Query Router
• Analyzes incoming queries to determine if they’re calculations or knowledge queries
• Can use the same model for the response generator or a lightweight model to reduce
overhead
• Implements rule-based fallbacks for robust classification
2. Document Retriever
• Connects to a persistent vector database (Chroma)
• Uses semantic embeddings to find relevant documents
• Returns contextually similar content for knowledge generation
3. Response Generator
• Creates answers based on retrieved documents
• Implements a two-stage approach with validation and improvement
• Adds appropriate disclaimers when information is insufficient
4. Validation Engine
• Evaluates answer quality using structured criteria
• Assigns a numerical score to each generated response
• Triggers enhancement processes when quality is insufficient
5. Interactive Interface
• Provides user-friendly interaction with clear quality indicators
• Supports model switching and verbosity control
• Offers guidance for improving query outcomes
Key Workflow
737
• Validate response quality
• If quality is low, attempt enhancement with improved prompt
• If still insufficient, add disclaimer about knowledge gaps
5. Return final answer with quality metrics
Query Routing
if is_calc:
# Use smaller, faster model for operation and number extraction
# ...extraction logic here...
return route_info
738
validation = validate_response(llm, query, answer)
validation_score = [Link]("score", 5)
{query}
{enhanced_context}
739
)
answer = answer + information_gap_note
740
741
Examples
Simple Calculation
742
First Pass Rag
743
Queries outside of the database scope:
Fine-tuning can adapt models to specific domains or tasks, improving performance for targeted
applications.
744
Preparing for Fine-Tuning
Fine-tuning on edge devices is typically impractical due to resource constraints. Instead, fine-
tune on a more powerful machine and deploy the result to the edge:
745
# This is a conceptual example
print(f"Fine-tuning model using data from {data_path}")
print(f"Fine-tuned model will be saved to {output_path}")
Supervised fine tuning (SFT) is a method to improve and customize pre-trained LLMs.
It involves retraining base models on a smaller dataset of instructions and answers. The
main goal is to transform a basic model that predicts text into an assistant that can follow
instructions and answer questions. SFT can also enhance the model’s overall performance,
add new knowledge, or adapt it to specific tasks and domains
Before considering SFT, it is recommended to try prompt engineering techniques like few-shot
746
prompting or retrieval augmented generation (RAG), as discussed previously. In practice,
these methods can solve many problems without fine-tuning. If this approach doesn’t meet
our objectives (regarding quality, cost, latency, etc.), then SFT becomes a viable option when
instruction data is available. SFT also offers benefits like additional control and customizability
to create personalized LLMs.
However, SFT has limitations. It works best when leveraging knowledge already present in
the base model. Learning completely new information, like an unknown language, can be
challenging and lead to more frequent hallucinations. For new domains unknown to the base
model, it is recommended that it be continuously pre-trained on a raw dataset first.
On the opposite end of the spectrum, instruct models (i.e., already fine-tuned models) can be
very close to our requirements. By providing chosen and rejected samples for a small set of
instructions (between 100 and 1000 samples), we can force the LLM to behave as we need.
The easiest way to finetune an SLM is by using Unsloth. The three most popular SFT tech-
niques are full fine-tuning, LoRA, and QLoRA.
For example, using this link, it is possible to find several notebooks with the steps to finetune
SLMs—for instance, the Gemma 3:1B.
The fine-tuned model can be saved on HF Hub or locally as GGUF, and to run a GGUF
model locally, we can use Ollama, as shown below:
cd ~
nano Modelfile
FROM ~/Downloads/[Link]
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 64
747
The model name here can be anything.
6. We can now use the model through as we have done with the llama3.2:3B in this chapter..
Conclusion
This chapter has explored comprehensive strategies for overcoming the inherent limitations of
Small Language Models in edge computing environments. By implementing techniques ranging
from optimized prompting strategies to sophisticated agent architectures and knowledge inte-
gration systems, we’ve demonstrated that it’s possible to significantly enhance the capabilities
of edge AI systems without requiring more powerful hardware or cloud connectivity.
The techniques presented—chain-of-thought prompting, task decomposition, function calling,
response validation, and RAG—form a toolkit that edge AI engineers can apply individually
or in combination to address specific challenges. Each approach offers unique advantages:
prompting techniques improve reasoning capabilities with minimal overhead, agent architec-
tures enable SLMs to perform actions beyond text generation, and RAG systems dramatically
expand an SLM’s knowledge without increasing model size.
Our practical implementations on the Raspberry Pi showcase that these enhancements are
not merely theoretical but can be deployed in real-world edge scenarios. From the simple
calculator agent to the more sophisticated knowledge router and RAG-enabled question an-
swering system, these examples provide templates that developers can adapt to their specific
application requirements.
The true power of these techniques emerges when they’re strategically combined. An agent
architecture with RAG capabilities, enhanced by chain-of-thought reasoning and validated
with a feedback loop, creates an edge AI system that approaches the capabilities of much
larger models while maintaining the advantages of edge deployment—privacy preservation,
reduced latency, and operation without internet connectivity.
As edge AI continues to evolve, these techniques will become increasingly important in bridging
the gap between the limited resources available on edge devices and the growing expectations
for AI capabilities. By thoughtfully applying these approaches, developers can create intelli-
gent systems that process data locally, respect user privacy, and operate reliably in diverse
environments.
The future of edge AI lies not necessarily in deploying ever-larger models but in developing
more innovative systems that combine efficient models with intelligent architectures, contextual
knowledge integration, and robust validation mechanisms. By mastering these techniques, edge
AI practitioners can create solutions that are not just technologically impressive but genuinely
useful and trustworthy in addressing real-world challenges.
748
Resources
The scripts used in this chapter can be found here: Advancing EdgeAI Scripts
#
Weekly Labs
749
Edge AI Engineering - Weekly Labs
Objectives:
Instructions:
Objectives:
Instructions:
750
2. Configure Jupyter Notebook for remote access:
Deliverable: A simple Python script that captures and displays an image from your camera
and a Screenshot showing a successful image capture
Objectives:
Instructions:
wget [Link]
751
Deliverable: Python script that successfully classifies sample images with MobileNet V2 and
a Screenshot showing a successful result
Objectives:
Instructions:
Deliverable: Structured dataset with at least 3 classes and 50 images per class
Objectives:
Instructions:
752
• Set image size to 160x160
• Use Transfer Learning for feature extraction
4. Generate features for all images
5. Train model using MobileNet V2
6. Analyze model performance (accuracy, confusion matrix)
7. Test model on validation data
Deliverable: Edge Impulse project link and screenshot of model performance metrics
Objectives:
Instructions:
Deliverable: Python script for real-time image classification with your custom model and a
Screenshot showing a successful result
Objectives:
753
• Understand object detection architecture
• Run pre-trained SSD-MobileNet model
• Process detection outputs
• Visualize detected objects
Instructions:
Deliverable: Python script that performs and visualizes object detection on test images and
a Screenshot showing a successful result.
Objectives:
Instructions:
754
Week 5: Custom Object Detection
Objectives:
Instructions:
Deliverable: Annotated dataset with at least 2 object classes and 100 total images
Objectives:
Instructions:
Deliverable: Edge Impulse project link with trained object detection model and performance
metrics
755
Week 6: Advanced Object Detection
Objectives:
Instructions:
Deliverable: Python application that compares SSD MobileNet vs. FOMO performance in
real-time
Objectives:
Instructions:
Deliverable: Python application for real-time object detection and counting using YOLO
756
Week 7: Object Counting Project
Objectives:
Instructions:
Objectives:
Instructions:
757
Deliverable: Integrated application combining multiple AI capabilities with visualization
dashboard
Objectives:
Instructions:
3. Install dependencies:
Deliverable: Screenshot showing system configuration with increased swap and temperature
monitor during stress test
Objectives:
758
Instructions:
1. Install Ollama:
• Load time
• Inference speed (tokens/sec)
• Memory usage
• Temperature
Objectives:
Instructions:
759
• Handle conversation context
3. Implement proper error handling
4. Create a simple interactive CLI application
Deliverable: Python script demonstrating Ollama library usage with conversation handling
Objectives:
Instructions:
Deliverable: Python application that uses function calling for structured interaction with
SLMs
760
Week 10: Retrieval-Augmented Generation
Objectives:
Instructions:
Deliverable: Python implementation of a basic RAG system with simple text documents
Objectives:
Instructions:
761
Deliverable: Optimized RAG implementation with specialized knowledge base and perfor-
mance analysis
Objectives:
Instructions:
Objectives:
Instructions:
762
1. Implement image captioning:
• Basic caption generation
• Detailed caption generation
2. Implement object detection:
• Bounding box visualization
• Multiple object detection
3. Implement visual grounding:
• Highlight specific objects based on text prompts
4. Create segmentation application
5. Measure the performance of each task
Deliverable: Python application demonstrating multiple vision tasks with Florence-2 and
performance analysis
Objectives:
Instructions:
763
3. Create a Python script to read sensor data
4. Implement LED control based on conditions
5. Create visualization of sensor data
Deliverable: Python application for reading sensor data and controlling actuators with visu-
alization
Objectives:
Instructions:
Deliverable: Jupyter Notebook with interactive widgets for sensor monitoring and actuator
control
Objectives:
Instructions:
764
1. Create a Python application that:
• Collects sensor data
• Formats data for SLM prompt
• Sends prompt to model
• Parses response
• Controls actuators based on response
2. Implement multiple analysis modes
3. Add error handling for SLM responses
Deliverable: Python application integrating SLMs with physical sensors and actuators
Objectives:
Instructions:
Deliverable: Complete IoT monitoring system with SLM integration and web interface
765
Week 14: Advanced Edge AI Techniques
Objectives:
Instructions:
Deliverable: Python implementation of agent architecture with tool usage and decision rout-
ing
Objectives:
Instructions:
766
5. Implement validation mechanisms
6. Compare the effectiveness of different strategies
Objectives:
Instructions:
Deliverable: Complete agentic RAG system with documentation and performance analysis
Objectives:
767
Instructions:
Hardware Requirements
• Raspberry Pi Zero 2W or Pi 5
• MicroSD card (32GB+)
• Camera module (USB webcam or Pi camera)
• Power supply
768
• Breadboard
Software Requirements
Development Environment
• Raspberry Pi OS (64-bit)
• Python 3.9+
• Jupyter Notebook
• SSH client
Generative AI
• Ollama
• Transformers
• Pytorch
• ChromaDB
• LangChain
• Pydantic
• Instructor
Physical Computing
• GPIO Zero
• Adafruit CircuitPython libraries
769
Assessment Criteria
1. Start Early: These labs build on each other. Falling behind makes later labs more
difficult.
2. Document As You Go: Take notes, screenshots, and document issues/solutions.
3. Optimize Resources: SLMs and VLMs require careful resource management.
4. Collaborate: Discuss approaches with classmates while ensuring individual work.
5. Backup Regularly: Create backups of your SD card after significant progress.
6. Measure Performance: Always benchmark and optimize your implementations.
7. Ask Questions: If you’re stuck, ask for help early rather than falling behind.
#
References & Author
770
References
To learn more:
Online Courses
• Harvard School of Engineering and Applied Sciences - CS249r: Tiny Machine Learning
• Professional Certificate in Tiny Machine Learning (TinyML) – edX/Harvard
• Introduction to Embedded Machine Learning - Coursera/Edge Impulse
• Computer Vision with Embedded Machine Learning - Coursera/Edge Impulse
• UNIFEI-IESTI01 TinyML: “Machine Learning for Embedding Devices”
Books
Projects Repository
771
TinyML4D
TinyML Made Easy, an eBook collection of a series of Hands-On tutorials, is part of the
TinyML4D, an initiative to make Embedded Machine Learning (TinyML) education avail-
able to everyone, explicitly enabling innovative solutions for the unique challenges Developing
Countries face.
772
About the author
Marcelo Rovai, a Brazilian living in Chile, is a recognized figure in engineering and technology
education. He holds the title of Professor Honoris Causa from the Federal University of Itajubá
(UNIFEI), Brazil. His educational background includes an Engineering degree from UNIFEI
and a specialization from the Polytechnic School of São Paulo University (POLI/USP). Further
773
enhancing his expertise, he earned an MBA from IBMEC (INSPER) and a Master’s in Data
Science from the Universidad del Desarrollo (UDD) in Chile.
With a career spanning several high-profile technology companies, including AVIBRAS
Airspace, AT&T, NCR, and IGT, where he served as Vice President for Latin Amer-
ica, he brings industry experience to his academic endeavors. He is a prolific writer on
electronics-related topics and shares his knowledge through open platforms like [Link].
In addition to his professional pursuits, he is dedicated to educational outreach, serving as
a volunteer professor at UNIFEI and engaging with the TinyML4D group and the EDGE
AIP– the Academia-Industry Partnership of EDGEAI Foundation as a Co-Chair, promoting
TinyML education in developing countries. His work underscores a commitment to leveraging
technology for societal advancement.
LinkedIn profile: [Link]
Lectures, books, papers, and tutorials: [Link]
774