Edge AI Engineering
Edge AI Engineering
Edge AI Engineering
Hands-on with the Raspberry Pi
Marcelo Rovai
2026-05-24
Table of contents
Preface 3
Acknowledgments 5
Introduction 6
Edge AI Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
Why Edge AI Matters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
The Raspberry Pi Advantage . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
What You’ll Learn . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Who This Book Is For . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Classification of AI Applications 11
Fixed Function AI vs. Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Fixed Function AI (Reactive) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
Generative AI (Proactive) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
Summary Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
The Edge AI Advantage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
Setup 15
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
Raspberry Pi Models (covered in this book) . . . . . . . . . . . . . . . . . . . . 17
Engineering Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
Hardware Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Raspberry Pi Zero 2W . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
Installing the Operating System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
The Operating System (OS) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
Installation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
Initial Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2
Remote Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
SSH Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
To shut down the Raspi via terminal: . . . . . . . . . . . . . . . . . . . . . . . . 26
Transfer Files between the Raspberry Pi and a computer . . . . . . . . . . . . . 26
Increasing Memory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
Increasing zram for the Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . 32
Increasing memory for the Rasp-Zero . . . . . . . . . . . . . . . . . . . . . . . . 33
Notes specific to Zero 2 W Trixie . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Installing a Camera . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Installing a Camera Module on the CSI port . . . . . . . . . . . . . . . . . . . 35
Installing a USB WebCam . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
Running the Raspi Desktop remotely . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
Updating and Installing Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
Model-Specific Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Raspberry Pi Zero (Raspi-Zero) . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Raspberry Pi 4 or 5 (Raspi-4 or Raspi-5) . . . . . . . . . . . . . . . . . . . . . . 52
Measuring Temperature and Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
Check CPU Temperature . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
Understand Power Measurement on Pi 5 . . . . . . . . . . . . . . . . . . . . . . 53
Script: Average Temperature and Power . . . . . . . . . . . . . . . . . . . . . . 54
Run a measurement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
Interpreting the Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
3
Custom Image Classification Project 88
Image Classification Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
The Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
Data Collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
Training the model with Edge Impulse Studio . . . . . . . . . . . . . . . . . . . . . . 98
Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
The Impulse Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
Image Pre-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
Model Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
Model Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
Trading off: Accuracy versus speed . . . . . . . . . . . . . . . . . . . . . . . . . 104
Model Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Live Image Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
Summary: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
4
Impulse Design, new Training and Testing . . . . . . . . . . . . . . . . . . . . . 185
Deploying the model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 188
Inference and Post-Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 190
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
5
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 259
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 260
6
Camera Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303
Image Capture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303
Performance Benchmarking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 306
Basic (Float32): mobilenet_v2.pte . . . . . . . . . . . . . . . . . . . . . . . . . 308
XNNPACK Backend (Flot32): mobilenet_v2_xnnpack.pte . . . . . . . . . . . 309
Quantization (INT8): mobilenet_v2_quantized_xnnpack.pte . . . . . . . . . . 310
Performance Comparison Table . . . . . . . . . . . . . . . . . . . . . . . . . . . 311
Exploring Custom Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 312
Exporting a Custom Trained Model . . . . . . . . . . . . . . . . . . . . . . . . 313
Running Custom Models on Raspberry Pi . . . . . . . . . . . . . . . . . . . . . 314
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317
Key Takeaways . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318
Performance Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Code Repository . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Official Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Books . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
7
Run Inference on MemryX Accelerator (MXA) . . . . . . . . . . . . . . . . . . 341
Decode the MXA Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 342
Comparing CPU vs. MXA Performance . . . . . . . . . . . . . . . . . . . . . . 343
Measuring Latency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 343
Testing with Larger Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 345
Clean Shutdown . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Folders Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Performance Comparison Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346
Key Observations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
When to Use the MX3? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
YOLOv8 Object Detection with MX3 Hardware Acceleration . . . . . . . . . . . . . 347
Model Export and Compilation . . . . . . . . . . . . . . . . . . . . . . . . . . . 349
Understanding YOLOv8 Output Format . . . . . . . . . . . . . . . . . . . . . . 352
Complete Inference Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 352
Going deeper in the functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 355
Making Inferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 358
Inference with a custom model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 360
Adjusting Confidence Threshold . . . . . . . . . . . . . . . . . . . . . . . . . . 364
Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Batch Processing (Optimization) . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Thermal Management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Confidence Threshold Tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Exploring MemryX eXamples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 365
Clone the MemryX eXamples Repository . . . . . . . . . . . . . . . . . . . . . 366
Troubleshooting Common Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 366
Device Not Detected . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 366
Compilation Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 368
Thermal Throttling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369
Python Version Conflicts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369
Low FPS / Poor Performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . 370
Import Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 370
Model Accuracy Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Next Steps and Extensions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Project Ideas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Advanced Topics to Explore . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
References and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Official Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Code and Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Background Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
Community and Support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
8
Text Generation with RNNs 374
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
What Are We Actually Building? . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Neural Network Architectures Background . . . . . . . . . . . . . . . . . . . . . . . . 375
The Human Brain Analogy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Recurrent Neural Networks (RNN) . . . . . . . . . . . . . . . . . . . . . . . . . 375
The Memory Problem and GRU Solution . . . . . . . . . . . . . . . . . . . . . 377
Dataset Preparation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 378
Data Preprocessing Steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 379
Tokenization and Vocabulary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
Character-Level Tokenization . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
Vocabulary Building Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
Creating the Character Dictionary . . . . . . . . . . . . . . . . . . . . . . . . . 381
Training Sequences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 382
The Sliding Window Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . 382
Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Creating Training Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
Character Embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 384
From Sparse to Dense Representation . . . . . . . . . . . . . . . . . . . . . . . 384
Learning Character Relationships . . . . . . . . . . . . . . . . . . . . . . . . . . 384
Visualization and Understanding . . . . . . . . . . . . . . . . . . . . . . . . . . 385
Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
RNN Architecture Components . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
Model Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386
Memory and Processing Flow . . . . . . . . . . . . . . . . . . . . . . . . . . . . 387
Why GRU over Basic RNN? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Training Process: Teaching the Model to Write . . . . . . . . . . . . . . . . . . . . . 388
The Learning Objective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Hardware and Time Requirements . . . . . . . . . . . . . . . . . . . . . . . . . 388
Monitoring Progress . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 388
Preventing Overfitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 389
Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 389
Training Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
Text Generation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
The Generation Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 390
Temperature Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
Generation Example (Temperature = 0.5) . . . . . . . . . . . . . . . . . . . . . 392
Example Output Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 392
Challenges and Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
Context Window Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
Character vs. Word Level Trade-offs . . . . . . . . . . . . . . . . . . . . . . . . 393
Coherence Challenges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 393
9
Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Potential Improvements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Connecting to Modern Language Models . . . . . . . . . . . . . . . . . . . . . . . . . 394
Scale Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Architectural Evolution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
Training Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
Summary: Why Modern Models Perform Better? . . . . . . . . . . . . . . . . . 395
Conclusion: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 396
Resourses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
10
Final Thoughts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416
11
Best Practices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 482
Function Calling Solution for Calculations . . . . . . . . . . . . . . . . . . . . . . . . 483
Define the Tool (Function Schema) . . . . . . . . . . . . . . . . . . . . . . . . . 483
Implement the Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 483
Project: Calculating Distances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 484
Running with other models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491
Adding images . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491
Retrievel Augmentation Generation (RAG) . . . . . . . . . . . . . . . . . . . . . . . 496
A simple RAG project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 497
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 504
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 504
12
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 548
Key Advantages of Florence-2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . 548
Trade-offs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 549
Best Use Cases . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 549
Future Implications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 550
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 550
13
Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 578
Accessing the GPIOs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 578
Pin Numbering Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
Safety First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
GPIO Zero Library . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 580
“Hello World”: Blinking an LED . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 581
Understanding the LED Circuit . . . . . . . . . . . . . . . . . . . . . . . . . . . 581
Testing with GPIO Zero . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 582
Installing all LEDs (the “actuators”) . . . . . . . . . . . . . . . . . . . . . . . . 584
Sensors Installation and setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 586
Button . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 586
Installing Adafruit CircuitPython . . . . . . . . . . . . . . . . . . . . . . . . . . 590
DHT22 - Temperature & Humidity Sensor . . . . . . . . . . . . . . . . . . . . . 592
Addendum: DHT22 on Raspberry Pi 5 with Debian Trixie . . . . . . . . . . . . 597
Installing the BMP280: Barometric Pressure & Altitude Sensor . . . . . . . . . 600
Measuring Weather and Altitude With BMP280 . . . . . . . . . . . . . . . . . 605
Playing with Sensors and Actuators . . . . . . . . . . . . . . . . . . . . . . . . . . . 608
Testing the Notebook setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
Initialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
GPIO Input and Output . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 611
Getting and displaying Sensor Data . . . . . . . . . . . . . . . . . . . . . . . . 615
Widgets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 617
Advanced GPIO Zero Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 618
PWM LED Brightness Control . . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Composite Devices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Other Useful GPIO Zero Devices . . . . . . . . . . . . . . . . . . . . . . . . . . 619
Interacting an SLM with the Physical world . . . . . . . . . . . . . . . . . . . . . . . 620
Other Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 627
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Key Achievements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Technical Insights . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Practical Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
Challenges and Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Future Enhancements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Final Thoughts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 630
14
SLM Basic Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 638
Act on Output (Actuators) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 642
Creating new functions for the Actuation: . . . . . . . . . . . . . . . . . . . . . 643
Prompting Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 646
Key Changes in the code: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 646
System Overview: Enhanced IoT Environmental Monitoring with SLM Control 647
The new code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 651
Code Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 654
TEST1: Temp above the threshold . . . . . . . . . . . . . . . . . . . . . . . . . 654
TEST2: Temp below the threshold . . . . . . . . . . . . . . . . . . . . . . . . . 656
TEST3: Alarm Button pressed . . . . . . . . . . . . . . . . . . . . . . . . . . . 657
Interacting with IoT Systems, using Natural Language Commands . . . . . . . . . . 657
What will change? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 658
How It will Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 659
Key Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 661
Running the System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 662
Special Commands . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Languages . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Advantages of This Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Error Handling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Tips for Best Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665
Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 665
Adding Data Logging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
Key Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 666
Running the System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 667
Example Queries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 667
Data Files . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
sensor_readings.csv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
command_history.csv . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
Tips . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
Flow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 669
Examples: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 670
Prompt Optimization and Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . 671
Quick Solution: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 672
New Code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 674
Key Changes Under the Hood: . . . . . . . . . . . . . . . . . . . . . . . . . . . 674
Using Pydantic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 675
2. Implementing Pydantic in the Code . . . . . . . . . . . . . . . . . . . . . . . 676
Next Steps . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 679
Conclusion 681
Resourses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 682
15
Advancing EdgeAI: Beyond Basic SLMs 683
Understanding SLM Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 684
1. Knowledge Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 684
2. Reasoning Limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
3. Inconsistent Outputs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
4. Domain Specialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
Techniques for Enhancing SLM at the Edge . . . . . . . . . . . . . . . . . . . . . . . 686
Optimizing Prompting Strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Chain-of-Thought Prompting . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Few-Shot Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Task Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 688
Building Agents with SLMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 689
General Knowledge Router . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 700
Improving Agent Reliability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703
1. Function Calling with Pydantic . . . . . . . . . . . . . . . . . . . . . . . . . 703
2. Response Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 705
Retrieval-Augmented Generation (RAG) . . . . . . . . . . . . . . . . . . . . . . . . . 707
Understanding RAG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 707
Implementing a Naive RAG System . . . . . . . . . . . . . . . . . . . . . . . . 708
Instalation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 709
Key Components of the Naive RAG System . . . . . . . . . . . . . . . . . . . . 709
Advantages of RAG for Edge AI . . . . . . . . . . . . . . . . . . . . . . . . . . 711
Optimizing RAG for Edge Devices . . . . . . . . . . . . . . . . . . . . . . . . . 712
Application: Enhanced Weather Station with RAG . . . . . . . . . . . . . . . . 712
Using the RAG System for Edge AI Engineering . . . . . . . . . . . . . . . . . 714
Testing Different Models and Chunk Sizes . . . . . . . . . . . . . . . . . . . . . 718
Advanced Agentic RAG System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 721
System Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 722
Key Workflow . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 723
Important Code Sections . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 724
Detailed Workflow Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 726
Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 728
Fine-Tuning SLMs for Edge Deployment . . . . . . . . . . . . . . . . . . . . . . . . . 730
Preparing for Fine-Tuning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 731
Setting Up a Fine-Tuning Process . . . . . . . . . . . . . . . . . . . . . . . . . 731
Real implementation: Supervised Fine-Tuning (SFT) . . . . . . . . . . . . . . . 732
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 734
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 734
16
Architecture Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 739
The .litertlm Format . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
Supported Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
LiteRT-LM vs Ollama: When to Use Each . . . . . . . . . . . . . . . . . . . . . . . . 741
Installation on Raspberry Pi 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Prerequisites . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Creating a Virtual Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . 742
Installing the Package . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 743
Installing the Hugging Face CLI (for model downloads) . . . . . . . . . . . . . 743
Downloading a Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 744
Quick Sanity Check with the CLI . . . . . . . . . . . . . . . . . . . . . . . . . . 744
Python API . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
The Engine and Conversation Objects . . . . . . . . . . . . . . . . . . . . . . . 745
1. Simple Single-Turn Query . . . . . . . . . . . . . . . . . . . . . . . . . . . . 746
2. Interactive Chat with Persistent History . . . . . . . . . . . . . . . . . . . . 747
3. Tool Calling (Function Calling) . . . . . . . . . . . . . . . . . . . . . . . . . 749
4. Streaming with Performance Metrics . . . . . . . . . . . . . . . . . . . . . . 752
Complete Example: Putting it All Together . . . . . . . . . . . . . . . . . . . . . . . 755
Going Further . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
17
Lab 14: Fixed-Function AI Integration (Optional) . . . . . . . . . . . . . . . . 768
Week 8: Introduction to Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . 769
Lab 15: Raspberry Pi Configuration for SLMs . . . . . . . . . . . . . . . . . . . 769
Lab 16: Ollama Installation and Testing . . . . . . . . . . . . . . . . . . . . . . 769
Week 9: SLM Python Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 770
Lab 17: Ollama Python Library . . . . . . . . . . . . . . . . . . . . . . . . . . . 770
Lab 18: Function Calling and Structured Outputs . . . . . . . . . . . . . . . . 771
Week 10: Retrieval-Augmented Generation . . . . . . . . . . . . . . . . . . . . . . . 772
Lab 19: RAG Fundamentals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 772
Lab 20: Advanced RAG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 772
Week 11: Vision-Language Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 773
Lab 21: Florence-2 Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 773
Lab 22: Vision Tasks with Florence-2 . . . . . . . . . . . . . . . . . . . . . . . 773
Week 12: Physical Computing Basics . . . . . . . . . . . . . . . . . . . . . . . . . . . 774
Lab 23: Sensor and Actuator Integration . . . . . . . . . . . . . . . . . . . . . . 774
Lab 24: Jupyter Notebook Integration . . . . . . . . . . . . . . . . . . . . . . . 775
Week 13: SLM-Physical Computing Integration . . . . . . . . . . . . . . . . . . . . . 775
Lab 25: Basic SLM Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 775
Lab 26: SLM-IoT Control System . . . . . . . . . . . . . . . . . . . . . . . . . . 776
Week 14: Advanced Edge AI Techniques . . . . . . . . . . . . . . . . . . . . . . . . . 777
Lab 27: Building Agents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 777
Lab 28: Advanced Prompting and Validation . . . . . . . . . . . . . . . . . . . 777
Week 15: Final Project Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778
Lab 29: Agentic RAG System . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778
Lab 30: Final Project . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 778
Hardware Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 779
Basic Setup (Weeks 1-7) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 779
Generative AI (Weeks 8-15) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 779
Physical Computing (Weeks 12-15) . . . . . . . . . . . . . . . . . . . . . . . . . 779
Software Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Development Environment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Computer Vision and DL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Generative AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Physical Computing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 780
Assessment Criteria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 781
Tips for Success . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 781
References 782
To learn more: . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
Online Courses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
Books . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
Projects Repository . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 782
18
AiEng4D 783
19
Preface
In the rapidly evolving technology landscape, the convergence of artificial intelligence and edge
computing is one of the most exciting frontiers. This intersection promises to revolutionize how
we interact with the world around us, bringing intelligence and decision-making capabilities
directly to the devices we use every day. At the heart of this revolution lies the Raspberry Pi,
a powerful yet accessible single-board computer (SBC) that has democratized computing and
now stands poised to do the same for edge AI.
This book, which serves as the official textbook for IESTI05 Edge AI Engineering at the
Federal University of Itajubá (UNIFEI) in Brazil, embodies both a passion for technology
and a conviction in its capacity to address real-world problems. While developed to support
UNIFEI’s engineering curriculum, the content is valuable for all learners, whether in academic
settings or pursuing independent study.
“Edge AI Engineering: Hands-on with the Raspberry Pi” is not just about theory or abstract
concepts. It’s about getting your hands dirty, writing code, training models, and seeing your
creations come to life. Each chapter blends foundational knowledge with practical application,
focusing on what’s possible with the Raspberry Pi platform.
From the compact Raspberry Pi Zero to the more powerful Pi 5, we explore how these incredible
devices can become the brains of intelligent systems—recognizing images, understanding speech,
detecting objects, and even running small language models. Each project serves as a stepping
stone, building your skills and confidence as you progress.
Beyond the technical skills, this book aims to instill something more valuable – a sense of
curiosity and possibility. The field of edge AI is still in its infancy, with new applications
and techniques emerging daily. By mastering the fundamentals presented here, you’ll be
well-equipped to explore these frontiers, perhaps even pushing the boundaries of what’s possible
on edge devices.
Whether you’re a student seeking to understand AI’s practical applications, a professional
looking to expand your skill set, or an enthusiast eager to add intelligence to your projects, we
hope this book serves as both a guide and an inspiration.
As you embark on this journey, remember that every expert was once a beginner. The learning
path is filled with challenges and moments of joy and discovery. Embrace both, and let your
creativity guide you.
20
Thank you for joining us on this exciting adventure into edge machine learning. Let’s begin
exploring what’s possible when we bring AI to the edge, one Raspberry Pi at a time.
Happy coding, and may your models always converge!
Prof. Marcelo Rovai
May, 2026
21
Acknowledgments
I extend my deepest gratitude to the entire AiEng4D Academic Network, comprised of distin-
guished professors, researchers, and professionals. Notable contributions from Marco Zennaro,
Ermanno Petrosemoli, Brian Plancher, José Alberto Ferreira, Jesus Lopez, Diego Mendez,
Shawn Hymel, Dan Situnayake, Pete Warden, and Laurence Moroney have been instrumental
in advancing our understanding of Embedded Machine Learning (TinyML) and Edge AI.
Special commendation is reserved for Professor Vijay Janapa Reddi of Harvard University. His
steadfast belief in the transformative potential of open-source communities, coupled with his
invaluable guidance and teachings, has been a beacon and a cornerstone of our efforts from the
beginning.
Acknowledging these individuals, we pay tribute to the collective wisdom and dedication that
have enriched this field and our work.
Google Nano Banana and OpenAI’s GPT were used to generate some of the images
in the book. Claude Sonnet and Perplexity helped with code and text reviews.
22
Introduction
Edge AI Engineering
Traditional AI deployment often relies on cloud infrastructure, which requires constant con-
nectivity and introduces latency. Edge AI addresses these limitations by bringing intelligence
directly to where data is generated and actions occur. This approach offers several compelling
advantages:
The Raspberry Pi, with its combination of affordability, processing capability, and extensive
GPIO options, provides an ideal platform for exploring Edge AI concepts. From the compact
Raspberry Pi Zero 2W to the more powerful Pi 5, these devices offer:
23
• Sufficient computational power for running optimized AI models
• A complete Linux-based operating system for straightforward development
• Extensive connectivity options for integrating with sensors and actuators
• A vibrant community and ecosystem of libraries and tools
• An accessible entry point for students, hobbyists, and professionals alike
This book takes a progressive approach to Edge AI engineering, starting with foundational
concepts and building toward more advanced applications:
1. Essential setup and configuration: Prepare your Raspberry Pi for Edge AI develop-
ment
2. Computer vision applications: Implement image classification and object detection
systems
3. Small Language Models (SLMs): Run and optimize language models directly on
your Raspberry Pi
4. Vision-Language Models: Explore multimodal AI with Florence-2
5. Physical computing integration: Connect AI systems with sensors and actuators
6. Advanced optimization techniques: Enhance model performance through methods
like RAG, agents, and function calling
Each chapter includes detailed explanations, step-by-step instructions, and practical projects
demonstrating real-world applications of Edge AI concepts.
Whether you’re a student exploring AI for the first time, an educator developing a curriculum,
a maker building innovative projects, or a professional seeking to expand your skills, this book
provides the knowledge and hands-on experience needed to successfully implement Edge AI
solutions on the Raspberry Pi platform.
Join us on this journey to the edge of AI innovation, where we’ll bridge theory and practice
through engaging, accessible projects that demonstrate the transformative potential of intelligent
edge computing.
24
About this Book
Several chapters (Labs) in this book also accompany the open-source book Machine Learning
Systems by Professor Vijay Janapa Reddi from Harvard, which we invite you to read.
“Edge AI Engineering: Hands-on with the Raspberry Pi” is designed as a practical, project-
based learning resource that bridges theoretical AI concepts with tangible implementations.
This book is part of the open-source Machine Learning Systems initiative, democratizing access
to AI education and applications.
Key Features
1. Progressive Learning Path: The book structure follows a natural progression from
basic to advanced concepts, beginning with foundational computer vision applications
and advancing to generative AI techniques.
25
2. Model-Specific Optimizations: Each chapter provides targeted guidance for specific
Raspberry Pi models, helping you maximize performance whether you’re using a Pi Zero
2W or a Pi 5.
3. Open-Source Foundation: We emphasize accessible tools and frameworks, including
Edge Impulse Studio, TensorFlow Lite, PyTorch, Transformers, and Ollama, ensuring
you can continue your learning journey with widely available resources.
4. Practical Problem-Solving: Rather than abstract exercises, each project addresses
real-world challenges that demonstrate Edge AI’s practical value.
5. Resource Optimization Techniques: Learn essential strategies for deploying AI on
resource-constrained devices, balancing performance needs with hardware limitations.
6. Cross-Domain Applications: Explore implementations spanning computer vision,
natural language processing, and physical computing, showcasing the versatility of Edge
AI.
Prerequisites
26
By completing this book, you’ll possess the skills to design, implement, and optimize Edge
AI applications across a wide range of use cases, leveraging the unique capabilities of the
Raspberry Pi platform to bring intelligence to the edge.
27
Classification of AI Applications
As we embark on our journey through Edge AI Engineering with the Raspberry Pi, it’s essential
to understand the fundamental classification of AI applications that form the structure of this
book. Our exploration is divided into two parts, each representing a different paradigm in
artificial intelligence implementation.
AI applications can be broadly categorized into two approaches that represent different capa-
bilities, interaction models, and implementation strategies:
Fixed Function AI, or Reactive AI, operates by analyzing specific inputs according to predeter-
mined patterns and rules and then producing consistent outputs for given scenarios. These
systems:
• Respond to specific triggers: They activate only when presented with particular
inputs.
• Follow defined patterns: Their behavior is predictable and consistent.
• Excel at structured tasks: They perform exceptionally well at classification, detection,
and pattern recognition
• Operate within boundaries: Their capabilities are limited to their specific program-
ming.
In the first part of this book (Chapters 2-4), we explore fixed-function AI through computer
vision applications:
These applications demonstrate how edge devices can deliver reliable, efficient AI in constrained
environments, focusing on specific, well-defined tasks.
28
Generative AI (Proactive)
Generative AI, also known as Proactive AI, represents a fundamental shift in capability. These
systems can:
The second part of this book (Chapters 5-9) explores Generative AI at the edge:
This progression from Fixed Function to Generative AI mirrors the evolution of artificial
intelligence itself—from specialized systems designed for specific tasks to more flexible, creative
systems capable of addressing a broader range of challenges.
Summary Table
Conclusion
29
The Edge AI Advantage
Both Fixed Function and Generative AI gain unique benefits when deployed at the edge:
30
31
Setup
32
This chapter will guide you through setting up the Raspberry Pi Zero 2 W (Raspi-Zero) and the
Raspberry Pi 5 (Raspi-5) models. We’ll cover hardware setup, operating system installation,
initial configuration, and tests.
The general instructions for the Raspi-5 also apply to the older Raspberry Pi
versions, such as the Raspi-3 and Raspi-4.
Introduction
The Raspberry Pi is a powerful and versatile single-board computer that has become an essential
tool for engineers across various disciplines. Developed by the Raspberry Pi Foundation, these
compact devices offer a unique combination of affordability, computational power, and extensive
GPIO (General Purpose Input/Output) capabilities, making them ideal for prototyping,
embedded systems development, and advanced engineering projects.
Key Features
1. Computational Power: Despite their small size, Raspberry Pis offer significant pro-
cessing capabilities, with the latest models featuring multi-core ARM processors and up
to 8GB of RAM.
2. GPIO Interface: The 40-pin GPIO header enables direct interaction with sensors,
actuators, and other electronic components, facilitating hardware-software integration.
3. Extensive Connectivity: Built-in Wi-Fi, Bluetooth, Ethernet, and multiple USB ports
enable a wide range of communication and networking projects.
4. Low-Level Hardware Access: Raspberry Pis provide access to interfaces such as I2C,
SPI, and UART, enabling detailed control and communication with external devices.
5. Real-Time Capabilities: With proper configuration, Raspberry Pis can be used for soft
real-time applications, making them suitable for control systems and signal processing
tasks.
6. Power Efficiency: Low power consumption enables battery-powered and energy-efficient
designs, especially in models like the Pi Zero.
33
Raspberry Pi Models (covered in this book)
Engineering Applications
1. Embedded Systems Design: Develop and prototype embedded systems for real-world
applications.
2. IoT and Networked Devices: Create interconnected devices and explore protocols
like MQTT, CoAP, and HTTP/HTTPS.
3. Control Systems: Implement feedback control loops, PID controllers, and interface
with actuators.
4. Computer Vision and AI: Utilize libraries like OpenCV and TensorFlow Lite for
image processing and machine learning at the edge.
5. Data Acquisition and Analysis: Collect sensor data, perform real-time analysis, and
create data logging systems.
6. Robotics: Build robot controllers, implement motion planning algorithms, and interface
with motor drivers.
7. Signal Processing: Perform real-time signal analysis, filtering, and DSP applications.
8. Network Security: Set up VPNs, firewalls, and explore network penetration testing.
This lab will guide you through setting up the most common Raspberry Pi models, so you can
get started on your machine learning project quickly. We’ll cover hardware setup, operating
system installation, and initial configuration, focusing on preparing your Pi for Machine
Learning applications.
34
Hardware Overview
Raspberry Pi Zero 2W
35
Raspberry Pi 5
• Processor:
– Pi 5: Quad-core 64-bit Arm Cortex-A76 CPU @ 2.4GHz
– Pi 4: Quad-core Cortex-A72 (ARM v8) 64-bit SoC @ 1.5GHz
• RAM: 2GB, 4GB, or 8GB options (8GB recommended for AI tasks)
• Wireless: Dual-band 802.11ac wireless (2.4 GHz and 5 GHz), Bluetooth 5.0
• Ports: 2 × micro HDMI ports, 2 × USB 3.0 ports, 2 × USB 2.0 ports, CSI camera port,
DSI display port
• Power: 5V/5A, 5V/3A limits peripherals to 600mA, via USB-C connector 27W USB-C
power supply
In the labs, we will use different names to address the Raspberry Pi: Raspi, Raspi-5,
Raspi-Zero, etc. Usually, “Raspi” or “Raspberry Pi” is used when the instructions
or comments apply to all models.
36
Installing the Operating System
An operating system (OS) is essential software that manages computer hardware and software
resources, providing standard services for computer programs. It is the core software that runs
on a computer, serving as an intermediary between hardware and application software. The
OS oversees the computer’s memory, processes, device drivers, files, and security protocols.
1. Key functions:
• Process management: Allocating CPU time to different programs
• Memory management: Allocating and freeing up memory as needed
• File system management: Organizing and keeping track of files and directories
• Device management: Communicating with connected hardware devices
• User interface: Providing a way for users to interact with the computer
2. Components:
• Kernel: The core of the OS that manages hardware resources
• Shell: The user interface for interacting with the OS
• File system: Organizes and manages data storage
• Device drivers: Software that allows the OS to communicate with hardware
The Raspberry Pi runs a specialized version of Linux designed for embedded systems. This
operating system, typically a variant of Debian called Raspberry Pi OS (formerly Raspbian), is
optimized for the Pi’s ARM-based architecture and limited resources.
Key features:
37
Installation
To use the Raspberry Pi, we will need an operating system. By default, Raspberry Pis check
for an operating system on any SD card inserted in the slot, so we should install an operating
system using Raspberry Pi Imager.
In November 2025, the Raspberry Pi Imager 2.0 was launched. It brings a new
wizard interface, the opportunity to pre-configure Raspberry Pi Connect, and
improved accessibility for screen readers and other assistive technologies.
Raspberry Pi Imager is a tool for downloading and writing images on macOS, Windows, and
Linux. It includes many popular operating system images for Raspberry Pi. We will also use
the Imager to preconfigure credentials and remote access settings.
Follow the steps to install the OS on your Raspberry Pi.
38
4. Choose the appropriate operating system:
• For Raspi-Zero: For example, you can select under Raspberry Pi OS (Other),
Raspberry Pi OS Lite (64-bit).
Due to the Raspberry Pi Zero’s limited SDRAM (512 MB), the recommended
OS is the 32-bit version. However, to run some machine learning models, such
as the YOLO from Ultralitics, we should use the 64-bit version. Although the
Raspi-Zero can run a desktop, we will choose the LITE version (no Desktop) to
reduce the RAM needed for regular operation.
• For Raspi-5: We can select the full 64-bit version, which includes a desktop:
Raspberry Pi OS (64-bit)
39
5. Select your microSD card as the storage device.
6. Click Next, then go to the Customization tab. The imager will guide you through setting
the hostname, the Raspberry Pi username and password, configuring WiFi, and enabling
SSH (Very important!).
7. Write the image to the microSD card.
In the examples here, we will use different hostnames depending on the device used:
raspi, raspi-5, raspi-Zero, etc. Please replace it with the one you’re currently using.
Initial Configuration
40
3. Please wait for the initial boot process to complete (it may take a few minutes).
You can find the most common Linux commands for the Raspberry Pi here or here.
Remote Access
SSH Access
The easiest way to interact with the Raspi-Zero is via SSH (“Headless”). You can use a
Terminal (MAC/Linux), PuTTy (Windows), or any other.
The Raspberry Pi and the notebook should be on the same WiFi network. Note
that the Raspberry Pi 5 supports dual-band 802.11ac Wi-Fi (2.4 GHz and 5 GHz),
whereas the Raspberry Pi Zero 2W supports only 2.4 GHz.
1. Find your Raspberry Pi’s IP address (for example, check your router).
2. On your computer, open a terminal and connect via SSH:
ssh username@[raspberry_pi_ip_address]
Alternatively, if you do not have the IP address, you can try, for example, ssh
mjrovai@[Link] , ssh mjrovai@[Link] , etc.: bash ssh username@[Link]
When you see the prompt:
mjrovai@rpi-5:~ $
41
You should confirm the Raspberry Pi IP address. On the terminal, you can use:
hostname -I
python3 --version
Once we use the latest Raspberry Pi OS (based on Debian Trixie), it should be: 3.13:
As of today (January 2026), some packages, such as ExecuTorch, officially support only Python
versions 3.10-3.12. Python 3.13.5 is too new and will likely cause compatibility issues. Since
Debian Trixie ships with Python 3.13 by default, we’ll need to install a compatible Python
version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python. We will do it in the appropriate Lab.
For now, we will keep the system Python.
If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.
42
To shut down the Raspi via terminal:
When you want to turn off your Raspberry Pi, there are better ideas than just pulling the
power cord. This is because the Raspi may still be writing data to the SD card, in which case
merely powering down may result in data loss or, even worse, a corrupted SD card.
For a safety shutdown, use the command line:
To avoid potential data loss and SD card corruption, wait a few seconds after
shutdown for the Raspberry Pi’s LED to stop blinking and go dark before removing
power. Once the LED goes out, it’s safe to power down.
Transferring files between the Raspberry Pi and our main computer can be done using a USB
drive, via the terminal (scp), or an FTP program over the network.
43
You can use any text editor. In the same terminal, an option is nano.
To copy the file named [Link] from your personal computer to a user’s home folder on your
Raspberry Pi, run the following command from the directory containing [Link], replacing
the <username> placeholder with the username you use to log in to your Raspberry Pi and the
<pi_ip_address> placeholder with your Raspberry Pi’s IP address:
Note that ~/ means we will move the file to the ROOT of our Raspberry Pi. You
can choose any folder in your Raspberry Pi. But you should create the folder before
you run scp, since scp won’t create folders automatically.
For example, let’s transfer the file [Link] to the ROOT of my Raspberry Pi Zero, which
has an IP of [Link]:
44
scp [Link] mjrovai@[Link]:~/
I use a different profile to differentiate the terminals. The action above occurs on your
computer. Now, let’s go to our Raspi (using SSH) and check if the file is there:
$ scp <username>@<pi_ip_address>:[Link] .
For example:
On the Raspi, let’s create a copy of the file with another name:
cp [Link] test_2.txt
45
scp mjrovai@[Link]:test_2.txt .
Transferring files via FTP, such as FileZilla FTP Client, is also possible and much easier to use.
Follow the instructions to install the program on your Desktop, then use the Raspberry Pi’s IP
address as the Host. For example:
s[Link]
Enter your Raspberry Pi username and password. Pressing Quickconnect opens two windows,
one for your host computer desktop (right) and another for the Raspberry Pi (left).
46
Increasing Memory
Using htop, a cross-platform interactive process viewer, we can easily monitor resources on our
Raspberry Pi in real time, including the list of processes, the running CPUs, and the memory
usage. To lunch hop, enter with the command on the terminal:
htop
47
Regarding memory, among the Raspberry Pi family, the Raspberry Pi Zero has the least SRAM
(500 MB), compared to 2GB to 16GB on the Raspberry Pi 4 or 5.
On any Raspberry Pi, it is possible to increase/modify the system’s available memory using
“Swap.” Swap memory, also known as swap space, is a technique used in computer operating
systems to temporarily store data from RAM (Random Access Memory) on the SD card when
the physical RAM is fully utilized. This allows the operating system (OS) to continue running
even when RAM is full, which can prevent system crashes or slowdowns.
Swap memory benefits devices with limited RAM, such as the Raspberry Pi Zero. Increasing
swap can help run more demanding applications or processes, but it’s essential to balance this
with the potential performance impact of frequent disk access.
We can check the swap memory using htopor by the command:
48
On the Debian Trixie, 2 MB of swap memory is configured by default on the Rasp-5
and 512MB on the Raspi-Zero
[Zram]
RamMultiplier=2
MaxSizeMiB=4096
49
Increasing memory for the Rasp-Zero
By default, the Rapi-Zero’s SWAP (zram) memory is 512MB, which may be insufficient for
running more complex and demanding Machine Learning applications (for example, YOLO).
So, let’s increase it to 2MB:
On Zero 2 W with Raspberry Pi OS Trixie, you configure rpi-swap by dropping a small config
file into /etc/rpi/[Link].d/. The same mechanism works on Pi 5 and Zero‑2.
Below are two common patterns: fixed swap file (on SD/SSD) and fixed zram size (in RAM).
#### Fixed‑size swap file (e.g., 2 GB on SD)
[File]
FixedSizeMiB=2048
4. Apply:
sudo systemctl restart rpi-swap
swapon --show
If we want to keep everything in compressed RAM (no SD wear), we can set a fixed zram size
instead:
50
1. Create the same drop‑in file:
sudo nano /etc/rpi/[Link].d/[Link]
[Zram]
FixedSizeMiB=2048
• Default Trixie on Zero‑2 uses rpi-swap with zram sized roughly to RAM×1, capped by
MaxSizeMiB (usually 2048), and no swap file unless you add it.
• If we create multiple drop‑ins in /etc/rpi/[Link].d/, they are merged; keep it
simple with one clear override file per board/use‑case.
To keep htop running, open another terminal window to continuously interact with
the Raspberry Pi.
Installing a Camera
The Raspberry Pi is an excellent device for computer vision applications that require a camera.
We can install a camera module connected to the Raspberry Pi CSI (Camera Serial Interface)
port, or a standard USB webcam on the micro-USB port using a USB OTG adapter (Raspi-Zero)
or directly on the USB port on the Raspi-5.
USB Webcams generally have inferior image quality compared to camera modules
that connect to the CSI port. They can not be controlled using the raspistill and
rasivid commands in the terminal, or the picamera recording package in Python.
Nevertheless, there may be reasons you want to connect a USB camera to your
Raspberry Pi, such as setting up multiple cameras with a single Raspberry Pi,
avoiding long cables, or simply because you have such a camera on hand.
51
Installing a Camera Module on the CSI port
There are now several Raspberry Pi camera modules. The original 5-megapixel model was
releasedin 2013, followed by the 8-megapixel Camera Module 2, released in 2016. The latest
camera model is the 12-megapixel Camera Module 3, released in 2023.
The original 5MP camera (Arducam OV5647) is no longer available from Raspberry Pi but
can be found from several alternative suppliers. Below is an example of such a camera on a
Raspberry Pi Zero.
Here is another example of a v2 Camera Module, which has a Sony IMX219 8-megapixel
sensor:
52
First, try to list the installed cameras:
rpicam-hello --list-cameras
Try to list the installed camera again. If Ok, you should see something like:
53
Any camera module will work on the Raspberry Pi, but for that, it is possible that the
[Link] file must be updated:
At the bottom of the file, for example, to use the 5MP Arducam OV5647 camera, add the
line:
dtoverlay=ov5647,cam0
Or for the v2 module, wich has the 8MP Sony IMX219 camera:
dtoverlay=imx219,cam0
Save the file (CTRL+O [ENTER] CRTL+X) and reboot the Raspi:
Sudo reboot
rpicam-hello --list-cameras
54
libcamerais an open-source software library that supports camera systems directly
on Linux for Arm processors. It minimizes the amount of proprietary code running
on the Broadcom GPU.
Let’s capture a JPEG image with a resolution of 640 x 480 for testing and save it to a file
named test_cli_camera.jpg
To view the saved file, we should use ls, which lists the contents of the current directory:
55
Alternatively, you can transfer it to your desktop using FileZilla.
sudo shutdown -h no
2. Connect the USB Webcam (USB Camera Module 30 fps, 1280x720) to your Raspberry
Pi (in this example, I am using a Raspberry Pi Zero, but the instructions work for all
Raspberry Pis).
56
3. Power on again and run the SSH
4. To check if your USB camera is recognized, run:
lsusb
57
5. To take a test picture with your USB camera, use:
fswebcam test_image.jpg
6. Since we are using SSH to connect to our Rapsi, we must transfer the image to our main
computer so we can view it. We can use FileZilla or SCP for this:
scp mjrovai@[Link]:~/test_image.jpg .
Replace “mjrovai” with your username and “raspi-zero” with Pi’s hostname.
7. If the image quality isn’t satisfactory, you can adjust various settings; for example, define
a resolution that is suitable for YOLO (640x640):
58
fswebcam -r 640x640 --no-banner test_image_yolo.jpg
59
And verified using lsusb
Video Streaming
For stream video (which is more resource-intensive), we can install and use mjpg-streamer:
First, install Git:
60
sudo apt install git
Now, we should install the necessary dependencies for mjpg-streamer, clone the repository, and
proceed with the installation:
We can then access the stream by opening a web browser and navigating to:
[Link] In my case: [Link]
We should see a webpage with options to view the stream. Click on the link that says “Stream”
or try accessing:
[Link]
61
Running the Raspi Desktop remotely
While we’ve primarily interacted with the Raspberry Pi via SSH terminal commands, we
can access the full graphical desktop environment remotely if we have installed the complete
Raspberry Pi OS (for example, Raspberry Pi OS (64-bit). This can be particularly useful
for tasks that benefit from a visual interface. To enable this functionality, we must set up a
VNC (Virtual Network Computing) server on the Raspberry Pi. Here’s how to do it:
62
• Connect to your Raspberry Pi via SSH.
• Run the Raspberry Pi configuration tool by entering:
sudo raspi-config
• Exit the configuration tool (use [Tab]), saving changes when prompted.
63
2. Install a VNC Viewer on Your Computer:
• Download and install a VNC viewer application on your main computer. Popular
options include RealVNC Viewer, TightVNC, or VNC Viewer by RealVNC. We will
install VNC Viewer by RealVNC.
3. Once installed, confirm the Raspberry Pi’s IP address. For example, on the terminal,
you can use:
hostname -I
64
• Enter your Raspberry Pi’s IP address and hostname.
• When prompted, enter your Raspberry Pi’s username and password.
65
5. The Raspberry Pi 5 Desktop should appear on your computer monitor.
66
6. Adjust Display Settings (if needed):
• Once connected, adjust the display resolution for optimal viewing. This can be done
through the Raspberry Pi’s desktop settings or by modifying the [Link] file.
• Let’s do it using the desktop settings. Reach the menu (the Raspberry Icon at the
left upper corner) and select the best screen definition for your monitor:
67
Updating and Installing Software
68
Rule of thumb: Use sudo apt install only for system dependencies and hardware interfaces.
Use pip install (without sudo) inside an activated virtual environment for everything else.
Inside the vent, PIP or PIP3 are the same.
Model-Specific Considerations
Remember to adjust your project requirements based on the specific Raspberry Pi model you’re
using. The Raspi-Zero is great for low-power, space-constrained projects, while the Raspi-4 or
5 models are better suited for more computationally intensive tasks.
Here’s a concise, copy‑pasteable tutorial you can share or publish.
In this section, we will explore how to measure (and monitor) CPU temperature and power
consumption on a Raspberry Pi 5 using only built‑in tools and a small shell script.
69
Check CPU Temperature
vcgencmd measure_temp
Example:
temp=47.2'C
cat /sys/class/thermal/thermal_zone0/temp
70
vcgencmd pmic_read_adc
Example (truncated):
3V3_SYS_A current(1)=0.06245952A
1V8_SYS_A current(2)=0.16102850A
...
3V3_SYS_V volt(13)=3.29687500V
1V8_SYS_V volt(14)=1.79687500V
...
• Applies a linear correction to estimate real board power, based on the RPi5‑power
calibration.
nano avg_temp_power.sh
Paste:
71
#!/bin/bash
# avg_temp_power.sh
# Measure average CPU temperature and power on Raspberry Pi 5
TEMP_FILE=$(mktemp)
PMIC_FILE=$(mktemp)
sum_pmic=0
for idx in "${!volts[@]}"; do
v=${volts[$idx]}
i_amp=${currents[$idx]}
sum_pmic=$(awk -v a="$sum_pmic" -v b="$v" -v c="$i_amp" \
'BEGIN{printf "%.8f", a + b*c}')
done
72
if (n > 1) {
mean = sum / n
std = sqrt( (sum2 - (sum*sum)/n) / (n-1) )
} else {
mean = sum
std = 0
}
return mean "|" std
}
BEGIN{
split(stats("'"${TEMP_FILE}"'"), t, "|")
t_mean = t[1]
t_std = t[2]
split(stats("'"${PMIC_FILE}"'"), p, "|")
p_pmic_mean = p[1]
p_pmic_std = p[2]
rm -f "${TEMP_FILE}" "${PMIC_FILE}"
Make it executable
chmod +x avg_temp_power.sh
Run a measurement
Idle baseline:
73
./avg_temp_power.sh
Then start a heavy workload (e.g., Llama 3.2:3B) in another terminal and run the script again
during inference. You should see higher temperature and power, often around 70–75 °C and
10–11 W with active cooling.
74
State Temp (°C) Estimated real power (W)
Idle CLI 40–50 3.0–3.5
Light desktop 50–60 4–5
Moderate CPU load 55–70 5–7
Heavy CPU/GPU 70–80 7–10
75
76
Image Classification Fundamentals
Figure 2: DALL·E prompt - “Create a Cartoon with style from the 50’s doing Image Classifi-
cation on a Raspberrry Pi - based on the image uploaded.”
77
Introduction
Image classification has found its way into numerous real-world applications, revolutionizing
various sectors:
Implementing image classification on edge devices such as the Raspberry Pi offers several
compelling advantages:
1. Low Latency: Processing images locally eliminates the need to send data to cloud servers,
significantly reducing response times.
2. Offline Functionality: Classification can be performed without an internet connection,
making it suitable for remote or connectivity-challenged environments.
3. Privacy and Security: Sensitive image data remains on the local device, addressing data
privacy concerns and compliance requirements.
4. Cost-Effectiveness: Eliminates the need for expensive cloud computing resources, espe-
cially for continuous or high-volume classification tasks.
78
5. Scalability: Enables distributed computing architectures in which multiple devices can
operate independently or in a network.
6. Energy Efficiency: Optimized models on dedicated hardware can be more energy-efficient
than cloud-based solutions, which is crucial for battery-powered or remote applications.
7. Customization: Deploying specialized or frequently updated models tailored to specific
use cases is more manageable.
We can create more responsive, secure, and efficient computer vision solutions by leveraging
the power of edge devices such as the Raspberry Pi for image classification. This approach
opens new possibilities for integrating intelligent visual processing across diverse applications
and environments.
In the following sections, we’ll explore how to implement and optimize image classification
on the Raspberry Pi, leveraging these advantages to build powerful, efficient computer vision
systems.
79
Setting up a Virtual Environment
source ~/tflite_env/bin/activate
deactivate
Verify installation
80
System vs pip Package Installation Rule
Rule of thumb: Use sudo apt install only for system dependencies and hard-
ware interfaces. Use pip install (without sudo) inside an activated virtual
environment for everything else. Inside the vent, PIP or PIP3 are the same.
The virtual environment will automatically include both system packages and
pip-installed packages thanks to the --system-site-packages flag.
Let’s set up Jupyter Notebook optimized for headless Raspberry Pi camera work and develop-
ment:
To run Jupyter Notebook, run the command (change the IP address for yours):
81
jupyter notebook --ip=[Link] --no-browser
On the terminal, you can see the local URL address to open the notebook:
You can access it from another device by entering the Raspberry Pi’s IP address and the
provided token in a web browser (you can copy the token from the terminal).
Define the working directory in the Raspi and create a new Python 3 notebook. For example:
82
cd Documents
mkdir Python
import time
import numpy as np
from PIL import Image
import [Link] as plt
from picamera2 import Picamera2
Load an image from the internet, for example (note that it is possible to run a command line
from the Notebook, using ! before the command:
!wget [Link]
img_path = "[Link]"
img = [Link](img_path)
83
Now, let’s use the camera to capture a local image:
# Initialize camera
picam2 = Picamera2()
[Link]()
84
[Link](2)
# Capture image
picam2.capture_file("class3_test.jpg")
print("Image captured: class3_test.jpg")
# Stop camera
[Link]()
[Link]()
And use a similar code as before to show it (adapting the img_path and title):
85
Installing LiteRT
We are interested in inference, which involves running trained models on a device to make
predictions from input data. To perform an inference with a model, we must run it through an
interpreter. For that, we will use LiteRT, Google’s on-device framework for high-performance
ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimiza-
tion.
LiteRT features advanced GPU/NPU acceleration, delivers superior ML & GenAI performance,
making on-device ML inference easier than ever.
For installation on the Raspi, let’s use the command:
If you are working on the Raspi-Zero with the minimum OS (No Desktop), you do not have a
user-pre-defined directory tree (you can check it with ls. So, let’s create one:
86
mkdir Documents
cd Documents/
mkdir TFLITE
cd TFLITE/
mkdir IMG_CLASS
cd IMG_CLASS
mkdir models
cd models
wget [Link]
wget [Link]
87
Verifying the Setup
Let’s test our setup by running a simple Python script on TFLITE/IMG_CLASS folder:
print("NumPy:", np.__version__)
print("Pillow:", Image.__version__)
We can create the Python script using nano on the terminal, saving it with CTRL+0 + ENTER +
CTRL+X
88
python setup_test.py
89
Making inferences with Mobilenet V2
In the last section, we set up the environment, including downloading a popular pre-trained
model, Mobilenet V2, trained on ImageNet’s 224x224 images (1.2 million) for 1,001 classes
(1,000 object categories plus 1 background). The model was converted to a compact 3.5MB
tflite format, making it suitable for the limited storage and memory of a Raspberry Pi.
In the IMG_CLASS working directory, let’s start a new notebook to follow all the steps to classify
one image:
Import the needed libraries:
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
The message means LiteRTsuccessfully enabled an optimized CPU backend (XNNPACK) for
our model, which is good and expected.
What XNNPACK is
• XNNPACK is a library of highly optimized operators (conv, FC, etc.) for running neural
networks on CPUs, especially ARM and x86.
90
• LiteRT can “delegate” supported ops to XNNPACK so they run using these faster kernels
instead of the default reference CPU implementation.
So, it means the interpreter has attached the XNNPACK delegate and will run all compatible
parts of the graph on the CPU using it.
• On devices like the Raspberry Pi, this usually results in lower inference latency at the
cost of slightly longer delivery time and a bit more RAM for packed weights.
• We are currently using CPU acceleration, not GPU, which is the standard/optimal path
for many TFLite/LiteRT models on Pi-class hardware.
input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()
Input details provide information on how the model should be fed an image. The shape of (1,
224, 224, 3) informs us that an image with dimensions (224x224x3) should be input one by one
(Batch Dimension: 1).
The output details indicate that the inference will produce an array of 1,001 integer values.
Those values result from image classification, where each value is the probability that the
corresponding label is associated with the image.
91
Let’s also inspect the dtype of the input details of the model
input_dtype = input_details[0]['dtype']
input_dtype
dtype('uint8')
This shows that the input image should be represented as raw pixels (0-255).
Let’s get a test image. We can either transfer it from our computer or download one for testing,
as we did before. Let’s first create a folder under our working directory:
mkdir images
cd images
wget [Link]
92
We can see the image size by running the command:
That shows that the image is an RGB image with a width and height of 1600 pixels each. To
use our model, we should reshape it to (224, 224, 3) and add a batch dimension of 1, as defined
in the input details: (1, 224, 224, 3). The inference result, as shown in the output details, will
be an array of size 1001, as shown below:
93
So, let’s reshape the image, add the batch dimension, and see the result:
input_data.dtype
dtype('uint8')
The input data dtype is ‘uint8’, which is compatible with the dtype expected for the model.
Using the input_data, let’s run the interpreter and get the predictions (output):
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
predictions = interpreter.get_tensor(output_details[0]['index'])[0]
The prediction is an array with 1001 elements. Let’s get the Top-5 indices where their elements
have high values:
top_k_results = 5
top_k_indices = [Link](predictions)[::-1][:top_k_results]
top_k_indices
94
The top_k_indices is an array with 5 elements: array([283, 286, 282])
So, 283, 286, 282, 288, and 479 are the image’s most probable classes. Having the index, we
must find to which class it belongs (such as car, cat, or dog). The text file downloaded with
the model includes a label for each index from 0 to 1,000. Let’s use a function to load the .txt
file as a list:
def load_labels(filename):
with open(filename, 'r') as f:
return [[Link]() for line in [Link]()]
And get the list, printing the labels associated with the indexes:
labels_path = "./models/[Link]"
labels = load_labels(labels_path)
print(labels[286])
print(labels[283])
print(labels[282])
print(labels[288])
print(labels[479])
As a result, we have:
Egyptian cat
tiger cat
tabby
lynx
carton
At least four of the top indices are related to felines. The prediction content is the probability
associated with each one of the labels. As we saw in the output details, those values are
quantized and should be dequantized:
95
The output (positive and negative numbers) shows that the output probably does not have a
Softmax. Checking the model documentation ([Link] MobileNet
V2 typically doesn’t include a softmax layer at the output. It usually ends with a 1x1 convolution
followed by average pooling and a fully connected layer. So, for getting the probabilities (0 to
1), we should apply Softmax:
print (probabilities[286])
print (probabilities[283])
print (probabilities[282])
print (probabilities[288])
print (probabilities[479])
0.265947
0.39499295
0.17906114
0.08961108
0.022443123
For clarity, let’s create a function to relate the labels to the probabilities:
for i in range(top_k_results):
print("\t{:20}: {}%".format(
labels[top_k_indices[i]],
(int(probabilities[top_k_indices[i]]*100))))
Let’s create a general function to give an image as input, and we get the Top-5 possible
classes:
96
def image_classification(img_path, model_path, labels, top_k_results=5):
# load the image
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
# Preprocess
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims(img, axis=0)
# Inference on Raspi
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {}%".format(
labels[top_k_indices[i]],
97
(int(probabilities[top_k_indices[i]]*100))))
Let’s modify the Python script used before to capture an image from the camera (size: 224x224),
saving it in the images folder:
def capture_image(image_path):
# Initialize camera
picam2 = Picamera2() # default is index 0
98
# Wait for camera to warm up
[Link](2)
# Capture image
picam2.capture_file(image_path)
print("Image captured: "+"image_path")
# Stop camera
[Link]()
[Link]()
img_path = './images/cam_img_test.jpg'
model_path = "./models/mobilenet_v2_1.0_224_quant.tflite"
labels = load_labels("./models/[Link]")
capture_image(img_path)
image_classification(img_path, model_path, labels, top_k_results=5)
99
Exploring a Model Trained from Zero
Let’s get a TFLite model trained from scratch. For that, we can follow the Notebook:
100
CNN to classify Cifar-10 dataset
In the notebook, we trained a model using the CIFAR10 dataset, which contains 60,000 images
from 10 classes of CIFAR (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and
truck). CIFAR has 32x32 color images (3 color channels) where the objects are not centered
and can have the object with a background, such as airplanes that might have a cloudy sky
behind them! In short, small but real images.
The CNN trained model (cifar10_model.keras) had a size of 2.0MB. Using the TFLite Converter,
the model [Link] became with 674MB (around 1/3 of the original size).
Conclusion:
This chapter has established a solid foundation for understanding and implementing image
classification on Raspberry Pi devices using Python and LiteRT. Throughout this journey, we
101
have explored the essential components that make edge-based computer vision both practical
and powerful.
We began by understanding the theoretical foundations of image classification and its real-
world applications across diverse sectors, from healthcare to environmental monitoring. The
advantages of running classification on edge devices like the Raspberry Pi—including low
latency, offline functionality, enhanced privacy, and cost-effectiveness—make it an attractive
solution for many practical applications.
The hands-on experience of setting up the development environment provided crucial insights
into the requirements and constraints of embedded systems. We successfully configured LiteRT,
installed essential Python libraries, and established a working directory structure that serves
as the foundation for computer vision projects.
Working with the pre-trained MobileNet V2 model demonstrated several key concepts:
This foundational knowledge prepares us for more advanced topics, including custom model
training and deployment. The skills developed here—understanding model architectures,
implementing inference pipelines, and working with embedded Python environments—are
transferable to a wide range of computer vision applications.
102
The chapter serves as a stepping stone toward building more sophisticated AI systems on edge
devices, demonstrating that powerful computer vision capabilities are accessible even on modest
hardware platforms when properly optimized and implemented.
Resources
• Dataset Example
• Setup Test Notebook on a Raspi
• Image Classification Notebook on a Raspi
• CNN to classify Cifar-10 dataset at CoLab
• Cifar 10 - Image Classification on a Raspi
• Python Scripts
103
104
Custom Image Classification Project
Figure 3: DALL·E prompt - A cover image for an ‘Image Classification’ chapter in a Raspberry
Pi tutorial, designed in the same vintage 1950s electronics lab style as previous covers.
The scene should feature a Raspberry Pi connected to a camera module, with the
camera capturing a photo of the small blue robot provided by the user. The robot should
be placed on a workbench, surrounded by classic lab tools like soldering irons, resistors,
and wires. The lab background should include vintage equipment like oscilloscopes
and tube radios, maintaining the detailed and nostalgic feel of the era. No text or
logos should be included.
105
Image Classification Project
In this chapter, we will develop a complete Image Classification project using the Edge Impulse
Studio. As we did with the MobiliNet V2, the trained and converted TFLite model will be
used for inference using a Python script.
Here is a typical ML workflow that we will use in our project:
The Goal
The first step in any ML project is to define its goal. In this case, it is to detect and classify
two specific objects present in one image. For this project, we will use two small toys: a robot
and a small Brazilian parrot (named Periquito). We will also collect images of a background
where those two objects are absent.
Data Collection
Once we have defined our Machine Learning project goal, the next and most crucial step is
collecting the dataset. We can use a phone for the image capture, but we will use the Raspi
here. Let’s set up a simple web server on our Raspberry Pi to view the QVGA (320 x 240)
captured images in a browser.
106
1. First, let’s install Flask, a lightweight web framework for Python:
pip install flask
2. Go to the working folder (IMG_CLASS) and create a new Python script combining image
capture with a web server. We’ll call it get_img_data.py:
app = Flask(__name__)
# Global variables
base_dir = "dataset"
picam2 = None
frame = None
frame_lock = [Link]()
capture_counts = {}
current_label = None
shutdown_event = [Link]()
def initialize_camera():
global picam2
picam2 = Picamera2()
config = picam2.create_preview_configuration(
main={"size": (320, 240)}
)
[Link](config)
[Link]()
[Link](2) # Wait for camera to warm up
def get_frame():
global frame
while not shutdown_event.is_set():
stream = [Link]()
picam2.capture_file(stream, format='jpeg')
with frame_lock:
107
frame = [Link]()
[Link](0.1) # Adjust as needed for smooth preview
def generate_frames():
while not shutdown_event.is_set():
with frame_lock:
if frame is not None:
yield (b'--frame\r\n'
b'Content-Type: image/jpeg\r\n\r\n' +
frame + b'\r\n')
[Link](0.1) # Adjust as needed for smooth streaming
def shutdown_server():
shutdown_event.set()
if picam2:
[Link]()
# Give some time for other threads to finish
[Link](2)
# Send SIGINT to the main process
[Link]([Link](), [Link])
108
</form>
</body>
</html>
''')
@[Link]('/capture')
def capture_page():
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture</title>
<script>
var shutdownInitiated = false;
function checkShutdown() {
if (!shutdownInitiated) {
fetch('/check_shutdown')
.then(response => [Link]())
.then(data => {
if ([Link]) {
shutdownInitiated = true;
[Link](
'video-feed').src = '';
[Link](
'shutdown-message')
.[Link] = 'block';
}
});
}
}
setInterval(checkShutdown, 1000); // Check
every second
</script>
</head>
<body>
<h1>Dataset Capture</h1>
<p>Current Label: {{ label }}</p>
<p>Images captured for this label: {{ capture_count
}}</p>
<img id="video-feed" src="{{ url_for('video_feed')
}}" width="640"
height="480" />
109
<div id="shutdown-message" style="display: none;
color: red;">
Capture process has been stopped.
You can close this window.
</div>
<form action="/capture_image" method="post">
<input type="submit" value="Capture Image">
</form>
<form action="/stop" method="post">
<input type="submit" value="Stop Capture"
style="background-color: #ff6666;">
</form>
<form action="/" method="get">
<input type="submit" value="Change Label"
style="background-color: #ffff66;">
</form>
</body>
</html>
''', label=current_label, capture_count=capture_counts.get(
current_label, 0))
@[Link]('/video_feed')
def video_feed():
return Response(generate_frames(),
mimetype='multipart/x-mixed-replace;
boundary=frame')
@[Link]('/capture_image', methods=['POST'])
def capture_image():
global capture_counts
if current_label and not shutdown_event.is_set():
capture_counts[current_label] += 1
timestamp = [Link]("%Y%m%d-%H%M%S")
filename = f"image_{timestamp}.jpg"
full_path = [Link](base_dir, current_label,
filename)
picam2.capture_file(full_path)
return redirect(url_for('capture_page'))
@[Link]('/stop', methods=['POST'])
110
def stop():
summary = render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Dataset Capture - Stopped</title>
</head>
<body>
<h1>Dataset Capture Stopped</h1>
<p>The capture process has been stopped.
You can close this window.</p>
<p>Summary of captures:</p>
<ul>
{% for label, count in capture_counts.items() %}
<li>{{ label }}: {{ count }} images</li>
{% endfor %}
</ul>
</body>
</html>
''', capture_counts=capture_counts)
return summary
@[Link]('/check_shutdown')
def check_shutdown():
return {'shutdown': shutdown_event.is_set()}
if __name__ == '__main__':
initialize_camera()
[Link](target=get_frame, daemon=True).start()
[Link](host='[Link]', port=5000, threaded=True)
python get_img_data.py
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
111
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your
Raspberry Pi’s IP address). For example: [Link]
This Python script creates a web-based interface for capturing and organizing image datasets
using a Raspberry Pi and its camera. It’s handy for machine learning projects that require
labeled image data.
Key Features:
1. Web Interface: Accessible from any device on the same network as the Raspberry Pi.
2. Live Camera Preview: This shows a real-time feed from the camera.
3. Labeling System: Allows users to input labels for different categories of images.
4. Organized Storage: Automatically saves images in label-specific subdirectories.
5. Per-Label Counters: Keeps track of how many images are captured for each label.
6. Summary Statistics: Provides a summary of captured images when stopping the
capture process.
Main Components:
1. Flask Web Application: Handles routing and serves the web interface.
2. Picamera2 Integration: Controls the Raspberry Pi camera.
3. Threaded Frame Capture: Ensures smooth live preview.
4. File Management: Organizes captured images into labeled directories.
Key Functions:
112
Usage Flow:
113
7. Click Stop Capture when finished to see a summary.
Technical Notes:
• The script uses threading to handle concurrent frame capture and web serving.
• Images are saved with timestamps in their filenames for uniqueness.
• The web interface is responsive and can be accessed from mobile devices.
Customization Possibilities:
Get around 60 images from each category (periquito, robot and background). Try to capture
different angles, backgrounds, and light conditions.
On the Raspi, we will end with a folder named dataset, which contains three sub-folders:
periquito, robot, and background, one for each class of images.
You can use Filezilla to transfer the created dataset to your main computer.
114
Training the model with Edge Impulse Studio
We will use the Edge Impulse Studio to train our model. Go to the Edge Impulse Page, enter
your account credentials, and create a new project:
Dataset
We will walk through four main steps using the EI Studio (or Studio). These steps are crucial
in preparing our model for use on the Raspi: Dataset, Impulse, Tests, and Deploy (on the Edge
Device, in this case, the Raspi).
Regarding the Dataset, it is essential to point out that our Original Dataset,
captured with the Raspi, will be split into Training, Validation, and Test. The Test
Set will be separated from the beginning and reserved for use only in the Test phase
after training. The Validation Set will be used during training.
1. Go to the Data acquisition tab, and in the UPLOAD DATA section, upload the files from
your computer in the chosen categories.
2. Leave to the Studio the splitting of the original dataset into train and test and choose
the label about
3. Repeat the procedure for all three classes. At the end, you should see your “raw data” in
the Studio:
115
The Studio allows you to explore your data, showing a complete view of all the data in your
project. You can clear, inspect, or change labels by clicking on individual data items. In our
case, a straightforward project, the data seems OK.
116
• Pre-process our data, which consists of resizing the individual images and determining
the color depth to use (be it RGB or Grayscale) and
• Specify a Model. In this case, it will be the Transfer Learning (Images) to fine-tune a
pre-trained MobileNet V2 image classification model on our data. This method performs
well even with relatively small image datasets (around 180 images in our case).
Transfer Learning with MobileNet offers a streamlined approach to model training, which is
especially beneficial for resource-constrained environments and projects with limited labeled
data. MobileNet, known for its lightweight architecture, is a pre-trained model that has already
learned valuable features from a large dataset (ImageNet).
By leveraging these learned features, we can train a new model for your specific task with fewer
data and computational resources and achieve competitive accuracy.
117
This approach significantly reduces training time and computational cost, making it ideal for
quick prototyping and deployment on embedded devices where efficiency is paramount.
Go to the Impulse Design Tab and create the impulse, defining an image size of 160 × 160 and
squashing them (squared form, without cropping). Select Image and Transfer Learning blocks.
Save the Impulse.
Image Pre-Processing
All the input QVGA/RGB565 images will be converted to 76,800 features (160 × 160 × 3).
118
Press Save parameters and select Generate features in the next tab.
Model Design
MobileNet is a family of efficient convolutional neural networks designed for mobile and
embedded vision applications. The key features of MobileNet are:
1. Lightweight: Optimized for mobile devices and embedded systems with limited computa-
tional resources.
2. Speed: Fast inference times, suitable for real-time applications.
119
3. Accuracy: Maintains good accuracy despite its compact size.
MobileNetV2, introduced in 2018, improves the original MobileNet architecture. Key features
include:
1. Inverted Residuals: Inverted residual structures are used where shortcut connections are
made between thin bottleneck layers.
2. Linear Bottlenecks: Removes non-linearities in the narrow layers to prevent the destruction
of information.
3. Depth-wise Separable Convolutions: Continues to use this efficient operation from Mo-
bileNetV1.
In our project, we will do a Transfer Learning with the MobileNetV2 160x160 1.0, which
means that the images used for training (and future inference) should have an input Size of
160 × 160 pixels and a Width Multiplier of 1.0 (full width, not reduced). This configuration
balances between model size, speed, and accuracy.
Model Training
120
return image, label
Exposure to these variations during training can help prevent your model from taking shortcuts
by “memorizing” superficial clues in your training data, meaning it may better reflect the deep
underlying patterns in your dataset.
The final dense layer of our model will have 0 neurons with a 10% dropout for overfitting
prevention. Here is the Training result:
The result is excellent, with a reasonable 35 ms of latency (for a Raspi-4), which should result
in around 30 fps (frames per second) during inference. A Raspi-Zero should be slower, and the
Raspi-5, faster.
If faster inference is needed, we should train the model using smaller alphas (0.35, 0.5, and
0.75) or even reduce the image input size, trading with accuracy. However, reducing the input
image size and decreasing the alpha (width multiplier) can speed up inference for MobileNet
V2, but they have different trade-offs. Let’s compare:
Pros:
121
• Significantly reduces the computational cost across all layers.
• Decreases memory usage.
• It often provides a substantial speed boost.
Cons:
• It may reduce the model’s ability to detect small features or fine details.
• It can significantly impact accuracy, especially for tasks requiring fine-grained recognition.
Pros:
Cons:
Comparison:
1. Speed Impact:
• Reducing input size often provides a more substantial speed boost because it reduces
computations quadratically (halving both width and height reduces computations
by about 75%).
• Reducing alpha provides a more linear reduction in computations.
2. Accuracy Impact:
• Reducing input size can severely impact accuracy, especially when detecting small
objects or fine details.
• Reducing alpha tends to have a more gradual impact on accuracy.
3. Model Architecture:
• Changing input size doesn’t alter the model’s architecture.
• Changing alpha modifies the model’s structure by reducing the number of channels
in each layer.
Recommendation:
1. If our application doesn’t require detecting tiny details and can tolerate some loss in
accuracy, reducing the input size is often the most effective way to speed up inference.
122
2. Reducing alpha might be preferable if maintaining the ability to detect fine details is
crucial or if you need a more balanced trade-off between speed and accuracy.
3. For best results, you might want to experiment with both:
• Try MobileNet V2 with input sizes like 160 × 160 or 92 × 92
• Experiment with alpha values like 1.0, 0.75, 0.5 or 0.35.
4. Always benchmark the different configurations on your specific hardware and with your
particular dataset to find the optimal balance for your use case.
Remember, the best choice depends on your specific requirements for accuracy, speed,
and the nature of the images you’re working with. It’s often worth experimenting
with combinations to find the optimal configuration for your particular use case.
Model Testing
Now, you should take the data set aside at the start of the project and run the trained model
using it as input. Again, the result is excellent (92.22%).
As we did in the previous section, we can deploy the trained model as .tflite and use Raspi to
run it using Python.
On the Dashboard tab, go to Transfer learning model (int8 quantized) and click on the download
icon:
123
Let’s also download the float32 version for comparison
Transfer the models from your computer to the Raspi (./models), for example, using FileZilla.
Also, capture some images for inference and save them in (./images), or use the images in the
./dataset folder.
Let’s remember what we did in the last chapter:
Activate the environment:
source ~/tflite_env/bin/activate
124
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
img_path = "./images/[Link]"
model_path = "./models/ei-raspi-img-class-int8-quantized-\
[Link]"
labels = ['background', 'periquito', 'robot']
Note that the models trained on the Edge Impulse Studio will output values with
index 0, 1, 2, etc., where the actual labels will follow an alphabetic order.
Load the model, allocate the tensors, and get the input and output tensor details:
One important difference to note is that the dtype of the input details of the model is now
int8, which means that the input values go from –128 to +127, while each pixel of our image
goes from 0 to 255. This means that we should pre-process the image to match it. We can
check here:
input_dtype = input_details[0]['dtype']
input_dtype
numpy.int8
125
img = [Link](img_path)
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
[Link]()
Checking the input data, we can verify that the input tensor is compatible with what is expected
by the model:
126
input_data.shape, input_data.dtype
Now, it is time to perform the inference. Let’s also calculate the latency of the model:
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert
# to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))
The model will take around 125ms to perform the inference in the Raspi-Zero, which is 3 to 4
times longer than a Raspi-5.
Now, we can get the output labels and probabilities. It is also important to note that the
model trained on the Edge Impulse Studio has a softmax activation function in its output
(different from the original Movilenet V2), and we can use the model’s raw output as the
“probabilities.”
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
127
print("\t{:20}: {:.2f}%".format(
labels[top_k_indices[i]],
probabilities[top_k_indices[i]] * 100))
Let’s modify the function created before so that we can handle different type of models:
# Preprocess
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_dtype = input_details[0]['dtype']
if input_dtype == np.uint8:
input_data = np.expand_dims([Link](img), axis=0)
elif input_dtype == np.int8:
scale, zero_point = input_details[0]['quantization']
128
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (
img_array / scale
+ zero_point
).clip(-128, 127).astype(np.int8)
input_data = np.expand_dims(img_array, axis=0)
else: # float32
input_data = np.expand_dims(
[Link](img, dtype=np.float32),
axis=0
) / 255.0
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time -
start_time
) * 1000 # Convert to milliseconds
# Obtain results
predictions = interpreter.get_tensor(output_details[0]
['index'])[0]
if apply_softmax:
# Apply softmax
exp_preds = [Link](predictions - [Link](predictions))
probabilities = exp_preds / [Link](exp_preds)
else:
probabilities = predictions
129
print("\n\t[PREDICTION] [Prob]\n")
for i in range(top_k_results):
print("\t{:20}: {:.1f}%".format(
labels[top_k_indices[i]],
probabilities[top_k_indices[i]] * 100))
print ("\n\tInference time: {:.1f}ms".format(inference_time))
And test it with different images and the int8 quantized model (160x160 alpha =1.0).
Let’s download a smaller model, such as the one trained for the Nicla Vision Lab (int8 quantized
model, 96x96, alpha = 0.1), as a test. We can use the same function:
The model lost some accuracy, but it is still OK once our model does not look for many details.
Regarding latency, we are aboutt ten times faster on the Raspi-Zero.
130
Live Image Classification
Let’s develop an app that captures images with the camera in real-time and displays their
classification.
Using the nano on the terminal, save the code below, such as img_class_live_infer.py.
app = Flask(__name__)
# Global variables
picam2 = None
frame = None
frame_lock = [Link]()
is_classifying = False
confidence_threshold = 0.8
model_path = "./models/ei-raspi-img-class-int8-quantized-\
[Link]"
labels = ['background', 'periquito', 'robot']
interpreter = None
classification_queue = Queue(maxsize=1)
def initialize_camera():
global picam2
picam2 = Picamera2()
config = picam2.create_preview_configuration(
main={"size": (320, 240)}
)
[Link](config)
[Link]()
[Link](2) # Wait for camera to warm up
def get_frame():
131
global frame
while True:
stream = [Link]()
picam2.capture_file(stream, format='jpeg')
with frame_lock:
frame = [Link]()
[Link](0.1) # Capture frames more frequently
def generate_frames():
while True:
with frame_lock:
if frame is not None:
yield (
b'--frame\r\n'
b'Content-Type: image/jpeg\r\n\r\n'
+ frame + b'\r\n'
)
[Link](0.1)
def load_model():
global interpreter
if interpreter is None:
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
return interpreter
img = [Link]((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims([Link](img), axis=0)\
.astype(input_details[0]['dtype'])
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
predictions = interpreter.get_tensor(output_details[0]
['index'])[0]
# Handle output based on type
output_dtype = output_details[0]['dtype']
132
if output_dtype in [np.int8, np.uint8]:
# Dequantize the output
scale, zero_point = output_details[0]['quantization']
predictions = ([Link](np.float32) -
zero_point) * scale
return predictions
def classification_worker():
interpreter = load_model()
while True:
if is_classifying:
with frame_lock:
if frame is not None:
img = [Link]([Link](frame))
predictions = classify_image(img, interpreter)
max_prob = [Link](predictions)
if max_prob >= confidence_threshold:
label = labels[[Link](predictions)]
else:
label = 'Uncertain'
classification_queue.put({
'label': label,
'probability': float(max_prob)
})
[Link](0.1) # Adjust based on your needs
@[Link]('/')
def index():
return render_template_string('''
<!DOCTYPE html>
<html>
<head>
<title>Image Classification</title>
<script
src="[Link]
</script>
<script>
function startClassification() {
$.post('/start');
$('#startBtn').prop('disabled', true);
$('#stopBtn').prop('disabled', false);
}
133
function stopClassification() {
$.post('/stop');
$('#startBtn').prop('disabled', false);
$('#stopBtn').prop('disabled', true);
}
function updateConfidence() {
var confidence = $('#confidence').val();
$.post('/update_confidence',
{confidence: confidence}
);
}
function updateClassification() {
$.get('/get_classification', function(data) {
$('#classification').text([Link] + ': '
+ [Link](2));
});
}
$(document).ready(function() {
setInterval(updateClassification, 100);
// Update every 100ms
});
</script>
</head>
<body>
<h1>Image Classification</h1>
<img src="{{ url_for('video_feed') }}"
width="640"
height="480" />
<br>
<button id="startBtn"
onclick="startClassification()">
Start Classification
</button>
<button id="stopBtn"
onclick="stopClassification()"
disabled>
Stop Classification
</button>
<br>
134
<label for="confidence">Confidence Threshold:</label>
<input type="number"
id="confidence"
name="confidence"
min="0" max="1"
step="0.1"
value="0.8"
onchange="updateConfidence()" />
<br>
<div id="classification">
Waiting for classification...
</div>
</body>
</html>
''')
@[Link]('/video_feed')
def video_feed():
return Response(
generate_frames(),
mimetype='multipart/x-mixed-replace; boundary=frame'
)
@[Link]('/start', methods=['POST'])
def start_classification():
global is_classifying
is_classifying = True
return '', 204
@[Link]('/stop', methods=['POST'])
def stop_classification():
global is_classifying
is_classifying = False
return '', 204
@[Link]('/update_confidence', methods=['POST'])
def update_confidence():
global confidence_threshold
confidence_threshold = float([Link]['confidence'])
return '', 204
135
@[Link]('/get_classification')
def get_classification():
if not is_classifying:
return jsonify({'label': 'Not classifying',
'probability': 0})
try:
result = classification_queue.get_nowait()
except [Link]:
result = {'label': 'Processing', 'probability': 0}
return jsonify(result)
if __name__ == '__main__':
initialize_camera()
[Link](target=get_frame, daemon=True).start()
[Link](target=classification_worker,
daemon=True).start()
[Link](host='[Link]', port=5000, threaded=True)
python img_class_live_infer.py
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Raspberry
Pi’s IP address). For example: [Link]
136
Here, you can see the app running on YouTube.
The code creates a web application for real-time image classification using a Raspberry Pi,
its camera module, and a TensorFlow Lite model. The application uses Flask to serve a web
interface where is possible to view the camera feed and see live classification results.
Key Components:
1. Flask Web Application: Serves the user interface and handles requests.
2. PiCamera2: Captures images from the Raspberry Pi camera module.
3. LiteRT: Runs the image classification model.
4. Threading: Manages concurrent operations for smooth performance.
Main Features:
Code Structure:
137
• Camera and frame management
• Classification control
• Model and label information
3. Camera Functions:
• initialize_camera(): Sets up the PiCamera2
• get_frame(): Continuously captures frames
• generate_frames(): Yields frames for the web feed
4. Model Functions:
• load_model(): Loads the TFLite model
• classify_image(): Performs inference on a single image
5. Classification Worker:
• Runs in a separate thread
• Continuously classifies frames when active
• Updates a queue with the latest results
6. Flask Routes:
• /: Serves the main HTML page
• /video_feed: Streams the camera feed
• /start and /stop: Controls classification
• /update_confidence: Adjusts the confidence threshold
• /get_classification: Returns the latest classification result
7. HTML Template:
• Displays camera feed and classification results
• Provides controls for starting/stopping and adjusting settings
8. Main Execution:
• Initializes camera and starts necessary threads
• Runs the Flask application
Key Concepts:
138
Usage:
Summary:
Image classification has emerged as a powerful and versatile application of machine learning,
with significant implications for various fields, from healthcare to environmental monitoring.
This chapter has demonstrated how to implement a robust image classification system on
edge devices like the Raspi-Zero and Raspi-5, showcasing the potential for real-time, on-device
intelligence.
We’ve explored the entire pipeline of an image classification project, from data collection and
model training using Edge Impulse Studio to deploying and running inferences on a Raspi.
The process highlighted several key points:
1. The importance of proper data collection and preprocessing for training effective models.
2. The power of transfer learning, allowing us to leverage pre-trained models like MobileNet
V2 for efficient training with limited data.
3. The trade-offs between model accuracy and inference speed, especially crucial for edge
devices.
4. The implementation of real-time classification using a web-based interface, demonstrating
practical applications.
The ability to run these models on edge devices like the Raspi opens up numerous possibilities
for IoT applications, autonomous systems, and real-time monitoring solutions. It allows for
reduced latency, improved privacy, and operation in environments with limited connectivity.
As we’ve seen, even with the computational constraints of edge devices, it’s possible to achieve
impressive results in terms of both accuracy and speed. The flexibility to adjust model
parameters, such as input size and alpha values, allows for fine-tuning to meet specific project
requirements.
Looking forward, the field of edge AI and image classification continues to evolve rapidly.
Advances in model compression techniques, hardware acceleration, and more efficient neural
network architectures promise to further expand the capabilities of edge devices in computer
vision tasks.
This project serves as a foundation for more complex computer vision applications and encour-
ages further exploration into the exciting world of edge AI and IoT. Whether it’s for industrial
139
automation, smart home applications, or environmental monitoring, the skills and concepts
covered here provide a solid starting point for a wide range of innovative projects.
Resources
• Dataset Example
• Python Scripts
• Edge Impulse Project
• Image Classification Project - Edge Impulse Notebook
140
141
Object Detection: Fundamentals
Figure 4: DALL·E prompt - A cover image for an ‘Object Detection’ chapter in a Raspberry Pi
tutorial, designed in the same vintage 1950s electronics lab style as previous covers.
The scene should prominently feature wheels and cubes, similar to those provided by
the user, placed on a workbench in the foreground. A Raspberry Pi with a connected
camera module should be capturing an image of these objects. Surround the scene
with classic lab tools like soldering irons, resistors, and wires. The lab background
should include vintage equipment like oscilloscopes and tube radios, maintaining the
detailed and nostalgic feel of the era. No text or logos should be included.
142
Introduction
Building on our exploration of image classification, we now turn to a more advanced computer
vision task: object detection. While image classification assigns a single label to an entire
image, object detection goes further by identifying and locating multiple objects within a single
image. This capability opens up many new applications and challenges, particularly in edge
computing and IoT devices like the Raspberry Pi.
Object detection combines classification and localization. It not only determines which objects
are present in an image but also pinpoints their locations, for example, by drawing bounding
boxes around them. This added complexity makes object detection a more powerful tool
for understanding visual scenes, but it also requires more sophisticated models and training
techniques.
In edge AI, where computational resources are constrained, implementing efficient object
detection models is crucial. The challenges we faced with image classification—balancing model
size, inference speed, and accuracy—are even more pronounced in object detection. However,
the rewards are also more significant, as object detection enables more nuanced and detailed
analysis of visual data.
Some applications of object detection on edge devices include:
As we put our hands into object detection, we’ll build on the concepts and techniques we
explored in image classification. We’ll examine popular object detection architectures designed
for efficiency, such as:
To learn more about object detection models, follow the tutorial A Gentle Intro-
duction to Object Recognition With Deep Learning.
143
Throughout this lab, we’ll cover the fundamentals of object detection and how it differs from
image classification. We’ll also learn how to train, fine-tune, test, optimize, and deploy popular
object detection architectures using a dataset created from scratch.
Object detection builds upon the foundations of image classification but extends its capabilities
significantly. To understand object detection, it’s crucial first to recognize its key differences
from image classification:
Image Classification:
Object Detection:
144
To visualize this difference, let’s consider an example:
This diagram illustrates the critical difference: image classification provides a single label for
the entire image, while object detection identifies multiple objects, their classes, and their
locations within the image.
1. Object Localization: This component identifies the location of objects within the image.
It typically outputs bounding boxes, rectangular regions encompassing each detected
object.
2. Object Classification: This component determines the class or category of each detected
object, similar to image classification but applied to each localized region.
• Multiple objects: An image may contain multiple objects of various classes, sizes, and
positions.
• Varying scales: Objects can appear at different sizes within the image.
• Occlusion: Objects may be partially hidden or overlapping.
145
• Background clutter: Distinguishing objects from complex backgrounds can be challenging.
• Real-time performance: Many applications require fast inference times, especially on edge
devices.
1. Two-stage detectors: These first propose regions of interest and then classify each region.
Examples include R-CNN and its variants (Fast R-CNN, Faster R-CNN).
2. Single-stage detectors: These predict bounding boxes (or centroids) and class probabilities
in a single forward pass through the network. Examples include YOLO (You Only Look
Once), EfficientDet, SSD (Single Shot Detector), and FOMO (Faster Objects, More
Objects). These are often faster and better suited to edge devices, such as the Raspberry
Pi.
Evaluation Metrics
• Intersection over Union (IoU) is a metric used to evaluate the accuracy of an object
detector. It measures the overlap between two bounding boxes: the Ground Truth
box (the manually labeled correct box) and the Predicted box (the box generated by
the object detection model). The IoU value is calculated by dividing the area of the
Intersection (the overlapping area) by the area of the Union (the total area covered by
both boxes). A higher IoU value indicates a better prediction.
146
• Mean Average Precision (mAP) is a widely used metric for evaluating the perfor-
mance of object detection models. It provides a single number that reflects a model’s
ability to accurately both classify and localize objects. The “mean” in mAP refers to
the average taken over all object classes in the dataset. The “average precision” (AP) is
calculated for each class, and then these AP values are averaged to get the final mAP
score. A high mAP score indicates that the model is excellent at identifying all objects
and placing a tight-fitting, accurate bounding box around them.
147
• Frames Per Second (FPS): Measures detection speed, crucial for real-time applications
on edge devices.
As we saw in the introduction, given an image or a video stream, an object detection model
can identify which of a known set of objects might be present and provide information about
their positions within the image.
You can test some common models online by visiting Object Detection - MediaPipe
Studio
On Kaggle, we can find the most common pre-trained TFLite models to use with the Raspberry
Pi, ssd_mobilenet_v1, and efficiendet. Those models were trained on the COCO (Common
Objects in Context) dataset, which contains over 200,000 labeled images across 91 categories.
Download the models and upload them to the ./models folder on the Raspberry Pi.
Alternatively, you can find the models and the COCO labels on GitHub.
148
For the first part of this lab, we will focus on a pre-trained 300x300 SSD-Mobilenet V1 model
and compare it with the 320x320 EfficientDet-lite0, also trained using the COCO 2017 dataset.
Both models were converted to a TensorFlow Lite format (4.2MB for the SSD Mobilenet and
4.6MB for the EfficientDet).
The model outputs up to ten detections per image, including bounding boxes,
class IDs, and confidence scores.
We should confirm the steps done on the last Hands-On Lab, Image Classification, as follows:
source ~/tflite/bin/activate
149
Creating a Working Directory:
Considering that we have created the Documents/TFLITE folder in the last Lab, let’s now create
the specific folders for this object detection lab:
cd Documents/TFLITE/
mkdir OBJ_DETECT
cd OBJ_DETECT
mkdir images
mkdir models
cd models
Let’s start a new notebook to follow all the steps to detect objects in an image:
Import the needed libraries:
import time
import numpy as np
import [Link] as plt
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
Download the model and labels from the folder models and save them in the models folder
under OBJ_DETECT.
Load the model and allocate tensors:
model_path = "./models/[Link]"
interpreter = Interpreter(model_path=model_path)
interpreter.allocate_tensors()
150
Get input and output tensors.
input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()
Input details will inform us how the model should be fed with an image. The shape of (1,
300, 300, 3) with a dtype of uint8 tells us that a non-normalized (pixel value range from
0 to 255) image with dimensions (300x300x3) should be input one by one (Batch Dimension:
1).
The output details include not only the labels (“classes”) and probabilities (“scores”), but
also the relative window positions of the bounding boxes (“boxes”), indicating where the object
is located in the image, and the number of detected objects (”num_detections”). The output
details also indicate that the model can detect up to 10 objects in the image.
151
So, for the above example, using the same cat image used with the Image Classification
Lab, looking for the output, we have a 76% probability of having found an object with a
class ID of 16 on an area delimited by a bounding box of [0.028011084, 0.020121813,
0.9886069, 0.802299]. Those four numbers are related to ymin, xmin, ymax, and xmax, the
box coordinates.
Considering that y ranges from the top (ymin) to the bottom (ymax) and x ranges from
left (xmin) to right (xmax), we have, in fact, the coordinates of the top-left corner and the
bottom-right one. With both edges and knowing the shape of the picture, it is possible to draw
a rectangle around the object, as shown in the figure below:
152
Next, we should find what class ID 16 means. Opening the file coco_labels.txt, we see that
each element has an associated index; inspecting index 16, we get, as expected, cat. The
probability is the value returned from the score.
Let’s now upload some images with multiple objects on them for testing.
img_path = "./images/cat_dog.jpeg"
orig_img = [Link](img_path)
153
[Link]("Original Image")
[Link]()
Based on the input details, let’s pre-process the image, changing its shape and expanding its
dimensions:
img = orig_img.resize((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
input_data = np.expand_dims(img, axis=0)
input_data.shape, input_data.dtype
The new input_data shape is(1, 300, 300, 3) with a dtype of uint8, which is compatible
with what the model expects.
Using the input_data, let’s run the interpreter, measure the latency, and get the output:
154
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))
With a latency of around 800 ms on the Raspi-Zero and 100 ms n the Raspi-5 , we can get four
distinct outputs:
boxes = interpreter.get_tensor(output_details[0]['index'])[0]
classes = interpreter.get_tensor(output_details[1]['index'])[0]
scores = interpreter.get_tensor(output_details[2]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[3]['index'])[0])
On a quick inspection, we can see that the model detected two objects with a score over 0.5:
for i in range(num_detections):
if scores[i] > 0.5: # Confidence threshold
print(f"Object {i}:")
print(f" Bounding Box: {boxes[i]}")
print(f" Confidence: {scores[i]}")
print(f" Class: {classes[i]}")
[Link](figsize=(12, 8))
[Link](orig_img)
for i in range(num_detections):
if scores[i] > 0.5: # Adjust threshold as needed
ymin, xmin, ymax, xmax = boxes[i]
155
(left, right, top, bottom) = (xmin * orig_img.width,
xmax * orig_img.width,
ymin * orig_img.height,
ymax * orig_img.height)
rect = [Link]((left, top), right-left, bottom-top,
fill=False, color='red', linewidth=2)
[Link]().add_patch(rect)
class_id = int(classes[i])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {scores[i]:.2f}',
color='red', fontsize=12, backgroundcolor='white')
The choice of the confidence threshold is crucial. For example, setting it to 0.2
will show false positives. A proper code should handle it.
156
EfficientDet
EfficientDet is not technically an SSD (Single Shot Detector) model, but it shares some
similarities and builds upon ideas from SSD and other object detection architectures:
1. EfficientDet:
• Developed by Google researchers in 2019
• Uses EfficientNet as the backbone network
• Employs a novel bi-directional feature pyramid network (BiFPN)
• It uses compound scaling to efficiently scale the backbone network and object
detection components.
2. Similarities to SSD:
• Both are single-stage detectors, meaning they perform object localization and
classification in a single forward pass.
• Both use multi-scale feature maps to detect objects at different scales.
3. Key differences:
• Backbone: SSD typically uses VGG or MobileNet, while EfficientDet uses Efficient-
Net.
• Feature fusion: SSD uses a simple feature pyramid, while EfficientDet uses the more
advanced BiFPN.
• Scaling method: EfficientDet introduces compound scaling for all components of the
network
4. Advantages of EfficientDet:
• Generally achieves better trade-offs between accuracy and efficiency than SSD and
many other object detection models.
• More flexible scaling enables a family of models with varying size-performance
trade-offs.
While EfficientDet is not an SSD model, it can be seen as an evolution of single-stage detection
architectures, incorporating more advanced techniques to improve efficiency and accuracy.
When using EfficientDet, we can expect outputs similar to those of SSD (e.g., bounding boxes
and class scores).
On GitHub, you can find another notebook exploring the EfficientDet model that
we did with SSD MobileNet.
157
Object Detection on a live stream
Object detection models can also detect objects in real-time using a camera. The captured
image should be the input for the trained and converted model. On the Raspberry Pi 4 or 5,
OpenCV can capture frames and display inference results on a desktop.
However, even without a desktop, it’s possible to create a live stream with a webcam to
detect objects in real time. For example, let’s start with the script developed for the Image
Classification app and adapt it for a Real-Time Object Detection Web Application Using
TensorFlow Lite and Flask.
Download the Python script object_detection_app.py from GitHub.
This app version should work for any TFLite/LiteRT models.
model_path = "./models/[Link]"
python object_detection_app.py
After starting, you should receive the message on the terminal (the IP is from my Raspberry):
158
* Running on [Link]
Press CTRL+C to quit
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Raspberry
Pi’s IP address). For example: [Link]
159
160
Let’s see a technical description of the key modules used in the object detection application:
1. LiteRT:
• Purpose: Efficient inference of machine learning models on edge devices.
• Why: LiteRT offers a smaller model size and optimized performance compared to
full TensorFlow, which is crucial for resource-constrained devices like the Raspberry
Pi. It supports hardware acceleration and quantization, further improving efficiency.
• Key functions: Interpreter for loading and running the model, get_input_details(),
and get_output_details() for interfacing with the model.
2. Flask:
• Purpose: Lightweight web framework for building backend servers.
• Why: Flask’s simplicity and flexibility make it ideal for rapidly developing and
deploying web applications. It’s less resource-intensive than larger frameworks
suitable for edge devices.
• Key components: route decorators for defining API endpoints, Response objects for
streaming video, render_template_string for serving dynamic HTML.
3. Picamera2:
• Purpose: Interface with the Raspberry Pi camera module.
• Why: Picamera2 is the latest library for controlling Raspberry Pi cameras, offering
improved performance and features over the original Picamera library.
• Key functions: create_preview_configuration() for setting up the camera,
capture_file() for capturing frames.
4. PIL (Python Imaging Library):
• Purpose: Image processing and manipulation.
• Why: PIL provides a wide range of image processing capabilities. It’s used here to
resize images, draw bounding boxes, and convert between image formats.
• Key classes: Image for loading and manipulating images, ImageDraw for drawing
shapes and text on images.
5. NumPy:
• Purpose: Efficient array operations and numerical computing.
• Why: NumPy’s array operations are much faster than pure Python lists, which is
crucial for efficiently processing image data and model inputs/outputs.
• Key functions: array() for creating arrays, expand_dims() for adding dimensions
to arrays.
6. Threading:
• Purpose: Concurrent execution of tasks.
161
• Why: Threading enables simultaneous frame capture, object detection, and web
server operation, which is crucial for maintaining real-time performance.
• Key components: Thread class creates separate execution threads, and Lock is used
for thread synchronization.
7. [Link]:
• Purpose: In-memory binary streams.
• Why: Allows efficient handling of image data in memory without needing temporary
files, improving speed and reducing I/O operations.
8. time:
• Purpose: Time-related functions.
• Why: Used for adding delays ([Link]()) to control frame rate and for perfor-
mance measurements.
9. jQuery (client-side):
• Purpose: Simplified DOM manipulation and AJAX requests.
• Why: It makes it easy to update the web interface dynamically and communicate
with the server without page reloads.
• Key functions: .get() and .post() for AJAX requests, DOM manipulation methods
for updating the UI.
1. Main Thread: Runs the Flask server, handling HTTP requests and serving the web
interface.
2. Camera Thread: Continuously captures frames from the camera.
3. Detection Thread: Processes frames using the LiteRT object detection model.
4. Frame Buffer: Shared memory space (protected by locks) storing the latest frame and
detection results.
This architecture enables efficient, real-time object detection while maintaining a responsive web
interface on a resource-constrained edge device, such as a Raspberry Pi. Threading and efficient
libraries, such as LiteRT and PIL, enable the system to process video frames in real-time, while
Flask and jQuery provide a user-friendly way to interact with them.
162
You can test the app with another pre-processed model, such as the EfficientDet, by changing
the app line:
model_path = "./models/lite-model_efficientdet_lite0_detection_metadata_1.tflite"
If we want to use the app with the SSD-MobileNetV2 model, trained in Edge
Impulse Studio with the “Box versus Wheel” dataset, the code should also be
adapted to the input details, as we explored in its notebook.
Conclusion
This lab has explored implementing object detection on edge devices such as the Raspberry
Pi, demonstrating the power and potential of running advanced computer vision tasks on
resource-constrained hardware. We examined the object detection models SSD-MobileNet and
EfficientDet, comparing their performance and trade-offs on edge devices.
The lab demonstrated a real-time object-detection web application, showing how these models
can be integrated into practical, interactive systems.
The ability to perform object detection on edge devices opens up numerous possibilities across
domains such as precision agriculture, industrial automation, quality control, smart home
applications, and environmental monitoring. By processing data locally, these systems can offer
reduced latency, improved privacy, and operation in environments with limited connectivity.
Looking ahead, potential areas for further exploration include: - Using a custom dataset
(labeled on Roboflow), walking through the process of training models using Edge Impulse
Studio and Ultralytics, and deploying them on Raspberry Pi. - To improve inference speed on
edge devices, explore various optimization methods, such as model quantization (TFLite int8)
and format conversion (e.g., to NCNN). - Implementing multi-model pipelines for more complex
tasks - Exploring hardware acceleration options for Raspberry Pi - Integrating object detection
with other sensors for more comprehensive edge AI systems - Developing edge-to-cloud solutions
that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the power
of AI directly into the physical world, opening up new frontiers in how we interact with and
understand our environment.
Resources
163
• Python Scripts
• Models
164
Custom Object Detection Project
165
Object Detection Project
In this chapter, we will develop a complete Object Detection project from data collection,
labelling, training, and deployment. As we did with the Image Classification project, the
trained and converted model will be used for inference.
We will use the same dataset to train 3 models: SSD-MobileNet V2, FOMO, and YOLO.
The Goal
All Machine Learning projects need to start with a goal. Let’s assume we are in an industrial
facility and must sort and count wheels and special boxes.
In other words, we should perform a multi-label classification, where each image can have three
classes:
166
• Wheel
Once we have defined our Machine Learning project goal, the next and most crucial step is
collecting the dataset. We can use a phone, the Raspi, or a mix to create the raw dataset (with
no labels). Let’s use the simple web app on our Raspberry Pi to view the QVGA (320 x 240)
captured images in a browser.
From GitHub, get the Python script get_img_data.py and open it in the terminal:
python3 get_img_data.py
• On the Raspberry Pi itself (if you have a GUI): Open a web browser and go to
[Link]
• From another device on the same network: Open a web browser and go to
[Link] (Replace <raspberry_pi_ip> with your Raspberry
Pi’s IP address). For example: [Link]
The
Python script creates a web-based interface for capturing and organizing image datasets using
a Raspberry Pi and its camera. It’s handy for machine learning projects that require labeled
image data, or not, as in our case here.
Access the web interface from a browser, enter a generic label for the images you want to
capture, and press Start Capture.
167
Note that the images to be captured will have multiple labels that should be defined
later.
Use the live preview to position the camera, then click Capture Image to save the images under
the current label (in this case, box-wheel).
168
When we have enough images, we can press Stop Capture. The captured images are saved in
the folder dataset/box-wheel:
169
Get around 60 images. Try to capture different angles, backgrounds, and light
conditions. FileZilla can transfer the raw dataset you created to your main computer.
Labeling Data
The next step in an Object Detect project is to create a labeled dataset. We should label the
raw dataset images, creating bounding boxes around each picture’s objects (box and wheel).
We can use labeling tools such as LabelImg, CVAT, Roboflow, or even the Edge Impulse Studio.
Once we have explored the Edge Impulse tool in other labs, let’s use Roboflow here.
We are using Roboflow (free version) here for two main reasons. 1) We can have an
auto-labeler, and 2) The annotated dataset is available in several formats and can
be used both on Edge Impulse Studio (we will use it for MobileNet V2 and FOMO
train) and on CoLab (YOLOv8 or YOLOv11 train), for example. An annotated
dataset created on Edge Impulse (Free account) cannot be used for training on
other platforms.
We should upload the raw dataset to Roboflow. Create a free account there and start a new
project, for example, (“box-versus-wheel”).
170
We will not go into great detail about the Roboflow process, as many tutorials are
already available.
Annotate
Once the project is created and the dataset is uploaded, you can use the “Auto-Label” tool to
generate annotations, or do it manually.
171
The Label Assist tool can be handy for the labeling process.
172
Note that you should also upload images with only a background, which should be saved w/o
any annotations using the Null Tool option.
173
Once all images are annotated, split them into training, validation, and test sets.
174
Data Pre-Processing
The last step in the dataset is preprocessing to generate a final training version. Let’s resize all
images to 320x320 and generate augmented versions of each image (augmentation) to create
new training examples from which our model can learn.
For augmentation, we will rotate the images (+/-15o ), crop, and vary the brightness and
exposure.
175
At the end of the process, we will have 153 images.
176
Now, you should export the annotated dataset in a format that Edge Impulse, Ultralitics, and
other frameworks/tools understand, for example, YOLOv8 (or v11). Let’s download a zipped
version of the dataset to our desktop.
177
Here, it is possible to review how the dataset was structured
178
There are 3 separate folders, one for each split (train/test/valid). For each of them,
there are 2 subfolders, images, and labels. The pictures are stored as image_id.jpg and
images_id.txt, where “image_id” is unique for every picture.
The labels file format will be class_id bounding box coordinates, where in our case, class_id
will be 0 for box and 1 for wheel. The numerical id (o, 1, 2…) will follow the alphabetical order
of the class name.
The [Link] file contains information about the dataset, such as the classes’ names (names:
['box', 'wheel']) following the YOLO format.
And that’s it! We are ready to start training using Edge Impulse Studio (as we will in the next
step), Ultralytics (as we will when discussing YOLO), or even training from scratch on CoLab
(as we did with the Cifar-10 dataset in the Image Classification lab).
179
Training an SSD MobileNet Model on Edge Impulse Studio
Go to Edge Impulse Studio, enter your credentials at Login (or create an account), and start
a new project.
Here, you can clone the project developed for this hands-on lab: Raspi - Object
Detection.
On the Project Dashboard tab, go down to Project info, and for Labeling method select
Bounding boxes (object detection)
In Studio, go to the Data acquisition tab, and in the UPLOAD DATA section, upload the raw
dataset from your computer.
We can use the Select a folder option, choosing, for example, the train folder on your
computer, which contains two sub-folders: images and labels. Select the Image label
format, “YOLO TXT”, upload it into the category Training, and press Upload data.
180
Repeat the process for the test data (upload both folders, test, and validation). At the end
of the upload process, you should end with the annotated dataset of 153 images split in the
train/test (84%/16%).
Note that labels will be stored at the labels files 0 and 1 , which are equivalent to
box and wheel.
The first thing to define when we enter the Create impulse step is to describe the target device
for deployment. A pop-up window will appear. We will select Raspberry 4, an intermediary
device between the Raspi-Zero and the Raspi-5.
This choice will not interfere with the training; it will only give us an idea about
the latency of the model on that specific target.
181
In this phase, you should define how to:
• Pre-processing consists of resizing the individual images. In our case, the images were
pre-processed on Roboflow, to 320x320 , so let’s keep it. The resize will not matter here
because the images are already squared. If you upload a rectangular image, squash it
(squared form, without cropping). Afterward, you could define if the images are converted
from RGB to Grayscale or not.
• Design a Model, in this case, “Object Detection.”
182
Preprocessing all dataset
In the section Image, select Color depth as RGB, and press Save parameters.
183
The Studio automatically moves to the next section, Generate features, where all samples will
be preprocessed, resulting in 480 objects: 207 boxes and 273 wheels.
184
The feature explorer shows that all samples exhibit a good separation after the feature
generation.
For training, we should select a pre-trained model. Let’s use the MobileNetV2 SSD FPN-
Lite (320x320 only).
185
It is a pre-trained object detection model that locates up to 10 objects in an image and outputs
a bounding box for each. The model is approximately 3.7 MB. It supports an RGB input at
320x320px.
• Epochs: 25
• Batch size: 32
• Learning Rate: 0.15.
For validation during training, 20% of the dataset (validation_dataset) will be spared.
186
As a result, the model achieves an overall precision score (based on COCO mAP) of 88.8%,
higher than the score on the test data (83.3%).
• TFLite model, which lets deploy the trained model as .tflite for the Raspberry Pi to
run it using Python.
• Linux (AARCH64), a binary for Linux (AARCH64), implements the Edge Impulse
Linux protocol, which lets us run our models on any Linux-based development board,
with SDKs such as Python. See the documentation for more information and setup
instructions.
Let’s deploy the TFLite model. On the Dashboard tab, go to Transfer learning model (int8
quantized) and click on the download icon:
187
Transfer the model from your computer to the Raspi folder./models and capture or get some
images for inference and save them in the folder ./images.
188
Inference and Post-Processing
The inference can be made as discussed in the Pre-Trained Object Detection Models Overview.
Let’s start a new notebook to follow all the steps to detect cubes and wheels in an image.
Import the needed libraries:
import time
import numpy as np
import [Link] as plt
import [Link] as patches
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
model_path = "./models/ei-raspi-object-detection-SSD-MobileNetv2-320x0320-\
[Link]"
labels = ['box', 'wheel']
Remember that the model will output the class ID as values (0 and 1), following an
alphabetic order regarding the class names.
Load the model, allocate the tensors, and get the input and output tensor details:
One crucial difference to note is that the dtype of the input details of the model is now int8,
which means that the input values go from -128 to +127, while each pixel of our raw image
goes from 0 to 256. This means that we should pre-process the image to match it. We can
check here:
input_dtype = input_details[0]['dtype']
input_dtype
numpy.int8
189
So, let’s open the image and show it:
190
And perform the pre-processing:
Checking the input data, we can verify that the input tensor is compatible with what is expected
191
by the model:
input_data.shape, input_data.dtype
Now, it is time to perform the inference. Let’s also calculate the latency of the model:
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to milliseconds
print ("Inference time: {:.1f}ms".format(inference_time))
The model will take around 600ms to perform the inference in the Raspi-Zero, which is around
5 times longer than a Raspi-5.
Now, we can get the output classes of objects detected, its bounding boxes coordinates, and
probabilities.
boxes = interpreter.get_tensor(output_details[1]['index'])[0]
classes = interpreter.get_tensor(output_details[3]['index'])[0]
scores = interpreter.get_tensor(output_details[0]['index'])[0]
num_detections = int(interpreter.get_tensor(output_details[2]['index'])[0])
for i in range(num_detections):
if scores[i] > 0.5: # Confidence threshold
print(f"Object {i}:")
print(f" Bounding Box: {boxes[i]}")
print(f" Confidence: {scores[i]}")
print(f" Class: {classes[i]}")
192
From the results, we can see that 4 objects were detected: two with class ID 0 (box)and two
with class ID 1 (wheel), what is correct!
Let’s visualize the result for a threshold of 0.5
threshold = 0.5
[Link](figsize=(6,6))
[Link](orig_img)
for i in range(num_detections):
if scores[i] > threshold:
ymin, xmin, ymax, xmax = boxes[i]
(left, right, top, bottom) = (xmin * orig_img.width,
xmax * orig_img.width,
ymin * orig_img.height,
ymax * orig_img.height)
rect = [Link]((left, top), right-left, bottom-top,
fill=False, color='red', linewidth=2)
[Link]().add_patch(rect)
class_id = int(classes[i])
class_name = labels[class_id]
[Link](left, top-10, f'{class_name}: {scores[i]:.2f}',
color='red', fontsize=12, backgroundcolor='white')
193
But what happens if we reduce the threshold to 0.3, for example?
194
We start to see false positives and multiple detections, where the model detects the same
object multiple times with different confidence levels and slightly different bounding boxes.
Commonly, sometimes, we need to adjust the threshold to smaller values to capture all objects,
avoiding false negatives, which would lead to multiple detections.
To improve the detection results, we should implement Non-Maximum Suppression (NMS),
which helps eliminate overlapping bounding boxes and keeps only the most confident detection.
For that, let’s create a general function named non_max_suppression(), with the role of
refining object detection results by eliminating redundant and overlapping bounding boxes.
195
It achieves this by iteratively selecting the detection with the highest confidence score and
removing other significantly overlapping detections based on an Intersection over Union (IoU)
threshold.
keep = []
while [Link] > 0:
i = order[0]
[Link](i)
xx1 = [Link](x1[i], x1[order[1:]])
yy1 = [Link](y1[i], y1[order[1:]])
xx2 = [Link](x2[i], x2[order[1:]])
yy2 = [Link](y2[i], y2[order[1:]])
return keep
How it works:
1. Sorting: It starts by sorting all detections by their confidence scores, highest to lowest.
2. Selection: It selects the highest-scoring box and adds it to the final list of detections.
3. Comparison: This selected box is compared with all remaining lower-scoring boxes.
4. Elimination: Any box that overlaps significantly (above the IoU threshold) with the
selected box is eliminated.
196
5. Iteration: This process repeats with the next highest-scoring box until all boxes are
processed.
Now, we can define a more precise visualization function that will take into consideration an
IoU threshold, detecting only the objects that were selected by the non_max_suppression
function:
# Apply NMS
keep = non_max_suppression(boxes_pixel, scores, iou_threshold)
[Link](image_np)
for i in keep:
if scores[i] > threshold:
ymin, xmin, ymax, xmax = boxes[i]
rect = [Link]((xmin * width, ymin * height),
(xmax - xmin) * width,
(ymax - ymin) * height,
linewidth=2, edgecolor='r', facecolor='none')
ax.add_patch(rect)
class_name = labels[int(classes[i])]
[Link](xmin * width, ymin * height - 10,
f'{class_name}: {scores[i]:.2f}', color='red',
fontsize=12, backgroundcolor='white')
[Link]()
Now we can create a function that will call the others, performing inference on any image:
197
def detect_objects(img_path, conf=0.5, iou=0.5):
orig_img = [Link](img_path)
scale, zero_point = input_details[0]['quantization']
img = orig_img.resize((input_details[0]['shape'][1],
input_details[0]['shape'][2]))
img_array = [Link](img, dtype=np.float32) / 255.0
img_array = (img_array / scale + zero_point).clip(-128, 127).\
astype(np.int8)
input_data = np.expand_dims(img_array, axis=0)
# Inference on Raspi-Zero
start_time = [Link]()
interpreter.set_tensor(input_details[0]['index'], input_data)
[Link]()
end_time = [Link]()
inference_time = (end_time - start_time) * 1000 # Convert to ms
print ("Inference time: {:.1f}ms".format(inference_time))
Now, running the code, having the same image again with a confidence threshold of 0.3, but
with a small IoU:
img_path = "./images/box_2_wheel_2.jpg"
detect_objects(img_path, conf=0.3,iou=0.05)
198
Training a FOMO Model at Edge Impulse Studio
The inference with the SSD MobileNet model worked well, but the latency was significantly
high. The inference varied from 0.5 to 1.3 seconds on a Raspi-Zero, which means around or
less than 1 FPS (1 frame per second). One alternative to speed up the process is to use FOMO
(Faster Objects, More Objects).
This novel machine learning algorithm lets us count multiple objects and find their location in
199
an image in real-time using up to 30x less processing power and memory than MobileNet SSD
or YOLO. The main reason this is possible is that while other models calculate the object’s
size by drawing a square around it (bounding box), FOMO ignores the size of the image,
providing only the information about where the object is located in the image through its
centroid coordinates.
In a typical object detection pipeline, the first stage is extracting features from the input image.
FOMO leverages MobileNetV2 to perform this task. MobileNetV2 processes the input
image to produce a feature map that captures essential characteristics, such as textures, shapes,
and object edges, in a computationally efficient way.
Once these features are extracted, FOMO’s simpler architecture, focused on center-point
detection, interprets the feature map to determine where objects are located in the image. The
200
output is a grid of cells, where each cell represents whether or not an object center is detected.
The model outputs one or more confidence scores for each cell, indicating the likelihood of an
object being present.
Let’s see how it works on an image.
FOMO divides the image into blocks of pixels using a factor of 8. For the input of 96x96,
the grid would be 12x12 (96/8=12). For a 160x160, the grid will be 20x20, and so on. Next,
FOMO will run a classifier through each pixel block to calculate the probability that there is a
box or a wheel in each of them and, subsequently, determine the regions that have the highest
probability of containing the object (If a pixel block has no objects, it will be classified as
background). From the overlap of the final region, the FOMO provides the coordinates (related
to the image dimensions) of the centroid of this region.
• Grid Resolution: FOMO uses a grid of fixed resolution, meaning each cell can detect if
an object is present in that part of the image. While it doesn’t provide high localization
201
accuracy, it makes a trade-off by being fast and computationally light, which is crucial
for edge devices.
• Multi-Object Detection: Since each cell is independent, FOMO can detect multiple
objects simultaneously in an image by identifying multiple centers.
Return to Edge Impulse Studio, and in the Experiments tab, create another impulse. Now,
the input images should be 160x160 (this is the expected input size for MobilenetV2).
On the Image tab, generate the features and go to the Object detection tab.
We should select a pre-trained model for training. Let’s use the FOMO (Faster Objects,
More Objects) MobileNetV2 0.35.
202
Regarding the training hyper-parameters, the model will be trained with:
• Epochs: 30
• Batch size: 32
• Learning Rate: 0.001.
For validation during training, 20% of the dataset (validation_dataset) will be spared. We will
not apply Data Augmentation for the remaining 80% (train_dataset) because our dataset was
already augmented during the labeling phase at Roboflow.
As a result, the model ends with an overall F1 score of 93.3% with an impressive latency of
8ms (Raspi-4), around 60X less than we got with the SSD MovileNetV2.
203
Note that FOMO automatically added a third label background to the two previously
defined boxes (0) and wheels (1).
On the Model testing tab, we can see that the accuracy was 94%. Here is one of the test
sample results:
204
In object detection tasks, accuracy is generally not the primary evaluation metric.
Object detection involves classifying objects and providing bounding boxes around
them, making it a more complex problem than simple classification. The issue is
that we do not have the bounding box, only the centroids. In short, using accuracy
as a metric could be misleading and may not provide a complete understanding of
how well the model is performing.
As we did in the previous section, we can deploy the trained model as TFLite or Linux
(AARCH64). Let’s do it now as Linux (AARCH64), a binary that implements the Edge
Impulse Linux protocol.
Edge Impulse for Linux models is delivered in .eim format. This executable contains our “full
impulse” created in Edge Impulse Studio. The impulse consists of the signal processing block(s)
and any learning and anomaly block(s) we added and trained. It is compiled with optimizations
for our processor or GPU (e.g., NEON instructions on ARM cores), plus a straightforward IPC
layer (over a Unix socket).
At the Deploy tab, select the option Linux (AARCH64), the int8model and press Build.
205
The model will be automatically downloaded to your computer.
206
On our Raspi, let’s create a new working area:
cd ~
cd Documents
mkdir EI_Linux
cd EI_Linux
mkdir models
mkdir images
The inference will be made using the Linux Python SDK. This library lets us run machine
learning models and collect sensor data on Linux machines using Python. The SDK is open
source and available on GitHub at edgeimpulse/linux-sdk-python.
Let’s set up a Virtual Environment for working with the Linux Python SDK
207
chmod +x [Link]
jupyter notebook
Let’s start a new notebook by following all the steps to detect cubes and wheels on an image
using the FOMO model and the Edge Impulse Linux Python SDK.
Import the needed libraries:
model_file = "[Link]"
model_path = "models/"+ model_file # Trained ML model from Edge Impulse
labels = ['box', 'wheel']
Remember that the model will output the class ID as values (0 and 1), following an
alphabetic order regarding the class names.
# Initialize model
model_info = [Link]()
208
The model_info will contain critical information about our model. However, unlike the TFLite
interpreter, the EI Linux Python SDK library will now prepare the model for inference.
So, let’s open the image and show it (Now, for compatibility, we will use OpenCV, the CV
Library used internally by EI. OpenCV reads the image as BGR, so we will need to convert it
to RGB :
209
Now we will get the features and the preprocessed image (cropped) using the runner:
And perform the inference. Let’s also calculate the latency of the model:
res = [Link](features)
Let’s get the output classes of objects detected, their bounding boxes centroids, and probabili-
ties.
The results show that two objects were detected: one with class ID 0 (box) and one with class
ID 1 (wheel), which is correct!
Let’s visualize the result (The threshold is 0.5, the default value set during the model testing
on the Edge Impulse Studio).
210
top = bbox['y']
width = bbox['width']
height = bbox['height']
211
Conclusion
This chapter has explored the implementation of a custom object detector on edge devices, such
as the Raspberry Pi, demonstrating the power and potential of running advanced computer
vision tasks on resource-constrained hardware. We’ve covered several vital aspects:
212
and FOMO, and compared their performance and trade-offs on edge devices.
2. Training and Deployment: Using a custom dataset of boxes and wheels (labeled on
Roboflow), we walked through the process of training models with Edge Impulse Studio
and Ultralytics and deploying them on a Raspberry Pi.
3. Optimization Techniques: To improve inference speed on edge devices, we explored
various optimization methods, such as model quantization (int8).
4. Performance Considerations: Throughout the lab, we discussed the balance between
model accuracy and inference speed, a critical consideration for edge AI applications.
As discussed earlier, the ability to perform object detection on edge devices opens up numerous
possibilities across domains, such as precision agriculture, industrial automation, quality control,
smart home applications, and environmental monitoring. By processing data locally, these
systems can offer reduced latency, improved privacy, and operation in environments with limited
connectivity.
Looking ahead, potential areas for further exploration include: - Implementing multi-model
pipelines for more complex tasks - Exploring hardware acceleration options for Raspberry Pi
- Integrating object detection with other sensors for more comprehensive edge AI systems -
Developing edge-to-cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the power
of AI directly into the physical world, opening up new frontiers in how we interact with and
understand our environment.
Resources
213
Computer Vision Applications with YOLO
In this chapter, we will explore YOLOv8 and v11. Ultralytics YOLO (v8 and v11) are versions
of the acclaimed real-time object detection and image segmentation model, YOLO. YOLOv8
and v11 are built on cutting-edge advances in deep learning and computer vision, offering
unparalleled speed and accuracy. Its streamlined design makes it suitable for a wide range
of applications and easily adaptable across hardware platforms, from edge devices to cloud
APIs.
214
Talking about the YOLO Model
The YOLO (You Only Look Once) model is a highly efficient, widely used object detection
algorithm known for its real-time performance. Unlike traditional object detection systems that
repurpose classifiers or localizers to perform detection, YOLO frames the detection problem as
a single regression task. This innovative approach enables YOLO to simultaneously predict
multiple bounding boxes and their class probabilities from full images during a single evaluation,
significantly boosting its speed.
Key Features:
• YOLO employs a single neural network to process the entire image. This network
divides the image into a grid and, for each grid cell, directly predicts bounding boxes
and associated class probabilities. This end-to-end training improves speed and
simplifies the model architecture.
2. Real-Time Processing:
• One of YOLO’s standout features is its ability to perform object detection in real-
time. Depending on the version and hardware, YOLO can process images at high
frames per second (FPS). This makes it ideal for applications requiring quick and
accurate object detection, such as video surveillance, autonomous driving, and live
sports analysis.
3. Evolution of Versions:
• Over the years, YOLO has undergone significant improvements, from YOLOv1
to the latest YOLOv12. Each iteration has introduced enhancements in accuracy,
speed, and efficiency. YOLOv8, for instance, incorporates advancements in net-
work architecture, improved training methodologies, and better support for various
hardware, ensuring a more robust performance.
• YOLOv11 offers substantial improvements in accuracy, speed, and parameter effi-
ciency compared to prior versions such as YOLOv8 and YOLOv10, making it one
of the most versatile and powerful real-time object detection models available as of
2025
215
4. Accuracy and Efficiency:
• While early versions of YOLO traded off some accuracy for speed, recent versions
have made substantial strides in balancing both. The newer models are faster and
more accurate, detecting small objects (such as bees) and performing well on complex
datasets.
• YOLO’s versatility has led to its adoption in numerous fields. It is used in traffic
monitoring systems to detect and count vehicles, security applications to identify
potential threats and agricultural technology to monitor crops and livestock. Its
application extends to any domain requiring efficient and accurate object detection.
7. Model Capabilities
YOLO models support multiple computer vision tasks:
216
Ultralitics YOLO Detect, Segment, and Pose models pre-trained on the COCO
dataset, and Classify on the ImageNet dataset.
Track mode is available for all Detect, Segment, and Pose models. The latest versions of
YOLO can also perform OBB, which stands for Oriented Bounding Box, a rectangular
box in computer vision that can rotate to match the orientation of an object within
an image, providing a much tighter and more precise fit than traditional axis-aligned
bounding boxes.
YOLO offers several model variants optimized for different use cases, for example. The
YOLOv8:
Installation
python --version
217
As of today (January 2026), Ultralytics officially supports only Python 3.9-3.12; Python
3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships with
Python 3.13 by default, we’ll need to install a compatible Python version alongside it.
One solution is to install Pyenv, so that we can easily manage multiple Python versions for
different projects without affecting the system Python.
If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.
Install pyenv
Configure Shell
# pyenv configuration
export PYENV_ROOT="$HOME/.pyenv"
[[ -d $PYENV_ROOT/bin ]] && export PATH="$PYENV_ROOT/bin:$PATH"
eval "$(pyenv init -)"
EOF
218
source ~/.bashrc
pyenv --version
cd Documents
mkdir YOLO
cd YOLO
# Verify
python --version # Should show Python 3.11.14
219
# Verify if we're using the correct Python
which python
python --version
deactivate
Installing Ultralytics/Yolo
cd Documents/
cd YOLO
mkdir models
mkdir images
And install the Ultralytics packages for local inference on the Raspberry Pi (inside the env)
1. Update the packages list, install and/or upgrade PIP to the latest:
sudo reboot
220
Testing the YOLO
After the Raspi booting, let’s activate the yolo env, go to the working directory,
cd /Documents/YOLO
source ~/yolo_env/bin/activate
And run inference on an image that will be downloaded from the Ultralytics website, using, for
example, the YOLOV8n model (the smallest in the family) at the Terminal (CLI):
Note that the first time we invoke a model, it will automatically be downloaded to
the current directory.
The inference result will appear in the terminal. In the image ([Link]), 4 persons, 1 bus,
and 1 stop signal were detected:
221
222
So, the Ultrayitics YOLO is correctly installed on our Raspberry Pi. Note that on the Raspberry
Pi Zero, an issue is the high latency for this inference, which takes several seconds, even with
the most compact model in the family (YOLOv8n).
The procedure is the same as we did with version v8. As a comparison, we can see that the
YOLOv11 is faster than the v8, but seems a little less precise, as it does not detect the “stop
sign” as the v8.
Deploying computer vision models on edge devices with limited computational power, such as
the Raspberry Pi Zero, can cause latency issues. One alternative is to use a format optimized
for optimal performance. This ensures that even devices with limited processing power can
handle advanced computer vision tasks well.
Of all the model export formats supported by Ultralytics, the NCNN is a high-performance
neural network inference computing framework optimized for mobile platforms. From the
beginning of the design, NCNN was deeply considerate of deployment and use on mobile phones,
and it did not have third-party dependencies. It is cross-platform and runs faster than all
known open-source frameworks (such as TFLite).
NCNN delivers the best inference performance when working with Raspberry Pi devices. NCNN
is highly optimized for mobile embedded platforms (such as ARM architecture).
Let’s move the downloaded YOLO models to the ./models folder and [Link] to
./images.
223
And convert our models and rerun the inferences:
The first inference, when the model is loaded, typically has a high latency; however,
from the second inference, it is possible to note that the inference time decreases.
We can now realize that neither model detects the “Stop Signal”, with YOLOv11 being the
fastest. The optimized models are more rapid but also less accurate.
To start, let’s call the Python Interpreter so we can explore how the YOLO model works, line
by line:
python
Now, we should call the YOLO library from Ultralitics and load the model:
224
from ultralytics import YOLO
model = YOLO('./models/yolov8n_ncnn_model')
img = './images/[Link]'
result = [Link](img, save=True, imgsz=640, conf=0.5, iou=0.3)
We can verify that the result is almost identical to the one we get running the inference at
the terminal level (CLI), except that the bus stop was not detected with the reduced NCNN
model. Note that the latency was reduced.
Let’s analyze the “result” content.
For example, we can see result[0].[Link], showing us the main inference result, which
is a tensor with a shape of (4, 6). Each line is one of the objects detected, being the first four
columns, the bounding boxes coordinates, the 5th, the confidence, and the 6th, the class (in
this case, 0: person and 5: bus):
We can access several inference results separately, as the inference time, and have it printed in
a better format:
225
inference_time = int(result[0].speed['inference'])
print(f"Inference Time: {inference_time} ms")
With Python, we can create a detailed output that meets our needs (See Model Prediction with
Ultralytics YOLO for more details). Let’s run a Python script instead of manually entering it
line by line in the interpreter, as shown below. Let’s use nano as our text editor. First, we
should create an empty Python script named, for example, yolov8_tests.py:
nano yolov8_tests.py
# Run inference
img = './images/[Link]'
result = [Link](img, save=False, imgsz=640, conf=0.5, iou=0.3)
226
And enter with the commands: [CTRL+O] + [ENTER] +[CTRL+X] to save the Python script.
Run the script:
python yolov8_tests.py
The result is the same as running the inference at the terminal level (CLI) and with the built-in
Python interpreter.
Calling the YOLO library and loading the model for inference for the first time
takes a long time, but the inferences after that will be much faster. For example,
the first single inference can take several seconds, but after that, the inference time
should be reduced to less than 1 second.
Inference Arguments
[Link]() accepts multiple arguments that can be passed at inference time to override
defaults:
Inference arguments:
227
Argument Type Default Description
source str Specifies the data source for inference. Can be
'ultralytics/assets'
an image path, video file, directory, URL, or
device ID for live feeds. Supports a wide
range of formats and sources, enabling flexible
application across different types of input.
conf float 0.25 Sets the minimum confidence threshold for
detections. Objects detected with confidence
below this threshold will be disregarded.
Adjusting this value can help reduce false
positives.
iou float 0.7 Intersection Over Union (IoU) threshold for
Non-Maximum Suppression (NMS). Lower
values result in fewer detections by
eliminating overlapping boxes, useful for
reducing duplicates.
imgsz int or 640 Defines the image size for inference. Can be a
tuple single integer 640 for square resizing or a
(height, width) tuple. Proper sizing can
improve detection accuracy and processing
speed.
rect bool True If enabled, minimally pads the shorter side of
the image until it’s divisible by stride to
improve inference speed. If disabled, pads the
image to a square during inference.
half bool False Enables half-precision (FP16) inference, which
can speed up model inference on supported
GPUs with minimal impact on accuracy.
device str None Specifies the device for inference (e.g., cpu,
cuda:0 or 0). Allows users to select between
CPU, a specific GPU, or other compute
devices for model execution.
batch int 1 Specifies the batch size for inference (only
works when the source is a directory, video file
or .txt file). A larger batch size can provide
higher throughput, shortening the total
amount of time required for inference.
max_det int 300 Maximum number of detections allowed per
image. Limits the total number of objects the
model can detect in a single inference,
preventing excessive outputs in dense scenes.
228
Argument Type Default Description
vid_stride int 1 Frame stride for video inputs. Allows skipping
frames in videos to speed up processing at the
cost of temporal resolution. A value of 1
processes every frame, higher values skip
frames.
stream_buffer
bool False Determines whether to queue incoming frames
for video streams. If False, old frames get
dropped to accommodate new frames
(optimized for real-time applications). If True,
queues new frames in a buffer, ensuring no
frames get skipped, but will cause latency if
inference FPS is lower than stream FPS.
visualize bool False Activates visualization of model features
during inference, providing insights into what
the model is “seeing”. Useful for debugging
and model interpretation.
augment bool False Enables test-time augmentation (TTA) for
predictions, potentially improving detection
robustness at the cost of inference speed.
agnostic_nmsbool False Enables class-agnostic Non-Maximum
Suppression (NMS), which merges overlapping
boxes of different classes. Useful in multi-class
detection scenarios where class overlap is
common.
classes list[int] None Filters predictions to a set of class IDs. Only
detections belonging to the specified classes
will be returned. Useful for focusing on
relevant objects in multi-class detection tasks.
retina_masksbool False Returns high-resolution segmentation masks.
The returned masks ([Link]) will match
the original image size if enabled. If disabled,
they have the image size used during
inference.
embed list[int] None Specifies the layers from which to extract
feature vectors or embeddings. Useful for
downstream tasks like clustering or similarity
search.
project str None Name of the project directory where
prediction outputs are saved if save is
enabled.
229
Argument Type Default Description
name str None Name of the prediction run. Used for creating
a subdirectory within the project folder,
where prediction outputs are stored if save is
enabled.
stream bool False Enables memory-efficient processing for long
videos or numerous images by returning a
generator of Results objects instead of loading
all frames into memory at once.
verbose bool True Controls whether to display detailed inference
logs in the terminal, providing real-time
feedback on the prediction process.
Visualization arguments:
230
Argument Type Default Description
show_conf bool True Displays the confidence score for each detection
alongside the label. Gives insight into the model’s
certainty for each detection.
show_boxes bool True Draws bounding boxes around detected objects.
Essential for visual identification and location of
objects in images or video frames.
line_width None or None Specifies the line width of bounding boxes. If None,
int the line width is automatically adjusted based on
the image size. Provides visual customization for
clarity.
Let’s set up Jupyter Notebook optimized for headless Raspberry Pi camera work and develop-
ment:
To run Jupyter Notebook, run the command (change the IP address for yours):
On the terminal, you can see the local URL address and its Token to open the notebook. Copy
and paste it into the Browser.
import time
import numpy as np
from PIL import Image
from ultralytics import YOLO
import [Link] as plt
Here we have all the necessary libraries, which we installed automatically when we installed
Ultralytics.
231
• Time: Performance measurement and benchmarking
• NumPy: Numerical computations and array operations
• PIL (Python Imaging Library): Image loading and manipulation
• Ultralytics YOLO: Core YOLO functionality
• Matplotlib: Visualization and plotting results
model_path= "./models/[Link]"
task = "detect"
verbose = False
• Model Selection: YOLOv11n (nano) is chosen for its balance of speed and accuracy
• Task Specification: We will select detect, which in fact is the default for the model.
But remember that YOLO supports multiple computer vision tasks, which will be explored
later.
• Verbose Control: output model information during model initialization
Performance Characteristics
source = [Link]("./images/[Link]")
From the inference results info, we can see that the first time an inference is run, the latency is
greater.
# First inference
0: 640x480 4 persons, 1 bus, 7528.3ms
# Second inference
0: 640x480 4 persons, 1 bus, 2822.1ms
232
The dramatic difference between the first inference (7.5s) and subsequent inferences (2.8s)
illustrates:
result = results[0]
# - boxes, keypoints, masks, names
# - orig_img, orig_shape, path
# - speed metrics
The Ultralytics plot() can be customized to show as the detection result, for example, only
the bounding boxes:
233
img = [Link](im_bgr[..., ::-1])
[Link](figsize=(6, 6))
[Link](img)
#[Link]('off') # This turns off the axis numbers
[Link]("YOLO Result")
[Link]()
234
235
Customization Options:
The plot() method in Ultralytics YOLO Results object accepts several arguments to control
what is visualized on the image, including boxes, masks, keypoints, confidences, labels, and
more. Common Arguments for plot()
Image Classification
As explored in previous chapters, the output of an image classifier is a single class label and a
confidence score. Image classification is useful when we need to know only what class an image
belongs to and don’t need to know where objects of that class are located or what their exact
shape is.
model_path= "./models/[Link]"
task = "clasification"
Note that a specific variation of the model, for image classification, will be downloaded. Now,
let’s do an inference, using the same bus image:
0: 224x224 minibus 0.57, police_van 0.34, trolleybus 0.04, recreational_vehicle 0.01, streetc
Speed: 5233.9ms preprocess, 3355.1ms inference, 28.2ms postprocess per image at shape (1, 3,
236
We can check the top5 inference results using Python:
classes = [Link].top5
classes
for id in classes:
print([Link][id])
minibus
police_van
trolleybus
recreational_vehicle
streetcar
probs = [Link]()
probs
[0.5710113048553467,
0.33745330572128296,
0.04209813103079796,
0.014150412753224373,
0.005880324635654688]
print([Link][[Link].top1],
round([Link](), 2))
minibus 0.57
Instance Segmentation
237
model_path= "./models/[Link]"
task = "segment"
Note that a specific variation of the model, for instance segmentation, will be downloaded.
Now, lt’s use another image for testing:
source = [Link]("./images/[Link]")
[Link](figsize=(6, 6))
[Link](source)
#[Link]('off') # This turns off the axis numbers
[Link]("Original Image")
[Link]()
238
results = [Link](source, save=False)
result = results[0]
Pose Estimation
model_path= "./models/[Link]"
task = "pose"
239
model = YOLO(model_path, task, verbose)
source = [Link]("./images/[Link]")
results = [Link](source, save=False)
result = results[0]
240
Training YOLO on a Customized Dataset
We will now develop a customized object detection project from the data collected and labelled
with Roboflow. The training and deployment will be done in Python using a CoLab and
Ultralytics functions.
We will use with YOLO, the same dataset previously used to train the SSD-
MobileNet V2 and FOMO models.
As a reminder, we are assuming we are in an industrial facility that must sort and count wheels
and special boxes.
241
Each image can have three classes:
The Dataset
Return to our “Boxe versus Wheel” dataset, labeled on Roboflow. On the Download Dataset,
instead of Download a zip to computer option done for training on Edge Impulse Studio, we
will opt for Show download code. This option will open a pop-up window with a code snippet
that should be pasted into our training notebook.
242
For training, let’s choose one model (let’s say YOLOv8) and adapt one of the publicly available
examples from Ultralytics, then run it on Google Colab. Below, you can find my adaptation:
243
3. Now, you can import the YOLO and upload your dataset to the CoLab, pasting the
Download code that we get from Roboflow. Note that our dataset will be mounted under
/content/datasets/:
4. It is essential to verify and change the file [Link] with the correct path for the images
(copy the path on each images folder).
names:
- box
- wheel
nc: 2
roboflow:
license: CC BY 4.0
project: box-versus-wheel-auto-dataset
url: [Link]
244
version: 5
workspace: marcelo-rovai-riila
test: /content/datasets/Box-versus-Wheel-auto-dataset-5/test/images
train: /content/datasets/Box-versus-Wheel-auto-dataset-5/train/images
val: /content/datasets/Box-versus-Wheel-auto-dataset-5/valid/images
5. Define the main hyperparameters that you want to change from default, for example:
MODEL = '[Link]'
IMG_SIZE = 640
EPOCHS = 25 # For a final project, you should consider at least 100 epochs
Figure 5: image-20240910111319804
The model took a few minutes to be trained and has an excellent result (mAP50 of 0.995). At the
end of the training, all results are saved in the folder listed, for example: /runs/detect/train/.
There, you can find, for example, the confusion matrix.
245
7. Note that the trained model ([Link]) is saved in the folder /runs/detect/train/weights/.
Now, you should validate the trained model with the valid/images.
8. Now, we should perform inference on the images left aside for testing
246
!yolo task=detect mode=predict model={HOME}/runs/detect/train/weights/[Link] conf=0.25 sourc
The inference results are saved in the folder runs/detect/predict. Let’s see some of them:
9. It is advised to export the train, validation, and test results for a Drive at Google. To do
so, we should mount the drive.
from [Link] import drive
[Link]('/content/gdrive')
and copy the content of /runs folder to a folder that you should create in your Drive, for
example:
!scp -r /content/runs '/content/gdrive/MyDrive/10_UNIFEI/Box_vs_Wheel_Project'
247
cd ..
python
We will import the YOLO library and define the model to use::
Now, let’s define an image and call the inference (we will save the image result this time to
external verification):
Let’s repeat for several images. The inference result is saved on the variable result, and the
processed image on runs/detect/predict8
Using FileZilla FTP, we can send the inference result to our Desktop for verification:
248
We can see that the inference result is excellent! The model was trained based on the smaller
base model of the YOLOv8 family (YOLOv8n). The issue is the latency, around 1 second (or
1 FPS on the Raspi-Zero). We can reduce this latency and convert the model to TFLite or
NCNN.
The model trained with YOLO11 has a latency of around 800 ms, similar to the
result of v8 with ncnn.
In the last section of the notebook, we can find inferences made with the trained YOLO11n
model on a Raspberry 5, which took around 400ms:
249
Conclusion
This chapter has explored the YOLO model and the implementation of a custom object detector
on a Raspberry Pi, demonstrating the power and potential of running advanced computer
vision tasks on resource-constrained hardware. We’ve covered several vital aspects:
250
AS discussed before, the ability to perform object detection on edge devices opens up numerous
possibilities across various domains, including precision agriculture, industrial automation,
quality control, smart home applications, and environmental monitoring. By processing data
locally, these systems can offer reduced latency, improved privacy, and operation in environments
with limited connectivity.
Looking ahead, potential areas for further exploration include: - Implementing multi-model
pipelines for more complex tasks - Exploring hardware acceleration options for Raspberry Pi
- Integrating object detection with other sensors for more comprehensive edge AI systems -
Developing edge-to-cloud solutions that leverage both local processing and cloud resources
Object detection on edge devices can create intelligent, responsive systems that bring the power
of AI directly into the physical world, opening up new frontiers in how we interact with and
understand our environment.
Resources
251
Counting objects with YOLO
252
Introduction
At the Federal University of Itajuba in Brazil, with the master’s student José Anderson Reis
and Professor José Alberto Ferreira Filho, we are exploring a project that delves into the
intersection of technology and nature. This tutorial will review our first steps and share our
observations on deploying YOLOv8, a cutting-edge machine learning model, on the compact
253
and efficient Raspberry Pi Zero 2W (Raspi-Zero). We aim to estimate the number of bees
entering and exiting their hive—a task crucial for beekeeping and ecological studies.
Why is this important? Bee populations are vital indicators of environmental health, and their
monitoring can provide essential data for ecological research and conservation efforts. However,
manual counting is labor-intensive and prone to errors. By leveraging the power of embedded
machine learning, or tinyML, we automate this process, enhancing accuracy and efficiency.
Figure 6: img
This tutorial will cover setting up the Raspberry Pi, integrating a camera module, optimizing
and deploying YOLOv8 for real-time image processing, and analyzing the data gathered.
For our project at the university, we are preparing to collect a dataset of bees at the entrance
of a beehive using the same camera connected to the Raspberry Pi. The images should be
collected every 10 seconds. With the Arducam OV5647, the horizontal Field of View (FoV)
is 53.5o , which means that a camera positioned at the top of a standard Hive (46 cm) will
capture all of its entrance (about 47 cm).
254
Dataset
The dataset collection is the most critical phase of the project and should take several weeks or
months. For this tutorial, we will use a public dataset: “Sledevic, Tomyslav (2023), “[Labeled
dataset for bee detection and direction estimation on beehive landing boards,” Mendeley Data,
V5, doi: 10.17632/8gb9r2yhfc.5”
The original dataset contains 6,762 images (1920 x 1080), and around 8% (518) of them have
no bees (only background). This is very important in Object Detection, where we should keep
around 10% of the dataset with only background (no objects to be detected).
255
The images contain from zero to up to 61 bees:
We downloaded the dataset (images and annotations) and uploaded it to Roboflow. There,
you should create a free account and start a new project, for example, (“Bees_on_Hive_land-
ing_boards”):
256
We will not enter details about the Roboflow process once many tutorials are
available.
Once the project is created and the dataset is uploaded, you should review the annotations
using the “Auto-Label” Tool. Note that all images with only a background should be saved
w/o any annotations. At this step, you can also add additional images.
257
Once all images are annotated, you should split them into training, validation, and testing.
258
Pre-Processing
The last step with the dataset is preprocessing to generate a final version for training. The
Yolov8 model can be trained with 640 x 640 pixels (RGB) images. Let’s resize all images and
generate augmented versions of each image (augmentation) to create new training examples
from which our model can learn.
For augmentation, we will rotate the images (+/-15o ) and vary the brightness and exposure.
259
This will create a final dataset of 16,228 images.
260
Now, you should export the annotateddataset in a YOLOv8 format. You can download a
zipped version of the dataset to your desktop or get a downloaded code to be used with a
Jupyter Notebook:
261
And that is it! We are prepared to start our training using Google Colab.
For training, let’s adapt one of the public examples available from Ultralitytics and run it on
Google Colab:
262
Critical points on the Notebook:
3. Now, you can import the YOLO and upload your dataset to the CoLab, pasting the
Download code that you get from Roboflow. Note that your dataset will be mounted
under /content/datasets/:
263
4. It is important to verify and change, if needed, the file [Link] with the correct path
for the images:
names:
- bee
nc: 1
roboflow:
license: CC BY 4.0
project: bees_on_hive_landing_boards
url: [Link]
version: 1
workspace: marcelo-rovai-riila
test: /content/datasets/Bees_on_Hive_landing_boards-1test/images
train: /content/datasets/Bees_on_Hive_landing_boards-1/train/images
val: /content/datasets/Bees_on_Hive_landing_boards-1/valid/images
5. Define the main hyperparameters that you want to change from default, for example:
MODEL = '[Link]'
IMG_SIZE = 640
EPOCHS = 25 # For a final project, you should consider at least 100 epochs
The model took 2.7 hours to train and has an excellent result (mAP50 of 0.984). At the end
of the training, all results are saved in the folder listed, for example: /runs/detect/train3/.
There, you can find, for example, the confusion matrix and the metrics curves per epoch.
264
7. Note that the trained model ([Link]) is saved in the folder /runs/detect/train3/weights/.
Now, you should validade the trained model with the valid/images.
8. Now, we should perform inference on the images left aside for testing
The inference results are saved in the folder runs/detect/predict. Let’s see some of them:
We can also perform inference with a completely new and complex image from another beehive
with a different background (the beehive of Professor Maurilio of our University). The results
were great (but not perfect and with a lower confidence score). The model found 41 bees.
265
9. The last thing to do is export the train, validation, and test results for your Drive at
Google. To do so, you should mount your drive.
from [Link] import drive
[Link]('/content/gdrive')
and copy the content of /runs folder to a folder that you should create in your Drive, for
example:
!scp -r /content/runs '/content/gdrive/MyDrive/10_UNIFEI/Bee_Project/YOLO/bees_on_hive_l
Using the FileZilla FTP, let’s transfer the [Link] to our Rasp-Zero (before the transfer, you
may change the model name, for example, bee_landing_640_best.pt).
The first thing to do is convert the model to an NCNN format:
266
As a result, a new converted model, bee_landing_640_best_ncnn_model is created in the
same directory.
Let’s create a folder to receive some test images (under Documents/YOLO/:
mkdir test_images
Using the FileZilla FTP, let’s transfer a few images from the test dataset to our Rasp-Zero:
python
As before, we will import the YOLO library and define our converted model to detect bees:
Now, let’s define an image and call the inference (we will save the image result this time to
external verification):
img = 'test_images/15_bees.jpg'
result = [Link](img, save=True, imgsz=640, conf=0.2, iou=0.3)
267
The inference result is saved on the variable result, and the processed image on
runs/detect/predict9
Using FileZilla FTP, we can send the inference result to our Desktop for verification:
268
let’s go over the other images, analyzing the number of objects (bees) found:
269
Depending on the confidence level, we may see some false positives or negatives. But in general,
with a model trained on a smaller base model in the YOLOv8 family (YOLOv8n) and converted
to NCNN, the results are pretty good, running on an Edge device such as the Rasp-Zero. Also,
note that the inference latency is around 730ms.
For example, by running the inference on [Link], we can find 40 bees. During
the test phase on Colab, 41 bees were found (we only missed one here).
Our final project should be very simple in terms of code. We will use the camera to capture an
image every 10 seconds. As we did in the previous section, the captured image should serve as
the input to the trained and converted model. We should count the bees in each image and
store the counts in a database (e.g., timestamp: number of bees).
We can do it with a single Python script, or use a Linux system timer, such as cron, to
periodically capture images every 10 seconds, and have a separate Python script process them
as they are saved. This method can be particularly efficient at managing system resources and
is more robust against potential delays in image processing.
270
Setting Up the Image Capture with cron
First, we should set up a cron job to use the rpicam-jpeg command to capture an image
every 10 seconds.
• Open the terminal and type crontab -e to edit the cron jobs.
• cron normally doesn’t support sub-minute intervals directly, so we should use a
workaround, such as a loop or a file watcher.
• Image Capture: This bash script captures images every 10 seconds using rpicam-
jpeg, a command in the raspijpeg tool. This command lets us control the camera
and capture JPEG images directly from the command line. This is especially useful
because we are looking for a lightweight, straightforward method to capture images
without requiring additional libraries like Picamera or external software. The script
also saves the captured image with a timestamp.
#!/bin/bash
# Script to capture an image every 10 seconds
while true
do
DATE=$(date +"%Y-%m-%d_%H%M%S")
rpicam-jpeg --output test_images/$[Link] --width 640 --height 640
sleep 10
done
Image Processing: The Python script continuously monitors the designated directory for
new images, processes each new image using the YOLOv8 model, updates the database with
the count of detected bees, and optionally deletes the image to conserve disk space.
Database Updates: The results, along with the timestamps, are saved in an SQLite database.
For that, a simple option is to use sqlite3.
In short, we need to write a script that continuously monitors the directory for new images,
processes them using a YOLO model, and then saves the results to a SQLite database. Here’s
how we can create and make the script executable:
271
#!/usr/bin/env python3
import os
import time
import sqlite3
from datetime import datetime
from ultralytics import YOLO
def setup_database():
"""
Establishes a database connection and creates the table
if it doesn't exist.
"""
conn = [Link](DB_PATH)
cursor = [Link]()
[Link]('''
CREATE TABLE IF NOT EXISTS bee_counts
(timestamp TEXT, count INTEGER)
''')
[Link]()
return conn
272
"""
Monitors the directory for new images and processes
them as they appear.
"""
processed_files = set()
while True:
try:
files = set([Link](IMAGES_DIR))
new_files = files - processed_files
for file in new_files:
if [Link]('.jpg'):
full_path = [Link](IMAGES_DIR, file)
process_image(full_path, model, conn)
processed_files.add(file)
[Link](1) # Check every second
except KeyboardInterrupt:
print("Stopping...")
break
def main():
conn = setup_database()
model = YOLO(MODEL_PATH)
monitor_directory(model, conn)
[Link]()
if __name__ == "__main__":
main()
We should consider keeping the script running even after closing the terminal; for that, we can
use nohup or screen:
273
nohup ./process_images.py &
or
screen -S bee_monitor
./process_images.py
Note that we capture images with their own timestamps and log a separate timestamp when the
inference results are saved to the database. This approach can be beneficial for the following
reasons:
#!/usr/bin/env python3
import sqlite3
def main():
db_path = 'bee_count.db'
conn = [Link](db_path)
274
cursor = [Link]()
query = "SELECT * FROM bee_counts"
[Link](query)
data = [Link]()
for row in data:
print(f"Timestamp: {row[0]}, Number of bees: {row[1]}")
[Link]()
if __name__ == "__main__":
main()
Besides bee counting, environmental data, such as temperature and humidity, are essential
for monitoring the bee-have health. Using a Rasp-Zero, it is straightforward to add a digital
sensor such as the DHT-22 to get this data.
Environmental data will be part of our final project. If you want to know more about connecting
sensors to a Raspberry Pi and, even more, how to save the data to a local database and send
275
it to the web, follow this tutorial: From Data to Graph: A Web Journey With Flask and
SQLite.
Conclusion
In this tutorial, we have thoroughly explored integrating the YOLOv8 model with a Raspberry
Pi Zero 2W to address the practical, pressing task of counting (or, better, “estimating”) bees
at a beehive entrance. Our project underscores the robust capability of embedding advanced
machine learning technologies within compact edge computing devices, highlighting their
potential impact on environmental monitoring and ecological studies.
This tutorial provides a step-by-step guide to deploying the YOLOv8 model in practice. We
demonstrate a tangible real-world application by optimizing it for edge computing, improving
efficiency and processing speed (using the NCNN format). This not only serves as a functional
solution but also as an instructional tool for similar projects.
276
The technical insights and methodologies shared in this tutorial are the basis for the complete
work to be developed at our university in the future. We envision further development, such as
integrating additional environmental sensing capabilities and refining the model’s accuracy and
processing efficiency. Implementing alternative energy solutions, such as the proposed solar
power setup, will enhance the project’s sustainability and applicability in remote or underserved
locations.
Resources
The Dataset paper, Notebooks, and PDF version are in the Project repository.
277
Image Classification with EXECUTORCH
Introduction
Image classification is a fundamental computer vision task that powers countless real-world
applications—from quality control in manufacturing to wildlife monitoring, medical diagnostics,
and smart home devices. In the edge AI landscape, the ability to run these models efficiently on
278
resource-constrained devices has become increasingly critical for privacy-preserving, low-latency
applications.
In the chapter Image Classification Fundamentals, we explored image classification with
TensorFlow Lite and demonstrated how to deploy efficient neural networks on the Raspberry
Pi. That tutorial covered the complete workflow from model conversion to real-time camera
inference, achieving excellent results with the MobileNet V2 architecture and a real dataset
(CIFAR-10).
This chapter takes a parallel approach using PyTorch EXECUTORCH—Meta’s modern
solution for edge deployment. Rather than replacing our TFLite knowledge, this chapter
expands your edge AI toolkit, giving us the flexibility to choose the right framework for our
specific needs.
What is EXECUTORCH?
EXECUTORCH is PyTorch’s official solution for deploying machine learning models on edge
devices, from smartphones and embedded systems to microcontrollers and IoT devices. Released
in 2023, it represents Meta’s commitment to bringing the entire PyTorch ecosystem to edge
computing.
Core Capabilities:
• Native PyTorch Integration: Seamless workflow from model training to edge deploy-
ment without switching frameworks
• Efficient Execution: Optimized runtime designed specifically for resource-constrained
devices
• Broad Portability: Runs on diverse hardware platforms (ARM, x86, specialized
accelerators)
• Flexible Backend System: Extensible delegate architecture for hardware-specific
optimizations
• Quantization Support: Built-in integration with PyTorch’s quantization tools for
model compression
279
2. Modern Architecture Built from the ground up for edge computing with contemporary
best practices, EXECUTORCH incorporates lessons learned from previous mobile deployment
frameworks.
3. Comprehensive Quantization Native support for various quantization techniques
(dynamic, static, quantization-aware training) enables significant model size reduction with
minimal accuracy loss.
4. Extensible Backend System The delegate system allows seamless integration with
hardware accelerators (XNNPACK for CPU optimization, QNN for Qualcomm chips, CoreML
for Apple devices, and more).
5. Active Development Backed by Meta with rapid iteration and strong community support,
ensuring the framework evolves with edge AI needs.
6. Growing Model Zoo Access to pretrained models specifically optimized for edge deploy-
ment, with consistent performance across devices.
Understanding when to choose each framework is crucial for effective edge deployment:
280
2. Team expertise and existing infrastructure
3. Specific hardware requirements
4. Project timeline and maturity needs
Install Python tools, camera libraries, and build dependencies for PyTorch:
rpicam-hello --list-cameras
281
We should see that the OV5647 cam is installed.
import numpy as np
from picamera2 import Picamera2
import time
# Initialize camera
picam2 = Picamera2()
config = picam2.create_preview_configuration(main={"size":(640,480)})
[Link](config)
[Link]()
# Capture image
picam2.capture_file("camera_capture.jpg")
print("Image captured: cam_test.jpg")
# Stop camera
282
[Link]()
[Link]()
python --version
If the Raspberry Pi OS is the legacy, the Python version should be 3.11, and it is
not necessary to install Pyenv.
Install pyenv
283
Configure Shell
# pyenv configuration
export PYENV_ROOT="$HOME/.pyenv"
[[ -d $PYENV_ROOT/bin ]] && export PATH="$PYENV_ROOT/bin:$PATH"
eval "$(pyenv init -)"
EOF
source ~/.bashrc
pyenv --version
284
cd Documents
mkdir EXECUTORCH
cd EXECUTORCH
# Verify
python --version # Should show Python 3.11.14
Verify installation:
285
PyTorch and EXECUTORCH Installation
For the Raspberry Pi Zero 2 W (32-bit ARM), we may need to build from
source or use lighter alternatives, which are not covered here.
# Install dependencies
./install_requirements.sh
286
Verifying the Setup
Let’s verify our setup with a test script. Create setup_test.py (for example, using nano):
import torch
import numpy as np
from PIL import Image
import executorch
print("=" * 50)
print("SETUP VERIFICATION")
print("=" * 50)
# Check versions
print(f"PyTorch version: {torch.__version__}")
print(f"NumPy version: {np.__version__}")
print(f"PIL version: {Image.__version__}")
print(f"EXECUTORCH available: {executorch is not None}")
# Test PIL
test_img = [Link]('RGB', (224, 224), color='red')
print(f"Created test PIL image: {test_img.size}")
Run it:
python setup_test.py
==================================================
SETUP VERIFICATION
==================================================
PyTorch version: 2.9.1+cpu
NumPy version: 2.2.6
PIL version: 12.1.0
287
EXECUTORCH available: True
Working directory:
cd Documents
cd EXECUTORCH
mkdir IMG_CLASS
cd IMG_CLASS
mkdir MOBILENET
cd MOBILENET
mkdir models images notebooks
wget "[Link] \
-O ./images/[Link]
Now, let’s create a test program where we should take into consideration:
288
and save it as img_class_test_torch.py:
import torch
import [Link] as transforms
from torchvision import models
from PIL import Image
import time
import json
import [Link]
import os
# Paths
MODEL_PATH = "models/mobilenet_v2.pth"
LABELS_PATH = "models/imagenet_labels.json"
IMAGE_PATH = "images/[Link]"
289
model.load_state_dict([Link](MODEL_PATH, map_location='cpu'))
[Link]()
with torch.no_grad():
output = model(batch)
# Get predictions
probabilities = [Link](output[0], dim=0)
top5_prob, top5_idx = [Link](probabilities, 5)
# Display results
print("\n" + "="*50)
print("CLASSIFICATION RESULTS")
print("="*50)
print(f"Inference Time: {inference_time:.2f} ms\n")
print("Top 5 Predictions:")
print("-"*50)
for i in range(5):
idx = top5_idx[i].item()
prob = top5_prob[i].item()
290
print(f"{i+1}. {labels[idx]:20s} - {prob*100:.2f}%")
print("="*50)
The result:
==================================================
CLASSIFICATION RESULTS
==================================================
Inference Time: 86.12 ms
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================
The inference was OK, taking 86ms (first time). We can also verify the size of the saved Torch
model
ls -lh ./models/mobilenet_v2.pth
Unlike TensorFlow Lite, where we downloaded pre-converted .tflite models, with EXECU-
TORCH, we typically export PyTorch models to the .pte (PyTorch EXECUTORCH) format
ourselves. This gives us full control over the export process.
291
Understanding the Export Process
Let’s export a MobileNet V2 model to EXECUTORCH basic format. Creating a Python script
as convert_mobv2_executorch.py
import torch
from torchvision import models
from [Link] import to_edge
from [Link] import export
# Paths
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
292
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2.pte"
print("\n" + "="*50)
print("MODEL SIZE COMPARISON")
print("="*50)
print(f"PyTorch model: {pytorch_size:.2f} MB")
print(f"ExecuTorch model: {executorch_size:.2f} MB")
293
print(f"Reduction: {((pytorch_size - executorch_size) \
/pytorch_size * 100):.1f}%")
print("="*50)
python export_mobv2_executorch.py
We will get:
==================================================
MODEL SIZE COMPARISON
==================================================
PyTorch model: 13.60 MB
ExecuTorch model: 13.58 MB
Reduction: 0.2%
==================================================
The basic ExecuTorch conversion doesn’t compress the model much - it’s mainly for runtime
efficiency. To get real size reduction, we need quantization, which we will explore later.
But first, let’s do an inference test using the converted model.
Runing the script mobv2_executorch.py:
import torch
import [Link] as transforms
from PIL import Image
import time
import json
from [Link].portable_lib import _load_for_executorch
# Paths
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2.pte"
294
LABELS_PATH = "models/imagenet_labels.json"
IMAGE_PATH = "images/[Link]"
# Load labels
print("Loading labels...")
with open(LABELS_PATH, 'r') as f:
labels = [Link](f)
# Get predictions
output_tensor = output[0] # ExecuTorch returns a list
probabilities = [Link](output_tensor[0], dim=0)
top5_prob, top5_idx = [Link](probabilities, 5)
# Display results
295
print("\n" + "="*50)
print("EXECUTORCH CLASSIFICATION RESULTS")
print("="*50)
print(f"Inference Time: {inference_time:.2f} ms\n")
print("Top 5 Predictions:")
print("-"*50)
for i in range(5):
idx = top5_idx[i].item()
prob = top5_prob[i].item()
print(f"{i+1}. {labels[idx]:20s} - {prob*100:.2f}%")
print("="*50)
As a result, we got a similar inference result, but a much higher latency (almost 2.5 seconds),
which was unexpected.
Loading labels...
Loading ExecuTorch model from models/mobilenet_v2.pte...
Loading image from images/[Link]...
Running ExecuTorch inference...
==================================================
EXECUTORCH CLASSIFICATION RESULTS
==================================================
Inference Time: 2445.78 ms
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================
That export path produces a generic ExecuTorch CPU graph with reference kernels and no
backend optimizations or fusions, so significantly higher latency than PyTorch is expected for
MobileNet_v2 on a Pi 5.
ExecuTorch is designed to shine when delegated to a backend (XNNPACK, OpenVINO, etc.),
where large subgraphs are lowered into highly optimized kernels. Without a delegate, most of
296
the graph runs on the generic portable path, which is known to be significantly slower than
PyTorch for many models.
So, let’s export the .pth model again with a CPU‑optimized backend (e.g., XNNPACK) and
run with that backend enabled; this alone should reduce latency when compared with the naïve
interpreter path.
Here’s the corrected conversion script with XNNPACK delegation (convert_mobv2_xn-
[Link]):
import torch
from torchvision import models
from [Link] import to_edge
from [Link] import export
from [Link].xnnpack_partitioner \
import XnnpackPartitioner
# Paths
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
EXECUTORCH_MODEL_PATH = "models/mobilenet_v2_xnnpack.pte"
297
print(" 4. Lowering to ExecuTorch...")
executorch_program = edge_program.to_executorch()
print("\n" + "="*50)
print("MODEL SIZE COMPARISON")
print("="*50)
print(f"PyTorch model: {pytorch_size:.2f} MB")
print(f"ExecuTorch+XNNPACK: {executorch_size:.2f} MB")
print("="*50)
Runing it we get:
==================================================
MODEL SIZE COMPARISON
==================================================
PyTorch model: 13.60 MB
ExecuTorch+XNNPACK: 13.35 MB
==================================================
We did not gain in terms of size, but let’s run the same inference script as before, with this
new converted model, to inspect the latency:
298
the result:
Loading labels...
Loading ExecuTorch model from models/mobilenet_v2_xnnpack.pte...
Loading image from images/[Link]...
Running ExecuTorch inference...
==================================================
EXECUTORCH CLASSIFICATION RESULTS
==================================================
Inference Time: 19.95 ms
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 47.44%
2. Egyptian Mau - 37.61%
3. lynx - 6.91%
4. tabby cat - 6.22%
5. plastic bag - 0.47%
==================================================
Now, the ExecuTorch runtime detects the backend automatically from the .pte file metadata.
We have achieved much faster inference: 20ms instead of 2445ms. This latency is, in fact,
several times faster than PyTorch.
Why XNNPACK is so fast:
This demonstrates:
Now we can add quantization to get an even smaller model size while maintaining (or even
increasing) this speed!
299
Model Quantization
Quantization reduces model size and can further improve inference speed. EXECUTORCH
supports PyTorch’s native quantization.
Quantization Overview
Quantization is a technique that reduces the precision of numbers used in a model’s computations
and stored weights—typically from 32-bit floats to 8-bit integers. This reduces the model’s
memory footprint, speeds up inference, and lowers power consumption, often with minimal loss
in accuracy.
Quantization is especially important for deploying models on edge devices such as wearables,
embedded systems, and microcontrollers, which often have limited compute, memory, and
battery capacity. By quantizing models, we can make them significantly more efficient and
better suited to these resource-constrained environments.
Quantization in ExecuTorch
ExecuTorch uses torchao as its quantization library. This integration allows ExecuTorch to
leverage PyTorch-native tools for preparing, calibrating, and converting quantized models.
Quantization in ExecuTorch is backend-specific. Each backend defines how models should be
quantized based on its hardware capabilities. Most ExecuTorch backends use the torchao PT2E
quantization flow, which works with models exported with [Link] and enables tailored
quantization for each backend.
For a quantized XNNPACK .pte we need a different pipeline: PT2E quantization (with
XNNPACKQuantizer), then lowering with XnnpackPartitioner before to_executorch(). Oth-
erwise, we will hit errors or get an undelegated model.
For the conversion, we need: (1) calibrate with real, preprocessed images, and (2) compute the
quantized .pte size after you actually write the file.
First, let us create a small calib_images/ folder (e.g., 50–100 natural images across a few
classes). A simple way is to reuse an existing dataset (e.g., CIFAR‑10) and save 50–100 images
into calib_images/ with an ImageNet‑style folder layout.
The script gen_calibr_images.py will: • Download CIFAR‑10. • Pick 10 classes × 10 images
each = 100 images. • Save them under calib_images/<class_name>/img_XXX.jpg.
import os
from pathlib import Path
import torch
from torchvision import datasets, transforms
from [Link] import save_image
300
# Where to store calibration images
OUT_ROOT = Path("calib_images")
OUT_ROOT.mkdir(parents=True, exist_ok=True)
idx = counts[cls_name]
out_path = class_dir / f"img_{idx:04d}.jpg"
save_image(img, out_path)
counts[cls_name] += 1
301
break
Let’s use the inference script convert_mobv2_xnnpack_int8.py, which is the same inference
script as before, with this new int8 converted model to inspect the latency:
import os
import torch
import [Link] as models
import [Link] as transforms
import [Link] as datasets
PYTORCH_MODEL_PATH = "models/mobilenet_v2.pth"
EXECUTORCH_QUANTIZED_PATH = "models/mobilenet_v2_quantized_xnnpack.pte"
CALIB_IMAGES_DIR = "calib_images" # <-- put some natural images here
302
# 2) Configure XNNPACK quantizer (global symmetric config)
qparams = get_symmetric_quantization_config(is_per_channel=True)
quantizer = XNNPACKQuantizer()
quantizer.set_global(qparams)
calib_dataset = [Link](CALIB_IMAGES_DIR,
transform=calib_transform)
calib_loader = [Link](
calib_dataset, batch_size=1, shuffle=True
)
et_program = to_edge_transform_and_lower(
exported_quant,
303
partitioner=[XnnpackPartitioner()],
).to_executorch()
pytorch_size = [Link](PYTORCH_MODEL_PATH)/(1024*1024)
quantized_size = [Link](EXECUTORCH_QUANTIZED_PATH)/(1024*1024)
print("\n" + "="*60)
print("MODEL SIZE COMPARISON")
print("="*60)
print(f"PyTorch (FP32): {pytorch_size:6.2f} MB")
print(f"ExecuTorch Quantized (INT8): {quantized_size:6.2f} MB")
print(f"Size reduction: {((pytorch_size - quantized_size) \
/ pytorch_size * 100):5.1f}%")
print(f"Savings: {pytorch_size - quantized_size:6.2f} MB")
print("="*60)
============================================================
MODEL SIZE COMPARISON
============================================================
PyTorch (FP32): 13.60 MB
ExecuTorch Quantized (INT8): 3.59 MB
Size reduction: 73.6%
Savings: 10.01 MB
============================================================
The quantized (int8) model achieved 74% size reduction: ~3.5 MB (similar to TFLite). Let’s
see about the inference latency, runing mobv2_xnnpack_int8.py.
Loading labels...
Loading ExecuTorch model from models/mobilenet_v2_quantized_xnnpack.pte...
Loading image from images/[Link]...
Running ExecuTorch inference (Quantized INT8)...
304
==================================================
EXECUTORCH QUANTIZED INT8 RESULTS
==================================================
Inference Time: 13.56 ms
Output dtype: torch.float32
Top 5 Predictions:
--------------------------------------------------
1. tiger cat - 51.01%
2. Egyptian Mau - 34.11%
3. lynx - 7.54%
4. tabby cat - 6.17%
5. plastic bag - 0.37%
==================================================
Slightly higher top‑1 probabilities in the INT8 model are normal and do not indicate
a problem by themselves. Quantization slightly changes the logits, and softmax can
become a bit “sharper” or “flatter” even when top‑1 remains correct.
NOTE
• Looking at Htop, we can see that only one of the Pi’s cores is at 100%. This indicates that
the shipped Python runtime currently runs our ExecuTorch/XNNPACK model effectively
single‑threaded on Pi.
• To exploit all four cores, the next step would be to move inference into a small C++
wrapper that sets the ExecuTorch threadpool size before executing the graph. With the
pure‑Python path, there is no clean public knob to change it yet. We will not explore it
here.
Now that we have our EXECUTORCH models, let’s explore them in more detail for image
classification using a Jupyter Notebook!
305
Setting up Jupyter Notebook
jupyter notebook
Access it from another device using the provided token in your web browser.
EXECUTORCH/MOBILENET/
��� convert_mobv2_executorch.py
��� convert_mobv2_xnnpack.py
��� convert_mobv2_xnnpack_int8.py
��� mobv2_executorch.py
��� mobv2_xnnpack.py
��� mobv2_xnnpack_int8.py
��� calib_images/
��� data/
��� models/
� ��� mobilenet_v2.pth # Float32 pytorch model
� ��� mobilenet_v2.pte # Float32 conv model
� ��� mobilenet_v2_xnnpack.pte # Float32 conv model
� ��� mobilenet_v2_quantized_xnnpack.pte # Quantized conv model
� ��� imagenet_labels.json # Labels
��� images/ # Test images
306
� ��� [Link]
� ��� camera_capture.jpg
��� notebooks/
��� image_classification_executorch.ipynb
Inside the folder ‘notebooks’, on the project space IMAGE_CLASS/MOBILENET, create a new
notebook: image_classification_executorch.ipynb.
print("=" * 50)
print("SETUP VERIFICATION")
print("=" * 50)
# Check versions
print(f"PyTorch version: {torch.__version__}")
print(f"NumPy version: {np.__version__}")
print(f"PIL version: {Image.__version__}")
print(f"EXECUTORCH available: {executorch is not None}")
# Test PIL
307
test_img = [Link]('RGB', (224, 224), color='red')
print(f"Created test PIL image: {test_img.size}")
We get:
==================================================
SETUP VERIFICATION
==================================================
PyTorch version: 2.9.1+cpu
NumPy version: 2.2.6
PIL version: 12.1.0
EXECUTORCH available: True
img_path = "../images/[Link]"
308
Image size: (1600, 1598)
Image mode: RGB
• python export_mobv2_executorch.py
imagenet_labels.json mobilenet_v2_quantized_xnnpack.pte
mobilenet_v2.pte mobilenet_v2_xnnpack.pte
mobilenet_v2.pth
The conversions were performed using the Python scripts in the previous sections.
309
# Load the EXECUTORCH model
model_path = "../models/mobilenet_v2.pte"
try:
model = _load_for_executorch(model_path)
print(f"Model loaded successfully from: {model_path}")
#print(f" Available methods: {model.method_names}")
except FileNotFoundError:
print(f"� Model not found: {model_path}")
print("\nPlease run the export script first:")
print(" python export_mobilenet.py")
# Download and save ImageNet labels (if you do not have it)
LABELS_PATH = "../models/imagenet_labels.json"
if not [Link](LABELS_PATH):
print("Downloading ImageNet labels...")
LABELS_URL = "[Link]
imagenet-simple-labels/master/[Link]"
with [Link](LABELS_URL) as url:
labels = [Link](url)
310
Check the labels:
Image Preprocessing
A preprocessing pipeline is needed because ExecuTorch only runs the exported core network; it
does not include the input normalization logic that MobileNet v2 expects, and the model will
give incorrect predictions if the input tensor is not in the exact format it was trained on.
What MobileNet v2 expects For typical PyTorch MobileNet v2 models (ImageNet‑pretrained):
• Input shape: 3‑channel RGB tensor of size. • Value range: floating-point values, usually in
float32 after dividing by 255. • Normalization: per‑channel mean/std (ImageNet) normalization,
e.g., mean=0.485, 0.456, 0.406, std=0.229, 0.224, 0.225.
These steps (resize, convert to tensor, normalize) are not “optional decorations”; they are part
of the functional definition of the model’s expected input distribution.
preprocess = [Link]([
[Link](256), # Resize to 256
[Link](224), # Center crop to 224x224
[Link](), # Convert to tensor [0, 1]
[Link]( # Normalize with ImageNet stats
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
Apply preprocessing
311
input_tensor = preprocess(img)
print(f" Input shape: {input_tensor.shape}")
print(f" Input dtype: {input_tensor.dtype}")
input_batch = input_tensor.unsqueeze(0)
Run Inference
For inference, we should run a forward pass of the model in inference mode (torch.no_grad()),
measure the time, and print basic information about the outputs.
torch.no_grad() is a context manager that disables gradient calculation inside its block.
During inference, we do not need gradients, so disabling them:
312
# Run inference
with torch.no_grad():
start_time = [Link]()
outputs = [Link]((input_batch,))
inference_time = [Link]() - start_time
type(outputs) tells us what container the model returned. Often this is a tuple or list when
working with exported/ExecuTorch‑style models, e.g., <class 'tuple'>.
That container may hold one or more tensors (e.g., logits, auxiliary outputs).
• outputs[0] accesses the first element of that container (usually the main output tensor),
and .shape prints its dimensions (For image classification, this is often batch_size,
num_classes).
Now we should take the model’s raw scores (logits) for a single image, convert them into proba-
bilities with softmax, select the top‑5 most likely classes, and print them nicely formatted.
• outputs[0][0] selects the first element in the batch, giving a 1D tensor of logits of
length num_classes.
• [Link](..., dim=0) applies the softmax function along that
1D dimension, turning logits into probabilities that sum to 1.
# Display results
print("\n" + "="*60)
313
print("TOP 5 PREDICTIONS")
print("="*60)
print(f"{'Class':<35} {'Probability':>10}")
print("-"*60)
for i in range(5):
label = labels[top5_indices[i]]
prob = top5_prob[i].item() * 100
print(f"{label:<35} {prob:>9.2f}%")
print("="*60)
============================================================
TOP 5 PREDICTIONS
============================================================
Class Probability
------------------------------------------------------------
tiger cat 12.85%
Egyptian cat 9.75%
tabby 6.09%
lynx 1.70%
carton 0.84%
============================================================
For simplicity and reuse across other tests, let’s create a reusable function that builds on what
was done so far.
Args:
img_path: Path to input image
model_path: Path to .pte model file
labels_path: Path to labels text file
top_k: Number of top predictions to return
show_image: Whether to display the image
314
Returns:
inference_time: Inference time in ms
top_indices: Indices of top k predictions
top_probs: Probabilities of top k predictions
"""
# Load image
img = [Link](img_path).convert('RGB')
# Display image
if show_image:
[Link](figsize=(4, 4))
[Link](img)
[Link]('off')
[Link]('Input Image')
[Link]()
# Load model
print(f"Model Path {model_path}")
model_size = [Link](model_path) / (1024 * 1024)
print(f"Model size: {model_size:6.2f} MB")
model = _load_for_executorch(model_path)
# Preprocess
preprocess = [Link]([
[Link](256),
[Link](224),
[Link](),
[Link](
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
input_tensor = preprocess(img)
input_batch = input_tensor.unsqueeze(0)
# Inference
with torch.no_grad():
start_time = [Link]()
315
outputs = [Link]((input_batch,))
inference_time = ([Link]() - start_time)*1000
# Process results
probabilities = [Link](outputs[0][0], dim=0)
top_prob, top_indices = [Link](probabilities, top_k)
# Load labels
with open(labels_path, 'r') as f:
labels = [Link](f)
# Display results
print(f"\nInference time: {inference_time:.2f} ms")
print("\n" + "="*60)
print(f"{'[PREDICTION]':<35} {'[Probability]':>15}")
print("-"*60)
for i in range(top_k):
label = labels[top_indices[i]]
prob = top_prob[i].item() * 100
print(f"{label:<35} {prob:>14.2f}%")
print("="*60)
316
We can also check what is retrurned fron the function
(2445.200204849243,
tensor([282, 285, 287, 281, 728]),
tensor([0.4744, 0.3761, 0.0691, 0.0622, 0.0047]))
317
# Test with the cat image
inf_time, indices, probs = classify_image_executorch(
img_path="../images/[Link]",
model_path="../models/mobilenet_v2_xnnpack.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)
318
# Test with the cat image
inf_time, indices, probs = classify_image_executorch(
img_path="../images/[Link]",
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)
Slightly higher probabilities in the INT8 model are normal and do not indicate a
problem by themselves. Quantization slightly changes the logits, and softmax can
become a bit “sharper” or “flatter” even when top‑1 remains correct.
319
Camera Integration
We essentially have two different Python worlds: system Python 3.13 (where the camera stack
is wired up) and our 3.11 virtual env (where ExecuTorch is installed). To run ExecuTorch on
live frames from the Pi camera, we need to bridge those worlds.
We will use a two‑process solution: capture in 3.13, infer in 3.11. For that, we should run
a small capture service under Python 3.13 that:
The 3.11 process (under venev) receives the frame, decodes it, runs the preprocessing pipeline
(resize, normalize), then calls ExecuTorch for inference..
Image Capture
Outside of the ExecuTorch env and folder, we will create a folder (CAMERA).
Documents/
��� EXECUTORCH/MOBILENET/ # Python 3.11
��� CAMERA/ # Python 3.13
��� camera_capture.py
��� camera_capture.jpg
320
import numpy as np
from picamera2 import Picamera2
import time
# Initialize camera
picam2 = Picamera2()
config = picam2.create_preview_configuration(main={"size":(640,480)})
[Link](config)
[Link]()
[Link](2)
# Capture image
picam2.capture_file("camera_capture.jpg")
print("Image captured: camera_capture.jpg")
# Stop camera
[Link]()
[Link]()
• /Documents/CAMERA/camera_capture.jpg
Looking from the notebook folder, the image path will be:
../../../../CAMERA/camera_capture.jpg
Let’s run the same function used with the test image:
321
# Test the quantized model with the captured image
inf_time, indices, probs = classify_image_executorch(
img_path="../../../../CAMERA/camera_capture.jpg",
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
labels_path="../models/imagenet_labels.json",
top_k=5
)
322
Performance Benchmarking
Let’s now define a function to run inference several times for each model and compare their
performance.
323
def benchmark_inference(model_path, num_runs=50):
"""
Benchmark model inference speed
"""
print(f"Benchmarking model: {model_path}")
print(f"Number of runs: {num_runs}\n")
# Load model
model = _load_for_executorch(model_path)
# Benchmark
print(f"Running benchmark...")
times = []
for i in range(num_runs):
start = [Link]()
with torch.no_grad():
_ = [Link]((dummy_input,))
[Link]([Link]() - start)
# Print statistics
print("\n" + "="*50)
print("BENCHMARK RESULTS")
print("="*50)
print(f" Mean: {[Link]():.2f} ms")
print(f" Median: {[Link](times):.2f} ms")
print(f" Std: {[Link]():.2f} ms")
print(f" Min: {[Link]():.2f} ms")
print(f" Max: {[Link]():.2f} ms")
print("="*50)
# Plot distribution
324
[Link](figsize=(12, 4))
# Histogram
[Link](1, 2, 1)
[Link](times, bins=20, edgecolor='black', alpha=0.7)
[Link]([Link](), color='red', linestyle='--',
label=f'Mean: {[Link]():.2f} ms')
[Link]('Inference Time (ms)')
[Link]('Frequency')
[Link]('Inference Time Distribution')
[Link]()
[Link](alpha=0.3)
# Time series
[Link](1, 2, 2)
[Link](times, marker='o', markersize=3, alpha=0.6)
[Link]([Link](), color='red', linestyle='--',
label=f'Mean: {[Link]():.2f} ms')
[Link]('Run Number')
[Link]('Inference Time (ms)')
[Link]('Inference Time Over Runs')
[Link]()
[Link](alpha=0.3)
plt.tight_layout()
[Link]()
return times
mobilenet_v2.pte
mobilenet_v2_xnnpack.pte
mobilenet_v2_quantized_xnnpack.pte
# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2.pte",
325
num_runs=50
)
# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2_xnnpack.pte",
num_runs=50
)
326
Quantization (INT8): mobilenet_v2_quantized_xnnpack.pte
# Run benchmark
benchmark_times = benchmark_inference(
model_path="../models/mobilenet_v2_quantized_xnnpack.pte",
num_runs=50
)
327
Performance Comparison Table
Mean Median
Model Configuration (ms) (ms) Std Dev (ms) File Size (MB) Latency
Float32 (basic) 2440 2440 2.17 13.58 +600×
Float32 + 11.24 10.84 1.67 13.35 ~3×
XNNPACK
INT8 + XNNPACK 3.91 3.69 0.55 3.59 1×
Key Observations:
328
2. Quantization Benefit: INT8 quantization, besides size reduction, adds additional
speedup beyond XNNPACK
3. Variability: Quantized model shows lower standard deviation, indicating more stable
performance
4. Size-Speed Tradeoff: 75% size reduction (14MB → 3.5MB) with 3× speed improvement
CIFAR-10 Dataset:
• 10 classes: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck
Figure 7: cifar10
329
• The images in CIFAR-10 are of size 3x32x32 (3-channel color images of 32x32 pixels in
size).
Let’s create a Project folder structure as below (some files are shown as they will appear
later)
EXECUTORCH/CIFAR-10/
��� export_cifar10_xnnpack.py
��� inference_cifar10_xnnpack.py
��� models/
� ��� cifar10_model_jit.pt # Float32 pytorch model
� ��� cifar10_xnnpack.pte # Float32 conv model
��� images/ # Test images
� ��� [Link]
��� notebooks/
��� CIFAR-10_Inference_RPI.ipynb
Let’s train a model from scratch on CIFAR-10. For that, we can run the Notebook below on
Google Colab:
cifar10_colab_training.ipynb
From the training, we will have the trained model:
cifar10_model_jit.pt, which should be saved on /models folder
Next, as we did before, we should export the PyTorch model to ExecuTorch, and let’s use
XNNPACK. Run the script: export_cifar10_xnnpack.py, as a result, we have:
330
Runing it, a converted model cifar10_xnnpack.pte will be saved in ./models/ folder.
Runing the script inference_cifar10_xnnpack.py, over the “cat” image, we can see that the
converted model is working fine:
python inference_cifar10_xnnpack.py ./images/[Link]
331
And runing 20 times….
332
Despite the exported model being OK, when we make an inference with the original PyTorch
model, in this case (a small model), we will find even lower latencies.
333
In short, our export script is conceptually the right pattern for ExecuTorch+XNNPACK on
Arm, but for this specific small CIFAR‑10 CNN, the overhead of ExecuTorch and partial
XNNPACK delegation on a Pi‑class device can easily make it slower than a well‑optimized
plain PyTorch JIT model.
Optionally, it is possible to explore those models with the notebook:
CIFAR-10_Inference_RPI_Updated.ipynb
Conclusion
This chapter adapted our image classification workflow from TensorFlow Lite to PyTorch
EXECUTORCH, demonstrating that the PyTorch ecosystem provides a powerful and modern
alternative for edge AI deployment on Raspberry Pi devices.
EXECUTORCH represents a significant evolution in edge AI deployment, bringing PyTorch’s
research-friendly ecosystem to production edge devices. While TensorFlow Lite remains
excellent and mature, having EXECUTORCH in your toolkit makes you a more versatile edge
AI practitioner.
The future of edge AI is multi-framework, multi-platform, and rapidly evolving. By mastering
both EXECUTORCH and TensorFlow Lite, you’re positioned to make informed technical
decisions and adapt as the landscape changes.
Remember: The best framework is the one that serves your specific needs. This
tutorial empowers you to make that choice confidently.
334
Key Takeaways
Technical Achievements:
EXECUTORCH Advantages:
Comparison with TFLite: Both frameworks achieve similar goals with different philoso-
phies:
The choice between them often comes down to your training framework and specific require-
ments.
Performance Considerations
On Raspberry Pi 4/5, you can expect: - Float32 models: 10-20ms per inference (MobileNet
V2)
335
Resources
Code Repository
Official Documentation
Quantization:
• PyTorch Quantization
• Quantization API Tutorial
Models:
• Torchvision Models
• Pretrained Model Deployment Guide
Hardware Resources
Books
336
Beyond CPU - Hardware Acceleration for Edge
AI
337
Introduction
Throughout this course, we’ve explored various approaches to deploying AI models at the edge.
We started with TensorFlow Lite running on the Raspberry Pi’s CPU, then moved to YOLO
and ExecuTorch with optimized backends like XNNPACK. While these software optimizations
significantly improve performance, they still rely on the general-purpose CPU to execute neural
network operations.
In this chapter, we’ll take the next step: dedicated hardware acceleration. We’ll use the
MemryX MX3 M.2 AI Accelerator Module—a specialized processor designed specifically for
neural network inference. The MX3 module contains four AI accelerator chips that can run
deep learning models with dramatically lower latency and power consumption compared to
CPU execution.
338
For learning more about AI Acceleration, please refer to MLSys book and how the
MemryX module works, read the Architecture Overview.
Our goal
By the end of this lab, we will have installed and configured the MX3 hardware on a Raspberry
Pi 5, set up the MemryX SDK and development environment, and gained a clear understanding
of the MX3 compilation and deployment workflow. We will also compile neural network models
for execution on the MX3 accelerator, compare their performance against CPU-based inference
while analyzing the trade-offs, and finally build a complete end-to-end inference pipeline using
the MemryX Python API.
Prerequisites
IMPORTANT NOTE: MemryX recommends the GeeekPi N04 M.2 2280 HAT
as an excellent choice for the Raspberry Pi 5. It delivers solid power and fits the
2280 MX3 M.2 form factor. Some hats can lead to instabilities, mainly due to PCIe
speed (Gen3). The Raspberry Pi 5 can have stability issues on Gen3.
339
The Raspberry Pi M.2 HAT+ is a good option. It works very well, despite the fact that we
should adapt the MX3 board to it (The MX3 is longer than the hat).
MemryX MX3 M.2 module with the heatsink installed
For heatsink installation, follow the video instructions: [Link]
It is essential to ensure we have sufficient cooling for the MemryX MX3 M.2 module, or we
may experience thermal throttling and reduced performance. The chips will throttle their
performance if they hit 100 °C.
During normal operation, the current MemryX MX3 temperature and throttle status can be
viewed at any time with:
cat /sys/memx0/temperature
340
Verification
After installing the hardware, turn on the Raspberry Pi and verify the system setup.
ls /dev/memx*
The lab temperature at the time of the above measurement was 25 °C.
Software Installation
cd Documents
mkdir MEMRYX
cd MEMRYX
python --version
341
If using the latest Raspberry Pi OS (based on Debian Trixie), it should be:
Python 3.13.5
Or, if the OS is the Legacy version:
Python 3.11.2
Important: As of January 2026, MemryX officially supports only Python 3.09 to 3.12.
Python 3.13.5 is too new and will likely cause compatibility issues. Since Debian Trixie ships
with Python 3.13 by default, we’ll need to install a compatible Python version alongside
it.
One solution is to install Pyenv, which allows us to easily manage multiple Python versions for
different projects without affecting the system Python.
If the Raspberry Pi OS is the legacy version, the Python version should be 3.11,
and it is not necessary to install Pyenv.
# Install dependencies
sudo apt install -y make build-essential libssl-dev zlib1g-dev \
libbz2-dev libreadline-dev libsqlite3-dev wget curl llvm \
libncursesw5-dev xz-utils tk-dev libxml2-dev libxmlsec1-dev \
libffi-dev liblzma-dev
# Install pyenv
curl [Link] | bash
# Add to ~/.bashrc
echo 'export PYENV_ROOT="$HOME/.pyenv"' >> ~/.bashrc
echo 'command -v pyenv >/dev/null || export PATH="$PYENV_ROOT/bin:$PATH"' \
>> ~/.bashrc
echo 'eval "$(pyenv init -)"' >> ~/.bashrc
# Reload shell
source ~/.bashrc
342
Set Python Version for Project
Once Pyenv and the selected Python version are installed, define it for the project direc-
tory:
• Drivers (memx-drivers): Kernel-level drivers for PCIe communication with the acceler-
ator hardware
• SDK (memx-accl): Python libraries, neural compiler, runtime, and benchmarking tools
This command downloads the repository’s GPG key for package verification and adds the
MemryX package repository:
343
Configure Platform Settings
Run the ARM setup utility to configure platform-specific settings. This opens a menu to
select the platform and apply the necessary configurations (e.g., enabling PCIe Gen 3.0 on the
Raspberry Pi 5):
sudo mx_arm_setup
Select the appropriate option for your hardware, and press <OK> in the next page:
sudo reboot
After rebooting, verify that the MemryX driver is installed by checking its version:
344
apt policy memx-drivers
Install Utilities
It’s best practice to use a virtual environment to avoid conflicts with system packages.
Create and activate a virtual environment:
345
Verify the neural compiler is installed:
mx_nc --version
Verification
Verify the complete installation by running the built-in “hello world” benchmark:
mx_bench --hello
With the benchmark results, our MemryX MX3 is properly installed and ready to use.
346
Our First Accelerated Model
Working with the MemryX MX3 follows a straightforward four-step workflow that differs from
traditional CPU-based inference:
347
Step 1: Select or Train a Model
Start with a pre-trained model or train your own. MemryX supports models from major
frameworks:
The model remains in its original format—no framework-specific conversions needed yet. For
this lab, we’re using MobileNetV2 from Keras Applications, but we could equally use a custom
model we have trained for a specific task, as we have seen before.
Supported Operations: The MX3 supports most common deep learning operators (convolu-
tions, pooling, activations, etc.). Check the supported operators if using custom architectures.
Unsupported operations will fall back to CPU, though this is rare for standard vision models.
The MemryX Neural Compiler (mx_nc) transforms the model into a DFP (Dataflow Package):
The MemryX Neural Compiler, mx_nc, is a command‑line tool that takes one or more
neural‑network models (Keras, TensorFlow, TFLite, ONNX, etc.) and compiles them into a
MemryX Dataflow Program (DFP) that can run on MemryX accelerators (MXA). Internally,
it does framework import, graph optimization (fusion/splitting, operator expansion, activation
approximation), resource mapping on the MXA cores, and finally emits the DFP used by the
runtime or simulator.
What mx_nc does:
• Compiles models into a single DFP file per compilation, then loads it onto one or more
MXA chips to run inference.
• Supports multi‑model, multi‑stream, and multi‑chip mapping, automatically distributing
models and layers across available MX3 devices for higher throughput.
• Handles mixed‑precision weights (per‑channel 4/8/16‑bit) while keeping activations in
floating point on the accelerator. By default, MemryX quantizes weights to INT8 precision
and activations to BFloat16.
• Can crop pre/post‑processing parts of the graph so the MXA focuses on the core CNN/ML
operators while the host CPU runs the cropped sections.
348
What happens during compilation?
1. Model parsing: Loads the model and extracts the computational graph
2. Graph optimization: Fuses operations, eliminates redundancies
3. Operator mapping: Maps each layer to MX3 hardware instructions
4. Dataflow scheduling: Determines optimal execution order for pipelined processing
5. Memory allocation: Assigns on-chip memory for all intermediate activations
6. Multi-chip distribution: If using multiple chips, partitions the workload
• Model specification:
– -m / --models – input model file(s) (e.g. .h5, .pb, .onnx, TFLite).
349
– --extensions – load Neural Compiler Extensions (.nce files or builtin names) to
add or patch graph handling (e.g., complex transformer subgraphs or unsupported
ops) without a new SDK release.
For the complete option list (including less common flags), run mx_nc -h or consult
the Neural Compiler page in the MemryX Developer Hub, which documents all
arguments and includes usage examples for single‑model, multi‑model, cropping,
and mixed‑precision flows.
Once compiled, the DFP file is portable across all MX3 hardware.
Before integrating into our application, we can verify performance with the benchmarking
tool:
The benchmarker: - Generates synthetic input data matching the model’s input shape - Runs
warm-up inferences to stabilize performance - Measures throughput (FPS), latency, and chip
utilization - Reports first-inference latency (includes loading overhead)
Why benchmark separately? Real-world applications involve preprocessing (image loading
and resizing) and postprocessing (parsing outputs). Benchmarking isolates pure inference
performance, letting to identify bottlenecks in our full pipeline.
Finally, integrate the accelerator into our Python application using the MemryX API:
350
# Run inference
output = [Link](input_data)
# Process results
# ...
# Clean up
[Link]()
The API handles all hardware communication, memory transfers, and scheduling. Our code
just provides input tensors and receives output tensors—the complexity is abstracted away.
# Step 3: Benchmark
mx_bench -d mobilenet_v2.dfp -f 1000
This workflow is remarkably consistent across models and use cases. Once we’ve done it for
one model, adapting to others is straightforward.
Key Takeaway: The MX3 workflow separates compilation (done once) from
inference (done repeatedly). This “compile-once, run-many” approach means the
optimization overhead is amortized over thousands or millions of inferences in
production.
351
Download and Compile MobileNetV2
In Keras Applications, we can find deep learning models that are provided with pre-trained
weights. These models can be used for prediction, feature extraction, and fine-tuning.
Let’s download MobileNetV2, which was used in previous labs:
mx_nc -v -m mobilenet_v2.h5
352
What is a DFP file?
The .dfp (Dataflow Package) file is MemryX’s proprietary compiled format. Unlike standard
model formats (H5, ONNX, etc.) that describe the network architecture, a DFP file contains:
The neural compiler (mx_nc) performs this transformation automatically, with no manual
tuning required. The compilation process: 1. Parses the input model (H5, ONNX, TFLite,
etc.) 2. Maps operators to MX3-supported operations 3. Optimizes the dataflow graph 4.
Allocates memory on-chip 5. Generates the DFP binary
This is why compilation takes a few minutes, but inference is blazingly fast—all the optimization
work happens once, upfront.
Benchmarking Performance
Now that the model is compiled, it’s time to deploy it and run a benchmark to test its
performance on the MXA hardware. We will run 1000 frames of random data through the
accelerator to measure performance metrics:
353
Let’s understand what these metrics mean:
• FPS (Frames Per Second): How many images the accelerator can process per second
(~1,200 FPS for MobileNetV2)
• Latency: Time for a single inference (shown as “Avg” in the output)
– Subsequent inferences: True steady-state performance (~2ms)
• Throughput: Total data processed per second
The benchmark runs with random input data, which is why we see consistent performance.
Real-world performance with actual images should be similar once the preprocessing pipeline
is optimized, but we have found bigger latency.
In true dataflow architecture, latency and FPS are not coupled in the traditional
sense; latency does not equal 1/FPS. Note that even though latency is ~2ms in
the above benchmarking results, FPS is not measured to be 1000 ms / 2 ms = 500
FPS; rather, the FPS from the benchmarking results is ~1160.
In MemryX’s dataflow architecture, the “usual” rule (latency = 1/FPS) only applies to
frame‑to‑frame latency, not to end‑to‑end in‑to‑out latency for a single frame. That is why we
see ~2 ms latency per frame, yet still measure around 1160 FPS in MX3 benchmarks.
Two different latencies MemryX explicitly distinguishes two metrics.
• Latency 1 (frame‑to‑frame latency): Time between consecutive outputs once the pipeline
is full. Its reciprocal is FPS.
354
• Latency 2 (full in‑to‑out latency): Time from when the first input frame enters the system
(host + MX3 pipeline) until its output appears. This can be larger, but it does not set
the FPS.
In a streaming, pipelined accelerator like MX3, multiple frames are in‑flight simultaneously, so
the pipeline “fills” once and then produces results at a steady cadence.
Now let’s build a complete inference application that processes real images and compares CPU
vs. MX3 performance.
mkdir models
mkdir images
Load an image from the internet, for example, a cat (for comparison, it is the same as used on
previous chapters):
wget "[Link] \
-O ./images/[Link]
355
Understanding Input Requirements
All neural networks expect input data in a specific format, determined during training. For
MobileNetV2 trained on ImageNet:
The preprocessing must match exactly what was used during training, or accuracy will suffer.
For inference, we will need the ImageNet labels. The following function checks if the file exists,
and if not, downloads it:
356
import os, json
from pathlib import Path
import requests
MODELS_DIR = Path("./models")
IMAGENET_JSON = MODELS_DIR / "imagenet_class_index.json"
IMAGENET_JSON_URL = (
"[Link]
imagenet_class_index.json"
)
def load_idx2label():
with open(IMAGENET_JSON, "r") as f:
class_idx = [Link](f)
idx2label = [class_idx[str(k)][1] for k in range(len(class_idx))]
return idx2label
Image Preprocessing
The image used for inference should be preprocessed in the same way as during model training.
[Link].mobilenet_v2.preprocess_input() takes an image of shape (224, 224)
and converts it to (1, 224, 224, 3):
import numpy as np
from PIL import Image
import tensorflow as tf
from tensorflow import keras
357
def load_and_preprocess_image(image_path):
img = [Link](image_path).convert("RGB").resize((224, 224))
arr = [Link](img).astype(np.float32)
arr = [Link].mobilenet_v2.preprocess_input(arr)
arr = np.expand_dims(arr, 0) # Add batch dimension
return arr
The processed image will serve as the model’s input tensor (x):
ensure_imagenet_labels()
idx2label_full = load_idx2label() # length 1000 for ImageNet
IMAGE_PATH = Path("./images/[Link]")
x = load_and_preprocess_image(IMAGE_PATH)
Move the models (the original and compiled) to the models folder and set up the paths:
MODELS_DIR = Path("./models")
DFP_PATH = MODELS_DIR / "mobilenet_v2.dfp"
KERAS_PATH = MODELS_DIR / "mobilenet_v2.h5"
accl = SyncAccl(dfp=str(DFP_PATH))
mxa_outputs = [Link](x)
We get a list/array of outputs. In this case, with a shape of (1, 1000) and a dtype of float32.
This output should be normalized to a NumPy array:
mxa_outputs = [Link](mxa_outputs)
if mxa_outputs.ndim == 3:
mxa_outputs = mxa_outputs[0]
358
Decode the MXA Results
Expected output:
MXA top-5:
# 282: tiger_cat (38.6%)
# 281: tabby (18.3%)
# 285: Egyptian_cat (15.2%)
# 287: lynx (3.9%)
# 478: carton (1.7%)
359
Comparing CPU vs. MXA Performance
We can also run the unconverted model (mobilenet_v2.h5) on the CPU, applying the code to
the same input tensor:
cpu_model = [Link].load_model(KERAS_PATH)
cpu_outputs = cpu_model.predict(x)
num_classes = cpu_outputs.shape[-1]
idx2label = idx2label_full if num_classes == len(idx2label_full) else None
Expected output:
CPU top-5:
# 282: tiger_cat (58.4%)
# 285: Egyptian_cat (12.9%)
# 281: tabby (11.6%)
# 287: lynx (3.4%)
# 588: hamper (1.3%)
Despite the probabilities not being identical, both models reach the same top prediction.
The slight differences are due to numerical precision variations between CPU and accelerator
implementations.
Measuring Latency
Note: The following sections break down the complete inference script into logical
components. The full working script is available separately and integrates all these
pieces together.
360
import time
# Warm-up run
_ = [Link](x)
# Timed inference
start = [Link]()
mxa_outputs = [Link](x)
mxa_latency = [Link]() - start
python run_inference_comp_mobilenetv2.py
Expected results:
361
The MobileNet V2 running with ExecuTorch/XNNPACK backend on a CPU has
around 20 ms of latency.
# Download ResNet50
python3 -c "import tensorflow as tf; \
[Link].ResNet50().save('resnet50.h5');"
# Compile
mx_nc -v -m resnet50.h5
362
The performance improvements are even more dramatic with larger models!
Clean Shutdown
[Link]()
Folders Structure
Documents/MEMRYX/
��� run_inference_comp_mobilenetv2.py # MobileNetV2 script
��� run_inference_comp_resnet50.py # ResNet50 script
��� images/
� ��� [Link] # Test image
��� models/
� ��� mobilenet_v2.h5 # Original model
� ��� mobilenet_v2.dfp # Compiled model
� ��� resnet50.h5 # Original model
� ��� [Link] # Compiled model
� ��� imagenet_class_index.json # Labels (auto-downloaded)
��� mx-env/
Here’s how the MX3 compares across different deployment approaches we’ve covered in this
course:
363
Approach Hardware MobileNetV2 Latency ResNet50 Latency Power (Active)
Execu- Raspberry ~20 ms NA ~
Torch/XN- Pi 5
NPACK
MemryX Dedi- ~10 ms ~10 ms ~
MX3 cated
accelera-
tor
Key Observations
In this part of the lab, we’ll deploy YOLOv8n (nano) for real-time object detection on the
Raspberry Pi 5 using the MemryX MX3 AI accelerator. We’ll cover the complete workflow
from model export to inference optimization.
364
But first, we should install ULTRALYTICS
The MemoryX team suggested:
The model and the image [Link] will be download and tested with the YOLOV8n:
4 persons, 1 bus, and one stop signal were detected in 522 ms.
365
Model Export and Compilation
YOLOv8 must be converted to ONNX format before compilation for the MX3:
366
#### Export to ONNX format
[Link](format='onnx', simplify=True)
We can use the MemryX Neural Compiler to generate the DFP file:
Key flags:
Output files:
367
2. Detection head (yolov8n_post.onnx): Bounding box decoding on CPU
368
Understanding YOLOv8 Output Format
Decoding Process
We should now create a script to run an object detector (YOLOv8 with a pre/post-processing
pipeline), print each detection (label, confidence, bounding box), and save a copy of the image
with the boxes drawn.
Configuration section
DFP_PATH = "./models/[Link]"
POST_MODEL_PATH = "./models/yolov8n_post.onnx"
IMAGE_PATH = "./images/[Link]"
CONF_THRESHOLD = 0.25
369
• DFP_PATH: path to the compiled model used for inference.
• CONF_THRESHOLD: minimum confidence score; detections below this are filtered out.
Running detection
Here we call a helper function detect_objects that encapsulates the heavy lifting:
• Runs inference.
• Returns:
– detections: list/array where each element is [x1, y1, x2, y2, conf,
class_id].
Printing results
370
print(f"\n{'='*60}")
print("Detection Results:")
print(f"{'='*60}")
for i, det in enumerate(detections):
x1, y1, x2, y2, conf, class_id = det
print(f" {i+1}. {COCO_CLASSES[int(class_id)]}: {conf:.3f}")
print(f" Box: [{int(x1)}, {int(y1)}, {int(x2)}, {int(y2)}]")
• The loop goes over each detection, unpacks the bounding box coordinates, confidence,
and class ID.
if len(detections) > 0:
output_path = IMAGE_PATH.rsplit('.', 1)[0] + '_detected.jpg'
annotated_image.save(output_path)
print(f"\nSaved: {output_path}")
• The output filename is built by taking the original name and appending _detected
before the extension (e.g., bus_detected.jpg).
• annotated_image.save(...) writes the image with drawn boxes and labels to disk.
print(f"\n{'='*60}")
print(f"Total: {len(detections)} objects")
print(f"Time: {inference_time:.2f} ms")
print(f"{'='*60}")
371
• Prints the inference time, which is useful to talk about performance (e.g., model size
vs. speed, hardware differences).
python yolov8_m3_detect.py
As a result, we can see that the models found 4 persons and 1 bus, missing only the stop signal.
Regarding latency, the MX3 runs inference about 11 times faster than a CPU-only
system.
Basically, the same accuracy result that we got on the YOLO chapter running
yolov11
372
• Load and normalize:
– Open the image with PIL, convert to RGB, and get its original size.
– Compute a scale ratio so the image fits into 640×640 without distortion (preserving
aspect ratio).
• Letterboxing:
– Resize the image to (new_w, new_h) = (int(w * ratio), int(h * ratio)).
– Paste it onto a 640×640 canvas filled with color (114, 114, 114) (same as
Ultralytics).
– Compute the padding offsets (pad_w, pad_h) so we can undo this later.
• Tensor conversion:
– Convert to numpy, normalize to [0,1], permute from HWC to CHW, and add a batch
dimension to get shape [1, 3, 640, 640], which matches YOLOv8’s expected
input.[1][2]
• For COCO YOLOv8n ONNX, the detection head outputs a tensor of shape (1, 84,
8400).
• 84 = 4 (bbox) + 80 (class scores). Each of the 8400 positions corresponds to one
candidate box.
Function Walkthrough:
• Transpose:
– From (1, 84, 8400) to (8400, 84) so each row is: [x_center, y_center,
width, height, class_0_score, ..., class_79_score].
• Best class per box:
– Take max_scores = [Link](class_scores, axis=1) and class_ids =
[Link](class_scores, axis=1) to select the most likely class and its
score for each of the 8400 candidates.
• Confidence filtering:
– Drop boxes whose max class score is below conf_threshold.
• Coordinate conversion:
373
– Convert from YOLO’s center-format (x, y, w, h) to corner-format (x1, y1, x2,
y2) to make drawing and IoU calculation simpler.
• NMS:
– Call apply_nms to remove overlapping boxes and keep only the best ones.
IoU:
area of intersection
IoU =
area of union
• compute_iou_batch does this between one box and many boxes at once using vectorized
Numpy operations.
NMS:
• apply_nms:
– Sort boxes by score descending.
– Repeatedly pick the highest-score box, compute its IoU with the remaining boxes,
and discard those whose IoU is above iou_threshold.
• The result is a list of indices for boxes that don’t overlap too much and represent unique
objects.
374
5. Drawing results (draw_detections)
• Preprocess once: call preprocess_image to get the model-ready tensor and the info
needed for rescaling.
• Create the accelerator:
– accl = AsyncAccl(dfp_path) loads the compiled Memryx DFP model.
– accl.set_postprocessing_model(post_model_path, model_idx=0) attaches
the ONNX post-processing graph.
• Streaming-style design:
– frame_queue is a queue of inputs; you put your tensor in it.
– generate_frame is a generator feeding frames into the accelerator.
– process_output is a callback that collects outputs into results.
– The code wires them with connect_input and connect_output, then waits for
completion with [Link]().
• Post-processing:
– Grab the first output, call decode_predictions, rescale boxes, and draw.
Making Inferences
Let’s change the script to easily handle different images and confidence threshold
(yolov8_m3_detect_v2.py). We should replace the hardcoded IMAGE_PATH with a
command-line argument:
375
import argparse
if __name__ == "__main__":
parser = [Link]()
parser.add_argument(
"-i", "--image",
type=str,
required=True,
help="Path to input image"
)
parser.add_argument(
"-c", "--conf",
type=float,
default=0.25,
help="Confidence threshold"
)
args = parser.parse_args()
# Configuration
DFP_PATH = "./models/[Link]"
POST_MODEL_PATH = "./models/yolov8n_post.onnx"
IMAGE_PATH = [Link]
CONF_THRESHOLD = [Link]
# Run detection
detections, annotated_image, inference_time = detect_objects(
DFP_PATH,
POST_MODEL_PATH,
IMAGE_PATH,
CONF_THRESHOLD
)
# Print results
print(f"\n{'='*60}")
print("Detection Results:")
print(f"{'='*60}")
for i, det in enumerate(detections):
x1, y1, x2, y2, conf, class_id = det
print(f" {i+1}. {COCO_CLASSES[int(class_id)]}: {conf:.3f}")
print(f" Box: [{int(x1)}, {int(y1)}, {int(x2)}, {int(y2)}]")
376
if len(detections) > 0:
output_path = IMAGE_PATH.rsplit('.', 1)[0] + '_detected.jpg'
annotated_image.save(output_path)
print(f"\nSaved: {output_path}")
print(f"\n{'='*60}")
print(f"Total: {len(detections)} objects")
print(f"Time: {inference_time:.2f} ms")
print(f"{'='*60}")
As we saw in the YOLO chapter, we are assuming we are in an industrial facility that must
sort and count wheels and special boxes.
377
Each image can have three classes:
We have captured a raw dataset using the Raspberry Pi Camera and labeled it with the
ROBOFLOW. The Yolo model was trained on a Google Colab using Ultralytics.
Using the FileZilla FTP, transfer a few images from the test dataset to .\images:
378
Let’s return to the ./MEMRYX/YOLO folder and using the Python Interpreter, to quickly do some
inferences:
python
We will import the YOLO library and define the model to use:
Now, let’s define an image and call the inference (we will save the image result this time to
external verification):
379
We can see that the model is working and that the latency was 168 ms.
Let’s now export the model first to ONNX and after to FFPls
, to run it in the MX3 device:
cd ./models
yolo export model=box_wheel_320_yolo.pt format=onnx
mx_nc -v --autocrop -m box_wheel_320_yolo.onnx
cd ..
Naturally we should enter with the new models ’names and instead of COCO_LA-
BELS, the script was changed to:
380
# dataset class names
CLASSES = [
'Box', 'Wheel'
]
Thant’s all!
Run it with:
The Result was great! And the latency (~38 ms) was 4 times lower than with the
CPU-only approach (even smaller than the model exported to NCNN, runing 100% at CPU -
80 ms).
381
Advanced Topics
accl = AsyncAccl(dfp_path)
accl.set_postprocessing_model(post_model_path)
Thermal Management
Model Selection
382
Clone the MemryX eXamples Repository
After cloning the repository, you’ll find several subdirectories with different categories of
applications:
• Preprocessing pipelines
• Multi-threaded inference
• Output visualization
• Performance optimization
• Multi-model orchestration
Exploring these examples is an excellent way to learn production-ready patterns for deploying
MemryX applications.
383
sudo raspi-config
# Navigate to: Advanced Options → PCIe Speed → Enable PCIe Gen 3
sudo reboot
• Ensure sufficient power: Use the official Raspberry Pi 27W power supplyCheck HAT
installation: Ensure the M.2 HAT is properly seated.
• Lower Frequency: Try running sudo mx_set_powermode with a lower frequency, such
as 200 or 300 MHz. Then restart mxa-manager for good measure with sudo service
mxa-manager restart
sudo mx_set_powermode
384
sudo service mxa-manager restart
cat /etc/memryx/[Link]
We can check it by the first field of the file, FREQ4. If the Raspberry Pi is set to the module’s
default operating frequency (500 MHz), we should see FREQ4C=500, indicating that the module
is set to a 500 MHz clock speed for 4-chip DFPs.
If decreasing the frequency solves the issue, then you can either keep the default
frequency for all DFPs at 300 MHz (or 400, 450, etc.), or you can raise it back to
500 MHz and use the C++ API’s set_operating_frequency function to change the
clock speed on a per-DFP basis.
Compilation Errors
import tensorflow as tf
import tf2onnx
model = [Link].load_model('model.h5')
onnx_model, _ = [Link].from_keras(model)
with open("[Link]", "wb") as f:
[Link](onnx_model.SerializeToString())
385
• Check the compilation log (-v flag) to identify which specific layer is causing issues
Thermal Throttling
• Verify heatsink installation: Ensure thermal paste is properly applied and heatsink is
firmly attached
• Improve airflow: Position the Raspberry Pi for better air circulation
• Check ambient temperature: Ensure the room temperature is reasonable (<30°C)
• Monitor continuously:
deactivate
rm -rf mx-env
python -m venv mx-env
source mx-env/bin/activate
pip install --upgrade pip wheel
pip install --extra-index-url [Link] memryx
386
Low FPS / Poor Performance
sudo raspi-config
# Advanced Options → PCIe Speed
htop
• Verify Frequency
By default, the frequency should be at 500 MHz. Smaller frequencies will reduce the FPS
(increase the latency)
cat /etc/memryx/[Link]
Import Errors
• Reinstall if necessary:
387
pip install --force-reinstall --extra-index-url \
[Link] memryx
import sys
print([Link])
• Check input normalization: Confirm the value range matches training (e.g., [0, 1] vs
[-1, 1])
• Test with known inputs: Use the validation dataset to verify accuracy
• Compare outputs numerically: Print raw logits/probabilities to identify differences
• Check for quantization effects: If using -q flag, try without quantization first
Project Ideas
388
• Compare against the models from previous labs
4. Power Efficiency Study
• Measure power consumption with a USB meter
• Compare CPU vs. MX3 energy per inference
• Calculate battery life for mobile applications
5. Multi-Stream Processing
• Process multiple camera streams simultaneously
• Demonstrate chip utilization across streams
• Build a multi-camera monitoring system
• Quantization: Experiment with 8-bit and 4-bit quantization for even better performance
• Model Zoo: Explore pre-optimized models in the MemryX Model Explorer
• Async API: Use AsyncAccl for non-blocking, concurrent processing
• Custom Operators: Learn to handle models with custom layers
• Multi-chip Scaling: Understand how workload distributes across the four accelerators
Conclusion
In this lab, we’ve explored hardware acceleration for edge AI using the MemryX MX3 accelerator.
We’ve learned:
The MX3 demonstrates that dedicated AI accelerators can deliver significant performance
improvements for edge applications, achieving FPS several times higher (up to 25x for ResNet-
50) than CPU inference while maintaining accuracy and providing deterministic latency.
As edge AI continues to evolve, hardware acceleration will become increasingly important for
real-time, power-efficient deployments. The skills we’ve developed in this lab—understanding
the compilation workflow, benchmarking methodologies, and performance optimization—will
transfer to other accelerator platforms as well.
389
References and Further Reading
Official Documentation
Background Reading
#
Generative AI (Proactive)
390
Text Generation with RNNs
Introduction
In this chapter, we will explore how to build a character-level text generation model using
Recurrent Neural Networks (RNNs), specifically inspired by the works of Jules Verne.
The “Jules Verne Bot” project will help to show us the fundamental concepts of sequence
modeling and text generation using deep learning techniques, as a preview of how the modern
LLMs work.
Project Overview:
391
• Goal: Create an AI model that generates text in the style of Jules Verne
• Architecture: RNN with GRU (Gated Recurrent Unit) layers
• Approach: Character-level text prediction
• Framework: TensorFlow/Keras
• Platform: Google Colab with Tesla T4 GPU
Imagine we could teach a computer to write like Jules Verne, the famous author of “Twenty
Thousand Leagues Under the Sea” and “Around the World in Eighty Days.” That’s precisely
what we’re doing with the Jules Verne Bot. This project creates an artificial intelligence system
that learns the patterns, style, and vocabulary from Jules Verne’s novels, then generates new
text that sounds like it could have come from his pen.
Think of it like this: if we read enough of someone’s writing, we start to recognize their style.
We notice they use certain phrases, prefer specific sentence structures, or have favorite topics.
Our neural network does something similar, but with mathematical precision. It analyzes
millions of characters from Verne’s works and learns to predict what character should come
next in any given sequence.
Before we dive into the technical details, let’s understand why we use neural networks for this
task and why we chose the specific type we did.
When you read a sentence like “The submarine descended into the dark…” your brain auto-
matically starts predicting what might come next. Maybe “depths” or “ocean” or “waters.”
Your brain does this because it has learned patterns from all the text you’ve ever read. Neural
networks work similarly, but they learn these patterns through mathematical calculations
rather than biological processes.
Before diving into our RNN implementation, let’s understand where RNNs fit in the neural
network ecosystem:
Key Neural Network Architectures:
392
• MLP (Multi-Layer Perceptron): Basic feedforward networks for general tasks, for
example, vibration analysis
• CNN (Convolutional Neural Networks): Specialized for image processing as Image
Classification tasks and spatial data
• RNN (Recurrent Neural Networks): Designed for sequential data like text and time
series
• GAN (Generative Adversarial Networks): Two networks competing for realistic
data generation, as images
• Transformers (Attention Networks): Modern architecture using attention mecha-
nisms, as in LLMs (Large Language Models, such as GPT)
We chose a Recurrent Neural Network (RNN) for this project because text has a crucial
property: order matters tremendously. The sequence “The cat sat on the mat” means something
completely different from “Mat the on sat cat the.” Regular neural networks process all inputs
simultaneously, like looking at a photograph. But for text, we need a network that processes
information sequentially, remembering what came before to understand what should go next.
In text generation, we aim to predict the most probable word to follow a sentence.
393
Think of reading a book. You don’t just look at all the words on a page simultaneously. You
read word by word, sentence by sentence, and your understanding builds as you progress. Each
new word is interpreted in the context of everything you’ve read before in that chapter. RNNs
work the same way.
Early RNNs had a significant problem: they couldn’t remember information for very long.
Imagine trying to understand a story where you could only remember the last few words you
read. You’d lose track of characters, plot points, and context very quickly.
This is where the Gated Recurrent Unit (GRU) comes in. Think of GRU as an improved
memory system with two special abilities:
Reset Gate: This decides when to “forget” old information. If the story switches to a new
scene or character, the reset gate helps the network forget irrelevant details from the previous
context.
Update Gate: This decides how much new information to incorporate. When encountering
important plot points or character names, the update gate helps the network remember these
crucial details for longer.
It’s like having a smart note-taking system that automatically decides what’s worth remembering
and what can be forgotten.
Why RNNs for Text Generation?
Recurrent Neural Networks are designed explicitly for sequential data processing. Key
characteristics:
394
• Sequential Processing: Process data one element at a time, making them ideal for text
• Variable Length Input: Can handle sequences of different lengths
• Parameter Sharing: Same weights applied across different time steps
Dataset Preparation
395
Our model is trained on a curated collection of 10 classic Jules Verne novels, downloaded from
public domain texts of the Gutenberg Project:
“A Journey to the Centre of the Earth” teaches the model about geological descriptions and
underground adventures. “Twenty Thousand Leagues Under the Sea” provides vocabulary
about marine life and submarine technology. “Around the World in Eighty Days” offers
geographical references and travel descriptions. Each book contributes unique vocabulary and
stylistic elements while maintaining Verne’s consistent voice.
The complete dataset contains 5,768,791 characters, of which 123 are unique. To put this
in perspective, that’s roughly equivalent to 3,000 pages of text. This provides our neural
network with ample material to learn from, enabling it to capture both common patterns and
unique expressions in Verne’s writing.
return text
396
Tokenization and Vocabulary
Character-Level Tokenization
Here’s where our approach differs from how humans typically think about text. While we
naturally think in words and sentences, our model processes text character by character. This
means it learns that certain letters frequently follow others, that spaces separate words, and
that punctuation marks signal sentence boundaries.
Why choose character-level processing? Consider the word “extraordinary,” which appears
frequently in Verne’s work. A word-level model would need to have seen this exact word during
training to use it. But a character-level model can generate this word by learning that ‘e’ often
starts words, ‘x’ can follow ‘e’, ‘t’ often follows ‘x’, and so on. This allows our model to create
new words or handle misspellings gracefully.
The downside is that character-level processing requires more computational steps to generate
the same amount of text. Generating “Hello world” requires 11 prediction steps instead of just
2. However, for our educational purposes, this trade-off provides valuable insights into how
language generation works at its most fundamental level.
Please see the following site for a great general visual explanation, from Andrej
Karpathy, The Unreasonable Effectiveness of Recurrent Neural Networks.
Computers work with numbers, not letters, so we need to convert our text into a numerical
representation. We start by finding every unique character in our dataset. This includes
not just letters A-Z and a-z, but also numbers, punctuation marks, spaces, and even special
characters that might appear in the original texts.
Our Jules Verne collection contains 123 unique characters. These include obvious ones like
letters and common punctuation, but also less common characters like accented letters from
French names or special typography marks from the original publications.
397
Creating the Character Dictionary
We create two dictionaries: one that converts characters to numbers (encoding) and another
that converts numbers back to characters (decoding). For example:
‘a’ might become 47, ‘b’ becomes 48, ‘c’ becomes 49, and so on. The space character might be
1, and the period might be 72. These assignments are arbitrary but consistent throughout our
project.
When we want to process the phrase “The sea”, we convert it to something like [84, 72, 69, 1,
83, 69, 47]. When the model generates numbers like [84, 72, 69, 1, 87, 47, 83], we convert them
back to “The was” (as an example).
398
It is possible to experiment with (sub-word level) tokenization using OpenAI’s
tokenizer tool at:
[Link]
Training Sequences
Our model learns by playing a sophisticated prediction game. We show it sequences of 120
characters and ask it to predict what the 121st character should be. Think of it like a
fill-in-the-blank exercise, but instead of missing words, we’re missing the next character.
For example, if our text contains “The submarine descended into the dark depths of the ocean”,
we might show the model “The submarine descended into the dark depths of the ocea” and ask
399
it to predict “n”. Then we slide our window forward by one character and show it “he submarine
descended into the dark depths of the ocean” and ask it to predict the next character.
Training Configuration
We chose 120 characters as our context window because it roughly corresponds to one
paragraph of text in English. This gives the model enough context to understand local patterns
(like completing words and phrases) while remaining computationally manageable. In practical
terms, 120 characters might look like:
“The Nautilus had been cruising in these waters for some time. Captain Nemo stood on the
bridge, observing the vast exp”
From this context, the model might predict “a” to complete “expanse” or “l” to form “explore”.
The longer the context window, the better the model can maintain coherence, but
the more computer memory and processing time it requires.
Training Example
This means our dataset of 5.8 million characters becomes millions of individual training
examples, each teaching the model about character sequence patterns.
400
def create_training_sequences(text, seq_length):
sequences = []
targets = []
Character Embeddings
Initially, each character is represented as a one-hot vector, which is mostly zeros with a single
one indicating which character it is. For 123 characters, this means each character is represented
by a vector with 123 elements, where 122 are zero and 1 is one. This is wasteful and doesn’t
capture any relationships between characters.
Character embeddings solve this problem by representing each character as a dense vector
of real numbers. Instead of 123 mostly-zero values, each character becomes 256 meaningful
numbers. These numbers are learned during training and end up encoding relationships between
characters.
Something fascinating happens during training: characters that behave similarly end up with
similar embedding vectors. Vowels tend to cluster together because they can often substitute
for each other in similar contexts. Consonants that frequently appear together (like ‘th’ or ‘ch’)
develop related embeddings.
The model learns that uppercase and lowercase versions of the same letter are related but
distinct. It discovers that digits form their own cluster since they appear in similar contexts
(dates, measurements, chapter numbers). Punctuation marks develop embeddings based on
their grammatical functions.
401
Visualization and Understanding
402
You can play with Word2Vec - Embedding Projector
Model Architecture
Our Jules Verne Bot consists of three main components, each serving a specific purpose in the
text generation pipeline.
Embedding Layer: This is our translation layer. It takes character indices (numbers like
47, 83, 72) and converts them into dense 256-dimensional vectors that capture character
relationships. Think of this as converting raw symbols into a format that captures meaning
and relationships.
GRU Layer: This is the brain of our operation. With 1024 hidden units, this layer processes
sequences and maintains memory about what it has seen. When processing the sequence “The
submarine descended”, the GRU maintains a hidden state that encodes information about the
submarine, the action of descending, and the overall maritime context.
Dense Output Layer: This is our decision-making layer. It takes the GRU’s 1024-dimensional
hidden state and converts it into 123 probabilities, one for each character in our vocabulary.
These probabilities represent the model’s confidence about what character should come next.
Model Summary
Model: "sequential_4"
_________________________________________________________________
Layer (type) Output Shape Param #
403
=================================================================
embedding_4 (Embedding) (1, 120, 256) 31,488
_________________________________________________________________
gru_3 (GRU) (1, 120, 1024) 3,938,304
_________________________________________________________________
dense_3 (Dense) (1, 120, 123) 126,075
=================================================================
Total params: 4,095,867 (15.62 MB)
Trainable params: 4,095,867 (15.62 MB)
Non-trainable params: 0 (0.00 B)
Our model has 4,095,867 parameters. These are the individual numbers that the model adjusts
during training to improve its predictions. To put this in perspective, each parameter is like a
tiny dial that affects how the model processes information.
The GRU layer contains most of these parameters (about 3.9 million) because it needs to learn
complex patterns about how characters relate to each other across different time steps. The
embedding layer has about 31,000 parameters (123 characters × 256 dimensions), and the
output layer has about 126,000 parameters.
When processing text, information flows through the model like this:
A character index enters the embedding layer and becomes a 256-dimensional vector. This
vector enters the GRU, which combines it with its current memory state to produce a new
1024-dimensional hidden state. This hidden state captures everything the model “knows” at
this point in the sequence.
The hidden state goes to the dense layer, which produces probability scores for each of the
123 possible next characters. The character with the highest probability becomes the model’s
prediction.
Crucially, the GRU’s hidden state becomes its memory for the next character prediction.
This creates a chain of memory that allows the model to maintain context across the entire
sequence.
404
Why GRU over Basic RNN?
GRU Advantages:
Training a neural network means adjusting its millions of parameters so it makes better
predictions. We use a loss function called sparse categorical crossentropy, which measures how
far off the model’s predictions are from the correct answers.
Think of it like teaching someone to play darts. Each throw (prediction) has a target (the
correct next character). The loss function measures how far each dart lands from the bullseye.
Training adjusts the player’s technique (the model’s parameters) to improve accuracy over
time.
We trained our model on a Tesla T4 GPU, which can perform thousands of calculations
simultaneously. This parallelization is crucial because each training step involves matrix
multiplications with millions of numbers. The training took 33 minutes for 30 complete passes
through the entire dataset.
To understand why we need a GPU, consider that training involves calculating gradients for all
4 million parameters, potentially thousands of times per second. A regular CPU would take
many hours to complete the same training that a GPU accomplishes in minutes.
Monitoring Progress
During training, we watch the loss decrease from about 1.9 to 0.9. This represents the model’s
improving ability to predict the next character. Early in training, the model makes essentially
random predictions. By the end, it has learned sophisticated patterns about English spelling,
grammar, and Jules Verne’s writing style.
The learning curve typically shows rapid improvement in the first few epochs as the model
learns basic patterns like common letter combinations. Later epochs show slower but steady
405
improvement as the model refines its understanding of more complex patterns like narrative
structure and thematic elements.
Preventing Overfitting
One challenge in training is overfitting, in which the model memorizes the training data rather
than learning generalizable patterns. We use techniques like monitoring validation loss and
potentially stopping training early if the model stops improving on unseen text.
Training Configuration
Hardware Setup:
Training Parameters:
406
Training Implementation
# Model compilation
[Link](
optimizer='adam',
loss='sparse_categorical_crossentropy',
metrics=['accuracy']
)
Text Generation
Once trained, our model becomes a text generation engine. We start with a seed phrase like:
“THE FLYING SUBMARINE”
and ask the model to continue the story. The process works character by character:
The model receives “THE FLYING SUBMARINE” and predicts the most likely next character
based on everything it learned from Jules Verne’s works. Maybe it predicts a space, starting a
new word. Then we feed “THE FLYING SUBMARINE” (with the space) back to the model
and ask for the next character.
This process continues indefinitely, with each new character becoming part of the context
for predicting the next one. The model might generate “THE FLYING SUBMARINE de-
scended into the mysterious depths…” as it draws upon patterns learned from Verne’s nautical
adventures.
407
Temperature Control
Here’s where we can control the model’s creativity through a parameter called temperature.
Temperature affects how the model chooses between different possible next characters.
With temperature set to 0.1, the model almost always picks the most probable next character.
This produces very predictable, conservative text that closely mimics the training data but
might be repetitive or boring.
With temperature set to 1.0, the model considers all possible next characters according to
their learned probabilities. This produces more varied and creative text, but sometimes makes
unusual choices that lead to interesting narrative directions.
With temperature above 1.5, the model becomes quite random, often producing text that starts
coherently but gradually becomes nonsensical as unlikely character combinations accumulate.
In short:
Implementation
text_generated = []
model.reset_states()
for i in range(num_generate):
predictions = model(input_eval)
predictions = [Link](predictions, 0)
# Apply temperature
predictions = predictions / temperature
predicted_id = [Link](predictions, num_samples=1)[-1,0].numpy()
408
return start_string + ''.join(text_generated)
This eBook is for the use of anyone anywhere in the United States and most
other parts of the earth and miserable eruptions. The solar rays should be
entirely under the shock of the intensity of the sea. We were all sorts. Are
we to prepare for our feelings?"
"Well, then, John, for I get to the Pampas, that we ought to obey the same
time. In the country of this latitude changed my brother, and the
_Nautilus_ floated in a sea which contained the rudder and
lower colour visibly. The loiter was a fatalint region the two
scientific discoverers. Several times turning toward the river, the cry
of doors and over an inclined plains of the Angara, with a threatening
water and disappeared in the midst of the solar rays.
The weather was spread and strewn with closed bottoms which soon appeared
that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives, and the prisoners
were successively assuming the sides of this agreement for fifteen days
with a threatening voice.
...
Let’s examine some generated text: “The weather was spread and strewn with closed bottoms
which soon appeared that the unexpected sheets of wind was soon and linen, and the whole
seas were again landed on the subject of the natives…”
This excerpt shows both the model’s strengths and limitations. It successfully captures Verne’s
descriptive style and maritime vocabulary (“seas,” “wind,” “natives”). The sentence structure
409
feels appropriately Victorian and elaborate. However, the meaning becomes confused with
phrases like “closed bottoms” and “sheets of wind was soon and linen.”
This illustrates the fundamental challenge of character-level generation: the model learns local
patterns (how words are spelled, common phrases) much better than global coherence (logical
narrative flow, consistent meaning).
Our 120-character context window creates a fundamental limitation. The model can only “see”
about one paragraph of previous text when making predictions. This means it might introduce
a character named Captain Smith, then 200 characters later introduce another character with
the same name, having “forgotten” the first introduction.
Humans writing stories maintain mental models of characters, plot lines, and world-building
details across entire novels. Our model’s memory effectively resets every 120 characters, making
long-term narrative consistency nearly impossible.
Character-level generation requires many more prediction steps than word-level generation.
Generating the phrase “extraordinary adventure” requires 22 character predictions instead
of just 2 word predictions. This makes character-level generation much slower and more
computationally expensive.
However, character-level generation offers unique advantages. The model can generate new
words it has never seen before by combining character patterns. It can handle misspellings,
made-up words, or technical terms more gracefully than word-level models that have fixed
vocabularies.
Coherence Challenges
Perhaps the biggest limitation is maintaining semantic coherence. The model might generate
grammatically correct text that makes no logical sense. It can describe “The submarine floating
in the air above the mountain peaks” because it has learned that submarines float and that
Verne often described mountains, but it hasn’t learned the physical constraint that submarines
float in water, not air.
410
This happens because the model learns statistical patterns without understanding meaning.
It knows that certain word combinations are common without understanding why they make
sense.
Summary
Potential Improvements
Scale Comparison
To appreciate how far language modeling has advanced, consider the scale differences between
our Jules Verne Bot and modern language models:
411
Our model has 4 million parameters and was trained on about 5.8 million characters (10 books).
GPT-3 has 175 billion parameters and was trained on 45 terabytes of text (roughly equivalent
to millions of books). That’s a difference of over 40,000 times more parameters and millions of
times more training data.
Modern small language models (SLMs) like Phi-3-mini still dwarf our model with 3.8 billion
parameters, but they represent more efficient designs that achieve impressive performance with
“only” 1,000 times more parameters than our model.
Architectural Evolution
The biggest advancement since RNNs is the Transformer architecture, which uses attention
mechanisms instead of recurrent processing. While RNNs process text sequentially (like
reading word by word), Transformers can examine all parts of a text simultaneously and learn
relationships between any two words, regardless of how far apart they are.
This solves the long-term memory problem that limits our RNN model. A Transformer can
maintain awareness of a character introduced in a hypothetical “chapter 1” while writing
“chapter 10”, something our 120-character context window makes impossible.
Training Efficiency
Modern models also benefit from more sophisticated training techniques. They’re pre-trained
on massive, diverse datasets to learn general language patterns, then fine-tuned on specific tasks.
They use techniques like instruction tuning, where they learn to follow human commands, and
reinforcement learning from human feedback, where they learn to generate text that humans
find helpful and appropriate.
1. Transformer Architecture
412
• Attention Mechanism: Can look at any part of the input sequence
• Parallel Processing: Much faster training and inference
• Better Long-range Dependencies: Maintains context over thousands of tokens
2. Scale
• More Data: Trained on vastly more diverse text
• More Parameters: Can memorize and generalize better
• More Compute: Allows for more sophisticated training techniques
3. Advanced Techniques
• Pre-training + Fine-tuning: Learn general language then specialize
• Instruction Tuning: Trained to follow human instructions
• RLHF: Reinforcement Learning from Human Feedback
Conclusion:
Building the Jules Verne Bot teaches us that creating artificial intelligence systems capable
of generating human-like text requires careful consideration of multiple components working
together. The embedding layer learns to represent characters meaningfully, the RNN layer
processes sequences and maintains memory, and the output layer makes predictions based on
learned patterns.
The project also illustrates the fundamental trade-offs in machine learning: between model
complexity and training speed, between creativity and coherence, between local accuracy and
global consistency. These trade-offs appear in every AI system, from simple character-level
generators to the most sophisticated language models.
Most importantly, this project demonstrates that impressive AI capabilities emerge from
relatively simple components combined thoughtfully. Our 4-million parameter model, while
limited compared to modern systems, genuinely learns to write in Jules Verne’s style through
nothing more than statistical pattern recognition and mathematical optimization.
The techniques we’ve explored, sequence processing, embedding learning, and generation
strategies, form the foundation for understanding any language model. Whether you encounter
RNNs, Transformers, or future architectures yet to be invented, the core concepts remain
consistent: learn patterns from data, encode meaning in mathematical representations, and
generate new content by predicting what should come next.
Understanding these fundamentals provides the foundation for working with, improving, or
creating the next generation of language models that will shape how humans and computers
communicate in the future.
413
Resourses
• Gutenberg Project
• The Unreasonable Effectiveness of Recurrent Neural Networks
• Word2Vec - Embedding Projector
• OpenAI’s tokenizer tool
• Generating Text with RNNs: The Jules Verne Bot - CoLab
414
Knowledge Distillation in Practice
415
Figure 8: Image created by DALLE-3
Knowledge distillation is a powerful technique in machine learning that enables the transfer
of knowledge from a large, complex model (the “teacher”) to a smaller, more efficient model
(the “student”). This process allows us to create compact models that maintain much of the
416
performance of their larger counterparts while being significantly faster and requiring fewer
computational resources.
In today’s AI landscape, models are becoming increasingly large and complex. While these
models achieve remarkable performance, they often require substantial computational resources,
making deployment challenging in resource-constrained environments such as mobile devices,
edge computing systems, or real-time applications. Knowledge distillation addresses this
challenge by:
The core concept of knowledge distillation revolves around the teacher-student relationship:
• Teacher Model: A large, well-trained model with high capacity and performance
• Student Model: A smaller, more efficient model trained to mimic the teacher’s behavior
• Knowledge Transfer: The process of transferring the teacher’s “dark knowledge” to
the student
417
Figure 9: Figure from “Knowledge Distillation: A Survey, Jianping Gou Baosheng Yu Stephen
J. Maybank Dacheng Tao, 2021
See paper: “Knowledge Distillation: A Survey, Jianping Gou, Baosheng Yu, Stephen
J. Maybank, Dacheng Tao, 2021
Theoretical Foundations
Traditional supervised learning uses “hard targets” - one-hot encoded labels that provide limited
information. For example, in MNIST digit classification, the label for the digit “5” would be
represented as [0, 0, 0, 0, 0, 1, 0, 0, 0, 0].
Knowledge distillation leverages “soft targets” - the probability distributions produced by
the teacher model. These soft targets contain richer information about the relationships
between classes. For instance, the teacher might output [0.01, 0.02, 0.01, 0.05, 0.1,
0.78, 0.02, 0.01, 0.0, 0.0] for a “5”, indicating that it’s most confident about “5” but
also considers “4” somewhat similar.
The temperature parameter (�) is crucial in knowledge distillation. It controls the “softness” of
the probability distribution by modifying the softmax function:
418
P(class_i) = exp(z_i / �) / Σ_j exp(z_j / �)
Where:
Effects of Temperature:
1. Distillation Loss (L_KD): Measures how well the student mimics the teacher’s soft
targets
2. Student Loss (L_CE): Traditional cross-entropy loss with hard targets
Where � is a weighting parameter that balances the two objectives, in our implementation, we
use � = 0.3, giving more weight to the soft targets from the teacher.
The distillation loss is typically computed using the KL divergence:
The �² factor compensates for the gradient scaling effect of temperature. This is crucial for
stable training.
419
Implementation with TensorFlow and MNIST
Dataset Overview
Our teacher model has substantial capacity with multiple convolutional layers, batch normal-
ization, and dropout for regularization:
def build_teacher_model():
model = [Link]([
layers.Conv2D(64, 3,
activation='relu',
padding='same',
input_shape=(28, 28, 1)),
[Link](),
layers.Conv2D(64, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),
[Link](256, activation='relu'),
[Link](),
[Link](0.5),
[Link](128, activation='relu'),
420
[Link](),
[Link](0.5),
[Link](10, activation='softmax')
])
return model
421
422
This architecture achieved 99.44 % accuracy on MNIST.
Our student model is intentionally much simpler, with fewer layers and significantly fewer
parameters:
def build_student_model():
model = [Link]([
layers.Conv2D(16, 3,
activation='relu',
padding='same',
input_shape=(28, 28, 1)),
layers.MaxPooling2D(2, 2),
layers.Conv2D(32, 3, activation='relu', padding='same'),
layers.MaxPooling2D(2, 2),
[Link](),
[Link](64, activation='relu'),
[Link](10, activation='softmax')
])
return model
The student model has fewer parameters than the teacher (105K versus 658K), or
6.2 times smaller.
423
Knowledge Distillation Implementation
In our implementation, we use a custom training loop that explicitly calculates both hard and
soft losses:
# Forward pass
with [Link]() as tape:
predictions = kd_student(x_batch, training=True)
Training Process
1. Train the Teacher: First, we train the complex teacher model using early stopping and
a reduced learning rate.
2. Train a Vanilla Student: We train a student model normally on the hard labels for
comparison.
3. Generate Soft Targets: We use the teacher to create softened probability distributions.
4. Train the Distilled Student: We train another student using our custom distillation
training loop.
5. Evaluate and Compare: We compare the performance, size, and speed of all three
models.
424
Results and Analysis
• Teacher Model: 99.44% accuracy, largest size, slowest inference (0.8632 seconds)
• Vanilla Student: 98.77% accuracy, smaller size, faster inference (0.3579 seconds)
• Distilled Student: 99.32% accuracy, same size as vanilla student, but better performance
(+0.55%) and faster inference than the Teacher (0.5467 seconds)
These results demonstrate the key benefit of knowledge distillation: the distilled student
achieves performance closer to the teacher while maintaining the efficiency benefits of the
smaller architecture.
We analyze the models on challenging examples where the teacher succeeds but the vanilla
student fails. This reveals how knowledge distillation enables students to handle complex
cases by learning the teacher’s “dark knowledge.”
425
Advanced Techniques
Feature-Based Distillation
426
loss += [Link](t_feat, s_feat_aligned)
return loss
Attention-Based Distillation
1. Layer Alignment: When using feature distillation, carefully align the feature dimensions
2. Feature Selection: Not all features are equally important; focus on the most informative
ones
3. Multi-Teacher Distillation: Combine knowledge from multiple teachers for better
results
4. Online Distillation: Train teacher and student simultaneously for mutual improvement
427
LLM Distillation Techniques
1. Sequence-Level Distillation
1. Teacher Model: The larger Llama 3.2 models (70B/8B parameters) serve as the teachers
2. Student Model: Llama 3.2 3B and 1B are the distilled student models. They were
created using pruning techniques, which systematically remove less meaningful connections
(weights) in the neural network.
428
Key Points
429
While our MNIST example demonstrates the application of knowledge distillation principles
using a simple dataset, these same principles can be applied directly to state-of-the-art language
models.
Meta’s Llama 3.2 family provides a perfect real-world example:
Parame- Size
Model ters Reduction Use Case
Llama 3.2 70 billion (Teacher) Data centers, high-performance applications
70B
Llama 3.2 8 billion ~9x Server deployment, high-end workstations
8B
Llama 3.2 3 billion ~23x Consumer laptops, desktop applications
3B
Llama 3.2 1 billion ~70x Edge devices, mobile applications, embedded
1B systems
The 3B and 1B models represent successful distillations of the larger models, preserving core
capabilities while dramatically reducing computational requirements. This demonstrates the
industrial importance of knowledge distillation techniques we’ve explored.
It is possible to note that while the underlying principles remain the same as our MNIST
example, industrial LLM distillation includes additional techniques:
Despite these additional complexities, the core concept remains the same: using a larger, more
capable model to guide the training of a smaller, more efficient one.
Large models like ChatGPT (175B+ parameters) can be distilled to create mobile-friendly
assistants (1-2B parameters) that maintain core capabilities while running locally on smart-
phones.
430
2. Domain-Specific Distillation
• Medical LLMs: Distill medical knowledge from large models to smaller, specialized
ones
• Legal Assistants: Create compact models focused on legal reasoning and terminology
• Educational Tools: Develop small models optimized for teaching specific subjects
The journey from research models to production deployment often involves distillation:
1. Teacher-Student Size Ratio: Aim for a 10-20x parameter reduction for significant
efficiency gains
2. Architectural Similarity: Maintain similar architectural patterns between teacher and
student
3. Bottleneck Identification: Ensure the student has adequate capacity at critical layers
Hyperparameter Selection
1. Temperature (�):
• For MNIST: 3-5 works well
• For complex tasks: 5-10 may be better
• If outputs are already soft: Lower temperatures (2-3) may suffice
2. Alpha Weighting (�):
431
• For simpler tasks: 0.3-0.5 (balanced approach)
• For complex reasoning: 0.1-0.3 (more emphasis on teacher’s knowledge)
• When teacher is extremely accurate: Lower � values work better
3. Training Duration:
• Distilled students often benefit from longer training (1.5-2x the epochs)
• Use early stopping with patience to avoid overfitting
1. Teacher Performance: Ensure the teacher actually outperforms the student (we fixed
this in our implementation)
2. Temperature Selection: If knowledge transfer is poor, experiment with different
temperatures
3. Loss Weighting: If the student ignores soft targets, reduce � to emphasize distillation
loss
4. Gradient Scaling: Always apply the �² correction factor to the soft loss
Performance Evaluation
1. Accuracy: Primary performance metric (should be closer to teacher than vanilla student)
2. Model Size: Parameter count and memory footprint (should match vanilla student)
3. Inference Speed: Time per prediction (should be significantly faster than teacher)
4. Challenging Cases: Performance on difficult examples (should be better than vanilla
student)
Conclusion
Key Takeaways
432
2. Dark Knowledge Matters: The soft probability distributions contain valuable infor-
mation beyond just the predicted class
3. Temperature and Alpha: These hyperparameters are crucial for effective knowledge
transfer
4. Practical Benefits: Smaller size, faster inference, and lower resource requirements make
AI more accessible
The principles we’ve demonstrated with MNIST directly scale to Large Language Models:
Final Thoughts
Knowledge distillation isn’t just an academic technique—it’s essential for practical AI deploy-
ment. The same principles that helped us compress our MNIST classifier can be scaled to
compress models with hundreds of billions of parameters. This universality makes knowledge
distillation an indispensable skill for engineering students entering the field of AI.
Resources
Notebook
433
434
Small Language Models (SLM)
Figure 10: Nano Banana prompt - A cartoon about SLMs runing SLMs as Llama and Gemma
with Ollama - Python on a Raspberry Pi
435
Introduction
Setup
We could use any Raspberry Pi model in the previous labs, but here, the choice must be the
Raspberry Pi 5 (Raspi-5). It is a robust platform that substantially upgrades the last version
4, equipped with the Broadcom BCM2712, a 2.4GHz quad-core 64-bit Arm Cortex-A76 CPU
featuring Cryptographic Extension and enhanced caching capabilities. It boasts a VideoCore
VII GPU, dual 4Kp60 HDMI® outputs with HDR, and a 4Kp60 HEVC decoder. Memory
options include 4GB, 8GB, and 16GB of high-speed LPDDR4X SDRAM, with 8GB as our
choice for running SLMs. It also features expandable storage via a microSD card slot and a
PCIe 2.0 interface for fast peripherals such as M.2 SSDs (Solid State Drives).
For real SSL applications, SSDs are a better option than SD cards.
By the way, as Alasdair Allan discussed, inferencing directly on the Raspberry Pi 5 CPU—with
no GPU acceleration—is now on par with the performance of the Coral TPU.
436
For more info, please see the complete article: Benchmarking TensorFlow and TensorFlow Lite
on Raspberry Pi 5.
We suggest installing an Active Cooler, a dedicated clip-on cooling solution for Raspberry Pi 5
(Raspi-5), for this lab. It combines an aluminum heatsink with a temperature-controlled blower
fan to keep the Raspi-5 operating comfortably under heavy loads, such as running SLMs.
437
The Active Cooler has pre-applied thermal pads for heat transfer and is mounted directly to
the Raspberry Pi 5 board using spring-loaded push pins. The Raspberry Pi firmware actively
manages it: at 60°C, the blower’s fan is turned on; at 67.5°C, the fan speed is increased; and
finally, at 75°C, the fan increases to full speed. The blower’s fan will spin down automatically
when the temperature drops below these limits.
Generative AI (GenAI)
Generative AI is an artificial intelligence system capable of creating new, original content across
various media such as text, images, audio, and video. These systems learn patterns from
existing data and use that knowledge to generate novel outputs that didn’t previously exist.
Large Language Models (LLMs), Small Language Models (SLMs), and multimodal
models can all be considered types of GenAI when used for generative tasks.
GenAI provides the conceptual framework for AI-driven content creation, with LLMs serving
as powerful general-purpose text generators. SLMs adapt this technology for edge computing,
while multimodal models extend GenAI capabilities across different data types. Together, they
represent a spectrum of generative AI technologies, each with its strengths and applications,
collectively driving AI-powered content creation and understanding.
438
Large Language Models (LLMs)
Large Language Models (LLMs) are advanced artificial intelligence systems that understand,
process, and generate human-like text. These models are characterized by their massive scale
in terms of the amount of data they are trained on and the number of parameters they contain.
Critical aspects of LLMs include:
1. Size: LLMs typically contain billions of parameters. For example, GPT-3 has 175 billion
parameters, while some newer models exceed a trillion parameters.
2. Training Data: They are trained on vast amounts of text data, often including books,
websites, and other diverse sources, amounting to hundreds of gigabytes or even terabytes
of text.
3. Architecture: Most LLMs use transformer-based architectures, which allow them to
process and generate text by paying attention to different parts of the input simultaneously.
4. Capabilities: LLMs can perform a wide range of language tasks without specific fine-
tuning, including:
• Text generation
• Translation
• Summarization
• Question answering
• Code generation
• Logical reasoning
5. Few-shot Learning: They can often understand and perform new tasks with minimal
examples or instructions.
6. Resource-Intensive: Due to their size, LLMs typically require significant computational
resources to run, often needing powerful GPUs or TPUs.
7. Continual Development: The field of LLMs is rapidly evolving, with new models and
techniques constantly emerging.
8. Ethical Considerations: The use of LLMs raises important questions about bias,
misinformation, and the environmental impact of training such large models.
9. Applications: LLMs are used in various fields, including content creation, customer
service, research assistance, and software development.
10. Limitations: Despite their power, LLMs can produce incorrect or biased information
and lack true understanding or reasoning capabilities.
439
We must note that we use large models beyond text, which we call multi-modal models. These
models integrate and process information from multiple types of input simultaneously. They
are designed to understand and generate content across various data types, such as text, images,
audio, and video.
Certainly. Let’s define open and closed models in the context of AI and language models:
Closed models, also called proprietary models, are AI models whose internal workings, code,
and training data are not publicly disclosed. Examples: GPT-3 and beyond (by OpenAI),
Claude (by Anthropic), Gemini (by Google).
Open models, also known as open-source models, are AI models whose code, architecture,
and, often, training data are publicly available. Examples: Gemma (by Google), LLaMA (by
Meta), and Phi (by Microsoft).
Open models are particularly relevant for running models on edge devices like Raspberry Pi as
they can be more easily adapted, optimized, and deployed in resource-constrained environments.
Still, it is crucial to verify their Licenses. Open models come with various open-source licenses
that may affect their use in commercial applications, while closed models have clear, albeit
restrictive, terms of service.
440
Small Language Models (SLMs)
In the context of edge computing on devices like Raspberry Pi, full-scale LLMs are typically
too large and resource-intensive to run directly. This limitation has driven the development of
smaller, more efficient models, such as the Small Language Models (SLMs).
SLMs are compact versions of LLMs designed to run efficiently on resource-constrained de-
vices such as smartphones, IoT devices, and single-board computers like the Raspberry Pi.
These models are significantly smaller in size and computational requirements than their
larger counterparts while still retaining impressive language understanding and generation
capabilities.
Key characteristics of SLMs include:
1. Reduced parameter count: Typically ranging from a few hundred million to a few
billion parameters, compared to two-digit billions in larger models.
2. Lower memory footprint: Requiring, at most, a few gigabytes of memory rather than
tens or hundreds of gigabytes.
3. Faster inference time: Can generate responses in milliseconds to seconds on edge
devices.
4. Energy efficiency: Consuming less power, making them suitable for battery-powered
devices.
5. Privacy-preserving: Enabling on-device processing without sending data to cloud
servers.
6. Offline functionality: Operating without an internet connection.
SLMs achieve their compact size through various techniques such as knowledge distillation,
model pruning, and quantization. While they may not match the broad capabilities of larger
models, SLMs excel in specific tasks and domains, making them ideal for targeted applications
on edge devices.
We will generally consider SLMs —language models with fewer than 5 to 8 billion
parameters — quantized to 4 bits.
441
Examples of SLMs include compressed versions of models like Meta Llama, Microsoft PHI,
and Google Gemma. These models enable a wide range of natural language processing tasks
directly on edge devices, from text classification and sentiment analysis to question answering
and limited text generation.
For more information on SLMs, the paper, LLM Pruning and Distillation in Practice: The
Minitron Approach, provides an approach applying pruning and distillation to obtain SLMs
from LLMs. And, SMALL LANGUAGE MODELS: SURVEY, MEASUREMENTS, AND
INSIGHTS, presents a comprehensive survey and analysis of Small Language Models (SLMs),
which are language models with 100 million to 5 billion parameters designed for resource-
constrained devices.
442
Ollama
The primary and most user-friendly tool for running Small Language Models (SLMs) directly
on a Raspberry Pi, especially the Pi 5, is Ollama. It is an open-source framework that allows us
to install, manage, and run various SLMs (such as TinyLlama, smollm, Microsoft Phi, Google
Gemma, Meta Llama, MoonDream, LLaVa, among others) locally on our Raspberry Pi for
tasks such as text generation, image captioning, and translation.
[Link] and the Hugging Face Transformers library are well supported for running
SLMs on Raspberry Pi. [Link] is particularly efficient for running quantized models
natively. At the same time, Hugging Face Transformers offers a broader range of models
and tasks, which are best suited to smaller architectures due to hardware limitations.
Note that Ollama runs [Link] under the hood.
For production deployments on edge devices, a good option is Google AI Edge’s LiteRT-
LM, which is covered in the chapter “LiteRT-LM: Production-Ready LLM Inference at
the Edge”.
443
1. Local Model Execution: Ollama enables running LMs on personal computers or edge
devices such as the Raspberry Pi 5, eliminating the need for cloud-based API calls.
2. Ease of Use: It provides a simple command-line interface for downloading, running, and
managing different language models.
3. Model Variety: Ollama supports a variety of LLMs, including Phi, Gemma, Llama,
Mistral, and other open-source models.
4. Customization: Users can create and share custom models tailored to specific needs or
domains.
5. Lightweight: Designed to be efficient and run on consumer-grade hardware.
6. API Integration: Offers an API that allows integration with other applications and
services.
7. Privacy-Focused: By running models locally, it addresses privacy concerns associated
with sending data to external servers.
8. Cross-Platform: Available for macOS, Windows, and Linux systems (our case, here).
9. Active Development: Regularly updated with new features and model support.
10. Community-Driven: Benefits from community contributions and model sharing.
To learn more about what Ollama is and how it works under the hood, you should see this
short video from Matt Williams, one of the founders of Ollama:
[Link]
Matt has an entirely free course about Ollama that we recommend: [Link]
be/9KEUFe4KQAI?si=D_-q3CMbHiT-twuy
Installing Ollama
Let’s set up and activate a Virtual Environment for working with Ollama:
444
As a result, an API will run in the background on [Link]:11434. From now on, we can
run Ollama via the terminal. For starting, let’s verify the Ollama version, which will also tell
us that it is correctly installed:
ollama -v
On the Ollama Library page, we can find the models Ollama supports. For example, by filtering
by Most popular, we can see Meta Llama, Google Gemma, Microsoft Phi, LLaVa, etc.
445
Meta Llama 3.2 1B/3B
Let’s install and run our first small language model, Llama 3.2 1B (and 3B). The Meta Llama,
3.2 collections of multilingual large language models (LLMs), is a collection of pre-trained and
instruction-tuned generative models in 1B and 3B sizes (text in/text out). The Llama 3.2
instruction-tuned text-only models are optimized for multilingual dialogue use cases, including
agentic retrieval and summarization tasks.
The 1B and 3B models were pruned from the Llama 8B, and then logits from the 8B and 70B
models were used as token-level targets (token-level distillation). Knowledge distillation was
used to recover performance (they were trained with 9 trillion tokens). The 1B model has 1,24B,
quantized to integer (Q8_0), and the 3B, 3.12B parameters, with a Q4_0 quantization, which
ends with a size of 1.3 GB and 2GB, respectively. Its context window is 131,072 tokens.
446
Install and run the Model
Running the model with the command before, we should have the Ollama prompt available for
us to input a question and start chatting with the LLM model; for example,
>>> What is the capital of France?
Almost immediately, we get the correct answer:
The capital of France is Paris.
Using the option --verbose when calling the model will generate several statistics about its
performance (The model will be polling only the first time we run the command).
447
Each metric gives insights into how the model processes inputs and generates outputs. Here’s
a breakdown of what each metric means:
• Total Duration (2.620170326s): This is the complete time taken from the start of
the command to the completion of the response. It encompasses loading the model,
processing the input prompt, and generating the response.
• Load Duration (39.947908ms): This duration indicates the time to load the model
or necessary components into memory. If this value is minimal, it can suggest that the
model was preloaded or that only a minimal setup was required.
• Prompt Eval Count (32 tokens): The number of tokens in the input prompt. In
NLP, tokens are typically words or subwords, so this count includes all the tokens that
the model evaluated to understand and respond to the query.
• Prompt Eval Duration (1.644773s): This measures the model’s time to evaluate or
process the input prompt. It accounts for the bulk of the total duration, implying that
understanding the query and preparing a response is the most time-consuming part of
the process.
• Prompt Eval Rate (19.46 tokens/s): This rate indicates how quickly the model
processes tokens from the input prompt. It reflects the model’s speed in terms of natural
448
language comprehension.
• Eval Count (8 token(s)): This is the number of tokens in the model’s response, which
in this case was, “The capital of France is Paris.”
• Eval Duration (889.941ms): This is the time taken to generate the output based
on the evaluated input. It’s much shorter than the prompt evaluation, suggesting that
generating the response is less complex or computationally intensive than understanding
the prompt.
• Eval Rate (8.99 tokens/s): Similar to the prompt eval rate, this indicates the speed
at which the model generates output tokens. It’s a crucial metric for understanding the
model’s efficiency in output generation.
This detailed breakdown can help understand the computational demands and performance
characteristics of running SLMs like Llama on edge devices like the Raspberry Pi 5. It shows
that while prompt evaluation is more time-consuming, the actual generation of responses is
relatively quicker. This analysis is crucial for optimizing performance and diagnosing potential
bottlenecks in real-time applications.
Loading and running the 3B model, we can see the difference in performance for the same
prompt;
The eval rate is lower, 5.3 tokens/s versus 9 tokens/s with the smaller model.
When question about
>>> What is the distance between Paris and Santiago, Chile?
The 1B model answered 9,841 kilometers (6,093 miles), which is inaccurate, and the 3B
model answered 7,300 miles (11,700 km), which is close to the correct (11,642 km).
Let’s ask for the Paris’s coordinates:
449
>>> what is the latitude and longitude of Paris?
Google Gemma
Google Gemma, is a collection of lightweight, state-of-the-art open models built from the same
technology that powers our Gemini models. Today, the Gemma family has the Gemma3 and
Gemma3n models.
450
Gemma3
We can, for example, install gemma3:latest, using ollama run gemma3:latest. This model
has 4.3B parameters, with a context length of 8,192 and an embedding length of 2,560. A
typical quantization schema is the Q4_K_M. This model has vision capabilities. Besides the
4B, we can also install the 1B parameter model.
Install and run the Model
Running the model with the command before, we should have the Ollama prompt available for
us to input a question and start chatting with the LLM model; for example,
>>> What is the capital of France?
Almost immediately, we get the correct answer:
The capital of France is **Paris**. It's a global center for art, fashion,
gastronomy, and culture. � Do you want to know anything more about Paris?
And its statistics.
We can see that Gemma 3:4B has roughly the same performance as Lama 3.2:3B,
despite having more parameters.
451
Other examples:
A good and accurate answer (a little more verbose than the Llama answers).
An advantage of the Gemma3 models are their vision capability, for example, we can ask it to
caption an image:
452
This is a very accurate model, but it still has high latency, as we can see in the example above
(more than 3 minutes to caption the image).
The Gemma 3 1B size models are text only and don’t support image input.
Gemma 3n
Gemma3 is a powerful, efficient open-source model that runs locally on phones, tablets, and
laptops. The models are listed with parameter counts, such as E2B and E4B, that are lower than
the total number of parameters contained in the models. The E prefix indicates these models
can operate with a reduced set of Effective parameters. This reduced-parameter operation can
be achieved using the flexible parameter technology built into Gemma 3n models, which helps
them run efficiently on lower-resource devices.
The parameters in Gemma 3n models are divided into four main groups: text, visual, audio,
and per-layer embedding (PLE) parameters. In the standard execution of the E2B model, over
5 billion parameters are loaded. However, by using parameter skipping and PLE caching, this
model can be operated with an effective memory load of just under 2 billion (1.91B) parameters,
as illustrated below:
453
Figure 13: Immage from Google: Gemma 3n diagram of parameter usage
Once installed, (using ollama run gemma3n:e2b), inspecting the model we get:
• Architecture: gemma3n
• Parameters: 4.5B
• Quantization Q4_K_M
• Capabilities: completion only
When we run it, we can see that, despite having 4.5B parameters, it is faster than Llama3.3:3B.
454
Gemma 4
Gemma 4 models are designed to deliver frontier-level performance across all sizes. They are
well-suited for reasoning, agentic workflows, coding, and multimodal understanding. Gemma 4
is released under a commercially permissive Apache 2.0 license, which is a huge milestone for
open generative AI models at the edge.
E2B and E4B models: A new level of intelligence for mobile and IoT devices
Engineered from the ground up for maximum compute and memory efficiency, these models
operate with an effective 2- and 4-billion-parameter footprint during inference to preserve RAM
and battery life. In close collaboration with our Google Pixel team and mobile hardware leaders
such as Qualcomm Technologies and MediaTek, these multimodal models run completely offline
with near-zero latency on edge devices such as phones, Raspberry Pi, and NVIDIA Jetson Orin
Nano. Android developers can now prototype agentic flows in the AICore Developer Preview
today for forward compatibility with Gemini Nano 4.
Gemma 4 E2B and E4B are too heavy to be deployed on a Raspberry Pi 5 with
Ollama, but it works very well with Google AI Edge’s LiteRT-LM, which is covered
in the chapter “LiteRT-LM: Production-Ready LLM Inference at the Edge”.
455
Microsoft Phi3.5 3.8B
Let’s pull now the PHI3.5, a 3.8B lightweight state-of-the-art open model by Microsoft.
The model belongs to the Phi-3 model family and supports 128K token context length and
the languages: Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German,
Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish,
Swedish, Thai, Turkish, and Ukrainian.
The model size, in terms of bytes, will depend on the specific quantization format used. The
size can go from 2-bit quantization (q2_k) of 1.4 GB (higher performance/lower quality) to
16-bit quantization (fp-16) of 7.6 GB (lower performance/higher quality).
Let’s run the 4-bit quantization (Q4_0), which will need 2.2 GB of RAM, with an intermediary
trade-off regarding output quality and performance.
You can use run or pull to download the model. What happens is that Ollama
keeps note of the pulled models, and once the PHI3 does not exist, before running
it, Ollama pulls it.
456
of Europe's leading economies with its capital playing a key role
...
In this case, the answer was still longer than we expected, with an eval rate of 2.25 tokens/s,
more than double that of Gemma and Llama.
Choosing the most appropriate prompt is one of the most important skills to be
used with LLMs, no matter its size.
When we asked the same questions about distance and Latitude/Longitude, we did not get
a good answer for a distance of 13,507 kilometers (8,429 miles), but it was OK for
coordinates. Again, it could have been less verbose (more than 200 tokens for each answer).
We can use any model as an assistant since their speed is relatively decent, but in October 2024,
the Llama2:3B or Gemma 3 are better choices. Try other models, depending on your needs.
� Open LLM Leaderboard can give you an idea about the best models in size, benchmark,
license, etc.
457
The best model to use is the one fit for your specific necessity. Also, take into
consideration that this field evolves with new models everyday.
MoonDream
Moondream is an open-source visual language model that understands images using simple
text prompts. It’s fast and wildly capable. It is 3 to 4 times faster than the Gemma3, for
example.
458
Moondream can interpret images much like a human would, enabling tasks like: • Captioning
and describing images in detail • Visual question answering (VQA) such as “What is the
color of the shirt?” • Object detection and pointing out locations • Counting and contextual
reasoning about visual scenes
459
Llava-Phi-3
Another multimodal model is the LLaVA-Phi-3, a fine-tuned LLaVA model from Phi 3 Mini
4k. It has strong performance benchmarks that are on par with the original LLaVA (Large
Language and Vision Assistant) model.
In terms of latency, it is a pair with Gemma3 and much slower than MoonDream
The response took around 30s, with an eval rate of 3.93 tokens/s! Not bad!
But let us know to enter with an image as input. For that, let’s create a directory for working:
cd Documents/
mkdir OLLAMA
cd OLLAMA
460
Let’s download a 640x320 image from the internet, for example (Wikipedia: Paris, France):
Using FileZilla, for example, let’s upload the image to the OLLAMA folder on the Raspi-5 and
name it image_test_1.jpg. We should have the whole image path (we can use pwd to get
it).
/home/mjrovai/Documents/OLLAMA/image_test_1.jpg
If you use a desktop, you can copy the image path by right-clicking the image.
461
Let’s enter with this prompt:
The result was great, but the overall latency was significant; almost 4 minutes to perform the
inference.
462
Ollama Commands
Ollama has several commands available, which can be accessed using/, such as /bye to exit,
or /setto toggle section variables as verbose/quiet to see variables as token/s or latency,
and think/nothink to turn on/off the thinking mode in models that “reason” as Qwen and
Gemma.
463
Inspecting Model parameters
It is possible to know more about each model downloaded by Ollama, using the command
ollama show <model>. For example:
464
We’re looking at a 3.2B‑parameter Llama‑family instruction model, quantized for efficient local
use. Let’s walk through each field and what it implies in practice.
High‑level identity
Capabilities
– tools: the model has been trained/fine‑tuned to call tools/functions (e.g., JSON
function calls) when the host runtime exposes them, so it can decide “I should call
a function to get data” vs just answering from its own weights.
This doesn’t mean the tools are built into the model; it means it knows how to
format tool calls when the surrounding system supports them.
Quantization: Q4_K_M
465
– Much smaller VRAM/RAM requirement, so a 3.2B model can run comfortably
on a Pi 5 + NPU/GPU or a modest desktop GPU.
– Slight drop in raw quality vs the original FP16/FP32 model, but for many tasks
(chat, simple coding, small reasoning tasks) it’s still very usable.
– Ideal for edge or low‑power devices.
• llama.context_length: 131072
Context length is 131,072 tokens (~131k):
– This is a very large context window (comparable to “128k context” models).
– It can ingest long documents, multiple files, or extended conversations without
immediately forgetting earlier content.
– Under the hood, this usually implies some form of advanced attention scaling (e.g.,
RoPE scaling, attention optimizations, maybe sparse/flash attention), but the key
user effect is: we can feed a lot of text at once.
• llama.embedding_length: 3072
The hidden/embedding dimension is 3072:
– Each token is represented as a 3072‑dimensional vector internally.
– This roughly sets the “width” of the model: wider models can represent richer
patterns but cost more per token.
– Often, this is also the size of embeddings we’d extract if using the model as an
encoder (assuming our runtime supports that mode).
– We should read the exact license text bundled with the model before using it in
commercial products, especially at scale.
Parameters
The only parameters explicitly set in this Modelfile are three stop strings:
• <|start_header_id|>
466
• <|end_header_id|>
• <|eot_id|>
No temperature, top_k, top_p, etc., are defined in the Model file. Those, therefore, use
Ollama’s defaults for this model at runtime.
• temperature 0.3,
• top_p 0.9and
• top_k 40
htop
During the time that the model is running, we can inspect the resources:
467
All four CPUs run at almost 100% of their capacity, and the memory used with the model
loaded is 3.24GB. Exiting Ollama, the memory goes down to around 377MB (with no desktop).
It is also essential to monitor the temperature. When running the Raspberry with a desktop,
you can have the temperature shown on the taskbar:
If you are “headless”, the temperature can be monitored continuously (for example, every 2
seconds) with the command:
If you are doing nothing, the temperature is around 50°C for CPUs running at 1%. During
inference, with the CPUs at 100%, the temperature can rise to almost 70°C. This is OK and
means the active cooler is working, keeping the temperature below 80°C / 85°C (its limit).
So far, we have explored SLMs’ chat capability using the command line on a terminal. However,
we want to integrate those models into our projects, so Python seems to be the right path. The
good news is that Ollama has such a library.
The Ollama Python library simplifies interaction with advanced LLM models, enabling more
sophisticated responses and capabilities, besides providing the easiest way to integrate Python
3.8+ projects with Ollama.
For a better understanding of how to create apps using Ollama with Python, we can follow
Matt Williams’s videos, as the one below:
[Link]
Installation:
In the terminal (and inside the virtual environment), run the command:
468
For better using, we will need a text editor or an IDE to create a Python script. If we
run Raspberry Pi OS on a desktop, several options, such as Thonny and Geany, are already
installed by default (accessed via [Menu][Programming]). You can download other IDEs, such
as Visual Studio Code, from [Menu][Recommended Software]. When the window pops up,
go to [Programming], select your option, and press [Apply].
469
A good option to run Python scripts during development is to use the Jupyter Notebook:
We can use the Jupyter Notebook using SSH on our host computer. Run the command below,
changing the IP address for the Raspi:
On the terminal, you can see the local URL address to open the notebook:
We can access it from another computer by entering the Raspberry Pi’s IP address and the
provided token in a web browser (we should copy it from the terminal).
In our working directory (Documents/OLLAMA/) on the Raspberry Pi, we will create a new
Python 3 notebook.
470
Let’s enter with a very simple script to verify the installed models:
import ollama
models = [Link]()
for model in models['models']:
print(model['model'])
gemma3:4b
riven/smolvlm:latest
gemma3n:e2b
moondream:latest
llama3.2:3b
Same as on the terminal using ollama show llama3.2: 3b, we can get model information
with the Python command [Link]('llama3.2:3b'):
info = [Link]('llama3.2:3b')
471
print("Supported Capabilities:", getattr(info, 'capabilities', None))
print("Date Modified (Local) :", getattr(info, 'modified_at', None))
print("License Short :", (str(getattr(info, 'license', None)).split('\n')[0]))
for key in [
'[Link]', '[Link]',
'llama.context_length', 'llama.embedding_length', 'llama.block_count'
]:
print(f" - {key}: {[Link](key)}")
As a result, we have:
Model/Format : None
Parameter Size : 3.2B
Quantization Level : Q4_K_M
Family : llama
Supported Capabilities : ['completion', 'tools']
Date Modified (Local) : 2025-09-15 15:39:48.153510-03:00
License Short : LLAMA 3.2 COMMUNITY LICENSE AGREEMENT
Key Architecture Details :
- [Link]: llama
- [Link]: Instruct
- llama.context_length: 131072
- llama.embedding_length: 3072
- llama.block_count: 28
• More layers typically improve the model’s ability to perform multi‑step reasoning and to
build hierarchical representations.
472
• For 3.2B parameters, 28 layers with width 3072 is a sensible trade‑off (moderately deep,
not extremely wide).
• Compared with a 7B model (~32–40 layers, 4k+ width), this will be less capable on very
complex reasoning or coding, but noticeably more capable than 1B‑class models.
Model/Format: None : This usually means the frontend (e.g., our UI or manager) didn’t
detect a specific “format template” (like “chat”, “instruct”, “plain”) beyond knowing it’s
Llama:
• The underlying file is probably something like a GGUF or similar binary weight file.
• Prompt formatting will depend on our runtime’s default “llama‑instruct” template; if
nothing is configured, we may need to wrap instructions manually (system/user style or
clear “You are an assistant…” preamble).
Ollama Generate
Let’s repeat one of the questions that we did before, but now using [Link]() from
the Ollama Python library. This API generates a response for the given prompt using the
provided model. This is a streaming endpoint, so there will be a series of responses. The final
response object will include statistics and additional data from the request.
response = [Link](
model="llama3.2:3b",
prompt="What is the capital of Brazil"
)
print(response['response'])
If you are running the code as a Python script, you should save it as, for example,
test_ollama.py. You can run it in the IDE or run it directly in the terminal. Also,
remember always to call the model and define it when running a stand-alone script.
python test_ollama.py
473
And run a new query:
MODEL="llama3.2:3b"
response = simple_query ("What is the capital of Peru?")
response
Let’s print the full response now. As a result, we will have the model response in a JSON
format:
import json
print([Link](response.__dict__, indent=2))
{
"model": "llama3.2:3b",
"created_at": "2026-02-26T18:58:30.730414614Z",
"done": true,
"done_reason": "stop",
"total_duration": 2307368739,
"load_duration": 343767495,
"prompt_eval_count": 32,
"prompt_eval_duration": 722764406,
"eval_count": 8,
"eval_duration": 1229567370,
"response": "The capital of Peru is Lima.",
"thinking": null,
"context": [
128006,
9125,
...
13
],
"logprobs": null
}
• response: the main output text generated by the model in response to our prompt.
474
– The capital of Peru is Lima.
• context: the token IDs representing the input and context used by the model. Tokens
are numerical representations of text that the language model uses to process it.
But what we want is the plain ‘response’ and, perhaps for analysis, the total duration of the
inference, so let’s change the code to extract it from the dictionary:
print(f"\n{response['response']}")
print(f"\nTotal Duration: {(response['total_duration']/1e9):.2f} seconds")
Now, we got:
print(f"eval_count: {response['eval_count']}")
print(f"eval_duration: {(response['eval_duration']/1e9):.2f} s")
print(f"eval_rate: {response['eval_count']/(response['eval_duration']/1e9):.2f} tokens/s")
eval_count: 8
eval_duration: 1.23 s
eval_rate: 6.51 tokens/s
475
Streaming with [Link]()
To stream the output from [Link]() in Python, set stream=True and iterate over
the generator to print each chunk as it’s produced. This enables real-time response streaming,
similar to chat models.
stream = [Link](
model='llama3.2:3b',
prompt='Tell me an interesting fact about Brazil',
stream=True
)
This approach is ideal for long or complex generations, making the user experience feel faster
and more interactive.
System Prompt
Add a system parameter to set overall instructions or behavior for the model (useful for role
assignment and tone control):
response = [Link](model='llama3.2:3b',
prompt='Tell about industry',
system='You are an expert on Brazil.',
stream=False)
print(response['response'])
476
Temperature and Sampling
Control creativity/randomness via temperature, and customize output style with extra settings
like top_p and num_ctx (context window size)
response = [Link](
model='llama3.2:3b',
prompt='Why the sky is blue?',
options={'temperature':0.1},
stream=True)
for chunk in response:
print(chunk['response'], end='', flush=True)
477
response = [Link](
messages=[
{"role": "user",
"content": "Poetically describe Paris in one short sentence"},
],
model='llama3.2:3b',
options={"temperature": 1.0} # Set. temp. to 1.0 for more creativity
)
print(response['message']['content'])
More Options
Besides temperature, we can control a lot via the options={…} parameter in [Link].
Key ones:
Sampling/style
Length/context
478
Putting it together in your code:
response = [Link](
model='llama3.2:3b',
prompt='Why is the sky blue?',
options={
'temperature': 0.1,
'top_p': 0.9,
'top_k': 40,
'repeat_penalty': 1.1,
'num_predict': 50, # Limit the answer
'num_ctx': 4096,
'seed': 42,
},
stream=True,
)
for chunk in response:
print(chunk['response'], end='', flush=True)
The sky appears blue because of a phenomenon called Rayleigh scattering, named after the
British physicist Lord Rayleigh, who first described it in the late 19th century.
top_k and top_p both control how random and diverse the model’s next tokens can be, but
they do it in different ways.
top_k: fixed shortlist size
• At each step, the model ranks all possible next tokens by probability.
• top_k = K means: keep only the K most likely tokens, set all others to probability
0, then sample from those K.
• Effects:
479
– Small k (1–20) → very focused, deterministic, “on‑rails”; less creative, fewer weird
tokens.
– Large k (50–200+) → more variety and creativity; higher chance of unusual or
off‑topic words.
• Extreme:
– top_k = 1 → greedy decoding: always pick the single most likely token, almost no
randomness.
• Starting from the top, you add tokens until their cumulative probability � p. Then
you sample only from that set.
• top_p = 0.9 means: “consider just enough top tokens to cover 90% of the probability
mass”.
• Effects:
– In confident situations (one token is clearly best), the shortlist may be very small
→ behavior similar to low k.
– In uncertain situations (probabilities spread out), more tokens enter the shortlist
→ more exploration.
• Typical:
– top_p � 0.9–0.95 is a common sweet spot for natural but not too wild text.
• Higher top_k / top_p → more diverse, more creative, but also more risk of nonsense.
We can use them together—for example, top_k=40, top_p=0.9—where top_k gives a hard
cap and top_p then trims that set by probability mass.
[Link]()
Another way to get our response is to use [Link](), which generates the next message
in a chat with a provided model. This is a streaming endpoint, so a series of responses will
occur. Streaming can be disabled using "stream": false. The final response object will also
include statistics and additional data from the request.
480
PROMPT_1 = 'What is the capital of France?'
In the above code, we are running two queries, and the second prompt considers the result of
the first one.
Here is how the model responded:
481
The above code works with two prompts. Let’s include a conversation variable to really provide
the chat with a memory:
# Question
prompt = "What is the capital of Brazil"
print(chat_with_memory(prompt))
Image Description:
As we did with the visual models and the command line to analyze an image, the same
can be done here with Python. Let’s use the same image of Paris, but now with the
[Link]():
482
MODEL = 'llava-phi3:3.8b'
PROMPT = "Describe this picture"
response = [Link](
model=MODEL,
prompt=PROMPT,
images= [img]
)
print(f"\n{response['response']}")
print(f"\n [INFO] Total Duration: {(res['total_duration']/1e9):.2f} seconds")
This image captures the iconic cityscape of Paris, France. The vantage point
is high, providing a panoramic view of the Seine River that meanders through
the heart of the city. Several bridges arch gracefully over the river,
connecting different parts of the city. The Eiffel Tower, an iron lattice
structure with a pointed top and two antennas on its summit, stands tall in the
background, piercing the sky. It is painted in a light gray color, contrasting
against the blue sky speckled with white clouds.
The buildings that line the river are predominantly white or beige, their uniform
color palette broken occasionally by red roofs peeking through. The Seine River
itself appears calm and wide, reflecting the city's architectural beauty in its
surface. On either side of the river, trees add a touch of green to the urban
landscape.
The image is taken from an elevated perspective, looking down on the city. This
viewpoint allows for a comprehensive view of Paris's beautiful architecture and
layout. The relative positions of the buildings, bridges, and other structures
create a harmonious composition that showcases the city's charm.
In summary, this image presents a serene day in Paris, with its architectural
marvels - from the Eiffel Tower to the river-side buildings - all bathed in soft
colors under a clear sky.
The model took about 4 minutes (256.45 s) to return with a detailed image description.
483
Let’s capture an image from the Raspberry Pi camera and get the description, now using the
MoonDream model:
import time
import numpy as np
import [Link] as plt
from picamera2 import Picamera2
from PIL import Image
def capture_image(image_path):
# Initialize camera
picam2 = Picamera2() # default is index 0
# Capture image
picam2.capture_file(image_path)
print("Image captured: "+image_path)
# Stop camera
[Link]()
[Link]()
Using the above code, we can capture an image, which can be displayed with:
def show_image(image_path):
img = [Link](image_path)
484
def image_description(img_path, model):
with open(img_path, 'rb') as file:
response = [Link](
model=model,
messages=[
{
'role': 'user',
'content': '''return the description of the image''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,
}
)
return response
Now, let’s put all togheter and capture an image from the camera:
IMG_PATH = "/home/mjrovai/Documents/OLLAMA/SST/capt_image.jpg"
MODEL = "moondream:latest"
apture_image(IMG_PATH)
show_image(IMG_PATH)
response = image_description(IMG_PATH, MODEL)
caption = response['message']['content']
print ("\n==> AI Response:", caption)
print(f"\n[INFO] ==> Total Duration: {
(response['total_duration']/1e9):.2f} seconds")
485
We got the description:
The image features a green table with various items on it. A white mug adorned with
black faces is prominently displayed, and there are several other mugs scattered around
the table as well. In addition to the mugs, there's also a microphone placed near them,
suggesting that this might be an office or workspace setting where someone could enjoy
their coffee while recording podcasts or audio content.
486
connected to a computer for work purposes. A mouse and a cell phone are also present on
the table, further emphasizing the technology-oriented nature of this scene.
We can now change the image_description function to ask “Who are the faces in the mug?”.
The answer:
==> AI Response:
The mug has a picture of the Beatles on it.
One alternative to running an SLM in Python using Ollama is to call the API directly. Let’s
explore some advantages and disadvantages of both methods.
Python Library:
response = [Link](
model=MODEL,
prompt=QUERY)
result = response['response']
import requests
import json
# Configuration
OLLAMA_URL = "[Link]
MODEL = MODEL
response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": QUERY,
487
"stream": False
}
)
response = [Link](
f"{OLLAMA_URL}/generate",
json={"model": MODEL,
"prompt": query,
"stream": False}
)
result = [Link]().get("response", "")
One clear advantage of the Python library is that it handles URL construction,
request formatting, and response parsing.
Error Handling
Python Library:
Connection Management
Python Library:
488
Features & Functionality
Python Library:
• Clean access to all Ollama features (generate, chat, embeddings, list models, pull, etc.)
• Streaming is simple: for chunk in [Link](..., stream=True)
• Type hints and better IDE support
Dependencies
Python Library:
Advanced Features
Python Library:
# Streaming
for chunk in [Link](model=MODEL, prompt=query, stream=True):
print(chunk['response'], end='')
# Chat history
response = [Link](
model=MODEL,
messages=[
489
{'role': 'user', 'content': 'Hello!'},
{'role': 'assistant', 'content': 'Hi there!'},
{'role': 'user', 'content': 'How are you?'}
]
)
# List models
models = [Link]()
# Pull models
[Link]('llama3.2:3b')
490
Performance Difference
Minimal difference in practice! The Python library uses httpx which is comparable to requests.
Both make the same underlying HTTP calls to Ollama.
Bottom line: For most use cases, the Python library is the better choice due to its
simplicity and built-in features. Use direct API calls only when you need specific
control or have constraints that prevent adding the dependency.
Going Further
The small LLM models tested worked well at the edge, both with text and with images, but, of
course, the last one had high latency. A combination of specific, dedicated models can lead
to better results; for example, in real cases, an Object Detection model (such as YOLO) can
provide a general description and count of objects in an image, which, once passed to an LLM,
can help extract essential insights and actions.
According to Avi Baum, CTO at Hailo,
In the vast landscape of artificial intelligence (AI), one of the most intriguing
journeys has been the evolution of AI on the edge. This journey has taken us from
classic machine vision to the realms of discriminative AI, enhancive AI, and now,
the groundbreaking frontier of generative AI. Each step has brought us closer to a
future where intelligent systems seamlessly integrate with our daily lives, offering
an immersive experience of not just perception but also creation at the palm of our
hand.
491
Conclusion
This chapter has demonstrated how a Raspberry Pi 5 can be transformed into a potent AI
hub capable of running large language models (LLMs) for real-time, on-site data analysis and
insights using Ollama and Python. The Raspberry Pi’s versatility and power, coupled with the
capabilities of lightweight LLMs like Llama 3.2 and MoonDream, make it an excellent platform
for edge computing applications.
The potential of running LLMs on the edge extends far beyond simple data processing, as in
this lab’s examples. Here are some innovative suggestions for using this project:
1. Smart Home Automation:
• Integrate SLMs to interpret voice commands or analyze sensor data for intelligent home
automation. This could include real-time monitoring and control of home devices, security
systems, and energy management, all processed locally without relying on cloud services.
• Deploy SLMs on Raspberry Pi in remote or mobile setups for real-time data collection
and analysis. This can be used in agriculture to monitor crop health, in environmental
studies for wildlife tracking, or in disaster response for situational awareness and resource
management.
3. Educational Tools:
492
• Create interactive educational tools that leverage SLMs to provide instant feedback,
language translation, and tutoring. This can be particularly useful in developing regions
with limited access to advanced technology and internet connectivity.
4. Healthcare Applications:
• Use SLMs for medical diagnostics and patient monitoring. They can provide real-time
analysis of symptoms and suggest potential treatments. This can be integrated into
telemedicine platforms or portable health devices.
6. Industrial IoT:
• Integrate SLMs into industrial IoT systems for predictive maintenance, quality control,
and process optimization. The Raspberry Pi can serve as a localized data processing
unit, reducing latency and improving the reliability of automated systems.
7. Autonomous Vehicles:
• Use SLMs to process sensory data from autonomous vehicles, enabling real-time decision-
making and navigation. This can be applied to drones, robots, and self-driving cars for
enhanced autonomy and safety.
• Implement SLMs to provide interactive and informative cultural heritage sites and
museum guides. Visitors can use these systems to get real-time information and insights,
enhancing their experience without internet connectivity.
• Use SLMs to analyze and generate creative content, such as music, art, and literature. This
can foster innovative projects in the creative industries and allow for unique interactive
experiences in exhibitions and performances.
493
Resources
494
SLM: Basic Optimization Techniques
Figure 14: DALL·E prompt - I am writing a tutorial using Raspberry Pi. I am talking about
Optimization Techniques, such as Function Calling and RAG, using Ollama and
SLMs. I want a landscape-format image for the tutorial cover (without title). Should
be a cartoon styled in 5Os
Introduction
Large Language Models (LLMs) have revolutionized natural language processing, but their
deployment and optimization come with unique challenges. One significant issue is the
tendency for LLMs (and more, the SLMs) to generate plausible-sounding but factually incorrect
495
information, a phenomenon known as hallucination. This occurs when models produce
content that appears coherent but lacks grounding in truth or real-world facts.
Other challenges include the immense computational resources required to train and run these
models, the difficulty of maintaining up-to-date knowledge within the model, and the need
for domain-specific adaptations. Privacy concerns also arise when handling sensitive data
during training or inference. Additionally, ensuring consistent performance across diverse tasks
and maintaining ethical use of these powerful tools present ongoing challenges. Addressing
these issues is crucial for the effective and responsible deployment of LLMs in real-world
applications.
The fundamental and more common techniques for enhancing LLM (and SLM) performance
and efficiency are Function (or Tool) Calling, Prompt engineering, Fine-tuning, and Retrieval-
Augmented Generation (RAG).
• Function (Tool) calling allows models to perform actions beyond generating text. By
integrating with external functions or APIs, SLMs can access real-time data, automate
tasks, and perform precise calculations—addressing the reliability issues that arise from
the model’s limitations in mathematical operations.
• Prompt engineering is at the forefront of LLM optimization. By carefully crafting
input prompts, we can guide models to produce more accurate and relevant outputs. This
technique involves structuring queries that leverage the model’s pre-trained knowledge
and capabilities, often incorporating examples or specific instructions to shape the desired
response.
• Retrieval-Augmented Generation (RAG) represents a powerful approach that’s ideal
for resource-constrained edge devices. This method combines the knowledge embedded
in pre-trained models with the ability to access external, up-to-date information without
requiring fine-tuning. By retrieving relevant data from a local knowledge base, RAG
significantly enhances accuracy and reduces hallucinations—all without the computational
overhead of model retraining.
• Fine-tuning, while more resource-intensive, offers a way to specialize LLMs for specific
domains or tasks. This process involves further training the model on carefully curated
datasets, allowing it to adapt its vast general knowledge to particular applications. Fine-
tuning can lead to substantial performance improvements, especially in specialized fields
or for unique use cases.
In this chapter, we’ll start focusing on two techniques that are particularly well-suited for edge
devices like the Raspberry Pi: Function Calling and RAG.
We will learn more in detail about optimization techniques for SLMs, in the chapter:
Advancing EdgeAI: Beyond Basic SLMs
496
Function Calling Introduction
So far, we can see that, with the model’s (“response”) answer to a variable, we can efficiently
work with it and integrate it into real-world projects. However, a big problem is that the model
can respond differently to the same prompt. Let’s say, as in the last examples, that we want
the model’s response to be only the name of a given country’s capital and its coordinates,
nothing more, even with very verbose models such as the Microsoft Phi. We can use the Ollama
function's calling to guarantee the same answers, which is perfectly compatible with the
OpenAI API.
In modern artificial intelligence, function calling with Large Language Models (LLMs) allows
these models to perform actions beyond generating text. By integrating with external functions
or APIs, LLMs can access real-time data, automate tasks, and interact with various systems.
For instance, instead of merely responding to a weather query, an LLM can call a weather API
to fetch the current conditions and provide accurate, up-to-date information. This capability
enhances the relevance and accuracy of the model’s responses, making it a powerful tool for
driving workflows and automating processes, thereby transforming it into an active participant
in real-world applications.
For more details about Function Calling, please see this video made by Marvin Prison:
[Link]
And on this link: HuggingFace Function Calling
123456*123456
The result would be: 15,241,383,936. No issues on it, but let’s ask a SLM to do the same
simple task:
import ollama
response = [Link](
model='llama3.2:3B',
messages=[{
"role": "user",
497
"content": "What is 123456 multiplied by 123456? Only give me the answer"
}],
options={"temperature": 0}
)
• LLMs work by predicting the next most likely token based on patterns in training data
• They don’t perform actual arithmetic operations
• They’re essentially “guessing” what a plausible answer looks like
2. Tokenization Issues
Numbers are broken into tokens in ways that don’t align with mathematical operations:
This makes it nearly impossible for the model to “see” the actual numbers properly for
computation.
The LLM will likely give something that “looks” like a big number but is mathematically
incorrect.
498
The Solution: Function Calling / Tool Use
Setting temperature=0 makes the output deterministic (same input → same output), but it
doesn’t make it correct. The model will confidently give the same wrong answer every time.
Best Practices
499
Bottom line: Use LLMs for natural language understanding and intent classifica-
tion, but delegate actual computations to proper tools/functions. This is the core
principle behind tool use and function calling in modern LLM applications!
multiply_tool = {
"type": "function",
"function": {
"name": "multiply_numbers",
"description": "Multiply two numbers together",
"parameters": {
"type": "object",
"required": ["a", "b"],
"properties": {
"a": {"type": "number", "description": "First number"},
"b": {"type": "number", "description": "Second number"}
}
}
}
}
Now, let’s create a function to handle the user query and calling for the tool when needed.
def answer_query(QUERY):
response = [Link](
'llama3.2:3B',
messages=[{"role": "user", "content": QUERY}],
500
tools=[multiply_tool]
)
Result: 15,241,383,936.00
Great! And now, can I use the same code to answer general questions? Let’s test it:
Result: 1,000,000.00
The result is wrong. So, the above approach works fine for using the tool, but to answer it
correctly (even without a tool), we should implement an “agentic approach”, which is a subject
for later (See the Chapter: Advancing EdgeAI: Beyond Basic SLMs)
Suppose we want an SLM to return the distance in km from the capital city of the country
specified by the user to the user’s current location. We can see that the first is not so simple:
it is not always enough to enter only the country’s name; the SLM can also give us a different
(and incorrect) answer every time.
501
OK, for trying to mitigate it, let’s create an app where the user enters a country’s name and
gets, as an output, the distance in km from the capital city of such a country and the app’s
location (for simplicity, we will use Santiago, Chile, as the app location).
Once the user enters a country name, the model should return the capital city’s name (as
a string) and its latitude and longitude (as floats). Using those coordinates, we can use
a simple Python library (haversine) to compute the great‑circle distance between the two
latitude/longitude points.
502
The idea of this project is to demonstrate a combination of language model interaction (IA)
and geospatial calculations using the Haversine formula (traditional computing).
First, let us install the Haversine library:
Now, we should create a Python script designed to interact with our model (LLM) to determine
the coordinates of a country’s capital city and calculate the distance from Santiago de Chile to
that capital.
Let’s go over the code:
Importing Libraries
import time
from haversine import haversine
from ollama import chat
• MODEL: Specifies the model being used, which is, in this example, the Lhama3.2.
• mylat and mylon: Coordinates of Santiago de Chile, used as the starting point for the
distance calculation.
This is the real Python function that Ollama will be allowed to call. It performs the following
steps: • Takes latitude, longitude, and city name as input arguments. • Uses the haversine
library to calculate the distance from Santiago to the target city. • Returns a JSON‑like
dictionary containing the computed distance and a human‑readable text summary.
503
In Ollama’s terminology, this is a tool — a callable external function that the LLM
may invoke automatically
tools = [
{
"type": "function",
"function": {
"name": "calc_distance",
"description": "Calculates the distance from Santiago, Chile to a \
given city's coordinates.",
"parameters": {
"type": "object",
"properties": {
"lat": {"type": "number", "description": "Latitude of the city"},
"lon": {"type": "number", "description": "Longitude of the city"},
"city": {"type": "string", "description": "Name of the city"}
},
"required": ["lat", "lon", "city"]
}
}
}
]
This JSON object describes the metadata and input schema of the tool so that the LLM knows:
• Name: which function to call. • Description: what purpose it serves. • Parameters: input
argument types and their descriptions. This schema mirrors the OpenAI function‑calling format
and is fully supported in Ollama � 0.4 .
Defining tools this way allows Ollama to validate arguments before sending a call
request back.
response = chat(
model=MODEL,
messages=[{
"role": "user",
"content": f"Find the decimal latitude and longitude of the capital of \
I am running a few minutes late; my previous meeting is running over. running a few\
minutes late; my previous meeting is running over.
{country},"
504
" then use the calc_distance tool to determine how far it is from \
Santiago de Chile."
}],
tools=tools
)
• The chat() function is called with the chosen model, a message, and the tools list.
• The prompt instructs the model first to identify the capital and its coordinates, and then
invoke the tool (calc_distance) with those values.
• Ollama returns a structured response that may include a tool_calls section, indicating
which tool to execute.
city = raw_args['city']
lat = float(raw_args['lat'])
lon = float(raw_args['lon'])
505
In this case:
NOTE: Sometimes the model returns parameter names that differ from what your function
expects. Specifically, Ollama occasionally returns argument objects like:
{"lat1": -33.33, "lon1": -70.51, "lat2": 48.8566, "lon2": 2.3522, "city":
"Paris"}
Instead of the schema-defined keys (lat, lon, city).
This happens because some LLMs (such as Llama 3.2 and Qwen 3) attempt to be “helpful”
by naming coordinates explicitly—lat1/lon1 for origin and lat2/lon2 for destination—even
when the schema only defines lat/lon.
To handle this, optionally a mapping-correction step can be added after decoding the tocall
arguments.
# Convert numbers
args["lat"] = float(args["lat"])
args["lon"] = float(args["lon"])
result = calc_distance(**args)
print(result["message"])
506
Timing and Diagnostic Output
This records how long the operation took from prompt submission to tool execution, useful for
benchmarking response performance.
Example Usage
If we enter different countries, for example, France, Colombia, and the United States, We can
note that we always receive the same structured information:
ask_and_measure("France")
ask_and_measure("Colombia")
ask_and_measure("United States")
If you run the code as a script, the result will be printed on the terminal:
507
The complete script can be found at: func_call_dist_calc.py and on the 10-Ollama_Function_Calling
notebook.
The models that will run with the described approach are the ones that can handle tools. For
example, Gemma 3 and 3n will not work.
An alternative is to use the Pydantic library to serialize the schema using model_json_schema().
Using the Pydantic library, models as Gemma can also be used, as explored in the:
20-Ollama_Function_Calling_Pydantic notebook
Adding images
Now it is time to wrap up everything so far! Let’s modify the script using Pydantic so that
instead of entering the country name (as a text), the user enters an image, and the application
(based on SLM) returns the city in the image and its geographic location. With that data, we
can calculate the distance as before.
508
For simplicity, we will implement this new code in two steps. First, the LLM will analyze the
image and create a description (text). This text will be passed on to another instance, where
the model will extract the information needed to pass along.
import time
from haversine import haversine
from ollama import chat
from pydantic import BaseModel, Field
We can see the image if you run the code on the Jupyter Notebook. For that, we also need to
import:
MODEL = 'gemma3:4b'
mylat = -33.33
mylon = -70.51
We can download a new image, for example, Machu Picchu from Wikipedia. On the Notebook
we can see it:
509
[Link](figsize=(8, 8))
[Link](img)
[Link]('off')
#[Link]("Image")
[Link]()
Now, let’s define a function that will receive the image and will return the decimal
latitude and decimal longitude of the city in the image, its name, and what
country it is located
def image_description(img_path):
with open(img_path, 'rb') as file:
response = chat(
model=MODEL,
messages=[
{
'role': 'user',
'content': '''return the decimal latitude and decimal longitude
of the city in the image, its name, and
what country it is located''',
'images': [[Link]()],
},
],
options = {
'temperature': 0,
510
}
)
#print(response['message']['content'])
return response['message']['content']
We can print the entire response for debug purposes. In this case, we can get
something as:
'{\n "city": "Machu Picchu",\n "country": "Peru",\n "lat": -
13.1631,\n "lon": -72.5450\n}\n'
Let’s define a Pydantic model (CityCoord) that describes the expected structure of the SLM’s
response. It expects four fields: country, city (city name), lat (latitude), and lon (longitude).
class CityCoord(BaseModel):
city: str = Field(..., description="Name of the city in the image")
country: str = Field(..., description="Name of the country where the city in the\ image i
lat: float = Field(..., description="Decimal Latitude of the city in the image")
lon: float = Field(..., description="Decimal Longitude of the city in the image")
The image description generated for the function will be passed as a prompt for the model
again.
response = chat(
model=MODEL,
messages=[{
"role": "user",
"content": image_description # image_description from previous model's run
}],
format=CityCoord.model_json_schema(), # Structured JSON format
options={"temperature": 0}
)
resp = CityCoord.model_validate_json([Link])
And so, we can calculate and print the distance, using haversine():
511
And we will get:
The image shows Machu Picchu, with lat:-13.16 and long: -72.55, located in Peru and about 2,2
Enter with the Machu Picchu image full patch as an argument. We will get the same previous
result.
The app is working fine with both models, with the Gemma being faster.
512
How about Paris?
Of course, there are many ways to optimize the code used here. Still, the idea is to explore
the considerable potential of function calling with SLMs at the edge, allowing those models to
integrate with external functions or APIs. Going beyond text generation, SLMs can access
real-time data, automate tasks, and interact with various systems.
In a basic interaction between a user and a language model, the user asks a question, which is
sent to the model as a prompt. The model generates a response based solely on its pre-trained
knowledge.
513
In a RAG process, there’s an additional step between the user’s question and the model’s
response. The user’s question triggers a retrieval process from a knowledge base.
Here are the steps to implement a basic Retrieval Augmented Generation (RAG):
• Determine the type of documents you’ll be using: The best types are documents
from which we can get clean and unobscured text. PDFs can be problematic because
they are designed for printing, not for extracting sensible text. To work with PDFs, we
should get the source document or use tools to handle it.
• Chunk the text: We can’t store the text as one long stream because of context size
limitations and the potential for confusion. Chunking involves splitting the text into
smaller pieces. Chunk text has many ways, such as character count, tokens, words,
paragraphs, or sections. It is also possible to overlap chunks.
514
• Create embeddings: Embeddings are numerical representations of text that capture
semantic meaning. We create embeddings by passing each chunk of text through a
particular embedding model. The model outputs a vector, the length of which depends
on the embedding model used. We should pull one (or more) embedding models from
Ollama, to perform this task. Here are some examples of embedding models available at
Ollama.
Generally, larger embedding sizes capture more nuanced information about the
input. Still, they also require more computational resources to process, and a
higher number of parameters should increase the latency (but also the quality
of the response).
• Store the chunks and embeddings in a vector database: We will need a way to
efficiently find the most relevant chunks of text for a given prompt, which is where a vector
database comes in. We will use Chromadb, an AI-native open-source vector database,
which simplifies building RAGs by creating knowledge, facts, and skills pluggable for
LLMs. Both the embedding and the source text for each chunk are stored.
• Build the prompt: When we have a question, we create an embedding and query the
vector database for the most similar chunks. Then, we select the top few results and
include their text in the prompt.
The goal of RAG is to provide the model with the most relevant information from our documents,
allowing it to generate more accurate and informative responses. So, let’s implement a simple
example of an SLM incorporating a particular set of facts about bees (“Bee Facts”).
Inside the ollama env, enter the command in the terminal for Chromadb instalation:
515
cd Documents/OLLAMA/
mkdir RAG-simple-bee
cd RAG-simple-bee/
import ollama
import chromadb
import time
EMB_MODEL = "nomic-embed-text"
MODEL = 'llama3.2:3B'
Initially, a knowledge base about bee facts should be created. This involves collecting relevant
documents and converting them into vector embeddings. These embeddings are then stored in
a vector database, allowing for efficient similarity searches later. Enter with the “document,” a
base of “bee facts” as a list:
documents = [
"Bee-keeping, also known as apiculture, involves the maintenance of bee \
colonies, typically in hives, by humans.",
"The most commonly kept species of bees is the European honey bee (Apis \
mellifera).",
...
516
We do not need to “chunk” the document here because we will use each element of
the list as a chunk.
Now, we will create our vector embedding database bee_facts and store the document in it:
client = [Link]()
collection = client.create_collection(name="bee_facts")
Now that we have our “Knowledge Base” created, we can start making queries, retrieving data
from it:
User Query: The process begins when a user asks a question, such as “How many bees are in
a colony? Who lays eggs, and how much? How about common pests and diseases?”
517
prompt = "How many bees are in a colony? Who lays eggs and how much? How about\
common pests and diseases?"
Query Embedding: The user’s question is converted into a vector embedding using the
same embedding model used for the knowledge base.
response = [Link](
prompt=prompt,
model=EMB_MODEL
)
Relevant Document Retrieval: The system searches the knowledge base using the query
embedding to find the most relevant documents (in this case, the 5 more probable). This is done
using a similarity search, which compares the query embedding to the document embeddings
in the database.
results = [Link](
query_embeddings=[response["embedding"]],
n_results=5
)
data = results['documents']
Prompt Augmentation: The retrieved relevant information is combined with the original
user query to create an augmented prompt. This prompt now contains the user’s question and
pertinent facts from the knowledge base.
Answer Generation: The augmented prompt is then fed into a language model, in this case,
the llama3.2:3b model. The model uses this enriched context to generate a comprehensive
answer. Parameters like temperature, top_k, and top_p are set to control the randomness and
quality of the generated response.
output = [Link](
model=MODEL,
prompt=f"Using this data: {data}. Respond to this prompt: {prompt}",
options={
"temperature": 0.0,
"top_k":10,
"top_p":0.5 }
)
518
Response Delivery: Finally, the system returns the generated answer to the user.
print(output['response'])
Based on the provided data, here are the answers to your questions:
results = [Link](
query_embeddings=[response["embedding"]],
n_results=n_results
)
data = results['documents']
519
print(output['response'])
Yes, bees are found in Brazil. According to the data, Brazil has more than 300
different bee species, and indigenous people in Brazil used bees for medicine and
food purposes. Additionally, reports from 1577 mention three native bees used by
indigenous people in Brazil.
[INFO] ==> The code for model: llama3.2:3b, took 22.7s to generate the answer.
By the way, if the model used supports multiple languages, we can use it (for example,
Portuguese), even if the dataset was created in English:
Sim, existem abelhas no Brasil! De acordo com o relato de Hans Staden, há três
espécies de abelhas nativas do Brasil que foram mencionadas: mandaçaia (Melipona
quadrifasciata), mandaguari (Scaptotrigona postica) e jataí-amarela (Tetragonisca
angustula). Além disso, o Brasil é conhecido por ter mais de 300 espécies
diferentes de abelhas, a maioria das quais não é agressiva e não põe veneno.
[INFO] ==> The code for model: llama3.2:3b, took 54.6s to generate the answer.
In the Chapter Advancing EdgeAI: Beyond Basic SLMs, we will learn how to
implement a Naive RAG System
520
Conclusion
Throughout this chapter, we’ve explored two fundamental optimization techniques that signifi-
cantly enhance the capabilities of Small Language Models (SLMs) running on edge devices like
the Raspberry Pi: Function Calling and Retrieval-Augmented Generation (RAG).
We began by addressing a critical limitation of language models—their inability to perform
accurate calculations. By implementing function calling, we demonstrated how to transform
SLMs from text generators into actionable agents that can interact with external tools and APIs.
Whether extracting structured data like geographic coordinates, fetching real-time weather
information, or performing precise mathematical operations, function calling bridges the gap
between natural language understanding and deterministic computation. The key principle
remains:
Use SLMs for intent classification and understanding, while delegating specific tasks
to specialized functions that guarantee accuracy.
Resources
• 10-Ollama_Function_Calling notebook
• 20-Ollama_Function_Calling_Pydantic
• 30-Function_Calling_with_images notebook
521
• 40-RAG-simple-bee notebook
• calc_distance_image python script
522
Vision-Language Models at the Edge
We will learn Vison-Language Models across tasks such as captioning, object detection,
grounding, and segmentation on a Raspberry Pi.
Figure 15: DALL·E prompt - A Raspberry Pi setup featuring vision tasks. The image shows a
Raspberry Pi connected to a camera, with various computer vision tasks displayed
visually around it, including object detection, image captioning, segmentation, and
visual grounding. The Raspberry Pi is placed on a desk, with a display showing
bounding boxes and annotations related to these tasks. The background should be a
home workspace, with tools and devices typically used by developers and hobbyists.
In this hands-on lab, we will continuously explore AI applications at the Edge, going from the
basic setup of the Florence-2, Microsoft’s state-of-the-art vision foundation model, to advanced
implementations on devices like the Raspberry Pi.
523
Why Florence-2 at the Edge?
Florence-2 is a vision-language model open-sourced by Microsoft under the MIT license, which
significantly advances vision-language models by combining a lightweight architecture with
robust capabilities. Thanks to its training on the massive FLD-5B dataset, which contains
126 million images and 5.4 billion visual annotations, it achieves performance comparable to
larger models. This makes Florence-2 ideal for deployment at the edge, where power and
computational resources are limited.
In this tutorial, we will explore how to use Florence-2 for real-time computer vision applications,
such as:
• Image captioning
• Object detection
• Segmentation
• Visual grounding
524
• Image Encoder: The image encoder is based on the DaViT (Dual Attention Vision
Transformers) architecture. It converts input images into a series of visual token embed-
dings. These embeddings serve as the foundational representations of the visual content,
capturing both spatial and contextual information about the image.
• Multi-Modal Transformer Encoder-Decoder: Florence-2’s core is the multi-modal
transformer encoder-decoder, which combines visual token embeddings from the image
encoder with textual embeddings generated by a BERT-like model. This combination
525
allows the model to simultaneously process visual and textual inputs, enabling a unified
approach to tasks such as image captioning, object detection, and segmentation.
The model’s training on the extensive FLD-5B dataset ensures it can effectively handle diverse
vision tasks without requiring task-specific modifications. Florence-2 uses textual prompts to
activate specific tasks, making it highly flexible and capable of zero-shot generalization. For
tasks like object detection or visual grounding, the model incorporates additional location
tokens to represent regions within the image, ensuring a precise understanding of spatial
relationships.
Technical Overview
Architecture
526
– Florence-2-Base: 232 million parameters
– Florence-2-Large: 771 million parameters
• Unified Representation: Handles multiple vision tasks through a single architecture
• DaViT Vision Encoder: Converts images into visual token embeddings
• Transformer-based Multi-modal Encoder-Decoder: Processes combined visual
and text embeddings
527
• Automated annotation pipeline using specialist models
• Iterative refinement process for high-quality labels
Key Capabilities
Zero-shot Performance
Fine-tuned Performance
Practical Applications
1. Content Understanding
• Automated image captioning for accessibility
• Visual content moderation
• Media asset management
2. E-commerce
• Product image analysis
• Visual search
• Automated product tagging
3. Healthcare
• Medical image analysis
• Diagnostic assistance
• Research data processing
4. Security & Surveillance
528
• Object detection and tracking
• Anomaly detection
• Scene understanding
Florence-2 stands out from other visual language models due to its impressive zero-shot
capabilities. Unlike models like Google PaliGemma, which rely on extensive fine-tuning to
adapt to various tasks, Florence-2 works right out of the box, as we will see in this lab. It can
also compete with larger models like GPT-4V and Flamingo, which often have many more
parameters but only sometimes match Florence-2’s performance. For example, Florence-2
achieves better zero-shot results than Kosmos-2 despite having over twice the parameters.
In benchmark tests, Florence-2 has shown remarkable performance in tasks like COCO cap-
tioning and referring expression comprehension. It outperformed models like PolyFormer and
UNINEXT in object detection and segmentation tasks on the COCO dataset. It is a highly
competitive choice for real-world applications where both performance and resource efficiency
are crucial.
Our choice of edge device is the Raspberry Pi 5 (Raspi-5). Its robust platform is equipped
with the Broadcom BCM2712, a 2.4GHz quad-core 64-bit Arm Cortex-A76 CPU featuring
Cryptographic Extension and enhanced caching capabilities. It boasts a VideoCore VII GPU,
dual 4Kp60 HDMI® outputs with HDR, and a 4Kp60 HEVC decoder. Memory options include
4GB and 8GB of high-speed LPDDR4X SDRAM, with 8GB being our choice to run Florence-2.
It also features expandable storage via a microSD card slot and a PCIe 2.0 interface for fast
peripherals such as M.2 SSDs (Solid State Drives).
We suggest installing an Active Cooler, a dedicated clip-on cooling solution for Raspberry Pi 5
(Raspi-5), for this lab. It combines an aluminum heatsink with a temperature-controlled blower
fan to keep the Raspi-5 operating comfortably under heavy loads, such as running Florense-2.
529
Environment configuration
1. Transformers:
• Florence-2 uses the transformers library from Hugging Face for model loading
and inference. This library provides the architecture for working with pre-trained
vision-language models, making it easy to perform tasks like image captioning, object
detection, and more. Essentially, transformers helps in interacting with the model,
processing input prompts, and obtaining outputs.
2. PyTorch:
• PyTorch is a deep learning framework that provides the infrastructure needed to
run the Florence-2 model, which includes tensor operations, GPU acceleration (if
a GPU is available), and model training/inference functionalities. The Florence-2
model is trained in PyTorch, and we need it to leverage its functions, layers, and
computation capabilities to perform inferences on the Raspberry Pi.
3. Timm (PyTorch Image Models):
530
• Florence-2 uses timm to access efficient implementations of vision models and pre-
trained weights. Specifically, the timm library is utilized for the image encoder
part of Florence-2, particularly for managing the DaViT architecture. It provides
model definitions and optimized code for common vision tasks and allows the easy
integration of different backbones that are lightweight and suitable for edge devices.
4. Einops:
• Einops is a library for flexible and powerful tensor operations. It makes it easy to
reshape and manipulate tensor dimensions, which is especially important for the
multi-modal processing done in Florence-2. Vision-language models like Florence-2
often need to rearrange image data, text embeddings, and visual embeddings to align
correctly for the transformer blocks, and einops simplifies these complex operations,
making the code more readable and concise.
• Transformers and PyTorch are needed to load the model and run the inference.
• Timm is used to access and efficiently implement the vision encoder.
• Einops helps reshape data, facilitating the integration of visual and text features.
All these components work together to help Florence-2 run seamlessly on our Raspberry Pi,
allowing it to perform complex vision-language tasks relatively quickly.
Considering that the Raspberry Pi already has its OS installed, let’s use SSH to reach it from
another computer:
ssh mjrovai@[Link]
hostname -I
[Link]
531
Updating the Raspberry Pi
First, ensure your Raspberry Pi is up to date:
Install Dependencies
Let’s set up and activate a Virtual Environment for working with Florence-2:
Install PyTorch
532
pip3 install setuptools numpy Cython
pip3 install requests
pip3 install torch torchvision --index-url [Link]
pip3 install torchaudio --index-url [Link]
533
jupyter notebook --ip=[Link] --no-browser
Running the above command on the SSH terminal, we can see the local URL address to open
the notebook:
The notebook with the code used on this initial test can be found on the Lab GitHub:
• 10-florence2_test.ipynb
We can access it on the remote computer by entering the Raspberry Pi’s IP address and the
provided token in a web browser ( copy the entire URL from the terminal).
From the Home page, create a new notebook [Python 3 (ipykernel) ] and copy and paste
the example code from Hugging Face Hub.
The code is designed to run Florence-2 on a given image to perform object detection. It
loads the model, processes an image and a prompt, and then generates a response to identify
and describe the objects in the image.
534
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("microsoft/Florence-2-base",
torch_dtype=torch_dtype,
trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base",
trust_remote_code=True)
prompt = "<OD>"
url = "[Link]
images/resolve/main/transformers/tasks/[Link]?download=true"
image = [Link]([Link](url, stream=True).raw)
generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
do_sample=False,
num_beams=3,
)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=False)[0]
print(parsed_answer)
535
1. Importing Required Libraries
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForCausalLM
• requests: Used to make HTTP requests. In this case, it downloads an image from a
URL.
• PIL (Pillow): Provides tools for manipulating images. Here, it’s used to open the
downloaded image.
• torch: PyTorch is imported to handle tensor operations and determine the hardware
availability (CPU or GPU).
• transformers: This module provides easy access to Florence-2 by using AutoProcessor
and AutoModelForCausalLM to load pre-trained models and process inputs.
model = AutoModelForCausalLM.from_pretrained("microsoft/Florence-2-base",
torch_dtype=torch_dtype,
trust_remote_code=True).to(device)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base",
trust_remote_code=True)
• Model Initialization:
536
– AutoModelForCausalLM.from_pretrained() loads the pre-trained Florence-2
model from Microsoft’s repository on Hugging Face. The torch_dtype is set
according to the available hardware (GPU/CPU), and trust_remote_code=True
allows the use of any custom code that might be provided with the model.
– .to(device) moves the model to the appropriate device (either CPU or GPU). In
our case, it will be set to CPU.
• Processor Initialization:
– AutoProcessor.from_pretrained() loads the processor for Florence-2. The pro-
cessor is responsible for transforming text and image inputs into a format the model
can work with (e.g., encoding text, normalizing images, etc.).
prompt = "<OD>"
• Prompt Definition: The string "<OD>" is used as a prompt. This refers to “Object
Detection”, instructing the model to detect objects on the image.
url = "[Link]
images/resolve/main/transformers/tasks/[Link]?download=true"
image = [Link]([Link](url, stream=True).raw)
• Downloading the Image: The [Link]() function fetches the image from the
specified URL. The stream=True parameter ensures the image is streamed rather than
downloaded completely at once.
• Opening the Image: [Link]() opens the image so the model can process it.
6. Processing Inputs
• Processing Input Data: The processor() function processes the text (prompt) and
the image (image). The return_tensors="pt" argument converts the processed data
into PyTorch tensors, which are necessary for inputting data into the model.
537
• Moving Inputs to Device: .to(device, torch_dtype) moves the inputs to the
correct device (CPU or GPU) and assigns the appropriate data type.
generated_ids = [Link](
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
do_sample=False,
num_beams=3,
)
538
• Post-Processing: processor.post_process_generation() is called to process the
generated text further, interpreting it based on the task ("<OD>" for object detection)
and the size of the image.
• This function extracts specific information from the generated text, such as bounding
boxes for detected objects, making the output more useful for visual tasks.
print(parsed_answer)
• Finally, print(parsed_answer) displays the output, which could include object detection
results, such as bounding box coordinates and labels for the detected objects in the image.
Result
539
By the Object Detection result, we can see that:
It seems that at least a few objects were detected. we can also implement a code to draw the
bounding boxes in the find objects:
540
x1, y1, x2, y2 = bbox
# Create a Rectangle patch
rect = [Link]((x1, y1), x2-x1, y2-y1, linewidth=1,
edgecolor='r', facecolor='none')
# Add the rectangle to the Axes
ax.add_patch(rect)
# Annotate the label
[Link](x1, y1, label, color='white', fontsize=8,
bbox=dict(facecolor='red', alpha=0.5))
Box (x0, y0, x1, y1): Location tokens correspond to the top-left and bottom-right
corners of a box.
And running
plot_bbox(image, parsed_answer['<OD>'])
We get:
541
Florence-2 Tasks
Florence-2 is designed to perform a variety of computer vision and vision-language tasks through
prompts. These tasks can be activated by providing a specific textual prompt to the model, as
we saw with <OD> (Object Detection).
Florence-2’s versatility comes from combining these prompts, allowing us to guide the model’s
behavior to perform specific vision tasks. Changing the prompt allows us to adapt Florence-2 to
different tasks without needing task-specific modifications in the architecture. This capability
directly results from Florence-2’s unified model architecture and large-scale multi-task training
on the FLD-5B dataset.
Here are some of the key tasks that Florence-2 can perform, along with example prompts:
542
1. Object Detection (OD)
• Prompt: "<OD>"
• Description: Identifies objects in an image and provides bounding boxes for each
detected object. This task is helpful for applications like visual inspection, surveillance,
and general object recognition.
2. Image Captioning
• Prompt: "<CAPTION>"
• Description: Generates a textual description for an input image. This task helps the
model describe what is happening in the image, providing a human-readable caption for
content understanding.
3. Detailed Captioning
• Prompt: "<DETAILED_CAPTION>"
• Description: Generates a more detailed caption with more nuanced information about
the scene, such as the objects present and their relationships.
4. Visual Grounding
• Prompt: "<CAPTION_TO_PHRASE_GROUNDING>"
• Description: Links a textual description to specific regions in an image. For example,
given a prompt like “a green car,” the model highlights where the red car is in the image.
This is useful for human-computer interaction, where you must find specific objects based
on text.
5. Segmentation
• Prompt: "<REFERRING_EXPRESSION_SEGMENTATION>"
• Description: Performs segmentation based on a referring expression, such as “the
blue cup.” The model identifies and segments the specific region containing the object
mentioned in the prompt (all related pixels).
• Prompt: "<DENSE_REGION_CAPTION>"
• Description: Provides captions for multiple regions within an image, offering a detailed
breakdown of all visible areas, including different objects and their relationships.
543
7. OCR with Region
• Prompt: "<OCR_WITH_REGION>"
• Description: Performs Optical Character Recognition (OCR) on an image and provides
bounding boxes for the detected text. This is useful for extracting and locating textual
information in images, such as reading signs, labels, or other forms of text in images.
• Prompt: "<OPEN_VOCABULARY_OD>"
• Description: The model can detect objects without being restricted to a predefined list
of classes, making it helpful in recognizing a broader range of items based on general
visual understanding.
• 20-florence_2.ipynb
Let’s use a couple of images created by Dall-E and upload them to the Rasp-5 (FileZilla can
be used for that). The images will be saved on a sub-folder named images :
dogs_cats = [Link]('./images/[Link]')
table = [Link]('./images/[Link]')
544
Let’s create a function to facilitate our exploration and to keep track of the latency of the
model for different tasks:
545
)
return parsed_answer
Caption
run_example(task_prompt='<CAPTION>',image=dogs_cats)
2. Table
run_example(task_prompt='<CAPTION>',image=table)
{'<CAPTION>': 'A wooden table topped with a plate of fruit and a glass of wine.'}
DETAILED_CAPTION
run_example(task_prompt='<DETAILED_CAPTION>',image=dogs_cats)
{'<DETAILED_CAPTION>': 'The image shows a group of cats and dogs sitting on top of a
lush green field, surrounded by plants with flowers, trees, and a house in the
background. The sky is visible above them, creating a peaceful atmosphere.'}
546
2. Table
run_example(task_prompt='<DETAILED_CAPTION>',image=table)
{'<DETAILED_CAPTION>': 'The image shows a wooden table with a bottle of wine and a
glass of wine on it, surrounded by a variety of fruits such as apples, oranges, and
grapes. In the background, there are chairs, plants, trees, and a house, all slightly
blurred.'}
MORE_DETAILED_CAPTION
run_example(task_prompt='<MORE_DETAILED_CAPTION>',image=dogs_cats)
{'<MORE_DETAILED_CAPTION>': 'The image shows a group of four cats and a dog in a garden.
The garden is filled with colorful flowers and plants, and there is a pathway leading up
to a house in the background. The main focus of the image is a large German Shepherd dog
standing on the left side of the garden, with its tongue hanging out and its mouth open,
as if it is panting or panting. On the right side, there are two smaller cats, one orange
and one gray, sitting on the grass. In the background, there is another golden retriever
dog sitting and looking at the camera. The sky is blue and the sun is shining, creating a
warm and inviting atmosphere.'}
2. Table
run_example(task_prompt='< MORE_DETAILED_CAPTION>',image=table)
{'<MORE_DETAILED_CAPTION>': 'The image shows a wooden table with a wooden tray on it. On
the tray, there are various fruits such as grapes, oranges, apples, and grapes. There is
also a bottle of red wine on the table. The background shows a garden with trees and a
house. The overall mood of the image is peaceful and serene.'}
We can note that the more detailed the caption task, the longer the latency and
the possibility of mistakes (like “The image shows a group of four cats and a dog in
a garden”, instead of two dogs and three cats).
547
OD - Object Detection
We can run the same previous function for object detection using the prompt <OD>.
task_prompt = '<OD>'
results = run_example(task_prompt,image=dogs_cats)
print(results)
Only by the labels ['cat,' 'cat,' 'cat,' 'dog,' 'dog'] is it possible to see that the main
objects in the image were captured. Let’s apply the function used before to draw the bounding
boxes:
plot_bbox(dogs_cats, results['<OD>'])
548
Let’s also do it with the Table image:
task_prompt = '<OD>'
results = run_example(task_prompt,image=table)
plot_bbox(table, results['<OD>'])
549
DENSE_REGION_CAPTION
It is possible to mix the classic Object Detection with the Caption task in specific sub-regions
of the image:
task_prompt = '<DENSE_REGION_CAPTION>'
results = run_example(task_prompt,image=dogs_cats)
plot_bbox(dogs_cats, results['<DENSE_REGION_CAPTION>'])
results = run_example(task_prompt,image=table)
plot_bbox(table, results['<DENSE_REGION_CAPTION>'])
550
CAPTION_TO_PHRASE_GROUNDING
With this task, we can enter with a caption, such as “a wine glass”, “a wine bottle,” or “a half
orange,” and Florence-2 will localize the object in the image:
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
551
[INFO] ==> Florence-2-base (<CAPTION_TO_PHRASE_GROUNDING>), took 15.7 seconds to execute
each task.
Cascade Tasks
We can also enter the image caption as the input text to push Florence-2 to find more objects:
task_prompt = '<CAPTION>'
results = run_example(task_prompt,image=dogs_cats)
text_input = results[task_prompt]
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(task_prompt, text_input,image=dogs_cats)
plot_bbox(dogs_cats, results['<CAPTION_TO_PHRASE_GROUNDING>'])
552
OPEN_VOCABULARY_DETECTION
task_prompt = '<OPEN_VOCABULARY_DETECTION>'
text = ["a house", "a tree", "a standing cat at the left",
"a sleeping cat on the ground", "a standing cat at the right",
"a yellow cat"]
for txt in text:
results = run_example(task_prompt, text_input=txt,image=dogs_cats)
bbox_results = convert_to_od_format(results['<OPEN_VOCABULARY_DETECTION>'])
plot_bbox(dogs_cats, bbox_results)
553
[INFO] ==> Florence-2-base (<OPEN_VOCABULARY_DETECTION>), took 15.1 seconds
to execute each task.
Note: Trying to use Florence-2 to find objects that were not found can leads to
mistakes (see exaamples on the Notebook).
We can also segment a specific object in the image and give its description (caption), such as
“a wine bottle” on the table image or “a German Sheppard” on the dogs_cats.
Referring expression segmentation results format: {'<REFERRING_EXPRESSION_SEGMENTATION>':
{'Polygons': [[[polygon]], ...], 'labels': ['', '', ...]}}, one object is repre-
sented by a list of polygons. each polygon is [x1, y1, x2, y2, ..., xn, yn].
Polygon (x1, y1, …, xn, yn): Location tokens represent the vertices of a polygon
in clockwise order.
554
from PIL import Image, ImageDraw, ImageFont
import copy
import random
import numpy as np
colormap = ['blue','orange','green','purple','brown','pink','gray','olive',
'cyan','red','lime','indigo','violet','aqua','magenta','coral','gold',
'tan','skyblue']
Parameters:
- image_path: Path to the image file.
- prediction: Dictionary containing 'polygons' and 'labels' keys.
'polygons' is a list of lists, each containing vertices
of a polygon.
'labels' is a list of labels corresponding to each polygon.
- fill_mask: Boolean indicating whether to fill the polygons with color.
"""
# Load the image
draw = [Link](image)
555
if fill_mask:
[Link](_polygon, outline=color, fill=fill_color)
else:
[Link](_polygon, outline=color)
task_prompt = '<REFERRING_EXPRESSION_SEGMENTATION>'
556
[INFO] ==> Florence-2-base (<REFERRING_EXPRESSION_SEGMENTATION>),
took 207.0 seconds to execute each task.
Region to Segmentation
With this task, it is also possible to give the object coordinates in the image to segment it. The
input format is '<loc_x1><loc_y1><loc_x2><loc_y2>', [x1, y1, x2, y2] , which is the
quantized coordinates in [0, 999].
For example, when running the code:
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(task_prompt, text_input="a half orange",image=table)
results
557
Using the bboxes rounded coordinates:
task_prompt = '<REGION_TO_SEGMENTATION>'
results = run_example(task_prompt,
text_input="<loc_343><loc_690><loc_531><loc_874>",
image=table)
output_image = [Link](table)
draw_polygons(output_image, results['<REGION_TO_SEGMENTATION>'], fill_mask=True)
Region to Texts
We can also give the region (coordinates and ask for a caption):
558
task_prompt = '<REGION_TO_CATEGORY>'
results = run_example(task_prompt, text_input="<loc_343><loc_690><loc_531>
<loc_874>",image=table)
results
{'<REGION_TO_CATEGORY>': 'orange<loc_343><loc_690><loc_531><loc_874>'}
The model identified an orange in that region. Let’s ask for a description:
task_prompt = '<REGION_TO_DESCRIPTION>'
results = run_example(task_prompt, text_input="<loc_343><loc_690><loc_531>
<loc_874>",image=table)
results
{'<REGION_TO_CATEGORY>': 'orange<loc_343><loc_690><loc_531><loc_874>'}
In this case, the description did not provide more details, but it could. Try another example.
OCR
With Florence-2, we can perform Optical Character Recognition (OCR) on an image, getting
what is written on it (task_prompt = '<OCR>' and also get the bounding boxes (location) for
the detected text (ask_prompt = '<OCR_WITH_REGION>'). Those tasks can help extract and
locate textual information in images, such as reading signs, labels, or other forms of text in
images.
Let’s upload a flyer from a talk in Brazil to Raspi. Let’s test works in another language, here
Portuguese):
flayer = [Link]('./images/[Link]')
# Display the image
[Link](figsize=(8, 8))
[Link](flayer)
[Link]('off')
#[Link]("Image")
[Link]()
559
Let’s examine the image with '<MORE_DETAILED_CAPTION>' :
The description is very accurate. Let’s get to the more important words with the task OCR:
task_prompt = '<OCR>'
run_example(task_prompt,image=flayer)
560
{'<OCR>': 'Machine LearningCafécomEmbarcadoEmbarcadosDemocratizando a
InteligênciaArtificial para Paises em25 de Setembro ás 17hDesenvolvimentoToda quarta-
feiraMarcelo RovalProfessor na UNIFIEI eTransmissão viainCo-Director do TinyML4D'}
task_prompt = '<OCR_WITH_REGION>'
results = run_example(task_prompt,image=flayer)
Let’s also create a function to draw bounding boxes around the detected words:
fill=color)
display(image)
output_image = [Link](flayer)
draw_ocr_bboxes(output_image, results['<OCR_WITH_REGION>'])
561
We can inspect the detected words:
results['<OCR_WITH_REGION>']['labels']
'</s>Machine Learning',
'Café',
'com',
'Embarcado',
'Embarcados',
'Democratizando a Inteligência',
'Artificial para Paises em',
'25 de Setembro ás 17h',
'Desenvolvimento',
'Toda quarta-feira',
'Marcelo Roval',
'Professor na UNIFIEI e',
'Transmissão via',
'in',
'Co-Director do TinyML4D']
562
Latency Summary
The latency observed for different tasks using Florence-2 on the Raspberry Pi (Raspi-5) varied
depending on the complexity of the task:
These latency times highlight the resource constraints of edge devices like the Raspberry Pi
and emphasize the need to optimize the model and the environment to achieve real-time
performance.
563
Running complex tasks can use all 8GB of the Raspi-5’s memory. For example, the
above screenshot during the Florence OD task shows 4 CPUs at full speed and over
5GB of memory in use. Consider increasing the SWAP memory to 2 GB.
Checking the CPU temperature with vcgencmd measure_temp , showed that temperature can
go up to +80oC.
Fine-Tunning
As explored in this lab, Florence supports many tasks out of the box, including captioning, object
detection, OCR, and more. However, like other pre-trained foundational models, Florence-2
may need domain-specific knowledge. For example, it may need to improve with medical or
satellite imagery. In such cases, fine-tuning with a custom dataset is necessary. The Roboflow
tutorial, How to Fine-tune Florence-2 for Object Detection Tasks, shows how to fine-tune
Florence-2 on object detection datasets to improve model performance for our specific use
case.
Based on the above tutorial, it is possible to fine-tune the Florence-2 model to detect boxes
and wheels used in previous labs:
It is important to note that after fine-tuning, the model can still detect classes that don’t
belong to our custom dataset, like cats, dogs, grapes, etc, as seen before).
The complete fine-tunning project using a previously annotated dataset in Roboflow and
executed on CoLab can be found in the notebook:
• 30-Finetune_florence_2_on_detection_dataset_box_vs_wheel.ipynb
564
Conclusion
Florence-2 offers a versatile and powerful approach to vision-language tasks at the edge,
providing performance that rivals larger, task-specific models, such as YOLO for object
detection, BERT/RoBERTa for text analysis, and specialized OCR models.
Thanks to its multi-modal transformer architecture, Florence-2 is more flexible than YOLO in
terms of the tasks it can handle. These include object detection, image captioning, and visual
grounding.
Unlike BERT, which focuses purely on language, Florence-2 integrates vision and language,
allowing it to excel in applications that require both modalities, such as image captioning and
visual grounding.
Moreover, while traditional OCR models such as Tesseract and EasyOCR are designed solely
for recognizing and extracting text from images, Florence-2’s OCR capabilities are part of a
broader framework that includes contextual understanding and visual-text alignment. This
makes it particularly useful for scenarios that require both reading text and interpreting its
context within images.
Overall, Florence-2 stands out for its ability to seamlessly integrate various vision-language
tasks into a unified model that is efficient enough to run on edge devices like the Raspberry Pi.
This makes it a compelling choice for developers and researchers exploring AI applications at
the edge.
1. Unified Architecture
• Single model handles multiple vision tasks vs. specialized models (YOLO, BERT,
Tesseract)
• Eliminates the need for multiple model deployments and integrations
• Consistent API and interface across tasks
2. Performance Comparison
• Object Detection: Comparable to YOLOv8 (~37.5 mAP on COCO vs. YOLOv8’s
~39.7 mAP) despite being general-purpose
• Text Recognition: Handles multiple languages effectively like specialized OCR models
(Tesseract, EasyOCR)
• Language Understanding: Integrates BERT-like capabilities for text processing while
adding visual context
3. Resource Efficiency
• The Base model (232M parameters) achieves strong results despite smaller size
565
• Runs effectively on edge devices (Raspberry Pi)
• Single model deployment vs. multiple specialized models
Trade-offs
1. Resource-Constrained Environments
• Edge devices requiring multiple vision capabilities
• Systems with limited storage/deployment capacity
• Applications needing flexible vision processing
2. Multi-modal Applications
• Content moderation systems
• Accessibility tools
• Document analysis workflows
3. Rapid Prototyping
• Quick deployment of vision capabilities
• Testing multiple vision tasks without separate models
• Proof-of-concept development
566
Future Implications
Florence-2 represents a shift toward unified vision models that could eventually replace task-
specific architectures in many applications. While specialized models maintain advantages in
specific scenarios, the convenience and efficiency of unified models like Florence-2 make them
increasingly attractive for real-world deployments.
The lab demonstrates Florence-2’s viability on edge devices, suggesting future IoT, mobile
computing, and embedded systems applications where deploying multiple specialized models
would be impractical.
Resources
• 10-florence2_test.ipynb
• 20-florence_2.ipynb
• 30-Finetune_florence_2_on_detection_dataset_box_vs_wheel.ipynb
567
Audio and Vision AI Pipeline
Introduction
In this chapter, we extend our SLM and SVL capabilities by creating a complete audio processing
pipeline that transforms voice or image input into intelligent vocal responses. We will learn to
integrate Speech-to-Text (STT), Small Language (or Visual) Models, and Text-to-Speech (TTS)
technologies to build conversational AI systems that run entirely on Raspberry Pi hardware.
This chapter bridges the gap between our existing computer vision knowledge and multimodal
AI applications, demonstrating how different AI components work together in real-world edge
deployments.
568
The Audio to Audio AI Pipeline Architecture
When we built computer vision systems earlier in the course, we processed visual data to
extract meaningful information. Audio AI systems follow a similar principle but work with
temporal audio signals instead of static images. The key insight is that speech processing
requires multiple specialized models working together, rather than a single end-to-end system.
Modern small models, such as Gemma 3n, can process audio directly and its prompt.
Today (September 2025), Gemma 3n can transcribe text from audio files using
Hugging Face Transformers, but it is not available with Ollama
Consider how humans process spoken language. We simultaneously parse the acoustic signal,
understand the linguistic content, reason about the meaning, and formulate responses. Our AI
pipeline mimics this process by breaking it into distinct, manageable components.
Our comprehensive audio AI pipeline comprises four main components, connected in sequence.
569
[Microphone] → [STT Model] → [SLM] → [TTS Model] → [Speaker]
Audio Text Text Audio
Audio captured by the microphone is processed through a Speech-to-Text model, which converts
sound waves into text transcriptions. This text becomes input for our Small Language Model,
which generates intelligent responses. Finally, a Text-to-Speech system converts the written
response back into spoken audio.
Each component has specific requirements and limitations. The STT model must handle
various accents and noise conditions. The SLM needs sufficient context to generate coherent
responses. The TTS system must produce speech that sounds natural. Understanding these
individual requirements helps us optimize the overall system performance.
Edge AI Considerations
Begin by identifying the audio capabilities of our system. The Raspberry Pi can work with
various audio input and output devices, but proper configuration is essential for reliable
operation.
Use the command arecord -l to list available recording devices. You should see output
showing your microphone’s card and device numbers. For USB microphones, this typically
appears as something like card 2: Microphone [USB Condenser Microphone], device 0:
USB Audio [USB Audio]. The critical information is the card number and device number,
which you’ll reference as hw:2,0 in the code.
570
Testing Basic Audio Functionality
Before writing Python code, we should verify that our audio setup works correctly at the
system level. Let’s record a short test file using:
arecord --device="plughw:2,0" --format=S16_LE --rate=16000 -c2 [Link]
571
Play back the recording with aplay [Link] to confirm that both capture and playback
work correctly (use [CTRL]+[C] to stop the recording or add a duration in seconds to the
command line).
The .WAV file can be played on another device (such as a computer) or on the
Raspberry Pi, as a speaker can be connected via USB or Bluetooth.
source ~/ollama/bin/activate
The PyAudio library requires system-level audio libraries; therefore, install them using sudo
apt-get.
Let’s create a working directory: Documents/OLLAMA/SST and verify the USB device index,
with the below script (verify_usb_index.py:
import pyaudio
p = [Link]()
for ii in range(p.get_device_count()):
print(ii, p.get_device_info_by_index(ii).get('name'))
572
As a result, we should get:
A lot of messages should appear. They are mostly ALSA and JACK warnings
about missing or undefined virtual/surround sound devices—they are common on
Raspberry Pi systems with minimal or headless sound configs and typically do
not impact basic USB microphone capture. If our USB Microphone appears as a
recording device (as it does: “hw:2,0”), we can safely ignore most of these unless
audio capture fails.
import pyaudio
import wave
FORMAT = pyaudio.paInt16
CHANNELS = 1
RATE = 16000 # 16 kHz
CHUNK = 1024
RECORD_SECONDS = 10
DEVICE_INDEX = 2 # replace this with your detected USB mic's index
WAVE_OUTPUT_FILENAME = "[Link]"
573
audio = [Link]()
print("Recording...")
frames = []
print("Finished recording.")
stream.stop_stream()
[Link]()
[Link]()
wf = [Link](WAVE_OUTPUT_FILENAME, 'wb')
[Link](CHANNELS)
[Link](audio.get_sample_size(FORMAT))
[Link](RATE)
[Link](b''.join(frames))
[Link]()
Understanding the audio configuration parameters helps prevent common problems. We use
16-bit PCM format (pyaudio.paInt16) because it provides good quality while remaining
computationally efficient. The 16kHz sampling rate balances audio quality with processing
requirements - most speech recognition models expect this rate.
The buffer size (CHUNK = 1024) affects latency and reliability. Smaller buffers reduce latency
but may cause audio dropouts on busy systems. Larger buffers increase latency but provide
more stable recording.
Let’s Playback to verify if we get it correctly:
574
Your browser does not support the audio element.
Traditional speech recognition systems, such as OpenAI’s Whisper, are highly accurate but
require substantial computational resources. Moonshine is specifically designed for edge devices,
using optimized model architectures and quantization techniques to achieve good performance
on resource-constrained hardware.
The ONNX (Open Neural Network Exchange) version of Moonshine provides additional
optimization benefits. ONNX Runtime includes hardware-specific optimizations that can
significantly improve inference speed on ARM processors, such as those found in Raspberry Pi
devices.
Moonshine offers different model sizes with clear trade-offs between accuracy and computational
requirements. The “tiny” model processes audio quickly but may struggle with difficult audio
conditions. The “base” model provides better accuracy but requires more processing time and
memory.
For initial development, we should start with the tiny model to ensure that the pipeline works
correctly. Once the complete system is functional, we can experiment with larger models to
find the optimal balance for our specific use case and hardware capabilities.
575
Implementation and Preprocessing
This specific installation method ensures compatibility with the ONNX runtime optimizations.
Let’s run the test script below (transcription_test.py):
import moonshine_onnx
text = moonshine_onnx.transcribe('[Link]', 'moonshine/tiny')
print(text[0])
As a result, we will get the corresponding text, which was recorded before:
The text output from your speech recognition system becomes input for your Small Language
Model. However, the characteristics of spoken language differ significantly from written text,
especially when filtered through speech recognition systems.
Spoken language tends to be more informal, may contain false starts and repetitions, and might
include transcription errors. Our SLM integration should account for these characteristics. On
a final implementation, we should consider preprocessing the STT output to clean up obvious
transcription errors or providing context to the SLM about the voice interaction nature of the
input.
576
Optimizing Prompts for Voice Interaction
Voice-based interactions have different expectations than text-based chats. Responses should
be concise since users must listen to the entire output. Avoid complex formatting or long lists
that work well in text but become cumbersome when spoken aloud.
We should design our system prompts to encourage responses appropriate for voice interaction.
For example, “Provide a brief, conversational response suitable for speaking aloud” can help
guide the SLM toward more appropriate output formatting.
Unlike single-query text interactions, voice conversations often involve multiple exchanges.
Implementing conversation context memory significantly enhances the user experience. However,
context management on edge devices requires careful consideration of memory usage.
Consider implementing a sliding window approach, where you maintain the last few exchanges
in memory but discard older context to prevent memory exhaustion, balancing context length
with available system resources.
Let’s create a function to handle this. For test, run slm_test.py:
import ollama
577
response = [Link](
model=model,
prompt=full_prompt
)
return response['response']
Text-to-speech systems face different challenges than speech recognition systems. While STT
must handle various input conditions, TTS must generate consistent, natural-sounding output
across diverse text inputs. The quality of TTS has a significant impact on the user experience
in voice interaction systems.
PIPER provides an excellent balance between voice quality and computational efficiency for
edge deployments. Unlike cloud-based TTS services, PIPER runs entirely locally, ensuring
privacy and eliminating network dependencies.
PIPER offers various voice models with different characteristics. The “low”, “medium”, and
“high” quality designations primarily refer to model size and computational requirements
rather than dramatic quality differences. For most applications, the low-quality models provide
acceptable voice output while running efficiently on Raspberry Pi hardware.
578
Install PIPER with pip install piper-tts, create a voices directory:
Download voice models from the Hugging Face repository. Each voice model requires both the
model file (.onnx) and a configuration file (.json). The configuration file contains model-specific
parameters essential for generating proper audio.
We should download both files for our chosen voice; for example, the English female “lessac”
voice provides clear, natural speech suitable for most applications.
import subprocess
import os
Args:
text (str): Text to convert to speech
output_file (str): Output WAV file path
"""
# Path to your voice model
model_path = "voices/en_US-[Link]"
try:
# Run PIPER command
process = [Link](
579
['piper', '--model', model_path, '--output_file', output_file],
stdin=[Link],
stdout=[Link],
stderr=[Link],
text=True
)
if [Link] == 0:
print(f"\nSpeech generated successfully: {output_file}")
return True
else:
print(f"Error: {stderr}")
return False
except Exception as e:
print(f"Error running PIPER: {e}")
return False
if text_to_speech_piper(txt):
print("You can now play the file with: aplay piper_output.wav")
else:
print("Failed to generate speech")
Runing the script, a piper_output.wav file will be generated, which is the text converted into
speech.
Your browser does not support the audio element.
To listen to the sound, we can run: aplay piper_output.wav.
580
Handling Long Text and Special Cases
TTS systems may struggle with very long input text or special characters. Implement text
preprocessing to handle these cases gracefully. Break long responses into shorter segments,
handle abbreviations and numbers appropriately, and filter out problematic characters that
might cause TTS failures.
Consider implementing text chunking for responses longer than a reasonable speaking length.
This prevents both TTS processing issues and user fatigue from overly long audio responses.
Integrating all components requires careful attention to error handling and resource management.
Each stage of the pipeline can fail independently, and robust systems must handle these failures
gracefully rather than crashing.
Design our integration with modularity in mind. Test each component independently before
combining them. This approach simplifies debugging and allows you to optimize individual
components separately.
We should also implement proper logging throughout our pipeline. When complex systems
fail, detailed logs help identify whether the issue occurs in audio capture, speech recognition,
language model processing, text-to-speech conversion, or audio playback.
581
Performance Optimization Strategies
Measure the timing of each pipeline component to identify bottlenecks. Typically, the SLM
inference takes the longest time, followed by TTS generation. Understanding these timing
characteristics helps prioritize optimization efforts.
Consider implementing concurrent processing where possible. For example, you might begin
TTS processing for the first part of an extended response while the SLM is still generating the
remainder. However, be cautious about memory usage when implementing parallel processing
on resource-constrained devices.
Edge devices have limited RAM, and loading multiple large models simultaneously can cause
memory pressure. Implement strategies to manage memory efficiently, such as loading models
only when needed or using model swapping for infrequently used components.
Monitor system memory usage during operation and implement safeguards to prevent memory
exhaustion. Consider implementing graceful degradation where your system switches to smaller,
more efficient models if memory becomes constrained.
Production-quality voice interaction systems must handle various failure modes gracefully.
Network interruptions, hardware disconnections, model loading failures, and unexpected input
conditions should not cause your system to crash.
Implement comprehensive error handling at each pipeline stage. When speech recognition
produces empty output, provide the user with meaningful feedback rather than processing
empty strings. When TTS fails, consider falling back to text display or simplified audio
feedback.
Design user feedback mechanisms that work within your voice interaction paradigm. Audio
beeps, LED indicators, or simple voice messages can communicate system status without
requiring visual displays.
Multi-stage systems present unique debugging challenges. When the overall system fails,
identifying the specific failure point requires systematic testing approaches.
582
Implement test modes that allow you to inject known inputs at each pipeline stage. This
capability enables you to isolate problems to specific components rather than repeatedly testing
the entire system.
Create diagnostic outputs that help understand system behavior. For example, displaying
transcription confidence scores, SLM response times, or TTS processing status helps identify
performance issues or quality problems.
Considering the previous points, let’s assemble all the essential components that work together:
audio capture, transcription, language model processing, and text-to-speech. We should combine
these into a complete voice pipeline that flows naturally from one step to the next.
The key insight here is that each of our developed scripts represents a stage in what’s called an
“audio processing pipeline”. Let’s walk through how we can connect these pieces.
Understanding the Pipeline Architecture
The pipeline follows a logical sequence: the voice becomes audio data, that audio is converted
into text, the text is processed into an AI response, the response is converted into speech audio,
and finally, that speech audio is transformed into sound that we can hear.
We should have a run_voice_pipeline() function in addition to the previous ones that acts
as a coordinator, ensuring each step completes successfully before proceeding to the next. If
any step fails, the entire pipeline stops gracefully rather than trying to continue with missing
data.
Key Integration Points
We should connect the scripts by ensuring the output of each function becomes the input
for the following function. For example, record_audio() creates “user_input.wav”, which
transcribe_audio() reads to produce text, which generate_response() processes to develop
an AI response, and so on.
The error handling at each step ensures that if our microphone isn’t working, or if the AI model
is busy, or if the voice model files are missing, we get clear feedback about what went wrong
rather than mysterious crashes.
Optimizations for Raspberry Pi
The Raspberry Pi has limited resources compared to a desktop computer, so we should include
several optimizations. The cleanup_temp_files() function prevents our storage from filling
up with temporary audio files. The audio configuration uses 16kHz sampling (which matches
Moonshine’s expectations) rather than CD-quality 44kHz, reducing processing overhead.
583
The continuous assistant mode includes a manual trigger (pressing Enter) rather than voice
activation detection, which saves CPU cycles that would otherwise be spent constantly
monitoring audio input.
584
We can start by testing individual components using detect_audio_devices() first to confirm
our USB microphone is still at index 2. Then, we can run run_voice_pipeline() with a
simple question to verify the complete flow works.
Once we are confident in single interactions, we can use continuous_voice_assistant()
for extended conversations. This mode lets us have back-and-forth exchanges with your AI
assistant, making it feel more like a natural conversation partner.
If you want to experiment with different SLM models, we need to change the model
parameter in generate_response().
Voice interaction systems have significant potential for educational applications and accessibility
improvements by designing interfaces that adapt to different user needs and capabilities. By
integrating small visual models, such as Moondream, with existing TTS pipelines, we can create
multimodal assistants that describe images for people with visual impairments, converting
visual content into detailed spoken descriptions of scenes, objects, and spatial relationships.
585
To use the camera, we must ensure that the NumPy version is compatible with the Raspberry
Pi system (1.24.2). If this version was changed (what it should be due to the Moonshine
installation), revert it to the 1.24.2 version.
Now, let’s apply what we have explored in previous chapters, along with what we learn in this
one, to create a simple code that converts an image captured by a camera into its corresponding
spoken caption by the speaker.
The code is straightforward and is intended solely to test the solution’s potential.
import os
import time
import subprocess
import ollama
from picamera2 import Picamera2
586
def capture_image(img_path):
# Initialize camera
picam2 = Picamera2()
[Link]()
# Capture image
picam2.capture_file(img_path)
print("\n==> Image captured: "+img_path)
# Stop camera
[Link]()
[Link]()
587
if not [Link](model_path):
print(f"Error: Model file not found at {model_path}")
return False
try:
# Run PIPER command
process = [Link](
['piper', '--model', model_path, '--output_file', output_file],
stdin=[Link],
stdout=[Link],
stderr=[Link],
text=True
)
if [Link] == 0:
print(f"\nSpeech generated successfully: {output_file}")
return True
else:
print(f"Error: {stderr}")
return False
except Exception as e:
print(f"Error running PIPER: {e}")
return False
def play_audio(filename="assistant_response.wav"):
try:
# Use aplay to play the audio file
result = [Link](['aplay', filename],
capture_output=True,
text=True)
if [Link] == 0:
print("\nAudio playback completed")
return True
else:
print(f"\nPlayback error: {[Link]}")
return False
588
except Exception as e:
print(f"\nError playing audio: {e}")
return False
IMG_PATH = "/home/mjrovai/Documents/OLLAMA/SST/capt_image.jpg"
MODEL = "moondream:latest"
capture_image(IMG_PATH)
caption = image_description(IMG_PATH, MODEL)
print ("\n==> AI Response:", caption)
text_to_speech_piper(caption)
play_audio()
589
Your browser does not support the audio element.
Troubleshooting
To work simultaneously with STT and the camera in the same environment, we should ensure
that all core scientific packages are version-aligned for the Raspberry Pi environment. To use
NumPy 1.24.2 on our Raspberry Pi, we must ensure that both our SciPy and Librosa versions
are compatible with that NumPy version.
So, to avoid an eventual [Link] import error, and other possible incompatibilities,
we should downgrade scipy (and librosa) to versions that support numpy 1.24.2. According to
official compatibility tables, scipy 1.11.x works with numpy 1.24.x. Librosa versions released
after 0.9.0 also provide better support for recent NumPy releases. However, some older versions
of Librosa are not compatible with NumPy 1.24.2 due to deprecated NumPy attributes. So, to
fix it, we should run the lines below:
Conclusion
This chapter introduced us to multimodal AI system development through an audio and vision
processing pipeline. We learned to integrate speech recognition, language models, and speech
synthesis into a cohesive system that runs efficiently on edge hardware.
We explored how to architect systems with multiple AI components, handle complex error
conditions, and optimize performance within resource constraints.
These skills are essential for the advanced topics in upcoming chapters, including RAG systems,
agent architectures, and the integration of physical computing. The system thinking approach
used here will be essential for future AI engineering work.
We should consider how the voice interaction capabilities we built might enhance other AI
systems. Many applications benefit from voice interfaces, and the foundation established
here can be adapted and extended for various use cases. We can, for example, transform our
audio pipeline into a smart home assistant by integrating physical computing elements. Voice
commands can trigger LED indicators, read sensor values, or control actuators connected to
our Raspberry Pi GPIO pins. Voice queries about environmental conditions can trigger sensor
readings, while voice commands can control connected devices.
590
This chapter extends our SLM and VLM work by adding input and/or output modalities beyond
text and images. The same language models we used previously now process voice-derived
input and generate responses for speech synthesis.
Consider how RAG systems from later chapters might integrate with voice interactions. Voice
queries could trigger document retrieval, with synthesized responses incorporating retrieved
information.
Resources
Python Scripts
591
592
Physical Computing with Raspberry Pi
593
Introduction
Physical computing creates interactive systems that sense and respond to the analog world.
While this field has traditionally focused on direct sensor readings and programmed responses,
we’re entering an exciting new era where Large Language Models (LLMs) can add sophisticated
decision-making and natural language interaction to physical computing projects.
In the Small Language Models (SLM) chapter, we learned how to run an LLM (or, more
precisely, an SLM) on a Single Board Computer (SBC) such as the Raspberry Pi. In this
chapter, we will go through the process of setting up a Raspberry Pi for physical computing,
with an eye toward future AI integration. We’ll cover:
We will also use a Jupyter notebook (programmed in Python) to interact with sensors and
actuators—an important and necessary first step toward the goal of integrating the Raspi with
an SLM.
The combination of Raspberry Pi’s versatility and the power of SLMs opens up exciting
possibilities for creating more intelligent and responsive physical computing systems.
The diagram below gives us an overview of the project:
594
Prerequisites
• Raspberry Pi (model 4 or 5)
• DHT22 Temperature and Relative Humidity Sensor
• BMP280 Barometric Pressure, Temperature and Altitude Sensor
• Colored LEDs (3x)
• Push Button (1x)
• Resistor 4K7 ohm (2x)
• Resistor 220 or 330 ohm (3x)
The Raspberry Pi’s GPIO (General Purpose Input/Output) pins allow us to connect electronic
components and control them with Python code. This opens up endless possibilities for creating
interactive projects, home automation systems, robotics, and more.
This chapter covers the modern GPIO Zero library for interactions with buttons and LEDs.
� IMPORTANT: [Link] does NOT support Raspberry Pi 5!
595
With the Raspberry Pi 5, we must use GPIO Zero or the newer lgpio library.
[Link] only works on Pi models 1-4, the Pi Zero, Pi Zero 2, and the Pi Zero
2W.
The Raspberry Pi has 40 pins on its header, but not all of them are GPIO pins. Some provide
power (3.3V and 5V), others are ground pins, and the rest are programmable GPIO pins.
In this chapter, we’ll use BCM numbering as it’s more commonly used in Python
programming.
Safety First
Before connecting any components, please read these important safety guidelines:
• Never connect 5V directly to GPIO pins - GPIO pins are 3.3V tolerant only
596
• Always use current-limiting resistors with LEDs - Without them, you risk damaging
the LED or GPIO pin
• Double-check your connections before powering on
• Disconnect power when making circuit changes
• Respect polarity - LEDs for example, have positive (long leg) and negative (short leg)
sides
A modern, high-level library that makes GPIO programming much simpler and more intuitive.
It uses object-oriented programming and includes built-in features like automatic pin cleanup,
device abstraction, and event detection.
It is essential to note that the GPIO Zero Library uses Broadcom (BCM) pin
numbering for GPIO pins, rather than physical (board) numbering. Any pin
marked “GPIO ” in the previous diagram can be used as a PIN. For example, if an
LED were attached to GPIO13, we would specify the PIN as 13 rather than 33 (the
physical one).
It was created by Ben Nuttall of the Raspberry Pi Foundation, Dave Jones, and other
contributors (GitHub).
Advantages of GPIO Zero:
Installation:
597
“Hello World”: Blinking an LED
Let’s start with the classic ‘Hello World’ of physical computing - making an LED blink!
To connect our RPi to the world, let’s first connect:
• Physical Pin 6 (GND) to GND Breadboard Power Grid (Blue -), using a black jumper
• Physical Pin 1 (3.3V) to +VCC Breadboard Power Grid (Red +), using a red jumper
Now, let’s connect an LED (red) using the physical pin 33 (GPIO13) connected to the LED
cathode (longer LED leg). Connect the LED anode to the breadboard GND using a 220 ohms
resistor to reduce the current drawn from the Raspberry Pi, as shown below:
An LED (Light-Emitting Diode) requires current to flow through it to produce light. However,
without a resistor, too much current can flow, damaging the LED or your Raspberry Pi. We
use a resistor to limit the current to a safe level.
598
Why 220Ω or 330Ω Resistors?
These values limit the current to approximately 10-15mA, which is safe for most standard
LEDs and GPIO pins. The exact value isn’t critical—anything from 220 Ω to 1 kΩ will work
fine.
We can use the built-in Python interpreter to test the LED. In the terminal, enter python, and
once in the interpreter, enter with the commands below:
python
>>> from gpiozero import LED
>>> led = LED(13)
>>> [Link]()
>>> [Link]
599
On the Raspberry, start at home and go to Documents.
cd Documents
Create a directory to save the scripts and install the libraries. Move to there:
mkdir GPIO
cd GPIO
We can use any text editor (such as Nano) to create and run the script. Save the file, for
example, as led_test.py, and then execute it using the terminal:
python led_test.py
Now, let’s blink the LED (the actual “Hello world”) when talking about physical computing.
To do that, we must also import another library: time. We need it to define how long the LED
will be ON and OFF. In the case below, the LED will blink every 1 second.
600
We can use any text editor (such as Nano) to create and run the script. Save the file, for
example, as [Link], and then execute it using the terminal:
python [Link]
The LEDs can be used as “actuators”; depending on the condition of a code running on our Pi,
we can command one of the LEDs to fire! We will install two more LEDs, in addition to the
red one already installed. Follow the diagram and install the yellow (on GPIO 19 ) and the
green (on GPIO 26).
601
For testing we can run a similar code as the used with the single red led, changing the pin
accordantly, for example.
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
602
sleep(5)
[Link]()
[Link]()
[Link]()
In this section, we will setup the Raspberry Pi to capture data from several different sensors:
Sensors and Communication type:
Button
Now, let’s learn how to read input from a button. This allows your Raspberry Pi to respond to
physical interactions!
603
Understanding Pull-Up and Pull-Down Resistors
When a button is not pressed, the GPIO pin is ‘floating’ - it’s not connected to anything and
can read random values. We use pull-up or pull-down resistors to set the pin’s default state.
• Pull-Down: Default state is LOW (0V), becomes HIGH when button pressed
• Pull-Up: Default state is HIGH (3.3V), becomes LOW when button pressed
The simplest way to run an external command is with a push button, and the GPIO Zero
Library makes it easy to include in the project. We do not need to think about Pull-up or
Pull-down resistors, etc. In terms of HW, the only thing to do is to connect one leg of our
push-button to any one of the Raspi GPIOs and the other one to GND, as shown in the
diagram:
604
• Push-Button leg1 to GPIO 20
• Push-Button leg2 to GND
605
On a Raspberry Pi Zero 2W, for example, we could use the [Link], whose code is:
[Link]([Link])
[Link](False)
BUTTON_PIN = 20
try:
while True:
if [Link](BUTTON_PIN) == [Link]:
print('Button pressed!')
else:
print('Button not pressed')
[Link](0.1)
except KeyboardInterrupt:
[Link]()
606
Installing Adafruit CircuitPython
The GPIO Zero library is an excellent hardware interfacing library for Raspberry Pi. It’s great
for digital in/out, analog inputs, servos, basic sensors, etc. However, it doesn’t cover SPI/I2C
sensors or drivers. By using CircuitPython via adafruit_blinka, we can take advantage of all
of the drivers and example code developed by Adafruit!
Note that we will keep using GPIO Zero for pins, buttons, and LEDs.
Enable Interfaces
Run these commands to enable the various interfaces such as I2C and SPI:
Install Blinka
Let’s enter the Ollama environment, alheady created to install Blinka:
source ~/ollama/bin/activate
ls /dev/i2c* /dev/spi*
607
Verify Blinka Version
import board
import digitalio
import busio
print("Hello, blinka!")
608
pin = [Link](board.D4)
print("Digital IO ok!")
print("done!")
python blinka_test.py
The first sensor to be installed will be the DHT22 for capturing air temperature and relative
humidity data.
Overview
The low-cost DHT temperature and humidity sensors are elementary and slow, but great
for logging basic data. They consist of a capacitive humidity sensor and a thermistor. A
bare chip inside performs the analog-to-digital conversion and outputs a digital signal con-
taining the temperature and humidity. The digital signal is relatively easy to read using any
microcontroller.
609
DHT22 Main characteristics:
Once we use the sensor at distances less than 20m, a 4K7 ohm resistor should be connected
between the Data and VCC pins. The DHT22 output data pin will be connected to Raspberry
GPIO 16. Check the electrical diagram, connecting the sensor to RPi pins as below:
Do not forget to Install the 4K7 ohm resistor between the VCC and Data pins.
610
Once the sensor is connected, we must install its library on our Raspberry Pi. First, we
should install the Adafruit CircuitPython library, which we have already done, and the
Adafruit_CircuitPython_DHT.
Create a new Python script as below and name it, for example, dht_test.py:
import time
import board
import adafruit_dht
dhtDevice = adafruit_dht.DHT22(board.D16)
611
while True:
try:
# Print the values to the serial port
temperature_c = [Link]
temperature_f = temperature_c * (9 / 5) + 32
humidity = [Link]
print(
"Temp: {:.1f} F / {:.1f} C Humidity: {}% ".format(
temperature_f, temperature_c, humidity
)
)
[Link](2.0)
Placing a finger on the sensor, we can see that both temperature and humidity begin to rise.
612
613
Addendum: DHT22 on Raspberry Pi 5 with Debian Trixie
The DHT22 protocol is timing-sensitive: it communicates over a single wire using pulses as
short as 26 µs to encode each bit. Reading those pulses correctly requires a GPIO backend
with microsecond-level precision.
The adafruit_dht library offers two backends, both broken on Pi 5 + Trixie:
The pigpio daemon — which was the traditional workaround — was removed from Debian
Trixie’s official repositories. Compiling it from source is possible but fragile, and its DMA-based
approach does not work with the RP1 chip anyway.
The Linux kernel ships a dht11 driver that works with both DHT11 and DHT22 sensors.
It runs in kernel space using GPIO interrupts, capturing edge transitions with nanosecond
timestamps — the only approach reliable enough on Pi 5.
Step 1 — Add the overlay to /boot/firmware/[Link]
Add the following line, replacing 16 with whichever GPIO pin your sensor data line uses:
614
dtoverlay=dht11,gpiopin=16
sudo reboot
cat /sys/bus/iio/devices/iio:device0/name
You should see output like dht11@16. Then do a quick read to confirm the sensor responds:
cat /sys/bus/iio/devices/iio:device0/in_temp_input
cat /sys/bus/iio/devices/iio:device0/in_humidityrelative_input
Values are returned in millidegrees and millipercent. A reading of 22900 means 22.9°C; 55000
means 55.0% RH.
The number after @ in dht11@10 or dht11@16 is the internal device-tree unit address,
which does not always match the BCM GPIO number directly on the Pi 5 / RP1 chip.
What matters is that the device exists and returns readings — ignore the suffix.
import time
DHT_IIO = "/sys/bus/iio/devices/iio:device0/"
def read_dht22():
try:
temp_c = int(open(DHT_IIO + "in_temp_input").read()) / 1000.0
humidity = int(open(DHT_IIO + "in_humidityrelative_input").read()) / 1000.0
615
return temp_c, humidity
except Exception as e:
print(f"Read error: {e}")
return None, None
while True:
temperature_c, humidity = read_dht22()
if temperature_c is not None:
temperature_f = temperature_c * (9 / 5) + 32
print(
"Temp: {:.1f} F / {:.1f} C Humidity: {:.1f}% ".format(
temperature_f, temperature_c, humidity
)
)
[Link](3.0)
Run it:
python dht_test_trixie.py
Expected output:
Raspberry Pi 4
What (Bullseye/Bookworm) Raspberry Pi 5 (Trixie)
DHT22 library adafruit-circuitpython- Kernel dht11 IIO driver
dht
DHT22 read [Link] /sys/bus/iio/devices/iio:device0/
GPIO backend [Link] or pigpio lgpio (explicit
LGPIOFactory)
pigpio Available in apt Removed from Trixie repos
[Link] no DHT entry needed dtoverlay=dht11,gpiopin=16
616
LIGHTBULB Why is the kernel driver more reliable anyway?
The dht11 kernel driver uses GPIO interrupts: every rising and falling edge on the data
line triggers a hardware interrupt that is timestamped in nanoseconds inside the kernel.
There is no scheduler jitter, no PCIe round-trip overhead, and no competition from other
user-space threads. Even on Pi 4, this approach is more robust than bit-banging from
Python — on Pi 5, it is the only approach that works.
Sensor Overview:
Environmental sensing has become increasingly important in various industries, from weather
forecasting to indoor navigation and consumer electronics. At the forefront of this technological
advancement are sensors like the BMP280 and BMP180 (deprected), which excel in measuring
temperature and barometric pressure with exceptional precision and reliability.
As its predecessor, the BMP180, the BMP280 is an absolute barometric pressure sensor,
which is especially feasible for mobile applications. Its diminutive dimensions and low power
consumption allow for its implementation in battery-powered devices such as mobile phones,
GPS modules, or watches. The BMP280 is based on Bosch’s proven piezo-resistive pressure
sensor technology featuring high accuracy and linearity as well as long-term stability and high
EMC robustness. Numerous device operation options guarantee the highest flexibility. The
device is optimized for power consumption, resolution, and filter performance.
Technical data
617
Parameter Technical data
Temperature coefficient offset (+25°…+40°C @ 1.5 Pa/K, equiv. to 12.6 cm/K
900hPa)
Interface I²C and SPI
618
Enabling I2C Interface
Go to RPi Configuration and confirm that the I2C interface is enabled. If not, enable it.
619
If everything has been installed and connected correctly, you can turn on your Rapspi and
start interpreting the BMP180’s information about the environment.
The first thing to do is to check if the Raspi sees your BMP280. Try the following in a
terminal:
sudo i2cdetect -y 1
Create a new Python script as below and name it, for example, bmp280_test.py:
import time
import board
import adafruit_bmp280
i2c = board.I2C()
bmp280 = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address = 0x76)
620
bmp280.sea_level_pressure = 1013.25
while True:
print("\nTemperature: %0.1f C" % [Link])
print("Pressure: %0.1f hPa" % [Link])
print("Altitude = %0.2f meters" % [Link])
[Link](2)
python [Link]
Note that pressure is presented in hPa. See the next section to better understand
this unit.
621
Measuring Weather and Altitude With BMP280
Let’s take some time to understand more about what we will get with the BMP readings.
622
The BMP280 (and its predecessor, the BMP180) was designed to measure atmospheric pressure
accurately. Atmospheric pressure varies with both weather and altitude.
What is Atmospheric Pressure?
Atmospheric pressure is a force that the air around you exerts on everything. The weight of the
gasses in the atmosphere creates atmospheric pressure. A standard unit of pressure is “pounds
per square inch” or psi. We will use the international notation, newtons per square meter,
called pascals (Pa).
This weight, pressing down on the footprint of that column, creates the atmospheric pressure
that we can measure with sensors like the BMP280. Because that cm-wide column of air weighs
about 1 kg, the average sea level pressure is about 101,325 pascals, or better, 1013.25 hPa
(1 hPa is also known as milibar - mbar). This will drop about 4% for every 300 meters you
ascend. The higher you get, the less pressure you’ll see because the column to the top of the
atmosphere is much shorter and weighs less. This is useful because you can determine your
altitude by measuring the pressure and doing math.
The air pressure at 3,810 meters is only half that at sea level.
The BMP280 outputs absolute pressure in hPa (mbar). One pascal is a minimal amount of
pressure, approximately the amount that a sheet of paper will exert resting on a table. You will
often see measurements in hectopascals (1 hPa = 100 Pa). The library here provides outputs
of floating-point values in hPa, equaling one millibar (mbar).
Here are some conversions to other pressure units:
Temperature Effects
Because temperature affects the density of a gas, density affects the mass of a gas, and mass
affects the pressure (whew), atmospheric pressure will change dramatically with temperature.
Pilots know this as “density altitude”, which makes it easier to take off on a cold day than a hot
one because the air is denser and has a more significant aerodynamic effect. To compensate for
temperature, the BMP280 includes a rather good temperature sensor and a pressure sensor.
To perform a pressure reading, you first take a temperature reading, then combine that with a
raw pressure reading to come up with a final temperature-compensated pressure measurement.
(The library makes all of this very easy.)
623
Measuring Absolute Pressure
If your application requires measuring absolute pressure, all you have to do is get a temperature
reading, then perform a pressure reading (see the test script for details). The final pressure
reading will be in hPa = mbar. You can convert this to a different unit using the above
conversion factors.
Note that the absolute pressure of the atmosphere will vary with both your altitude
and the current weather patterns, both of which are useful things to measure.
Weather Observations
The atmospheric pressure at any given location on Earth (or anywhere with an atmosphere)
isn’t constant. The complex interaction between the earth’s spin, axis tilt, and many other
factors result in moving areas of higher and lower pressure, which in turn cause the variations
in weather we see every day. By watching for changes in pressure, you can predict short-term
changes in the weather. For example, dropping pressure usually means wet weather or a storm
is approaching (a low-pressure system is moving in). Rising pressure usually means clear
weather is coming (a high-pressure system is moving through). But remember that atmospheric
pressure also varies with altitude. The absolute pressure in my home, Lo Barnechea, in Chile
(altitude 960m), will always be lower than that in San Francisco (less than 2 meters, almost
sea level). If weather stations just reported their absolute pressure, it would be challenging
to compare pressure measurements from one location to another (and large-scale weather
predictions depend on measurements from as many stations as possible).
To solve this problem, weather stations continuously remove the effects of altitude from their
reported pressure readings by mathematically adding the equivalent fixed pressure to make it
appear that the reading was taken at sea level. When you do this, a higher reading in San
Francisco than in Lo Barnechea will always be because of weather patterns and not because of
altitude.
Sea Level Pressure Calculation
The See Level Pressure can be calculated with the formula:
Where,
po = SeaLevel Pressure
p = Atmospheric Pressure
L = Temperature Lapse Rate
h = Altitude
To = Sea Level Standard Temperature
g = Earth Surface Gravitational Acceleration
624
M = Molar Mass Of Dry Air
R = Universal Gas Constant
Having the absolute pressure in Pa, you check the sea level pressure using the
Calculator.
Or calculating in Python, where the altitude is the real altitude in meters where the sensor
is located.
Determining Altitude
Since pressure varies with altitude, you can use a pressure sensor to measure altitude (with a
few caveats). The average pressure of the atmosphere at sea level is 1013.25 hPa (or mbar).
This drops off to zero as you climb towards the vacuum of space. Because the curve of this
drop-off is well understood, you can compute the altitude difference between two pressure
measurements (p and p0) by using a specific equation. The BMP280 gives the measured
altitude using [Link].
The above explanation was based on the BMP 180 Sparkfun tutorial.
In this section, using the Jupyter Notebook, we will read sensors and act on actuators directly
on the Pi.
On the terminal, start the Jupyter notebook server with the command (change the IP address
with the one for your Raspi):
625
You will need the Token; you can copy it from the terminal as shown above.
http:localhost:8888
The first time you connect, you’ll need the token that appears in the Pi terminal
when you start the notebook server.
When you start your Pi and want to use Jupyter Notebook, type the “Jupyter
Notebook” command on your terminal and keep it running. This is very important!
If you need to use the terminal for another task, such as running a program, open
a new Terminal window.
To stop the server and close the “kernels” (the Jupyter notebooks), press [Ctrl] + [C].
626
Testing the Notebook setup
Let’s create a new notebook (Kernel: Python 3). Open dht_test.py, copy the code, and paste
it into the notebook. That’s it. We can see the temperature and humidity values appearing on
the cell. To interrupt the execution, go to the [stop] button at the top menu.
OK, this means we can access the physical world from our notebook! Let’s create a more
structured code for dealing with sensors and actuators.
Initialization
# time library
import time
import datetime
627
import adafruit_dht
DHT22Sensor = adafruit_dht.DHT22(board.D16)
# LEDs
from gpiozero import LED
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
[Link]()
[Link]()
[Link]()
# Push-Button
from gpiozero import Button
button = Button(20)
628
buttonSts = button.is_pressed
ledRedSts = ledRed.is_lit
ledYlwSts = ledYlw.is_lit
ledGrnSts = ledGrn.is_lit
[Link]()
[Link]()
[Link]()
629
If you press the push-button, its status will also be shown:
[Link]()
[Link]()
[Link]()
630
We can create a function to simplify turning LEDs on and off:
getGpioStatus()
PrintGpioStatus()
631
Getting and displaying Sensor Data
First, we should create a function to read the BMP280 and calculate the pressure value at sea
level, once the sensor only gives us the absolute pressure based on the actual altitude:
temp = [Link]
pres = [Link]
alt = [Link]
presSeaLevel = pres / pow(1.0 - real_altitude/44330.0, 5.255)
Entering the BMP280 real altitude where it is located, run the code:
bmp280GetData(960)
• Temperature of 26.9 oC
• Absolute Pressure of 906.73 hPa
632
• Measured Altitude (from Pressure) of 927 m
• Sea Level converted Pressure: 1,017.29 hPa
Now, we will generate a unique function to get the BMP280 and the DHT data, including a
timestamp:
tempDHT = [Link]
humDHT = [Link]
Runing them:
633
real_altitude = 960 # real altitude of where the BMP280 is installed
getSensorData(real_altitude)
printData()
Results:
Using Python, we can command the actuators (LEDs) and read the sensors and GIPOs status
at this stage. This is important, for example, to generate a data log to be read by an SLM in
the future.
� IMPORTANT: The Notebook Kernel should end after using to liberate GPIOs
The problem is that Jupyter Notebook is still holding onto the GPIO pins
even after our code finishes running. When we try to run it from the terminal,
those pins are already claimed by the Jupyter process. So, before running a code
that deals with GPIO in the terminal, In Jupyter:
Widgets
pywidgets, or jupyter-widgets orwidgets, are interactive HTML widgets for Jupyter note-
books and the IPython kernel. Notebooks come alive when interactive widgets are used. We
can gain control of our data and visualize changes in them.
Widgets are eventful Python objects that have a representation in the browser, often as a
control like a slider, text box, etc. We can use widgets to build interactive GUIs for our
project.
634
In this lab, for example, we will use a slide bar to control the state of actuators in real time,
such as by turning on or off the LEDs. Widgets are great for adding more dynamic behavior
to Jupyter Notebooks.
Installation
To use Widgets, we must install the Ipywidgets library using the commands:
# widget library
from ipywidgets import interactive
import ipywidgets as widgetsfrom
[Link] import display
And running the below line, we can control the LEDs in real-time:
This interactive widget is very easy to implement and very powerful. You can learn
more about Interactive on this link: Interactive Widget.
GPIO Zero includes many advanced features that make complex projects much easier to build.
635
PWM LED Brightness Control
led = PWMLED(13)
Composite Devices
while True:
[Link]()
sleep(2)
[Link]()
sleep(1)
[Link]()
[Link]()
[Link]()
sleep(3)
[Link]()
[Link]()
sleep(1)
[Link]()
636
• Motor: from gpiozero import Motor
• Servo: from gpiozero import Servo
• DistanceSensor: from gpiozero import DistanceSensor
• MotionSensor: from gpiozero import MotionSensor
• Robot: from gpiozero import Robot
This section demonstrates in a simple way how to integrate a Small Language Model (SLM)
with the sensors and LEDs we have set up. The diagram below shows how data flows from
sensors through processing and AI analysis to control the actuators and ultimately provide
user feedback.
637
We will use the Transformers library from Hugging Face for model loading and inference. This
library provides the architecture for working with pre-trained language models, helping interact
with the model, processing input prompts, and obtaining outputs.
Installation
Let’s create a simple SLM test in the Jupyter Notebook that checks if the model loads
and measures inference time. The model used here is the TinyLLama 1.1B. We will ask a
straightforward question:
As a result, besides the SLM answer, we will also measure the latency.
Run this script:
import time
from transformers import pipeline
import torch
model='TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T'
generator = pipeline('text-generation',
model=model,
device=device)
load_time = [Link]() - start_time
print(f"Model loading time: {load_time:.2f} seconds")
# Test prompt
test_prompt = "The weather today is"
638
num_return_sequences=1,
temperature=0.7)
inference_time = [Link]() - start_time
As we can see, the SLM works, but the latency is very high (+3 minutes). It is OK because this
particular test is on a Raspberry Pi 4. With a Raspberry Pi 5, the result would be better.:
The Raspi uses around 1GB of memory (model + process) and all four cores to process the
answer. The model alone needs around 800MB.
639
Now, let us create a code showing a basic interaction pattern where the SLM can respond to
sensor data and interact with the LEDs.
Install the Libraries:
import time
import datetime
import board
import adafruit_dht
import adafruit_bmp280
from gpiozero import LED, Button
from transformers import pipeline
Initialize sensors
DHT22Sensor = adafruit_dht.DHT22(board.D16)
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)
model='TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T'
generator = pipeline('text-generation',
model=model,
device='cpu')
Support Functions
Now, let’s create support functions for readings from all sensors and control the LEDs:
640
def get_sensor_data():
"""Get current readings from all sensors"""
try:
temp_dht = [Link]
humidity = [Link]
temp_bmp = [Link]
pressure = [Link]
return {
'temperature_dht': round(temp_dht, 1) if temp_dht else None,
'humidity': round(humidity, 1) if humidity else None,
'temperature_bmp': round(temp_bmp, 1),
'pressure': round(pressure, 1)
}
except RuntimeError:
return None
def process_conditions(sensor_data):
"""Process sensor data and control LEDs based on conditions"""
if not sensor_data:
control_leds(red=True) # Error condition
return
temp = sensor_data['temperature_dht']
humidity = sensor_data['humidity']
641
Generating an SLM’s response
So far, the LEDs reaction is only based on logic, but let’s also use the SLM to “analyse” the
sensors condition, generating a response based on that:
def generate_response(sensor_data):
"""Generate response based on sensor data using SLM"""
if not sensor_data:
return "Unable to read sensor data"
return response
Main Function
And now, let’s create a main() function to wait for the user to, for example, press a button
and, capture the data generated by the sensors, delivering some observation or recommendation
from the SLM:
def main_loop():
"""Main program loop"""
print("Starting Physical Computing with SLM Integration...")
print("Press the button to get a reading and SLM response.")
try:
while True:
if button.is_pressed:
# Get sensor readings
sensor_data = get_sensor_data()
642
# Process conditions and control LEDs
process_conditions(sensor_data)
if sensor_data:
# Get SLM response
response = generate_response(sensor_data)
except KeyboardInterrupt:
print("\nShutting down...")
control_leds(False, False, False) # Turn off all LEDs
Test Result
The sensors are read after the user presses the button to trigger a reading, and LEDs are
controlled based on conditions. Sensor data is formatted into a prompt for the SLM to generate
a response analyzing the current conditions. The results are displayed in the terminal, and the
LED indicators are shown.
This simple code integrates a Small Language Model (TinyLlama model (1.1B parameters)
with our physical computing setup, providing raw sensor data and intelligent responses from
the SLM about the environmental conditions.
643
We can extend this first test to more sophisticated and valuable uses of the SLM integration,
for example: adding:
Other Models
We can use other SLMs in a Raspberry Pi that have distinct ways of handling them. For
example, many modern models use GGUF formats, and to use them, we need to install
llama-cpp-python, which is designed to work with GGUF models.
Also, as we saw in a previous lab, Ollama is a great way to download and test SLMs on the
Raspberry Pi.
644
Conclusion
Key Achievements
Throughout this tutorial, we’ve successfully: - Set up a complete physical computing environ-
ment using Raspberry Pi - Integrated multiple environmental sensors (DHT22 and BMP280)
- Implemented visual feedback through LED actuators - Created interactive controls using
push buttons - Integrated a Small Language Model (TinyLLama 1.1B) for intelligent analysis -
Developed a foundation for AI-enhanced environmental monitoring
Technical Insights
Hardware Integration
The combination of digital (DHT22) and I2C (BMP280) sensors demonstrated different commu-
nication protocols and their implementations. This multi-sensor approach provides redundancy
and comprehensive environmental monitoring capabilities. The LED actuators and push-button
interface created a responsive and interactive system that bridges the digital and physical
worlds.
Software Architecture
AI Integration Learnings
The integration of TinyLLama 1.1B revealed several important insights: - Small Language
Models can effectively run on edge devices like Raspberry Pi - Natural language processing can
enhance sensor data interpretation - Real-time analysis is possible, though with some latency
considerations - The system can provide human-readable insights from complex sensor data
Practical Applications
645
Challenges and Solutions
2. Data Integration:
• Developed robust sensor data validation
• Created effective data preprocessing pipelines
• Implemented error handling for sensor failures
3. AI Integration:
• Designed effective prompting strategies
• Managed inference latency
• Balanced accuracy with response time
Future Enhancements
Final Thoughts
This chapter demonstrates that integrating physical computing with AI is feasible and practical
on readily accessible hardware such as the Raspberry Pi. Combining sensors, actuators, and
AI creates a powerful platform for developing intelligent environmental monitoring and control
systems.
While the current implementation focuses on environmental monitoring, the principles and
techniques can be adapted to various applications. The modular nature of hardware and
software components allows for customization and expansion based on specific needs.
Integrating small language models into physical computing opens new possibilities for creating
more intuitive and intelligent IoT devices. As edge AI capabilities evolve, projects like this
will become increasingly important in developing the next generation of smart devices and
systems.
646
Remember that this is just the beginning. Our foundation can be extended in countless ways
to create more sophisticated and capable systems. The key is to build on these basics while
balancing functionality, reliability, and resource usage.
Resources
• GPIOs - Scripts
• Sensors - Scripts
• Notebooks
647
Experimenting with SLMs for IoT Control
Introduction
This chapter explores the implementation of Small Language Models (SLMs) in IoT control
systems, demonstrating the possibility of creating a monitoring and control system using edge
AI. We’ll integrate these models with physical sensors and actuators, creating an intelligent
IoT system capable of natural language interaction. While this implementation shows
the potential of integrating AI with physical systems, it also highlights current limitations and
areas for improvement.
This chapter builds on the concepts introduced in “Small Language Models (SLMs)”
and “Physical Computing with Raspberry Pi.”
648
The Physical Computing chapter laid the groundwork for interfacing with hardware compo-
nents using the Raspberry Pi’s GPIO pins. We’ll revisit these concepts, focusing on connecting
and interacting with sensors (DHT22 for temperature and humidity, BMP280 for temperature
and pressure, and a push-button for digital inputs), as well as controlling actuators (LEDs) in
a more sophisticated setup.
We will progress from a simple IoT system to a more advanced platform that combines real-time
monitoring, historical data analysis, and natural language processing (NLP).
649
• Simple observation and reporting of system state
• Demonstration of SLM’s ability to interpret sensor data
3. Active Control Implementation
• Direct LED control based on SLM decisions
• Temperature threshold monitoring
• Emergency state detection via button input
• Real-time system state analysis
4. Natural Language Interaction
• Free-form command interpretation
• Context-aware responses
• Multiple SLM model support
• Flexible query handling
5. Data Logging and Analysis
• Continuous system state recording
• Trend analysis and pattern detection
• Historical data querying
• Performance monitoring
Let’s begin by setting up our hardware and software environment, building upon the foundation
established in our previous labs.
Setup
Hardware Setup
Connection Diagram
650
• Raspberry Pi 5 (with an OS installed, as detailed in previous labs)
• DHT22 temperature and humidity sensor
• BMP280 temperature and pressure sensor
• 3 LEDs (red, yellow, green)
• Push button
• 330Ω resistors (3)
• Jumper wires and breadboard
651
Software Prerequisites
Let’s create a Python script ([Link]) to handle the sensors and actuators. This script
will contain functions to be called from other scripts later:
import time
import board
import adafruit_dht
import adafruit_bmp280
from gpiozero import LED, Button
DHT22Sensor = adafruit_dht.DHT22(board.D16)
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25
ledRed = LED(13)
ledYlw = LED(19)
ledGrn = LED(26)
button = Button(20)
def collect_data():
try:
temperature_dht = [Link]
humidity = [Link]
temperature_bmp = [Link]
pressure = [Link]
button_pressed = button.is_pressed
return temperature_dht, humidity, temperature_bmp, pressure, button_pressed
except RuntimeError:
return None, None, None, None, None
def led_status():
ledRedSts = ledRed.is_lit
652
ledYlwSts = ledYlw.is_lit
ledGrnSts = ledGrn.is_lit
return ledRedSts, ledYlwSts, ledGrnSts
while True:
ledRedSts, ledYlwSts, ledGrnSts = led_status()
temp_dht, hum, temp_bmp, press, button_state = collect_data()
[Link](2)
653
IMPORTANT: Updated [Link] for Pi 5 + Trixie
When integrating the DHT22 into the full monitoring script alongside LEDs and BMP280, two
changes are needed: replace adafruit_dht with the sysfs reader, and explicitly set gpiozero’s
pin factory to LGPIOFactory (required on Pi 5).
import time
import board
import adafruit_bmp280
from gpiozero import LED, Button, Device
from [Link] import LGPIOFactory
# Required on Raspberry Pi 5
Device.pin_factory = LGPIOFactory()
def read_dht22():
try:
temp_c = int(open(DHT_IIO + "in_temp_input").read()) / 1000.0
humidity = int(open(DHT_IIO + "in_humidityrelative_input").read()) / 1000.0
return temp_c, humidity
except Exception:
return None, None
654
i2c = board.I2C()
bmp280Sensor = adafruit_bmp280.Adafruit_BMP280_I2C(i2c, address=0x76)
bmp280Sensor.sea_level_pressure = 1013.25
def collect_data():
temperature_dht, humidity = read_dht22()
try:
temperature_bmp = [Link]
pressure = [Link]
except Exception:
temperature_bmp, pressure = None, None
button_pressed = button.is_pressed
return temperature_dht, humidity, temperature_bmp, pressure, button_pressed
def led_status():
return ledRed.is_lit, ledYlw.is_lit, ledGrn.is_lit
Now, let’s create a new script, slm_basic_analysis.py, which will be responsible for analysing
the hardware components’ status, according to the following diagram:
655
The diagram shows the basic analysis system, which consists of:
1. Hardware Layer:
• Sensors: DHT22 (temperature/humidity), BMP280 (temperature/pressure)
• Input: Emergency button
• Output: Three LEDs (Red, Yellow, Green)
2. [Link]:
• Handles all hardware interactions
• Provides two main functions:
– collect_data(): Reads all sensor values
– led_status(): Checks current LED states
3. slm_basic_analysis.py:
• Creates a descriptive prompt using sensor data
• Sends prompt to SLM (for example, the Llama 3.2 1B)
• Displays analysis results
• In this step we will not control the LEDs (observation only)
Okay, let’s implement the code, starting for importing the Ollama library and the functions to
monitor the HW (from the previous script):
import ollama
from monitor import collect_data, led_status
656
ledRedSts, ledYlwSts, ledGrnSts = led_status()
temp_dht, hum, temp_bmp, press, button_state = collect_data()
Now, the heart of out code, we will generate the Prompt, using the data captured on the
previous variables:
prompt = f"""
You are an experienced environmental scientist.
Analyze the information received from an IoT system:
Where,
- The button, not pressed, shows a normal operation
- The button, when pressed, shows an emergency
- Red LED when is on, indicates a problem/emergency.
- Yellow LED when is on indicates a warning situation.
- Green LED when is on, indicates system is OK.
"""
Now, the Prompt will be passed to the SLM, which will generate a response:
MODEL = 'llama3.2:3b'
PROMPT = prompt
response = [Link](
model=MODEL,
prompt=PROMPT
)
The last stage will be show the real monitored data and the SLM’s response:
657
print(f"\nSmart IoT Analyser using {MODEL} model\n")
In this initial experiment, the system successfully collected sensor data (temperatures of
26.3°C and 26.1°C from DHT22 and BMP280, respectively, 40.2% humidity, and 908.84hPa
pressure) and processed this information through the SLM, which produced a coherent response
recommending the activation of the yellow LED due to elevated temperature conditions.
658
The model’s ability to interpret sensor data and provide logical, rule-based decisions shows
promise. Still, the simplistic nature of the current implementation (using basic thresholds and
binary LED outputs) suggests significant room for improvement through more sophisticated
prompting strategies, historical data integration, and the implementation of safety mechanisms.
Also, the result is probabilistic, meaning it should change after execution.
Let’s use the output generated and use it to actuate on the LEDs, our “actuators”:
We can add a new function, parse_llm_response(), to return a command to the the LEDs
based on the SLM’s response:
def parse_llm_response(response_text):
"""Parse the LLM response to extract LED control instructions."""
response_lower = response_text.lower()
red_led = 'activate red led' in response_lower
yellow_led = 'activate yellow led' in response_lower
green_led = 'activate green led' in response_lower
return (red_led, yellow_led, green_led)
659
red, yellow, green = parse_llm_response(response['response'])
Would be:(False, True, False), which can be the input to the function control_leds(red,
yellow, green)'
660
def output_actuator(response, MODEL):
print(f"\nSmart IoT Actuator using {MODEL} model\n")
By updating the system and calling the two functions in sequence, we can have the full cycle of
analysis by the SLM and the actuation on the output LEDs:
Let’s press the button, call the functions, and see what happens:
661
We can see that, despite the button being pressed, the SLM did not consider it and misinterpreted
the temperature value as ABOVE the threshold, not below. Also, despite the fact that we
asked for a straight answer about which LED to turn on, the model lost time with analysis,
This issue relates to the prompt we wrote. Let’s cover it in the next section.
662
Prompting Engineering
Looking at the answers we got with the previous implementation, which are not always correct,
we can see that the main issue is unreliable text parsing. Using JSON to parse the answer
should be much better!
663
3. Changed to JSON format - The prompt now asks for a structured JSON response:
{"red_led": true, "yellow_led": false, "green_led": false}
4. Updated parser - Now parses JSON instead of searching for text strings. Includes error
handling and fallback to safe state (all LEDs off) if parsing fails.
Let’s revise the previous code, with a better prompt and parcing funcion, now based on JSON.
This project demonstrates how a Small Language Model (SLM) can make intelligent decisions
in an IoT system. The system monitors environmental conditions using sensors and controls
LED indicators based on the SLM’s analysis.
664
New System Architecture
665
– Green LED: Normal operation
This layer provides the bridge between hardware and software with three key functions:
• collect_data(): Reads all sensor values (temperature, humidity, pressure) and button
state
• led_status(): Checks the current state of all LEDs
• control_leds(red, yellow, green): Controls which LED is turned on based on
boolean values
This is where the SLM makes decisions. The workflow follows these steps:
a) Data Collection & Preparation
b) Prompt Generation
c) SLM Inference
666
{
"red_led": false,
"yellow_led": false,
"green_led": true
}
d) Response Processing
e) Actuation
667
Configurable Threshold: The temperature threshold is a parameter that makes the system
adaptable to different environments without code changes. Can be updated by the user.
Clear Priority Rules: The system follows explicit priority rules that the SLM understands:
Now, this architecture showcases how Small Language Models can be integrated into IoT
systems to provide intelligent, context-aware decision-making while maintaining simplicity and
reliability.
import ollama
import json
from monitor import collect_data, led_status, control_leds
SENSOR DATA:
- DHT22 Temperature: {temp_dht:.1f}°C
- BMP280 Temperature: {temp_bmp:.1f}°C
- Humidity: {hum:.1f}%
- Pressure: {press:.2f}hPa
- Button: {"PRESSED" if button_state else "NOT PRESSED"}
668
CURRENT ANALYSIS:
- Button status: {"PRESSED" if button_state else "NOT PRESSED"}
- DHT22 temp ({temp_dht:.1f}°C) is {"OVER" if temp_dht > TEMP_THRESHOLD
else "AT OR BELOW"} threshold ({TEMP_THRESHOLD}°C)
- BMP280 temp ({temp_bmp:.1f}°C) is {"OVER" if temp_bmp > TEMP_THRESHOLD
else "AT OR BELOW"} threshold ({TEMP_THRESHOLD}°C)
Based on these rules, respond with ONLY a JSON object (no other text):
{{"red_led": true, "yellow_led": false, "green_led": false}}
Only ONE LED should be true, the other two must be false.
"""
def parse_llm_response(response_text):
"""Parse the LLM JSON response to extract LED control instructions."""
try:
# Clean the response - remove any markdown code blocks if present
response_text = response_text.strip()
if response_text.startswith('```'):
# Extract JSON from markdown code block
lines = response_text.split('\n')
response_text = '\n'.join(lines[1:-1])
if len(lines) > 2
else response_text
# Parse JSON
data = [Link](response_text)
red_led = [Link]('red_led', False)
yellow_led = [Link]('yellow_led', False)
green_led = [Link]('green_led', False)
return (red_led, yellow_led, green_led)
except ([Link], KeyError) as e:
print(f"Error parsing JSON response: {e}")
print(f"Response was: {response_text}")
# Fallback to safe state (all LEDs off)
return (False, False, False)
669
print(f" - DHT22 ==> Temp: {temp_dht:.1f}°C, Humidity: {hum:.1f}%")
print(f" - BMP280 => Temp: {temp_bmp:.1f}°C, Pressure: {press:.2f}hPa")
print(f" - Button {'pressed' if button_state else 'not pressed'}")
670
Code Flow Diagram
Definitions
# Model to be used
MODEL = 'llama3.2:3b'
671
Calling the program
slm_analyse_act(MODEL, TEMP_THRESHOLD)
Result
672
TEST2: Temp below the threshold
slm_analyse_act(MODEL, TEMP_THRESHOLD)
673
TEST3: Alarm Button pressed
slm_analyse_act(MODEL, TEMP_THRESHOLD)
Now, let’s transform our IoT monitoring setup into an interactive assistant that accepts
natural language commands and queries. Instead of autonomously monitoring conditions, the
674
system will now respond to our requests in real-time.
675
How It will Work
1. Interactive Loop
User Input → Sensor Reading → SLM Analysis → LED Control → Status Display → Wait for Next Inp
2. Dual-Purpose Response
{
"message": "Helpful text response to the user",
"leds": {
"red_led": false,
"yellow_led": true,
"green_led": false
}
}
3. Command Types
A. Information Queries
Ask questions about sensor readings - LEDs remain unchanged.
Examples:
Response format:
{
"message": "The current temperature is 21.5°C from DHT22 and 22.3°C from BMP280.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false} // keeps current state
}
676
B. Direct LED Commands
Tell the system which LEDs to turn on/off.
Examples:
Response format:
{
"message": "Yellow LED turned on.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false}
}
C. Conditional Commands
Commands that depend on sensor readings or button state.
Examples:
Response format:
{
"message": "Temperature is 21.5°C, which is above 20°C. Yellow LED turned on.",
"leds": {"red_led": false, "yellow_led": true, "green_led": false}
}
D. Toggle/Switch Commands
Commands that change LED states based on current conditions.
Examples:
Response format:
677
{
"message": "Button is pressed. LED states switched.",
"leds": {"red_led": true, "yellow_led": false, "green_led": false} // inverted from curren
}
E. Analysis Queries
Ask the SLM to analyze sensor data.
Examples:
Response format:
{
"message": "Based on pressure of 910.18hPa and humidity of 28.8%, conditions are dry. Rain
"leds": {"red_led": false, "yellow_led": false, "green_led": true} // keeps current state
}
Key Functions
create_interactive_prompt()
parse_interactive_response()
678
display_system_status()
interactive_mode()
python slm_act_leds_interactive.py
Tests
============================================================
IoT Environmental Monitoring System - Interactive Mode
Using Model: llama3.2:3b
============================================================
679
- Will it rain based on current conditions?
- Type 'status' to see system status
- Type 'exit' or 'quit' to stop
============================================================
You: status
============================================================
SYSTEM STATUS
============================================================
DHT22 Sensor: Temp = 21.5°C, Humidity = 28.8%
BMP280 Sensor: Temp = 22.3°C, Pressure = 910.18hPa
Button: NOT PRESSED
LED Status:
Red LED: � OFF
Yellow LED: � ON
Green LED: � OFF
============================================================
You: exit
Exiting interactive mode. Goodbye!
680
Special Commands
Languages
One advantage of using SLMs is that, once multilingual models are used (as Llama 3,2 or
Gemma), the user can choose the better language for them, independent of the language used
during coding.
Error Handling
• If sensor data cannot be read, the system will notify you and wait for the next command
• If the SLM response cannot be parsed, the system keeps LEDs in their current state
681
Tips for Best Results
1. Be specific: “Turn on the yellow LED” works better than “turn on the light”
2. Use conditions clearly: “If temperature is above 20°C” is clearer than “when it’s hot”
3. Ask for status: Use the status command frequently to verify system state
4. One LED at a time: Unless you specifically say “all LEDs”, the system defaults to one
LED on
Flow Diagram
The development of our Interactive IoT-SLM system can also be followed using the
Jupyter Notebook: SLM_IoT.ipynb.
682
Adding Data Logging
Now, we will develop an enhanced version that adds data logging, analysis, and historical query
capabilities. The system automatically logs all sensor readings and commands, allowing it to
analyze trends and query historical data using natural language.
Key Features
683
3. Statistical Analysis
• Min/max/average calculations
• LED state change tracking
• Button press counting
• Trend analysis
python slm_act_leds_with_logging.py
Example Queries
Real-time:
Historical:
Built-in commands:
684
Data Files
sensor_readings.csv
timestamp,temp_dht,humidity,temp_bmp,pressure,
button_pressed,red_led,yellow_led,green_led
command_history.csv
timestamp,user_command,slm_response,red_led,yellow_led,green_led
Tips
The complete scripts for the datalogger version arehere: data_logger.py and
slm_act_leds_with_logging.py
685
Flow Diagram
686
Examples:
687
Prompt Optimization and Efficiency
688
print(f"eval_rate: {response['eval_count']/(response['eval_duration']/1e9):.2f} \
tokens/s")
We will get:
Based on the prompt_eval_duration value, the PROMPT is our main bottleneck, so we must
make it as concise as possible. The SLM must process all of the context before it can generate
the first output token.
Quick Solution:
• Condense System Status: Remove unnecessary descriptive text and present the status
information in a compact, structured format.
– Example (Before):
CURRENT SYSTEM STATUS:
- DHT22: Temperature 25.5°C, Humidity 45.1%
- BMP280: Temperature 25.6°C, Pressure 1012.34hPa
- Button: NOT PRESSED
- Red LED: OFF
- Yellow LED: OFF
- Green LED: ON
– Example (After):
STATUS: DHT22=22.5°C/65.0% BMP280=22.3°C/1013.25hPa Button=OFF LEDs:R=OFF/Y=ON/G=OFF
689
• Reduce Examples: While examples are crucial for instruction-following, eliminate
redundancy. Keep only the most diverse and representative examples. The current
prompt is very long due to verbose examples and instructions. Focus on the single-shot
example that shows the required JSON output format.
• Simplify Instructions: Make the instructions as direct and short as possible. Use
keywords instead of full sentences where clarity is maintained.
• System message: It defines the assistant’s behavior and should sent once at initialization,
not at PROMPT
SYSTEM_MESSAGE = """You are an IoT assistant controlling an environmental
monitoring system with LEDs.
RULES:
Always respond with valid JSON containing both "message" and "leds" fields."""
690
Model Pre-loading
New Code
Original Functions:
New functions:
691
The Prompt evaluation time was reduced drastically!
Using Pydantic
Using Pydantic is a robust way to improve the reliability, efficiency, and maintainability of
our system, mainly since we rely on the LLM to output precise JSON.
Here’s how Pydantic can help and what you would need to do:
692
1. How Pydantic Reduces Latency and Improves Reliability
Pydantic doesn’t directly speed up the model’s token generation, but it can indirectly reduce
latency and eliminate error-handling overhead by enabling cleaner, faster parsing and
more robust communication.
• JSON Schema: Pydantic can generate a JSON Schema from your Python classes.
We can include this schema directly in our prompt, which acts as an unambiguous,
machine-readable instruction for the SLM. This often leads to fewer errors in the model’s
output, reducing the need for costly retries or complex string manipulation.
693
Step 1: Define the Pydantic Models
class LEDControl(BaseModel):
"""LED control configuration."""
red_led: bool = Field(description="Red LED state (on/off)")
yellow_led: bool = Field(description="Yellow LED state (on/off)")
green_led: bool = Field(description="Green LED state (on/off)")
class AssistantResponse(BaseModel):
"""Complete assistant response with message and LED control."""
message: str = Field(description="Helpful response to the user")
leds: LEDControl = Field(description="LED control configuration")
def parse_interactive_response(response_text):
"""Parse the interactive SLM response using Pydantic (guaranteed valid)."""
try:
# Parse directly into Pydantic model - guaranteed valid JSON structure
694
data = AssistantResponse.model_validate_json(response_text)
except Exception as e:
print(f"Error parsing response: {e}")
print(f"Response was: {response_text}")
return "Error: Could not parse SLM response.", (False, False, False)
With such modifications, the final latency was reduced from 90 to around 60 seconds
Note that when we sent two different commands to turn on the LEDs, the new one did not
turn off the previous one. This is due to the change in the rules. If we want, we can modify it
to match whatever we wish to.
695
The final code can be found on GitHub: slm_act_leds_interactive_pydantic.py
and in notebook SLM_IoT.ipynb
In short, using Pydantic over tradicional approuch with JSON, we have as benefits:
Next Steps
This chapter involved experimenting with simple applications and verifying the feasibility of
using an SLM to control IoT devices. The final result is far from usable in the real world, but
it can serve as a starting point for more interesting applications. Below are some observations
and suggestions for improvement:
696
• Some simple commands could be handled without SLM intervention. We can do it
programmatically.
• Consider implementing a proper state machine for LED control to ensure consistent
behavior.
• Implement more sophisticated trend analysis using statistical methods.
• Add support for more complex queries combining multiple data points.
697
Conclusion
This chapter has demonstrated the progressive evolution of an IoT system from basic sensor
integration to an intelligent, interactive platform powered by Small Language Models. Through
our journey, we’ve explored several key aspects of combining edge AI with physical computing:
Key Achievements
1. SLM Reliability
• Probabilistic nature of responses
• Consistency issues in decision making
698
• Need for better validation and verification
2. System Performance
• Response time considerations
• Resource usage on edge devices
• Efficiency of data logging and analysis
3. Architectural Constraints
• Simple state management
• Basic error handling
• Limited data validation
Final Thoughts
While this implementation demonstrates the potential of combining SLMs with IoT systems,
it also highlights the exciting possibilities and challenges ahead. Though experimental, the
system we’ve built provides a solid foundation for understanding how edge AI can enhance IoT
applications. As SLMs evolve and improve, their integration with physical computing systems
will likely become more robust and practical for real-world applications.
This chapter has shown that, despite current limitations, SLMs can provide intelligent, natural-
language interfaces to IoT systems, opening new possibilities for human-machine interaction in
the physical world.
The future of IoT systems is shaped by intelligent, edge-based solutions that combine AI’s
power with the practicality of physical computing.
Resourses
699
Advancing EdgeAI: Beyond Basic SLMs
Figure 17: Image from author, with a Raspberry Pi from ImageFX - prompt - Create a cartoon
image with a single Raspberry Pi with a white background
700
Understanding SLM Limitations
Small Language Models, while impressive in their ability to run on edge devices, face several
key limitations:
1. Knowledge Constraints
SLMs have limited knowledge based on their training data, often outdated and incomplete.
Unlike their larger counterparts, they cannot store the vast information needed for comprehensive
expertise across all domains.
Let’s run the below example to verify this limitation.
import ollama
response = [Link](
model="llama3.2:1b",
prompt="Who won the 2024 Summer Olympics men's 100m sprint final?"
)
print(response['response'])
The output of the previous code will likely show hallucination or admission of not knowing, as
in the case below:
This constraint could be solved simply by having an Agent search the Internet for the answer
or using Retrieval-Augmented Generation (RAG), as we will see later.
701
2. Reasoning Limitations
Complex reasoning tasks often exceed the capabilities of SLMs, which struggle with multi-
step logical deductions, mathematical computations, and a nuanced understanding of context.
Agents can be used to mitigate such limitations.
For example, let’s reuse the previous code and ask to the SLM to multiply two numbers :
import ollama
response = [Link](
model="llama3.2:3b",
prompt="Multiply 123456 by 123456"
)
print(response['response'])
The response is wrong; once the multiplication result should be 15,241,383,936. This is
expected once the language models are not suitable for mathematical computations. Still, we
can use an “agent” to determine whether a user asks for multiplication or a general question.
We will learn how to create an agent later.
3. Inconsistent Outputs
SLMs may produce inconsistent responses to the same query, making them unreliable for
critical applications requiring deterministic outputs. Several enhancements, such as Function
Calling and Response Validation, can improve reliability.
4. Domain Specialization
SLMs perform worse than specialized models in domain-specific tasks like visual recognition or
time-series analysis. Fine-tuning can adapt models to specific domains or tasks, improving
performance for targeted applications.
702
Techniques for Enhancing SLM at the Edge
Small Language Models (SLMs) offer remarkable capabilities for edge devices, but various
techniques can significantly enhance their effectiveness. Here, we present a comprehensive
framework for optimizing SLMs on resource-constrained devices like the Raspberry Pi, organized
from fundamental to advanced approaches.
We will divide those technics into 3 segments:
– Chain-of-Thought Prompting
– Few-Shot Learning
– Task Decomposition
The true power of these techniques emerges when they’re strategically combined:
1. Agent Architecture with RAG: Create agents that can access both tools and knowl-
edge bases
2. Validation-Enhanced RAG: Apply response validation to ensure RAG outputs are
accurate
3. Fine-Tuned Routers: Use specialized fine-tuned models to handle routing decisions
4. Chain-of-Thought with Function Calling: Combine reasoning traces with structured
outputs
703
• Function calling to structure sensor data analysis
• Response validation to verify recommendations
• Task decomposition to handle complex multi-part weather analysis
Chain-of-Thought Prompting
Chain-of-thought prompting encourages SLMs to break down complex problems into step-by-
step reasoning, leading to more accurate results:
Problem:
{problem}
"""
response = [Link](model="llama3.2:3b", prompt=prompt)
return response["response"]
Few-Shot Learning
Few-shot learning provides examples within the prompt, helping SLMs understand the expected
response format and reasoning pattern:
def classify_sentiment(text):
prompt = f"""
Task: Classify the sentiment of the text as positive, negative, or neutral.
Examples:
704
Text: "I love this product, it works perfectly!"
Sentiment: positive
Text: "This is the worst experience I've ever had."
Sentiment: negative
Text: "The package arrived on time."
Sentiment: neutral
Text: "{text}"
Sentiment:
"""
response = [Link](model="llama3.2:1b", prompt=prompt)
return response['response'].strip()
This approach is particularly effective for classification tasks and standardized outputs.
Task Decomposition
For complex tasks, breaking them into smaller subtasks helps SLMs manage complexity:
def analyze_product_review(review):
# Step 1: Extract main points
points_prompt = f"Extract the main points from this product review: {review}"
points_response = [Link](model="llama3.2:1b", prompt=points_prompt)
main_points = points_response['response']
# Final synthesis
final_prompt = f"""
Create a concise analysis of this product review based on:
Main points: {main_points}
705
Overall sentiment: {sentiment}
Improvement suggestions: {improvements}
"""
final_response = [Link](model="llama3.2:1b",
prompt=final_prompt)
return final_response['response']
This technique distributes cognitive load across multiple simpler prompts, enabling SLMs to
handle tasks that might otherwise exceed their capabilities.
To address some of these limitations, we can develop agents that leverage SLMs as part of a
more extensive system with additional capabilities.
Let’s think about the multiplication problem that we faced before. A minimal agent can be
used for that.
An agent is a system that uses an AI Model as its core reasoning engine to:
706
For example, if it is a multiplication, we can use a Python function as a “tool” to calculate it,
as shown in the diagram:
707
This example shows a minimal agentic workflow: the model decides whether to use
a tool for multiplication or answer normally.
1. User Input: The user types a query like “What is 7 times 8?” or “What is the capital
of France?”
2. Process Query: The process_query() function handles the input and decides what to
do with it.
708
3. Classification: The ask_ollama_for_classification() function sends the user’s
query to the SLM (using Ollama) with a prompt asking it to classify whether the query
is requesting multiplication or asking a general question.
4. Decision: Based on the SLM’s classification:
• If it’s a multiplication request, the SLM also extracts the numbers, and we use our
multiply() function.
• If it’s a general question, we send the original query to the SLM for a direct answer.
5. Response: The system returns either the multiplication result or the SLM’s answer to
the general question.
Here’s a Python script that creates a simple agent (or router) between multiplication operations
and general questions as described:
import requests
import json
# Configuration
OLLAMA_URL = "[Link]
MODEL = "llama3.2:3b" # You can change this to any model you have installed
VERBOSE = True
def ask_ollama_for_classification(user_input):
"""
Ask Ollama to classify whether the query is a multiplication request or a \
general question.
"""
classification_prompt = f"""
Analyze the following query and determine if it's asking for multiplication \
or if it's a general question.
Query: "{user_input}"
If it's asking for multiplication, respond with a JSON object in this format:
{{
"type": "multiplication",
"numbers": [number1, number2]
}}
709
}}
try:
if VERBOSE:
print(f"Sending classification request to Ollama")
response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": classification_prompt,
"stream": False
}
)
if response.status_code == 200:
response_text = [Link]().get("response", "").strip()
if VERBOSE:
print(f"Classification response: {response_text}")
710
except Exception as e:
if VERBOSE:
print(f"Error connecting to Ollama: {str(e)}")
return {"type": "general_question"}
def ask_ollama(query):
"""
Send a query to Ollama for general question answering.
"""
try:
if VERBOSE:
print(f"Sending query to Ollama")
response = [Link](
f"{OLLAMA_URL}/generate",
json={
"model": MODEL,
"prompt": query,
"stream": False
}
)
if response.status_code == 200:
return [Link]().get("response", "")
else:
return f"Error: Received status code {response.status_code} \
from Ollama."
except Exception as e:
return f"Error connecting to Ollama: {str(e)}"
def process_query(user_input):
"""
Process the user input by first asking Ollama to classify it,
then either performing multiplication or sending it back as a
711
general question.
"""
# Let Ollama classify the query
classification = ask_ollama_for_classification(user_input)
if VERBOSE:
print("Ollama classification:", classification)
if [Link]("type") == "multiplication":
numbers = [Link]("numbers", [0, 0])
if len(numbers) >= 2:
return multiply(numbers[0], numbers[1])
else:
return "I understood you wanted multiplication, but couldn't \
extract the numbers properly."
else:
return ask_ollama(user_input)
def main():
"""
Main function to run the agent interactively.
"""
print("Ollama Agent (Type 'exit' to quit)")
print("-----------------------------------")
while True:
user_input = input("\nYou: ")
response = process_query(user_input)
print(f"\nAgent: {response}")
# Example usage
if __name__ == "__main__":
# Set to True to see detailed logging
VERBOSE = True
main()
When we run the script, we can see that, first, the SLM chooses multiplication, passing the
712
numbers entered by the user to the “tool,” which, in this case, is the multiply() function. As
a result, we got 15,241,383,936, which it is correct.
Let’s now enter with another question that has no relation with arithmetic, for example: What
is the capital of Brazil? In this case, the SLM will decide that the query is a general
question and pass it on to the SLM to answer it.
This simple agent (or router) demonstrates the fundamental concept of using an SLM to make
decisions about processing different types of user inputs. It shows both the power of SLMs for
natural language understanding and their limitations in structured tasks.
713
Limitations and Considerations
This agent seems to resolve our problem, but it has several limitations that are common when
working with SLMs:
1. JSON Parsing Issues: SLMs don’t always perfectly format JSON responses as requested.
The code includes error handling for this.
2. Classification Reliability: The SLM might not always correctly classify the query,
especially with ambiguous questions.
3. Number Extraction: The SLM might extract numbers incorrectly or miss them entirely.
4. Error Handling: Robust error handling is essential when working with SLMs because
their outputs can be unpredictable.
5. Latency: Significant latency is involved in making multiple calls to the SLM. For example,
for the above simple agent, the latency was about 50s when using the llama3.2:3B on a
Raspberry Pi 5.
Here, you can see the SLM latency (simple query) per device (in tokens/s):
In my simple tests, the 1B models struggled to classify the tasks correctly. The the
3B and 4B models worked fine
Improvements
1. Expand Capabilities: Add support for more operations (addition, subtraction, division).
2. Better Error Handling: Improve fallback mechanisms when the SLM fails to extract
numbers or classify correctly.
3. Model Preloading: Initialize the model at startup to reduce latency.
4. Adding Regex Fallbacks: Use regular expressions as a fallback to extract numbers
when the SLM fails.
5. Context Preservation: Maintain conversation context for multi-turn interactions.
A more robust script can be used with the above improvements. The diagram shows how it
would work:
714
Figure 18: image-20250322113135882
1. Initialization:
• The system starts by initializing both models in parallel threads
• This prevents cold starts and reduces latency
2. Query Processing Flow:
715
• User input is first sent to a classification step
• A model (llama3.2:3B) determines if it’s a calculation or a general question (we can
choose a different model here).
• If it’s a calculation:
– The system extracts the operation type and numbers
– Numbers are converted from strings to floats
– The appropriate calculation is performed
– Results are formatted with comma separators (e.g., 1,234,567.89)
• If it’s a general question:
– The query is sent to the main model (llama3.2:3b) for answering (we can choose
a different model here)
3. Optimizations (highlighted in the subgraph):
• Persistent HTTP session for connection reuse
• Keep-alive parameter to prevent model unloading
• Simplified classification prompt for faster processing
• Using a smaller model for the classification task
• Rule-based fallback logic if the model classification fails
This approach maintains the intelligent classification capability while significantly reducing
execution time compared to the original implementation.
Runing the script [Link], we get correct results with reduced latency
of about 60%.
716
General Knowledge Router
Remember when we asked our SLM: Who won the 2024 Summer Olympics men's 100m
sprint final? We could not receive an answer because the modes were trained with
information previously in late 2023.
To solve this issue, let’s build a more advanced agent to classify whether it should use its
knowledge to answer a question or fetch updated information from the Internet. This addresses
a key limitation of Small Language Models: their knowledge cutoff date.
The general architecture of our agent will be similar to the calculator, but now, we will use a
web search API as a tool.
717
This agent addresses a critical limitation of SLMs - their knowledge cutoff date - by determining
when to use the model’s built-in knowledge versus when to search for up-to-date information
from the web.
How it works:
Uses SLM for Classification: Relies entirely on the SLM to determine whether a query
needs web search or can be answered from the model’s knowledge.
Provides Date Context: This section supplies the current date to help the SLM make
informed decisions about whether information is outdated.
Integrates Tavily Search: Uses Tavily’s powerful search API to find relevant information
for queries that need external data.
Handles Timeouts: Includes fallback mechanisms when the model takes too long to respond.
Maintains Source Attribution: Clearly indicates to the user whether the answer comes
from the model’s knowledge or web search.
Let’s run the script: [Link]
But first, we should install the required libraries:
718
Runing the script and entering with the same questions that could not be answered before,
we now have: Noah Lyles won the men's 100m sprint final at the 2024 Summer
Olympics. He set a new personal best time of 9.79 seconds. This victory
marked the United States' first win in the event since 2004.
When the user enters a common-knowledge question, the agent will send it directly to the SLM.
For example, if the user asks, "Who is Albert Einstein?", we get:
719
Improving Agent Reliability
There are several ways to enhance an agent’s reliability. One is to implement effective, approved,
structured function calling, which makes agents’ responses more consistent and predictable.
In the SLM chapter, we explored function calling when we created an app where the user enters
a country’s name and gets, as an output, the distance in km from the capital city of such a
country and the app’s location.
720
Once the user enters a country name, the model will return the name of its capital city (as a
string) and the latitude and longitude of such city (in float). Using those coordinates, the app
used a simple Python library (haversine) to calculate the distance between those 2 points.
The critical library used was Pydantic (and instructor), a robust data validation and settings
management library engineered by Python to enhance the robustness and reliability of our
codebase. In short, Pydantic helps ensure that the model’s response will always be consistent.
Function calling can improve an agent’s reliability by ensuring structured outputs and clear
tool selection logic. Here’s a generic template about how we can implement it :
import time
from haversine import haversine
from pydantic import BaseModel, Field
from ollama import chat
class CityCoord(BaseModel):
city: str = Field(..., description="Name of the city")
lat: float = Field(..., description="Decimal Latitude of the city")
lon: float = Field(..., description="Decimal Longitude of the city")
resp = CityCoord.model_validate_json([Link])
721
end_time = time.perf_counter() # End timing
elapsed_time = end_time - start_time # Calculate elapsed time
# Test
calc_dist('france')
calc_dist('colombia')
calc_dist('united states')
2. Response Validation
• Response Relevancy: Determines if the LLM output addresses the input informatively
and concisely.
• Prompt Alignment: Check if the LLM output follows instructions from the prompt
template.
• Correctness: Assesses factual accuracy based on ground truth.
• Hallucination Detection: Identifies fake or made-up information in LLM outputs.
Adding validation prevents incorrect or harmful responses, and here, we can test it with a
simple script:
import ollama
import json
722
def validate_response(query, response):
"""Validate that the response is appropriate for the query"""
validation_prompt = f"""
User query: {query}
Generated response: {response}
try:
validation = [Link](
model="llama3.2:3b",
prompt=validation_prompt
)
result = [Link](validation['response'])
return result
except Exception as e:
print(f"Error during validation: {e}")
return {"valid": False, "reason": "Validation error", "score": 0}
# Test
query = "What is the Raspberry Pi 5?"
response = "It is a pie created with raspberry and cooked in an oven"
validation = validate_response(query, response)
print(validation)
723
Retrieval-Augmented Generation (RAG)
RAG systems enhance Small Language Models (SLMs) by providing relevant information from
external sources before generation. This is particularly valuable for edge devices with limited
model sizes, as it allows them to access knowledge beyond their training data without increasing
the model size.
Understanding RAG
In a basic interaction between a user and a language model, the user asks a question, which is
sent as a prompt to the model. The model generates a response based solely on its pre-trained
knowledge. In a RAG process, there’s an additional step between the user’s question and the
model’s response. The user’s question triggers a retrieval process from a knowledge base.
724
The RAG process consists of these key steps:
1. Query Processing: When a user asks a question, the system converts it into an
embedding (a numerical representation).
2. Document Retrieval: The system searches a knowledge base for documents with similar
embeddings.
3. Context Enhancement: Relevant documents are retrieved and combined with the
original query.
4. Generation: The SLM generates a response using both the query and the retrieved
context.
725
• Accepts user queries
• Retrieves relevant documents based on query similarity
• Combines documents with the query in a prompt
• Generates a response using the SLM
Instalation
Let’s examine how these components work together to implement a RAG system on edge
devices.
1. Document Processing
726
def create_vectorstore():
# Load documents from PDFs and URLs
docs_list = []
# [Document loading code]
# Persist to disk
[Link]()
This function processes our documents (chunk size of 300 with an overlap of 30), creating a
searchable knowledge base. Notice we’re using OllamaEmbeddings with the nomic-embed-text
model, which can run efficiently on edge devices like the Raspberry Pi.
print(f"Question: {question}")
print("Retrieving documents...")
docs = [Link](question)
docs_content = "\n\n".join(doc.page_content for doc in docs)
print(f"Retrieved {len(docs)} document chunks")
print("Generating answer...")
727
client = Client()
rag_prompt = client.pull_prompt("rlm/rag-prompt")
end_time = [Link]()
latency = end_time - start_time
print(f"Response latency: {latency:.2f} seconds using model: {local_llm}")
return answer
This function retrieves relevant documents based on the query and combines them with a
specialized RAG prompt to generate a response. The RAG prompt is particularly important
as it tells the model how to use the context documents to answer the question.
1. SLM Integration
We’re using Ollama to run the SLM locally on our edge device, in this case using the 3B
parameter version of Llama 3.2.
1. Knowledge Extension: RAG allows small models to access knowledge beyond their
training data, effectively extending their capabilities without increasing model size.
2. Reduced Hallucination: By providing factual context, RAG significantly reduces the
likelihood of SLMs generating incorrect information.
3. Up-to-date Information: Unlike the fixed knowledge in a model’s weights, RAG
knowledge bases can be updated regularly with new information.
4. Domain Specialization: RAG can make general SLMs perform like domain specialists
by providing domain-specific knowledge bases.
728
5. Resource Efficiency: RAG allows smaller models (which require less memory and
computation) to achieve performance comparable to much larger models.
When implementing RAG on resource-constrained edge devices like the Raspberry Pi, consider
these optimizations:
1. Chunk Size: Smaller chunks (300-500 tokens) reduce memory usage during retrieval
and generation.
2. Retrieval Limits: Limit the number of retrieved documents (k=3 to 5) to reduce context
size.
3. Embedding Model Selection: Choose lightweight embedding models like nomic-
embed-text (137M parameters) or all-minilm (23M parameters).
4. Persistent Storage: As shown in our examples, using persistent storage prevents
recomputing embeddings every time that the RAG system is initiated.
5. Query Optimization: Implement query preprocessing to improve retrieval accuracy
while reducing computational load.
def optimize_query(query):
"""Optimize the query for better retrieval results"""
# Remove filler words, focus on key terms
stop_words = {"and", "or", "the", "a", "an", "in", "on", "at", "to",
"for", "with"}
terms = [term for term in [Link]().split() if term not in stop_words]
return " ".join(terms)
Building on our advanced weather station (see the chapter “Experimenting with SLMs for IoT
Control”), we can, for example, integrate RAG to provide more contextual responses about
weather conditions and historical patterns:
729
def weather_station_with_rag(retriever, model="llama3.2:3b"):
# Get current sensor readings
temp_dht, humidity, temp_bmp, pressure, button_state = collect_data()
730
prompt = f"""
Current Weather Station Data:
- Temperature (DHT22): {temp_dht:.1f}°C
- Humidity: {humidity:.1f}%
- Pressure: {pressure:.2f}hPa
Reference Information:
{context}
return [Link]
This function enhances our weather station by providing context-aware responses incorporating
current sensor readings and relevant information from our knowledge base. This is only an
example. To use it, we should have “Weather Reference Data,” which we do not currently have.
Instead, let’s create a general RAG system specializing in Edge AI Engineering.
For our RAG system, we will create a database with all chapters alheady written for the
EdgeAI Engineering book (chapters as URLs) and a PDF Wevolver 2025 Edge AI Technology
Report.
731
"[Link]
"[Link]
/counting_objects_yolo.html",
"[Link]
"[Link]
"[Link]
/RPi_Physical_Computing.html",
"[Link]
]
Using the RAG system is straightforward. First, ensure you’ve created the vector database:
732
# Start the interactive query interface
python [Link]
Example interactions:
Generating answer...
ANSWER:
==================================================
EdgeAI refers to the application of artificial intelligence (AI) at the edge
of a network, typically in real-time applications such as IoT sensors, industrial
robots, and smart cameras. The Edge AI ecosystem includes edge devices, edge
servers, and cloud platforms that work together to enable low-latency AI
inferencing and processing of data on-site without relying on continuous cloud connectivity.
in AI by leveraging energy-efficient, affordable, and scalable solutions for machine
learning and advanced edge computing.
==================================================
733
Those responses demonstrate how RAG enhances the SLM’s response with specific information
from our knowledge base about Edge AI applications on Raspberry Pi. One issue that should
be addressed is the latency.
To reduce latency, we can use for embedding, the all-minilm model which is much smaller
(23M parameters vs. 137M for nomic-embed-text) and creates 384-dimensional embeddings
instead of 768, significantly reducing computation time.
Also, smaller chunks can be helpful but have some disadvantages. For example, let’s say that
we can use a small chunk size (100 tokens with 50 overlap). Here are some considerations:
Advantages
1. Memory Efficiency: Smaller chunks require less memory during retrieval and processing,
which is beneficial for resource-constrained devices like the Raspberry Pi.
2. More Granular Retrieval: Smaller chunks can potentially provide more precise matches
to specific questions, especially for targeted queries about very specific details.
3. Reduced Context Window Usage: SLMs have limited context windows; smaller
chunks allow you to include more distinct pieces of information while staying within these
limits.
734
Disadvantages
1. Loss of Context: 100 tokens is approximately 75-80 words, which is often insufficient to
capture complete concepts or explanations. Many paragraphs and technical descriptions
require more space to convey their full meaning.
2. Increased Vector Store Size: More chunks mean more embeddings to store, potentially
increasing the overall size of your vector database.
3. Fragmented Information: With such small chunks, related information will be split
across multiple chunks, making it harder for the model to synthesize coherent answers.
A good practice would be to experiment with different chunk sizes and embedding models and
measure:
We can create a simple benchmarking function to have one embedding model defined test the
best chunk size:
def benchmark_chunk_sizes(document_list,
query_list,
sizes=[(100, 50), (300, 30), (500, 50), (1000, 100)]):
"""Test different chunk sizes and measure performance"""
results = {}
# Split documents
start_time = [Link]()
doc_splits = text_splitter.split_documents(document_list)
split_time = [Link]() - start_time
735
# Create embeddings and store
embedding_function = OllamaEmbeddings(model="nomic-embed-text")
temp_db_path = f"temp_db_{chunk_size}_{overlap}"
start_time = [Link]()
vectorstore = Chroma.from_documents(
documents=doc_splits,
collection_name="benchmark",
embedding=embedding_function,
persist_directory=temp_db_path
)
db_time = [Link]() - start_time
# Create retriever
retriever = vectorstore.as_retriever(k=3)
# Test queries
query_times = []
for query in query_list:
start_time = [Link]()
docs = [Link](query)
query_time = [Link]() - start_time
query_times.append(query_time)
# Store results
results[(chunk_size, overlap)] = {
"num_chunks": len(doc_splits),
"splitting_time": split_time,
"db_creation_time": db_time,
"avg_query_time": sum(query_times) / len(query_times),
"max_query_time": max(query_times),
"min_query_time": min(query_times)
}
# Clean up temporary DB
[Link](temp_db_path)
return results
Regarding the query side, some optimations can also reduce the latency at the edge. Let’s
modify the previous script, with:
1. Direct Ollama API Calls: Bypasses the LangChain abstraction layer for embedding
736
and LLM generation to reduce overhead.
2. Embedding Caching: Uses lru_cache to prevent recalculating embeddings for repeated
queries.
3. Preloading Models: Initializes models at startup to avoid cold-start latency.
4. Optimized Retriever Settings: Uses minimal k-value (2) and adds a score threshold
to filter out irrelevant matches.
5. Reduced Dependency Usage: Removes unnecessary imports and simplifies the
pipeline.
6. Concurrent Processing: Uses ThreadPoolExecutor for batch document embedding
(when needed).
7. Early Termination: Checks for empty document results before running the LLM.
8. Simplified Prompt: Uses a more concise prompt template focused on getting direct
answers.
9. Fixed Seed: Uses a consistent seed for the LLM to reduce variability in response times.
The direct Ollama API approach removes several layers of abstraction in the
LangChain implementations.
We can see latency improvements from 2 minutes down to approximately 50-110 seconds,
depending on the complexity of the queries.
737
In the next section, we’ll explore how RAG can be combined with our agent architecture to
create even more powerful edge AI systems.
Note that any tool could be used here; the calculator is only a simple example to
demonstrate the concept.
738
When the user asks a question, the system first determines if it needs to use a tool or the RAG
approach. For knowledge queries, the RAG system enhances the response with information
from the database. The system then validates the answer quality, and if it’s not sufficient, tries
again with an improved prompt. In cases where questions fall outside the database’s scope, the
system will clearly inform the user rather than attempting to generate potentially misleading
answers.
System Architecture
739
The system functions through several key components:
1. Query Router
• Analyzes incoming queries to determine if they’re calculations or knowledge queries
• Can use the same model for the response generator or a lightweight model to reduce
overhead
• Implements rule-based fallbacks for robust classification
2. Document Retriever
• Connects to a persistent vector database (Chroma)
• Uses semantic embeddings to find relevant documents
• Returns contextually similar content for knowledge generation
3. Response Generator
• Creates answers based on retrieved documents
• Implements a two-stage approach with validation and improvement
• Adds appropriate disclaimers when information is insufficient
4. Validation Engine
• Evaluates answer quality using structured criteria
• Assigns a numerical score to each generated response
• Triggers enhancement processes when quality is insufficient
5. Interactive Interface
• Provides user-friendly interaction with clear quality indicators
• Supports model switching and verbosity control
• Offers guidance for improving query outcomes
Key Workflow
740
• Validate response quality
• If quality is low, attempt enhancement with improved prompt
• If still insufficient, add disclaimer about knowledge gaps
5. Return final answer with quality metrics
Query Routing
if is_calc:
# Use smaller, faster model for operation and number extraction
# ...extraction logic here...
return route_info
741
# Validate the response quality
validation = validate_response(llm, query, answer)
validation_score = [Link]("score", 5)
{query}
{enhanced_context}
742
"complete answer."
)
answer = answer + information_gap_note
743
744
Examples
Simple Calculation
745
More Complex Queries
746
Queries outside of the database scope:
Fine-tuning can adapt models to specific domains or tasks, improving performance for targeted
applications.
747
Preparing for Fine-Tuning
748
print(f"Fine-tuning model using data from {data_path}")
print(f"Fine-tuned model will be saved to {output_path}")
Supervised fine tuning (SFT) is a method to improve and customize pre-trained LLMs.
It involves retraining base models on a smaller dataset of instructions and answers. The
main goal is to transform a basic model that predicts text into an assistant that can follow
instructions and answer questions. SFT can also enhance the model’s overall performance, add
new knowledge, or adapt it to specific tasks and domains
Before considering SFT, it is recommended to try prompt engineering techniques like few-shot
prompting or retrieval augmented generation (RAG), as discussed previously. In practice,
749
these methods can solve many problems without fine-tuning. If this approach doesn’t meet
our objectives (regarding quality, cost, latency, etc.), then SFT becomes a viable option when
instruction data is available. SFT also offers benefits like additional control and customizability
to create personalized LLMs.
However, SFT has limitations. It works best when leveraging knowledge already present in
the base model. Learning completely new information, like an unknown language, can be
challenging and lead to more frequent hallucinations. For new domains unknown to the base
model, it is recommended that it be continuously pre-trained on a raw dataset first.
On the opposite end of the spectrum, instruct models (i.e., already fine-tuned models) can be
very close to our requirements. By providing chosen and rejected samples for a small set of
instructions (between 100 and 1000 samples), we can force the LLM to behave as we need.
The easiest way to finetune an SLM is by using Unsloth. The three most popular SFT techniques
are full fine-tuning, LoRA, and QLoRA.
For example, using this link, it is possible to find several notebooks with the steps to finetune
SLMs—for instance, the Gemma 3:1B.
The fine-tuned model can be saved on HF Hub or locally as GGUF, and to run a GGUF
model locally, we can use Ollama, as shown below:
cd ~
nano Modelfile
FROM ~/Downloads/[Link]
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 64
6. We can now use the model through as we have done with the llama3.2:3B in this chapter..
750
Conclusion
This chapter has explored comprehensive strategies for overcoming the inherent limitations of
Small Language Models in edge computing environments. By implementing techniques ranging
from optimized prompting strategies to sophisticated agent architectures and knowledge inte-
gration systems, we’ve demonstrated that it’s possible to significantly enhance the capabilities
of edge AI systems without requiring more powerful hardware or cloud connectivity.
The techniques presented—chain-of-thought prompting, task decomposition, function calling,
response validation, and RAG—form a toolkit that edge AI engineers can apply individually
or in combination to address specific challenges. Each approach offers unique advantages:
prompting techniques improve reasoning capabilities with minimal overhead, agent architectures
enable SLMs to perform actions beyond text generation, and RAG systems dramatically expand
an SLM’s knowledge without increasing model size.
Our practical implementations on the Raspberry Pi showcase that these enhancements are not
merely theoretical but can be deployed in real-world edge scenarios. From the simple calculator
agent to the more sophisticated knowledge router and RAG-enabled question answering system,
these examples provide templates that developers can adapt to their specific application
requirements.
The true power of these techniques emerges when they’re strategically combined. An agent
architecture with RAG capabilities, enhanced by chain-of-thought reasoning and validated with
a feedback loop, creates an edge AI system that approaches the capabilities of much larger
models while maintaining the advantages of edge deployment—privacy preservation, reduced
latency, and operation without internet connectivity.
As edge AI continues to evolve, these techniques will become increasingly important in bridging
the gap between the limited resources available on edge devices and the growing expectations
for AI capabilities. By thoughtfully applying these approaches, developers can create intelli-
gent systems that process data locally, respect user privacy, and operate reliably in diverse
environments.
The future of edge AI lies not necessarily in deploying ever-larger models but in developing
more innovative systems that combine efficient models with intelligent architectures, contextual
knowledge integration, and robust validation mechanisms. By mastering these techniques, edge
AI practitioners can create solutions that are not just technologically impressive but genuinely
useful and trustworthy in addressing real-world challenges.
Resources
The scripts used in this chapter can be found here: Advancing EdgeAI Scripts
751
752
LiteRT-LM: Production-Ready LLM Inference
at the Edge
753
Introduction
In the previous chapters, we explored Ollama as our primary gateway to running Small Language
Models (SLMs) on the Raspberry Pi 5. Ollama is excellent for rapid experimentation: one
command downloads and serves a model, and the Python library wraps everything in a clean,
familiar API. However, Ollama was designed primarily for desktop and server machines; it
runs [Link] under the hood, and its runtime is not purpose-built for the tight memory and
latency budgets of IoT deployments.
This chapter introduces LiteRT-LM, Google AI Edge’s production-ready, open-source inference
framework for deploying Large Language Models directly on edge devices. While Ollama is great
for getting started, LiteRT-LM offers a different set of tradeoffs that make it particularly well-
suited for scenarios where startup latency, memory footprint, and cross-platform consistency
matter — or simply when you want your Raspberry Pi to run the same model binary that
powers Google’s own Pixel Watch and Chromebook Plus.
• Explain the architecture of LiteRT-LM and how it differs from Ollama / [Link].
• Install the litert-lm Python package on a Raspberry Pi 5.
• Download and run a Gemma 4 model in the .litertlm format from Hugging Face.
• Build a simple query function, an interactive chat loop, a basic tool-calling pipeline, and
a streaming inference loop with performance metrics — all using LiteRT-LM’s Python
API.
754
What is LiteRT-LM?
Google’s on-device ML story has a long history. TensorFlow Lite (TFLite) was released in
2017 as a stripped-down runtime for deploying neural networks on mobile and embedded devices.
It popularized quantization-aware training, delegate-based hardware acceleration (GPU, DSP,
Edge TPU), and the flat .tflite model format — all ideas that have since become standard
practice across the industry.
In 2024, Google rebranded TensorFlow Lite as LiteRT (Lite RunTime), reflecting a broader
vision: a unified on-device ML runtime that covers not just classification and detection models,
but the full spectrum of modern AI workloads, including generative models. LiteRT-LM is
the part of this ecosystem dedicated to Large Language Models.
755
Architecture Overview
756
The execution pipeline follows a two-phase pattern familiar from any autoregressive LLM:
1. Prefill — the full prompt is tokenized and processed to fill the KV cache.
2. Decode — tokens are generated one at a time, each conditioned on the cached context.
The important insight for Edge AI is that LiteRT-LM manages the KV cache for
you. You never need to manually rebuild message history turn by turn (as you do
with raw REST APIs). As long as you stay inside the same Conversation object,
the model retains full context.
LiteRT-LM uses its own model packaging format: .litertlm. A .litertlm file is a single
archive that bundles together:
This is analogous to Ollama’s .gguf format, but with one key difference: .litertlm files are
specifically optimized for the LiteRT executor pipeline. You cannot use GGUF files with
LiteRT-LM, and you cannot use .litertlm files with Ollama. Pre-converted models are hosted
on the LiteRT Community on Hugging Face.
Supported Models
The table below shows some of the models available at the time of writing. Check the LiteRT-LM
overview page for the latest additions.
757
Quantiza- Con- RAM
Model Parameters tion text (approx.) Best For
Qwen3-0.6B 0.6B INT4 4096 ~0.5 GB Fastest option
Gemma 4 architecture note: The “E” in Gemma 4 E2B stands for Effective
parameters. Gemma 4 uses a MatMul-free attention design combined with Per-
Layer Embeddings (PLE), which achieves frontier-level quality with substantially
fewer active parameters than traditional dense transformers. The E2B model
benchmarks above Gemma 3 27B on several reasoning tasks despite being 12×
smaller.
Both Ollama and LiteRT-LM are excellent tools for running SLMs on Raspberry Pi. They
solve the same fundamental problem but with different priorities. Understanding the tradeoffs
will help you choose the right tool for each project.
Choose Ollama when you need the widest model selection, the simplest API, or OpenAI-
compatible endpoints for existing code.
Choose LiteRT-LM when you are targeting a production IoT deployment, want a consistent
runtime from prototype (Raspberry Pi) to mobile (Android/iOS), or need tighter memory
control and a framework that is actively hardened for edge constraints.
758
Important benchmark note: On the Raspberry Pi 5 with the Gemma 4 E2B
model, you can expect roughly 4–8 tokens/second at INT4 precision using the
CPU backend. This is comparable to Ollama with Gemma 3 2B, confirming that
both frameworks extract similar throughput from the RPi5’s Cortex-A76 cores.
The practical difference shows up on mobile hardware with GPU/NPU backends,
where LiteRT-LM achieves massive speedups that [Link]-based tools cannot
match.
Installation on Raspberry Pi 5
Prerequisites
Confirm that your Raspberry Pi 5 is running the 64-bit (aarch64) version of Raspberry Pi OS.
LiteRT-LM requires this architecture.
uname -m
# Expected: aarch64
Also, confirm your Python version. LiteRT-LM’s Python package targets Python 3.10–3.13:
python3 --version
It is good practice to use a dedicated virtual environment for LiteRT-LM experiments, especially
since litert-lm may conflict with torch or tensorflow packages when installed in the same
environment.
Starting criating a working directory
mkdir -p ~/Documents/LITERT
cd ~/Documents/LITERT
759
python3 -m venv .venv
source .venv/bin/activate
This single command pulls the pre-built wheel for aarch64 Linux and its dependencies (including
the XNNPACK-accelerated LiteRT runtime). There is no separate server to start — LiteRT-LM
runs in-process, embedded inside your Python script.
Models are hosted on Hugging Face. The huggingface_hub CLI is the most reliable way to
download .litertlm files.
760
Downloading a Model
We will use Gemma 4 E2B (gemma-4-E2B-it), the sweet-spot model for Raspberry Pi 5.
It runs in approximately 1.5 GB of working memory, leaving headroom for your application
code.
mkdir -p ~/Documents/LITERT/models
huggingface-cli download \
litert-community/gemma-4-E2B-it-litert-lm \
[Link] \
--local-dir ~/Documents/LITERT/models
~/Documents/LITERT/models
��� [Link]
LiteRT-LM ships with a command-line runner. Before writing any Python, confirm the model
works:
litert-lm run \
~/Documents/LITERT/models/[Link] \
--prompt="What is the capital of Brazil?"
761
First-run latency: Loading a 1.5 GB model from an SD card takes 10–20 seconds.
Using an NVMe SSD via the PCIe slot reduces this to 2–4 seconds. For production
IoT applications, an SSD is strongly recommended.
Python API
We can create Python scripts with a text editor, such as nano, or install a Jupyter Notebook:
Now let’s explore the Python API systematically, building from a single query to a streaming
chat loop with performance metrics. All examples assume the following preamble:
import litert_lm
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"
The two core objects you will use in every LiteRT-LM script are:
Both objects are Python context managers (use them with with statements), which ensures
clean resource deallocation.
762
��������������������������������������������
� Engine (model weights loaded once) �
� �������������������������������������� �
� � Conversation A (KV cache A) � �
� �������������������������������������� �
� �������������������������������������� �
� � Conversation B (KV cache B) � �
� �������������������������������������� �
��������������������������������������������
A single Engine can host multiple concurrent Conversation objects, each with an independent
state. This is useful for multi-user applications or for maintaining separate conversation threads
for different tasks.
The simplest pattern: open an Engine, create a Conversation, send one message, read the
response.
import litert_lm
litert_lm.set_min_log_severity(litert_lm.[Link])
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"
Output:
The response object is a dictionary with a content list, where each item has a type field
("text") and a text field containing the generated string. This structure is intentionally similar
to the Anthropic and OpenAI messages API, making the pattern familiar.
763
Wrapping it as a Reusable Function
Output:
print(simple_query("And Peru"))
Output:
Please provide more context! "And Peru?" is a very open-ended question. To give you a helpful
764
import litert_lm
litert_lm.set_min_log_severity(litert_lm.[Link])
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"
Example session:
>>> quit
[Exiting chat]
Notice how "and Peru?" (a context-dependent follow-up) is correctly resolved to Peru’s capital.
The model understood the implicit reference because the KV cache preserved the prior turn.
With the simple_query function from the previous section, the same follow-up would fail.
The key API difference here is send_message_async vs send_message. The async variant
returns a generator that yields partial chunks as tokens are produced, enabling streaming
output that makes interactions feel responsive. Each chunk has the same structure as the full
response from send_message.
765
3. Tool Calling (Function Calling)
As we saw in the SLM: Basic Optimization Techniques chapter, LLMs are unreliable calculators.
Let’s reproduce that demonstration with LiteRT-LM:
...123,456² = 15,238,827,536
print(f"{123456*123456:,}") # → 15,241,383,936
LiteRT-LM supports native function-calling protocols for Gemma models. At the Python
API level (still in active development), the most robust and universally compatible approach
is JSON-based prompt engineering — the same technique we used with Ollama in the
previous chapter. The idea is identical: use a system instruction to constrain the model’s
output to a specific JSON schema, then parse and dispatch that JSON in Python.
766
Step 1 — Define the Tool Function
Args:
a: The first number.
b: The second number.
"""
return f'{a * b:,}'
LiteRT-LM does not yet have a separate system role in its Python API. The standard
workaround is to send the system instruction as the first user message before any real query.
The conversation’s KV cache retains this instruction for all subsequent turns.
import json
SYSTEM_INSTRUCTION = """
You are a helpful assistant with access to a single tool: multiply_numbers(a, b).
{
"tool": "multiply_numbers",
"a": <number>,
"b": <number>
}
767
with litert_lm.Engine(MODEL_PATH) as engine:
with engine.create_conversation() as conversation:
# Prime the model with the system instruction
conversation.send_message(SYSTEM_INSTRUCTION)
Output:
1. The LLM handles natural language understanding — extracting intent and operands.
768
2. Python handles actual computation — always correct, always deterministic.
Swapping the tool: You can replace multiply_numbers with any Python function
— a GPIO sensor read, a file lookup, a distance calculation, a REST API call. The
model doesn’t care how the function is implemented internally; it only needs to
emit the correct JSON. This makes the pattern reusable across all kinds of Edge
AI + IoT applications.
If we run the tool calling loop again with a new query, such as “What is the capital of
Guatemala?”, the answer will be a normal response from the model:
[Could not parse tool call or invalid format] Expecting value: line 1 column 1 (char 0)
For production deployments, knowing your inference speed is essential. The following code
adds Time to First Token (TTFT), total latency, and tokens/second measurements to
the chat loop.
TTFT measures how long the user waits before any output appears — it is dominated by the
prefill phase. Tokens/second measures decode throughput. Together, they give a complete
picture of perceived responsiveness.
import time
import re
import litert_lm
litert_lm.set_min_log_severity(litert_lm.[Link])
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"
769
def chunk_to_text(chunk) -> str:
"""Extract plain text from a LiteRT-LM response chunk."""
parts = []
if isinstance(chunk, dict):
for item in [Link]("content", []):
if isinstance(item, dict) and [Link]("type") == "text":
[Link]([Link]("text", ""))
return "".join(parts)
start = time.perf_counter()
first_token_time = None
pieces = []
end = time.perf_counter()
print()
output = "".join(pieces)
output_tokens = rough_token_count(output)
total_latency = end - start
ttft = (first_token_time - start) if first_token_time else None
decode_time = (end - first_token_time) if first_token_time else None
770
tps = (output_tokens / decode_time) if decode_time and decode_time > 0 else 0.0
Example session:
>>> y la de Bolivia?
The capital of Bolivia is **Sucre**.
However, it's important to note that **La Paz** is the administrative capital and the seat of
[TTFT: 1.427s | Total: 11.427s | OutTok(est): 54 | Tk/s(est): 5.40]
Typical RPi 5
Metric Value Meaning
TTFT 1-2 s (SSD) Time until first visible output. Dominated by prefill
(prompt processing).
Total 3-11 s Full wall-clock time from send to end-of-stream.
latency
OutTok(est) varies Estimated number of output tokens (word-level
approximation).
Tk/s(est) 5-7 Decode throughput in tokens per second.
771
Why is TTFT several seconds? On the first question in a session, TTFT
includes model-loading overhead. On subsequent turns, it shrinks because the
Engine is already warm. Longer prompts increase TTFT proportionally, since
prefill scales with prompt length.
#!/usr/bin/env python3
"""
litert_lm_chat.py — Interactive streaming chat with Gemma 4 E2B on RPi5.
Usage:
source ~/litert-env/bin/activate
python3 litert_lm_chat.py
"""
import re
import time
import litert_lm
MODEL_PATH = "/home/mjrovai/Documents/LITERT/models/[Link]"
litert_lm.set_min_log_severity(litert_lm.[Link])
772
def main():
print(f"Loading model: {MODEL_PATH}")
print("Type 'quit' or 'exit' to stop.\n")
end = time.perf_counter()
print("\n")
output = "".join(pieces)
output_tokens = rough_token_count(output)
total_latency = end - start
ttft = (first_token_time - start) if first_token_time else None
decode_time = (end - first_token_time) if first_token_time else None
tps = (output_tokens / decode_time) if decode_time and decode_time > 0 else 0
773
if ttft is not None:
print(
f" � [TTFT: {ttft:.2f}s | Total: {total_latency:.2f}s | "
f"~{output_tokens} tokens | {tps:.1f} tk/s]\n"
)
if __name__ == "__main__":
main()
During the inference, it is possible to see that the memory consumption is low (2.1GB of RAM
+ ~1GB of Swap):
774
Runing a simple query on the Raspi, using the same model but with Ollama, we can see that mem-
Going Further
Multi-modal inputs. Gemma 4 is natively multi-modal. LiteRT-LM can accept image buffers
alongside text prompts, opening the door to vision-language tasks directly on Raspberry Pi —
using the same .litertlm file.
775
Feeding tool results back into the conversation. The tool-calling example stops after the
function is called. A richer pattern sends the function’s result back as a follow-up message,
allowing the model to formulate a natural-language reply that incorporates the correct answer.
Multiple simultaneous conversations. A single Engine can host multiple Conversation
objects at the same time — useful for multi-user local assistants in, say, a classroom setting,
without loading model weights twice.
Upgrading to Gemma 4 E4B. If your Raspberry Pi 5 has 8 GB of RAM, the E4B model
(~3.7 GB working memory) offers noticeably better reasoning quality, especially for longer
contexts (with higher latency).
Cross-platform consistency. A .litertlm binary tested on a Raspberry Pi 5 will run
without modification on an NVIDIA Jetson Orin Nano or a Qualcomm-based Android device
— with the additional benefit of GPU/NPU acceleration on those targets. This is the key
architectural advantage of LiteRT-LM: one model file, one runtime, every platform.
Conclusion
LiteRT-LM occupies a distinct niche in the edge AI toolbox. It is not a replacement for
Ollama — it is a complement. Ollama excels at broad model selection, simplicity, and desktop
development. LiteRT-LM excels at production-quality, cross-platform deployment where the
same runtime binary needs to travel from prototyping (Raspberry Pi) to production (Android
phone, wearable, embedded system).
In this chapter, we covered the complete development arc: understanding LiteRT-LM’s ar-
chitecture (Engine → Conversation → XNNPACK executor); installing the Python package
and downloading a Gemma 4 .litertlm model; and building four progressively more complex
patterns — single query, persistent chat, tool calling, and streaming with metrics.
The most important mental model to take away is the Engine / Conversation separation:
weights are loaded once (expensive), and conversation state is maintained automatically per
session (free). This design makes LiteRT-LM well-suited for always-on IoT applications where
the model is loaded at boot and conversations flow naturally for as long as the device is
running.
Resources
776
• Gemma 4 E2B model: [Link]/litert-community/gemma-4-E2B-it-litert-lm
• LiteRT-Optimized INT8 LLM for Raspberry Pi (ICCV Workshop 2025):
[Link]
• Google AI Edge Gallery (Android/iOS demo app): [Link]/google-ai-
edge/gallery
• Study Notebook for this chapter: [Link]
#
Weekly Labs
777
Edge AI Engineering - Weekly Labs
Objectives:
Instructions:
Objectives:
Instructions:
778
2. Configure Jupyter Notebook for remote access:
pip install jupyter
jupyter notebook --generate-config
jupyter notebook --ip=[Link] --no-browser
Deliverable: A simple Python script that captures and displays an image from your camera
and a Screenshot showing a successful image capture
Objectives:
Instructions:
wget [Link]
Deliverable: Python script that successfully classifies sample images with MobileNet V2 and
a Screenshot showing a successful result
779
Lab 4: Custom Dataset Creation
Objectives:
Instructions:
Deliverable: Structured dataset with at least 3 classes and 50 images per class
Objectives:
Instructions:
780
5. Train model using MobileNet V2
6. Analyze model performance (accuracy, confusion matrix)
7. Test model on validation data
Deliverable: Edge Impulse project link and screenshot of model performance metrics
Objectives:
Instructions:
Deliverable: Python script for real-time image classification with your custom model and a
Screenshot showing a successful result
Objectives:
781
Instructions:
Deliverable: Python script that performs and visualizes object detection on test images and
a Screenshot showing a successful result.
Objectives:
Instructions:
782
Week 5: Custom Object Detection
Objectives:
Instructions:
Deliverable: Annotated dataset with at least 2 object classes and 100 total images
Objectives:
Instructions:
Deliverable: Edge Impulse project link with trained object detection model and performance
metrics
783
Week 6: Advanced Object Detection
Objectives:
Instructions:
Deliverable: Python application that compares SSD MobileNet vs. FOMO performance in
real-time
Objectives:
Instructions:
Deliverable: Python application for real-time object detection and counting using YOLO
784
Week 7: Object Counting Project
Objectives:
Instructions:
Objectives:
Instructions:
785
Deliverable: Integrated application combining multiple AI capabilities with visualization
dashboard
Objectives:
Instructions:
3. Install dependencies:
sudo apt update
sudo apt install build-essential python3-dev
Deliverable: Screenshot showing system configuration with increased swap and temperature
monitor during stress test
Objectives:
786
Instructions:
1. Install Ollama:
• Load time
• Inference speed (tokens/sec)
• Memory usage
• Temperature
Objectives:
Instructions:
787
• Handle conversation context
3. Implement proper error handling
4. Create a simple interactive CLI application
Deliverable: Python script demonstrating Ollama library usage with conversation handling
Objectives:
Instructions:
Deliverable: Python application that uses function calling for structured interaction with
SLMs
788
Week 10: Retrieval-Augmented Generation
Objectives:
Instructions:
Deliverable: Python implementation of a basic RAG system with simple text documents
Objectives:
Instructions:
789
Deliverable: Optimized RAG implementation with specialized knowledge base and perfor-
mance analysis
Objectives:
Instructions:
Objectives:
Instructions:
790
1. Implement image captioning:
• Basic caption generation
• Detailed caption generation
2. Implement object detection:
• Bounding box visualization
• Multiple object detection
3. Implement visual grounding:
• Highlight specific objects based on text prompts
4. Create segmentation application
5. Measure the performance of each task
Deliverable: Python application demonstrating multiple vision tasks with Florence-2 and
performance analysis
Objectives:
Instructions:
791
3. Create a Python script to read sensor data
4. Implement LED control based on conditions
5. Create visualization of sensor data
Deliverable: Python application for reading sensor data and controlling actuators with
visualization
Objectives:
Instructions:
Deliverable: Jupyter Notebook with interactive widgets for sensor monitoring and actuator
control
Objectives:
Instructions:
792
1. Create a Python application that:
• Collects sensor data
• Formats data for SLM prompt
• Sends prompt to model
• Parses response
• Controls actuators based on response
2. Implement multiple analysis modes
3. Add error handling for SLM responses
Deliverable: Python application integrating SLMs with physical sensors and actuators
Objectives:
Instructions:
Deliverable: Complete IoT monitoring system with SLM integration and web interface
793
Week 14: Advanced Edge AI Techniques
Objectives:
Instructions:
Deliverable: Python implementation of agent architecture with tool usage and decision
routing
Objectives:
Instructions:
794
5. Implement validation mechanisms
6. Compare the effectiveness of different strategies
Objectives:
Instructions:
Deliverable: Complete agentic RAG system with documentation and performance analysis
Objectives:
795
Instructions:
Hardware Requirements
• Raspberry Pi Zero 2W or Pi 5
• MicroSD card (32GB+)
• Camera module (USB webcam or Pi camera)
• Power supply
796
• Breadboard
Software Requirements
Development Environment
• Raspberry Pi OS (64-bit)
• Python 3.9+
• Jupyter Notebook
• SSH client
Generative AI
• Ollama
• Transformers
• Pytorch
• ChromaDB
• LangChain
• Pydantic
• Instructor
Physical Computing
• GPIO Zero
• Adafruit CircuitPython libraries
797
Assessment Criteria
1. Start Early: These labs build on each other. Falling behind makes later labs more
difficult.
2. Document As You Go: Take notes, screenshots, and document issues/solutions.
3. Optimize Resources: SLMs and VLMs require careful resource management.
4. Collaborate: Discuss approaches with classmates while ensuring individual work.
5. Backup Regularly: Create backups of your SD card after significant progress.
6. Measure Performance: Always benchmark and optimize your implementations.
7. Ask Questions: If you’re stuck, ask for help early rather than falling behind.
#
References & Author
798
References
To learn more:
Online Courses
• Harvard School of Engineering and Applied Sciences - CS249r: Tiny Machine Learning
• Professional Certificate in Tiny Machine Learning (TinyML) – edX/Harvard
• Introduction to Embedded Machine Learning - Coursera/Edge Impulse
• Computer Vision with Embedded Machine Learning - Coursera/Edge Impulse
• UNIFEI-IESTI01 TinyML: “Machine Learning for Embedding Devices”
Books
Projects Repository
799
AiEng4D
800
About the author
Marcelo Rovai, a Brazilian living in Chile, is a recognized figure in engineering and technology
education. He holds the title of Professor Honoris Causa from the Federal University of Itajubá
(UNIFEI), Brazil. His educational background includes an Engineering degree from UNIFEI
and a specialization from the Polytechnic School of São Paulo University (POLI/USP). Further
enhancing his expertise, he earned an MBA from IBMEC (INSPER) and a Master’s in Data
Science from the Universidad del Desarrollo (UDD) in Chile.
801
With a career spanning several high-profile technology companies, including AVIBRAS Airspace,
AT&T, NCR, and IGT, where he served as Vice President for Latin America, he brings industry
experience to his academic endeavors. He is a prolific writer on electronics-related topics and
shares his knowledge through open platforms like [Link].
In addition to his professional pursuits, he is dedicated to educational outreach, serving as a
volunteer professor at UNIFEI and engaging with the AIEng4D (formerly TinyML4D) group,
and the EDGE AIP– the Academia-Industry Partnership of EDGEAI Foundation as a Co-Chair,
promoting TinyML education in developing countries. His work underscores a commitment to
leveraging technology for societal advancement.
LinkedIn profile: [Link]
Lectures, books, papers, and tutorials: [Link]
802