Communications IoT et protection de la vie privée
Communications IoT et protection de la vie privée
Samuel Pélissier
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Département de la Formation par la Recherche
et des Études Doctorales (FEDORA)
Bâtiment INSA direction, 1er étage
37, av. J. Capelle
69621 Villeurbanne Cédex
fedora@[Link]
L’INSA Lyon a mis en place une procédure de contrôle systématique via un outil de
détection de similitudes (logiciel Compilatio). Après le dépôt du manuscrit de thèse,
celui-ci est analysé par l’outil. Pour tout taux de similarité supérieur à 10%, le manuscrit
est vérifié par l’équipe de FEDORA. Il s’agit notamment d’exclure les auto-citations, à
condition qu’elles soient correctement référencées avec citation expresse dans le
manuscrit.
Par ce document, il est attesté que ce manuscrit, dans la forme communiquée par la
personne doctorante à l’INSA Lyon, satisfait aux exigences de l’Etablissement concernant
le taux maximal de similitude admissible.
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Département FEDORA – INSA Lyon - Ecoles Doctorales
1. ScSo : Histoire, Géographie, Aménagement, Urbanisme, Archéologie, Science politique, Sociologie, Anthropologie
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Contents
3 Background 36
3.1 LoRaWAN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
3.1.1 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
3.1.2 Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
3.1.3 Frame structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.1.4 LoRaWAN identifiers . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.1.5 Message types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
3.1.6 Join process and keys generation . . . . . . . . . . . . . . . . . . . . . 40
3.1.7 Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2 A machine learning primer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2.1 Step-by-step overview . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2.2 Unsupervised vs supervised learning . . . . . . . . . . . . . . . . . . . 42
3.2.3 Common pitfalls . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
3.2.4 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
5 CONTENTS
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
6 CONTENTS
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
7 CONTENTS
V Conclusion 115
9 Contributions and Future Directions 116
9.1 Contributions summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116
9.2 Perspectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
9.2.1 Short-term . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
9.2.2 Long-term . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
List of Figures
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
9 LIST OF FIGURES
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
List of Tables
10
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Remerciements
Trois ans auparavant, je n’aurais pu imaginer une fraction de ce qu’il m’a été possible de
vivre. Au-delà de la richesse en apprentissages et en émotions, cette thèse n’aurait pas abouti
sans l’aide, la patience et la bienveillance de nombreuses personnes.
Je souhaiterais tout d’abord remercier mes encadrants, Mathieu Cunche et Vincent Roca,
pour leurs précieux conseils et leur soutien indéfectible ; merci d’avoir fait ce pari et offert
la chance de découvrir la recherche dans un tel environnement. Merci à Didier Donsez pour
ses innombrables idées et son éternel enthousiasme, ainsi qu’à Abhishek, qui m’a tant appris
en si peu de temps. Merci à Jan et Thomas pour les pauses eau-café, les projets d’arrosage
fonctionnels-mais-étranges, et le camping. Plus généralement, merci à toute l’équipe Privatics
pour son accueil chaleureux et les discussions passionnantes.
Je n’aurais jamais pu débuter l’aventure, et encore moins la continuer, sans ma famille.
Merci à mes parents, mes frères, mes grands-parents et M. pour votre amour au quotidien.
Panou, j’aurais aimé pouvoir te montrer le manuscrit terminé ; tu aurais sûrement suggéré
de le traduire en français ou en espagnol.
Enfin, merci à mes ami·e·s J., M. et A. pour m’avoir aidé à naviguer trois années étranges
et rebondissantes. Merci au cercle des imbéciles anonymes et au chaudron, certainement
piliers du bon déroulement de cette thèse, et sans qui j’aurais probablement pu tout faire en
deux ans environ.
11
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Résumé
Les dernières décennies ont été témoins de l’émergence et de la prolifération d’objets con-
nectés, communément appelés Internet des Objets (IdO). Cet écosystème divers correspond à
une large gamme de dispositifs spécialisés, allant de la caméra IP aux capteurs détectant les
fuites d’eau, chacun conçu pour répondre à des objectifs et des contraintes de consommation
d’énergie, de puissance de calcul ou de coût. Le développement rapide de nombreuses tech-
nologies et leur connexion en réseau s’accompagne de la génération d’un important volume
de données, soulevant des préoccupations en matière de vie privée, en particulier dans des
domaines sensibles tels que la santé ou les maisons connectées.
Dans cette thèse, nous exploitons les techniques d’apprentissage automatique (machine
learning) pour explorer les problèmes liés à la vie privée des objets connectés via leurs pro-
tocoles réseau. Tout d’abord, nous étudions les attaques possibles contre LoRaWAN, un
protocole longue distance et à faible coût d’énergie. Nous explorons la relation entre deux
identifiants du protocole et montrons que leur séparation théorique peut être contrecarrée en
utilisant les métadonnées produites lors de la connexion au réseau. En nous appuyant sur
une approche multi-domaines (contenu, temps, radio), nous démontrons que ces métadonnées
permettent à un attaquant d’identifier les objets connectés de manière unique malgré le chiffre-
ment du trafic, ouvrant la voie au traçage ou à la ré-identification.
Nous explorons ensuite les possibles contre-mesures, en analysant systématiquement les
données utilisées lors de ces attaques et en proposant des techniques pour les obfusquer ou
réduire leur pertinence. Nous démontrons que seule une approche combinée offre une réelle
protection. Par ailleurs, nous proposons et évaluons diverses solutions de pseudonymes tem-
poraires adaptées aux contraintes de LoRaWAN, en particulier la consommation énergétique.
Enfin, nous adaptons notre méthodologie d’apprentissage automatique à DNS, un pro-
tocole largement déployé dans l’IdO grand public. À nouveau basées sur les métadonnées,
notre attaque permet d’identifier les objets connectés, malgré le chiffrement du flux DNS-
over-HTTPS. Explorant les contre-mesures potentielles, nous observons un non-respect des
standards liés au padding, entraı̂nant la compromission partielle de la vie privée des utilisa-
teurs.
Plus généralement, nos travaux mettent en évidence que les efforts déployés par les pro-
tocoles IdO tels que LoRaWAN pour protéger la vie privée sont insuffisants. Des évolutions
potentiellement profondes sont nécessaires pour correctement répondre à ces enjeux.
Mots-clefs : Vie privée; Internet des Objets; Objets connectés; Sécurité; Protocoles; Ra-
dio; Réseaux sans fil; Métadonnées; Empreinte; Re-identification; Inférence d’activités; Lo-
RaWAN; LPWAN; DNS; DNS-over-HTTPS.
12
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
List of publications
Accepted
International conferences
• [165] Samuel Pélissier, Jan Aalmoes, Abhishek Kumar Mishra, Mathieu Cunche, Vin-
cent Roca, and Didier Donsez. Privacy-preserving pseudonyms for LoRaWAN. Pro-
ceedings of the 17th ACM Conference on Security and Privacy in Wireless and Mobile
Networks (WiSec), 2024.
• [166] Samuel Pélissier, Mathieu Cunche, Vincent Roca, and Didier Donsez. Device re-
identification in LoRaWAN through messages linkage. Proceedings of the 15th ACM
Conference on Security and Privacy in Wireless and Mobile Networks (WiSec), 2022.
Code and data available: [Link]
ductibility.
National conferences
• Samuel Pélissier, Jan Aalmoes, Abhishek Kumar Mishra, Mathieu Cunche, Vincent
Roca, and Didier Donsez. Privacy-preserving pseudonyms for LoRaWAN. 14ème Atelier
sur la Protection de la Vie Privée (APVP), 2024.
• Samuel Pélissier, Mathieu Cunche, Vincent Roca, and Didier Donsez. Introducing
privacy-preserving identifiers in LoRaWAN. Poster LPWAN Days, 2023.
• Samuel Pélissier, Mathieu Cunche, Vincent Roca, and Didier Donsez. Linkage attacks
and countermeasures in LoRaWAN. 13ème Atelier sur la Protection de la Vie Privée
(APVP), 2023.
• Samuel Pélissier, Mathieu Cunche, Vincent Roca, and Didier Donsez. Device re-
identification in LoRaWAN through messages linkage. 12ème Atelier sur la Protection
de la Vie Privée (APVP), 2022.
Under review
• [167] Samuel Pélissier, Abhishek Kumar Mishra, Mathieu Cunche, Vincent Roca, and
Didier Donsez. Introducing Multi-Domain Fingerprints for Lorawan Device Linkage.
Submitted to the International Journal for the Computer and Telecommunications In-
dustry (Computer Communications).
• Samuel Pélissier, Gianluca Anselmi, Abhishek Kumar Mishra, Anna Maria Mandalari,
Mathieu Cunche. Does Your Smart Home Tell All? Assessing the Impact of DNS En-
cryption on IoT Device Identification. Submitted to the 20th International Conference
on Wireless and Mobile Computing, Networking and Communications (WiMob), 2024.
• Abhishek Kumar Mishra, Samuel Pélissier, and Mathieu Cunche. Fingerprinting con-
nected Wi-Fi devices using per-network MAC addresses. Submitted to the Workshop
on Security and Artificial Intelligence (SECAI), 2024.
13
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Part I
14
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 1
Introduction
Contents
1.1 Challenges and contributions . . . . . . . . . . . . . . . . . . . . 15
1.2 Outline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
1.3 Funding . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
During the past decades, we have witnessed the emergence of connected devices, commonly
referred to as the Internet-of-Things (IoT). Unlike general-purpose, consumer-grade comput-
ers, IoT is not a monolithic concept with a singular, clear definition. Instead, it represents a
diverse ecosystem corresponding to a wide range of specialized devices, each designed to meet
specific objectives and constraints, whether it is energy efficiency, computational power, or
affordability. For instance, IoT encompasses devices as varied as wired IP cameras deployed in
smart home environments and wireless water leakage sensors installed in industrial settings.
Despite their disparate functionalities and deployment scenarios, both examples fall under
the umbrella of IoT devices.
By definition, these devices are connected to their surroundings through networks of var-
ious topologies and technologies. This connectivity leads to the generation of specialized
data, much of which raises privacy concerns, especially in sensitive domains such as health
monitoring or activity tracking in smart homes.
In parallel to the expanding capabilities and ubiquitous deployment of IoT, machine learn-
ing techniques have proved to be powerful tools for automating the analysis of large volumes
of data and showcasing privacy attacks. By applying these methods to the network traffic
generated by IoT devices, we study network protocols and uncover new insights into their
vulnerabilities to various privacy attacks.
In this thesis, we leverage such tools to explore privacy issues in IoT devices through the
lens of their communications. With a specific focus on the highly constrained LoRaWAN
protocol, we study the trade-offs between privacy and energy consumption. We complete this
analysis by adapting our machine learning methodology to encrypted DNS, a battle-tested
and widely deployed protocol in consumer-grade IoT devices.
15
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
16 CHAPTER 1. INTRODUCTION
Imperfect protection
mechanisms Linking LoRaWAN
identifiers (contribution #1)
What are the privacy
challenges of the
LoRaWAN protocol?
Fingerprinting LoRaWAN
Identifiable patterns in devices (contribution #2)
communication
studied [74, 131, 196], with major features lacking attention from the community. By focusing
on identifiers leveraged by the protocol, we identify two possible privacy threats during the
lifetime of a LoRaWAN device: imperfect protection mechanisms in the join process, and
identifiable patterns in the actual transmission.
16
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
17 CHAPTER 1. INTRODUCTION
• We design a set of mitigations for each feature previously identified as posing a threat
against privacy, and respecting the constraints of the protocol. This includes padding,
delay, and radio power modulation.
• Through careful evaluation on the same datasets used for conducting attacks, we mea-
sure the individual and combined impact of such countermeasures.
• We study their inherent overhead and outline how both their limited effect on attacks
performance and restrictions of the LoRaWAN protocol restrict their potential viability.
• Given the constraints of LoRaWAN, we tailor two solutions for the protocol and propose
a renewal strategy to rotate pseudonyms.
• We thoroughly assess their performance and intrinsic limitations via a set of theoretical
and simulation-based evaluations.
• Finally, we reflect on possible adaptations in other Low Power Wide Area Network
protocols under similar constraints.
Contribution #5: Identifying IoT devices via encrypted DNS traffic (Chapter 8).
• We analyze potential contenders for encrypted DNS in IoT devices and choose to focus
on DNS-over-HTTPS (DoH).
• Only leveraging metadata from the network traffic, we evaluate machine learning meth-
ods and find that device identification is both possible and reliable despite encryption.
• We study possible countermeasures and analyze various padding strategies via simula-
tions.
• Coincidentally, we find that multiple DNS resolvers do not respect relevant standard,
failing to correctly protect the privacy of their users.
1.2 Outline
This thesis is structured as follows. Chapter 2 presents a state of the art regarding privacy
in the IoT, both from an offensive and defensive perspective. Additional elements on the
LoRaWAN protocol and machine learning are summarized in Chapter 3.
The first part of this thesis proposes some answers to Q1 What are the privacy challenges
of the LoRaWAN protocol?. Chapter 4 deals with linking theoretically unrelated identifiers
during the join process, allowing an eavesdropper to associate a device identity with its
17
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
18 CHAPTER 1. INTRODUCTION
activity. We design a second attack in Chapter 5, and fingerprint devices based on their
communication patterns.
The second part explores the second research question, Q2 How can we address privacy
challenges in LoRaWAN?, studying possible countermeasures. Chapter 6 applies various
noises to machine learning features in order to reduce their susceptibility to linkage and
identification attacks. A more radical and impactful solution is presented in Chapter 7, with
the design of rotating pseudonyms for LoRaWAN.
The third - and last - part includes Chapter 8, answering Q3 Can this methodology be
adapted to other IoT protocols?, by adjusting the machine learning methodology to identify
IoT devices via encrypted DNS, despite its encryption.
Finally, Chapter 9 concludes this thesis, summarizing our contributions and hinting at
both short and long-term perspectives.
1.3 Funding
This thesis was carried out within the CITI laboratory (INSA Lyon & Inria), in the Privatics
team (Inria). It was funded by the ANR-BMBF PIVOT project (ANR-20-CYAL-0002), and
supported by the H2020 SPARTA project, as well as the IoT SPIE ICS - INSA Lyon chair.
18
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 2
State of the art
Contents
2.1 Security in wireless communications for the IoT . . . . . . . . . 19
2.2 Privacy for the IoT . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.2.1 From human to data . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.2.2 Privacy regulations . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.2.3 Guidelines and standards . . . . . . . . . . . . . . . . . . . . . . . 22
2.3 Attacks on IoT privacy . . . . . . . . . . . . . . . . . . . . . . . . 22
2.3.1 Threat model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
2.3.2 Activity inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.3.3 Radio-based positional tracking . . . . . . . . . . . . . . . . . . . . 24
2.3.4 Protocol-based tracking . . . . . . . . . . . . . . . . . . . . . . . . 26
2.3.5 Device fingerprinting . . . . . . . . . . . . . . . . . . . . . . . . . . 27
2.4 Countermeasures against IoT privacy attacks . . . . . . . . . . 30
2.4.1 Privacy metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
2.4.2 Dummy traffic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
2.4.3 Noise-based perturbations . . . . . . . . . . . . . . . . . . . . . . . 31
2.4.4 Modified radio medium . . . . . . . . . . . . . . . . . . . . . . . . 33
2.4.5 Pseudonyms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
2.4.6 Pseudonym rotations . . . . . . . . . . . . . . . . . . . . . . . . . . 35
2.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
Building upon the thesis outline and research questions presented in the previous section,
we delve deeper into the multifaceted realm of the IoT privacy. Through this exploration, we
aim to offer an overview of various techniques, both offensive and defensive, related to the
specific context of IoT privacy, while setting reasonable boundaries to such a broad scope. We
begin by highlighting the inherent challenges and security vulnerabilities of IoT ecosystems.
Next, we explore the nuanced concept of privacy in the context of IoT, exploring its various
dimensions and implications. Subsequently, we dissect the prevalent attacks targeting IoT
privacy, outlining a generic eavesdropping threat model and the main methodologies employed
by adversaries. Finally, we analyze the countermeasures and mitigation strategies deployed
to combat these privacy attacks.
19
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
20 CHAPTER 2. STATE OF THE ART
(MitM) [41, 154], MAC spoofing attacks [41], and many others [159]. Emerging IoT protocols
such as LoRaWAN suffer from similar weaknesses, from the process allowing devices to join
a network, to data transmission [100, 154]. For instance, a bit flipping attack is possible by
design as the payload authentication is not checked by the application server [100]. Although
newer versions tackle most of these vulnerabilities, a few still remain, such as denial-of-service
attacks [212, 99].
A significant portion of security research focuses on data integrity and confidentiality. Con-
fidentiality, achieved through encryption and access controls, plays a crucial role in preventing
unauthorized access to sensitive information. It is important to note that privacy cannot exist
without encryption and the confidentiality measures it provides. However, privacy extends
beyond mere confidentiality by encompassing metadata generation and individuals’ rights to
control their personal data, regardless of security mechanisms. As explored further in Sec-
tion 2.2, it includes broader ethical and legal considerations regarding data collection, use,
and sharing.
For instance, when capturing traffic from a device using the address A and known to be
a glucose monitor [55], the identity corresponds to A and the activity to monitoring blood
glucose levels. We can also infer the location based on radio signals (see Section 2.3.3). Of
20
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
21 CHAPTER 2. STATE OF THE ART
course, such a baseline could be leveraged to infer additional information: a glucose monitor
is probably used by someone with diabetes, and the attacker could know who uses the device
with address A.
In a broader sense, we examine the privacy of IoT users through the lens of their in-
formational identities [218]. Instead of a single identifier (e.g. name only), the identity is
considered as a set of multidimensional information forming a profile (e.g. behavior, prefer-
ences) [83, 184].
a goal that member states should achieve, without defined means to achieve said results. It is implemented
nationally as laws and regulations. On the other hand, a European regulation is effective as law in all member
states at once (e.g. GDPR).
21
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
22 CHAPTER 2. STATE OF THE ART
22
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
23 CHAPTER 2. STATE OF THE ART
Attacks on
IoT privacy
Location Identity
Activity inference
& tracking & device fingerprinting
transmits via
Device Medium Receiver
eavesdrops
Attacker
instance via their own gateway or off-the-shelf hardware, can eavesdrop on communications
between devices and receivers.
Our threat model is thus defined as follows. First, every other part of the infrastructure is
trusted [222], including device, gateways, and servers physical and software integrity. Second,
contrary to scenarios detailed by Torres et al. [206], the attacker is only passive. For instance,
they cannot alter existing frames, inject new ones in the communication, jam it, or replay
previous exchanges. This is a reasonable assumption as messages generally benefit from
integrity protection and counters are used to detect replays (see Section 3.1.3). Third, we
assume the cryptographic layer is secured: it is impossible to access the content of encrypted
payloads, and the integrity as well as authentication are robust. Following sections present
various attacks under these assumptions.
23
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
24 CHAPTER 2. STATE OF THE ART
Such attacks also concern lower level and more IoT specific protocols such as ZigBee [186]
or LoRaWAN [131]. Interesting information can be gleaned not only from the traffic con-
tent but also from its structure: Leu et al. demonstrate that if a real-life event is directly
linked to network messages, a single frame is enough to infer a device state (e.g. a garage
door opening) [131]. Although Copos et al. extract features from clear-text protocols such
as DNS [66], following works only leverage metadata such as length and timing of mes-
sages [27, 131]. For instance, Acar et al. provide a detailed analysis on smart home activity
inference via encrypted traffic, from device activation to movements in the house [15]. They
distinguish between time-independent, non-sequential events (e.g. device power cycling) and
time-dependent, sequential events (e.g. movement between locations), shedding light on the
intricacies of activity inference in smart home environments.
24
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
25 CHAPTER 2. STATE OF THE ART
Position 1
Position 2
... ...
Position n
is located somewhere that was not previously surveyed. Hernández et al. estimate the lo-
cation of non-mapped fingerprints for indoor Wi-Fi positioning through continuous dataset
thanks to support vector machines [98]. From an attacker’s point of view, this approach al-
lows reducing by half the number of surveyed positions used to train the model while keeping
the same mean distance error [98].
As seen in Figure 2.4a, triangulation requires two receivers with directional antennas.
The position of S is computed through trigonometry using the Angle of Arrival (AoA) at
both A and B, and the distance d(A, B). Triangulation based on AoA requires less hardware
compared to other methods, but its effectiveness is optimized in scenarios with clear line-
of-sight [172]. As the distance between the device and receivers increases, the accuracy of
AoA-based triangulation tends to decrease [172].
Figure 2.4b presents trilateration based on 3 receivers with non-directional antennas using
Time of Arrival (ToA). The distance between each receiver and S is obtained via subtraction
of a timestamp included in the frame sent by S and the timestamp of arrival at A, B, and
C. To precisely compute the position, time synchronization is required between device and
receivers, which can be challenging [79, 135]. Such a technique falls outside the scope of
our threat model, as it requires a close synchronization with the device, that should not be
possible for a passive attacker.
Multilateration expands such a concept by including more than 3 receivers [124]. Another
improvement also involves computing the distance based on relative time instead of absolute
25
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
26 CHAPTER 2. STATE OF THE ART
time. Figure 2.5 illustrates how the Time Difference of Arrival (TDoA) can be computed at
each receiver to find the position of S, without time synchronization (i.e. without collaboration
from S). Fargas and Petersen leverage such a method in LoRaWAN with 4 receivers, obtaining
an outdoor precision of around 100 meters [79]. Their work is improved by Podevijn et al.:
they evaluate various trajectories (walking, cycling, and driving) in a larger network, reaching
a median accuracy of 75 meters when enriching TDoA with the underlying road map [172].
Although such figures are impressive, they may not be sufficient when tracking devices in
dense environments where multiple devices are deployed within meters of each other. Instead,
alternative solutions based on the protocol itself can be explored.
26
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
27 CHAPTER 2. STATE OF THE ART
highlighting that interlayer synchronization, i.e. changing all identifiers and counters at once,
is required to correctly protect from protocol-based tracking [56].
Although most related works ask “do these two fingerprints come from the same device? ”,
the others choose a different path. Instead of computing a distance between two fingerprints,
they let a machine learning model determine to which device (i.e. class) the fingerprint
corresponds. This advantageously avoids computing a distance, but is less robust: adding a
new device to the network requires to re-train the machine learning model. Such an approach
is abusively denoted as “ML-based” distance in Table 2.1 for completeness.
Vectors of values : Multiple works rely on vectors of raw values, without any modification
of information extracted from network traces [17, 162], such as F = v. For instance, Aernouts
et al. fingerprint the location of LoRaWAN devices based on their signal power, with values
directly associated with GPS positions [17]. When dealing with more than one feature, the
fingerprint is a concatenation of multiple vectors: Pang et al. identify devices via various
features extracted from their wireless traces, including a vector of packet lengths and a
vector of network identifiers (SSID) [162]. For n features, the fingerprint thus corresponds to:
∀v, F = f (v1 , v2 , ..., vn ) = (v1 , v2 , ..., vn ).
27
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
28 CHAPTER 2. STATE OF THE ART
Markov chains : IoT devices have distinct communication patterns (cf. Section [Link]).
For example, a sensor can be programmed to send a report every hour, or every day at
midnight, with occasional alerts in between if a threshold is exceeded. Previous fingerprint
representation do not fully exploit such a stateful behavior.
Although this aspect has seldom been investigated in wireless networks, a few works
consider Web browsing as a stochastic process and model it as Markov chains [121, 161, 189].
They all model TLS communication as a series of states, constructing homogeneous3 Markov
chains from browsing traces, with the fingerprint corresponding to a stochastic matrix P =
{pi,j }, where pi,j is the probability of transition from state i to state j. These chains, built
using known websites, are then compared to unknown visits to infer users’ browsing habits
despite encryption.
More precisely, initial work by Korczynski and Duda considers a single state based on the
TLS message type (e.g. Application Data, Client Hello, etc.), with an enter state, various
transition states, and an exit state [121]. Following research by Shen et al. and Pan et al.
improves the solution in two ways: 1) a state is now a combination of TLS message type and
message length, and 2) their Markov chains are not first order, but second order, meaning that
the next state is not only influenced by the current state, but also the previous one [161, 189].
While more accurate, this approach generates complex chains with many states. Increasing
the order and number of states in Markov chains results in high computational complexity
due to the underlying representation as transition matrices.
3 The transition probabilities do not vary across time.
28
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
29 CHAPTER 2. STATE OF THE ART
Image-based : A subset of radio fingerprinting works focus on the physical layer using
techniques previously developed for image recognition. This approach is based on the envi-
ronment differences between devices, as well as the unique hardware generating identifying
variations in the radio signal [197]. It involves converting radio frequency signals into visual
representations, including spectrograms illustrating signal frequency over time [113, 188].
These representations are then analyzed using image recognition techniques, such as convo-
lutional neural networks (CNNs). A highly visual example is provided by Jiang et al. [113],
reproduced in Figure 2.6. By drawing the differential constellation trace figure4 of various
LoRa devices (Figure 2.6a), they are able to uniquely identify them (Figure 2.6b).
(a) Example of differential constellation (b) Clustering for 6 LoRa devices based on
trace figure. differential constellation traces.
These works are currently limited by the medium quality, with their accuracy depending
on good Signal-to-Noise Ratio (SNR) [188] while some physical layer protocols, such as Long
Range (LoRa), are designed to operate effectively with low SNR. Although our work primarily
focuses on higher-level network protocols and fundamental radio information such as signal
power, it is essential to acknowledge physical layer fingerprinting as a potential threat to
privacy.
29
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
30 CHAPTER 2. STATE OF THE ART
Works leveraging Markov chains depend on the Maximum Likelihood Criterion [121, 189].
It seeks to find the parameter values that make the observed data the most probable under
the assumed statistical model. In the Web browsing fingerprinting scenario, it helps identify
the website that is most likely to be the source of the captured encrypted traffic.
In text classification and document clustering, the KL divergence outperforms the cosine
similarity [103], which in turn outclasses the Jaccard index [132] and Euclidean distance [103].
However, results highly depend on the studied dataset [103, 132]. To the best of our knowl-
edge, no comparison of these various distance have been published in the context of network
traffic fingerprinting.
Attacks on
IoT privacy
Location Identity
Activity inference
& tracking & device fingerprinting
Dummy traffic + noise-based perturbations
+ pseudonym rotations
Please note that each mitigation is presented in isolation, but combining multiple coun-
termeasures is often required for higher impact. For example, padding alone (included in
“noise-based perturbations” in Figure 2.7) may not be sufficient to counter DNS fingerprint-
ing [53]. Furthermore, practicing data minimization whenever possible is mandatory: an
attacker cannot exploit information they do not have [27]. Finally, for attacks based on clear-
text fields leaking information (see Section [Link]), hiding their content through encryption
remains the best countermeasure.
30
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
31 CHAPTER 2. STATE OF THE ART
being targeted. For instance, different metrics are employed for location privacy compared
to protecting against activity inference. Second, privacy metrics are influenced by the threat
model, which encompasses the goals and capabilities of potential attackers [219].
From the eighty privacy metrics Wagner and Eckhoff cite, authors usually pick a few
depending on their use-case. For instance, Becker et al. work on pseudonyms in BLE and
define an anonymity set (i.e. the set of users that an attacker cannot distinguish from the
targeted user) and the maximum tracking time (ie. the cumulative time during which an user
is uniquely identified) [44]. More generally, machine learning-based privacy research often
focuses on the change of accuracy of models trained with or without countermeasures [22, 53].
31
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
32 CHAPTER 2. STATE OF THE ART
1. Linear padding: Packets are padded to the nearest multiple of 128 (or to the Maximum
Transmission Unit (MTU), usually 1500 bytes).
2. Packet Random MTU padding: Packets are randomly padded, up to the MTU.
3. Pad to MTU : Packets are padded to the MTU.
4. Exponential padding: Packets are padded to the nearest power of 2 (or to the MTU).
5. Mice-Elephants padding: Packets are padded to 128 if smaller, or to the MTU if bigger.
This is relevant when packets are generally small, with a few large outliers.
The RFC 8467 [145] on DNS padding cites the first 3 strategy and advocates only for the
first (named Block-Length Padding), while proposing Random-Block-Length Padding, which
pads to a randomly selected multiple of 128.
At a high level, two main strategies emerge from the remaining literature: padding to
the maximum length possible or randomly padding within a specified interval. These involve
trade-offs between privacy and performance: padding to the MTU provides perfect protec-
tion, but requires significant resources. Pinheiro et al. observe that real-world packet lengths
are mostly under 100 bytes long, suggesting that uniform padding up to the MTU is ineffi-
cient [171]. They propose a hybrid approach: padding packets to the nearest hundred up to
300 and randomly beyond 301 to the MTU. This method reduces resource consumption while
maintaining privacy. Additionally, they explore adaptive padding based on network activity,
padding more during idle periods to obscure information and less during active network us-
age. Engelberg and Wool evaluate the approach and find that while it effectively hides the
specific device being used in the network, it does not prevent detecting the number of active
devices [77].
Most of the literature focuses on high-throughput network protocols (e.g. HTTP [76] or
DNS [145]) allowing to pad 128-bytes long blocks of data. However, Shafqat et al. outline
that this approach may not be feasible in constrained protocols, such as ZigBee for which
they propose a random padding of 0 to 3 bytes [186] (also see Section 3.1.7).
Finally, Alshehri et al. explore the choice of distribution for random padding selection,
aiming to achieve Differential Privacy [22]. While Laplace noise seemed promising, its imple-
mentation requires reducing half of the packet lengths. Instead, they demonstrate that uni-
formly random noise satisfies (ϵ, δ) Approximate Differential Privacy, a more lenient form [75].
Parameters ϵ = 0 and δ = ∆s n are selected, with n the maximum number of bytes, ensuring
optimal privacy protection (minimal ϵ). Here, ∆s denotes the maximum difference in packet
size: higher potential padding correlates with stronger privacy guarantees.
32
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
33 CHAPTER 2. STATE OF THE ART
used to identify devices [173]. They address this issue by introducing a random 10ms delay
before responding to a request. However, their analysis combines this countermeasure with
dummy traffic, making it challenging to determine which measure has the most significant
impact.
Unlike padding, delaying packets can significantly disrupt the application flow, poten-
tially causing timeouts or incorrect arrival order if the delay is too high [139]. Srinivasan
et al. emphasize the importance of adjusting delay values for each application, as some can
handle extended delays of several minutes, while others cannot tolerate any delay at all (e.g.
emergency applications) [197].
2.4.5 Pseudonyms
Instead of relying on stable identifiers, modern devices often opt for random and ephemeral
link-layer addresses to enhance privacy, often named pseudonyms. Pseudonyms are intimately
linked with the devices capabilities they are deployed on and the protocol requirements for
address space (i.e. the number of bits available). As seen in Section 2.2.3, this is still an
open question for the 802.11 family, and work is required to bridge the gap between resource-
intensive pseudonym solutions and constrained IoT devices. We propose a tailored solution
to this problem in Chapter 7.
In this section, we present various pseudonym schemes for 802.11/Wi-Fi [30, 90], sensor
networks [152, 231], BLE [92], and vehicular networks [138, 169]. The solutions discussed
here only support mutual authentication, where both the device and central authority can
authenticate the message’s source. Hence, we consider device anonymity without two-way
verification as out of scope. For instance, we exclude randomized MAC addresses in Android
Wi-Fi module [88] or group signatures in vehicular networks [169], for which the receiver can
not authenticate the sender.
33
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
34 CHAPTER 2. STATE OF THE ART
(AP) they are connected to. It is used to encrypt MAC addresses along a sequence number
si by both the devices and AP. Such encrypted number si is also appended at the end of
the frame. Then, the receiving device or AP matches the encrypted pseudonym with its own
pre-generated values.5
If a frame is lost, the sender still updates si , meaning the device and AP are now desyn-
chronized. Upon receiving a frame containing a pseudonym not matching any pre-generated
value, the AP tries to decrypt si with each pre-shared key. If si is found, the pseudonym can
be decrypted and both parties are synchronized again.
This approach has been updated by Greenstein et al. [90]. Their solution, named Shroud,
is different in two aspects. First, they encrypt the sequence number via AES-128, selecting
the whole block as identifier, instead of appending a sequence value at the end of the frame.
Second, they generate multiple encrypted pseudonyms at once, contrary to one at a time.
Thus, the receiver can look up a larger set of pre-generated pseudonyms when packet loss
occurs. Although this approach requires slightly more initial computation power and memory
storage, Greenstein et al. leverage hash tables instead of lists to reduce complexity and do not
require a costly process to re-synchronize, as long as enough pseudonyms are pre-generated.
[Link] HASHA
Zhang and Zhang propose an alternative design to protect device location in sensor net-
works [231]. First, devices transmit their original address (i.e., MAC) through beacon frames,
resulting in each device maintaining a list of all nearby identifiers. Second, each pseudonym
is generated using a rolling key, starting with an arbitrary fixed value. The temporary
pseudonym T P of the first data message is generated by both devices using a fixed-value key
k and the original address OA of both source S and destination D as: T Ps = HMAC(k, OAs ).
Finally, devices update the key by XORing it with the payload of the previous message. All
messages are acknowledged to make sure the keys are correctly synchronized between two
devices.
Keyless Entry systems [86]. In their case, the mechanism is used to prevent replays.
34
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
35 CHAPTER 2. STATE OF THE ART
2.5 Conclusion
Privacy, influenced by cultural, sociological, legal and technological factors, remains a com-
plex concept, even within the reduced scope of IoT devices. Our work aligns with evolving
regulations that emphasize data protection, security, and transparency to uphold privacy
in the IoT ecosystem. As guidelines and standards continue to evolve alongside legal and
technological advancements, our research hopes to inspire privacy-preserving solutions and
contribute to more precise guidelines.
While this section sheds light on various eavesdropping-based attacks in the IoT landscape,
there remains ample opportunity for further exploration, particularly in resource-constrained
protocols like LPWANs. We also outline that existing tools such as fingerprinting are battle-
tested but rarely systematically compared.
Countermeasures against IoT privacy attacks span a diverse array of strategies, reflecting
the multifaceted nature of the threats they mitigate. However, implementing these counter-
measures often entails navigating a delicate balance between enhancing privacy and minimiz-
ing the impact on network performance, ranging from increased bandwidth usage to potential
message loss. Such trade-offs are even more apparent in highly constrained network protocols
such as LoRaWAN.
35
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 3
Background
Contents
3.1 LoRaWAN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
3.1.1 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
3.1.2 Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
3.1.3 Frame structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.1.4 LoRaWAN identifiers . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.1.5 Message types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
3.1.6 Join process and keys generation . . . . . . . . . . . . . . . . . . . 40
3.1.7 Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2 A machine learning primer . . . . . . . . . . . . . . . . . . . . . . 41
3.2.1 Step-by-step overview . . . . . . . . . . . . . . . . . . . . . . . . . 41
3.2.2 Unsupervised vs supervised learning . . . . . . . . . . . . . . . . . 42
3.2.3 Common pitfalls . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
3.2.4 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
This chapter provides additional elements required for the rest of the thesis. It details
applications, architecture and various technical features of the LoRaWAN protocol, followed
by an introduction to the fundamentals of machine learning and its pitfalls.
3.1 LoRaWAN
In this section, we introduce Long Range Wide Area Network (LoRaWAN), a Low Power
Wide Area Network (LPWAN) protocol. Standardized for the first time in 2015 by the
LoRa Alliance, its adoption grew rapidly and LoRaWAN is now deployed by 181 operators
worldwide [67]. Although multiple versions evolve in parallel, we focus on LoRaWAN v1.1 as
it provides the most mature solution [65, 140]. Likewise, the protocol is available in various
frequency bands, but we are interested in the ones deployed in Europe: EU863-870MHz and
EU433MHz [64].
Although we refrain from rewriting the entire LoRaWAN specification [65], we provide
an overview of relevant concepts related to networks and privacy. Thus, we first present
existing LoRaWAN application, before delving into its network architecture. We introduce
the frame structure and relevant fields for our thesis, including identifiers. We then detail the
join process along with keys generation and message types. Finally, we present the multiple
constraints under which the protocol operates.
3.1.1 Applications
Anything connected to a low data rate sensor, sending at most a few messages per hour, can
leverage LoRaWAN for data extraction or occasional notifications. Following in the steps
36
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
37 CHAPTER 3. BACKGROUND
of large-scale IoT networks, LoRaWAN now connects 300 million devices to a wide range of
applications [67].
From urban environments to rural landscapes, LoRaWAN finds utility in various sec-
tors such as smart cities, agriculture, environmental monitoring, logistics, healthcare, and
industrial automation [96, 199]. In smart cities, LoRaWAN facilitates intelligent infrastruc-
ture management, including waste management [175], street lighting, and parking [74]. In
agriculture, it aids in precision farming [192], crop monitoring, and livestock tracking [114].
Environmental monitoring applications leverage LoRaWAN for air and water quality monitor-
ing [137, 204], weather stations, and wildlife conservation [34, 38]. Furthermore, LoRaWAN
plays a pivotal role in asset tracking, supply chain management, and fleet monitoring in logis-
tics [187]. In healthcare, it enables remote patient monitoring, medication adherence tracking,
and telemedicine initiatives [147].
3.1.2 Architecture
As seen in Figure 3.1, a classic LoRaWAN architecture includes various components [65].
Radio (LoRa)
TCP/IP
TCP/IP
Join Server
At the core are End-Devices, which include sensors or actuators deployed in the field
to collect or act on data. These End-Devices transmit information to gateways acting
as intermediaries between the End-Device and the network infrastructure. Gateways thus
receive data from End-Devices via the LoRa radio modulation and forward it through a
classic TCP/IP link. We note that LoRa is different from LoRaWAN : the radio modulation
is only one of the physical layers hosting the MAC layer protocol. In our work, we focus on
LoRaWAN.
Following the gateways, a set of servers manages the incoming messages (uplink) and,
if necessary, responds to End-Devices (downlink). Many LoRaWAN implementations use a
single node to host the servers for deployment simplicity. In any case, they can be functionally
separated as three entities:
• The Network Server is responsible for managing the communication between End-
Devices and gateways. It handles network-related tasks, including message deduplica-
tion, radio parameter optimization, and MAC layer commands. The NS also enforces
security mechanisms, such as authentication and encryption to protect the integrity and
confidentiality of the data transmitted over the network. Finally, it acts as a routing
interface between End-Devices and the other servers. For instance, it routes uplink
application payloads to the designated Application Server (AS) and facilitates the join
process with the Join Server (JS) (see Section 3.1.6).
• The Join Server handles the onboarding process for new End-Devices joining the net-
work. It manages the authentication and key generation process between End-Devices
and the NS (see Section 3.1.6).
• The Application Server receives and processes data from End-Devices forwarded by
the NS. It interfaces with external systems or applications, integrating LoRaWAN data
37
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
38 CHAPTER 3. BACKGROUND
MAC header DevAddr FCtrl FCnt FOpts FPort Frame Payload MIC
Both uplink (End-Device to server) and downlink (server to End-Device) messages follow
the same structure [65, Sec. 4]. They are encapsulated in a MAC payload, along with a MAC
header, which includes the Message Type (uplink, downlink, etc.) and the protocol Major
Version (v1.0, v1.1). The following fields are relevant to our work (from left to right in
Figure 3.2):
• DevAddr: The Device Address is used to identify End-Devices in the network (see
Section [Link]).
• FCnt: Each uplink and downlink message is ordered via a separate Frame Counter,
respectively FCntUp and FCntDown. Starting at 0 after the join process (see Section 3.1.6)
and sent unencrypted, the values are initially computed on 32 bits but only the 16 least-
significant bits are actually transmitted.
Although the FCntUp is a unique value, the FCntDown corresponds to two counters
incremented independently. The NFCntDown is increased for each downlink messages on
port 0 (or when the field is missing), while AFCntDown is incremented for each downlink
messages with a port different from 0. Both End-Devices and NS track FCnt values to
re-order messages or detect duplicate messages.
• FOPts: Frame options are used to send MAC commands for network management and
configuration tasks, including device status request, data rate adjustments, and security
parameters management.
• FPort: Similarly to the TCP/IP model, LoRaWAN leverages a Frame Port to direct
traffic to specific applications. For instance, one port may serve maintenance functions
while another is selected to transmit sensor data.
• Frame Payload: Containing application data or MAC commands, the payload is
encrypted using AES-128 CCM* (see Section 3.1.6). Its maximum length depends on
a radio parameter controlling the number of bits sent per second, the DataRate (DR).
Table 3.1 details possible length and DR values [64, Sec. 2.1.6].
• MIC: The Message Integrity Code enables the verification of the frame integrity by
computing an AES-128 CMAC over various fields (see Section 3.1.6). It may also be
used for resolution of identifier conflicts in some cases (see Section [Link]).
The payload and FOpts are encrypted (see Section 3.1.6) but all other fields are sent in
clear-text and thus available to any eavesdropper.
38
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
39 CHAPTER 3. BACKGROUND
Table 3.1: Maximum payload length based on the DataRate (EU863-870Hz band).
[Link] DevEUI
Each LoRaWAN device is identified by a globally unique EUI-64 identifier named DevEUI.
Allocated by the manufacturer or owner, the DevEUI remains unchanged throughout the
device lifespan. Similarly to a BLE or Wi-FI MAC addresses, the first 3 bytes corresponds
to the Organizationally Unique Identifier (OUI) of the End-Device, identifying the device’s
vendor, manufacturer or organization [5], and registered to the Institute of Electrical and
Electronics Engineers (IEEE) [108].
[Link] DevAddr
Every time an End-Device joins the network, the NS allocates a new 32-bit temporary identi-
fier called DevAddr. It is then used for the duration of the session. Similarly to IP addresses,
this identifier conveys both routing and identity information. More precisely the DevAddr is
divided into two subfields:
• The AddrPrefix is fixed and identifies the network an End-Device belongs to.
• The NwkAddr is randomly generated by the NS and identifies the device within the
network it belongs to (designated by the AddrPrefix).
Due to the diverse requirements of network operators, LoRaWAN supports 8 network types
(type 0 to type 7), each offering an increasing range of available addresses. As indicated in
Table 3.2, End-Devices are addressed over 7 to 25 bits, corresponding to 27 to 225 unique
identifiers per network.
Multiple End-Devices might share a single DevAddr. This happens due to configuration er-
rors or address allocation optimization [97]. In such instances, End-Devices are differentiated
through the Message Integrity Code (MIC). The receiver computes the MIC using various
potential keys until one of them produces a value equal to the one received, thereby singling-
out the device.1 In case of an End-Device receiving the message, it only has to compute the
MIC with its own key.
//[Link]/TheThingsNetwork/lorawan-stack/blob/301af52a0e64d67f90576ededa59836bca19a03b/pkg
/networkserver/grpc_gsns.go#L931
39
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
40 CHAPTER 3. BACKGROUND
a type defined through the Message Type (MType) field in the MAC header and summarized
below:
• Join Request: an End-Device asks to join the network.
• Rejoin Request: an End-Device asks to re-join the network, by changing a part or all
of the session context (inducing a DevAddr update).
• Join Accept: the NS validates a connection to the network (answering both Join and
Rejoin requests).
• Data up: uplink data frame, can require acknowledgment (Confirmed) or (Uncon-
firmed) from the NS.
• Data down: downlink data frame, can require acknowledgment (Confirmed) or (Un-
confirmed) from the End-Device.
[Link] Over-The-Air-Activation
In OTAA, End-Devices dynamically obtain the DevAddr and cryptographic materials during
the join process. Figure 3.3 illustrates this sequence of events.
LoRa: Uplink
IP: Uplink
DevAddr IP: Uplink
DevAddr
DevAddr
Figure 3.3: Join procedure using OTAA (red open lock: this specific field is unencrypted;
green closed lock: these specific fields are encrypted).
First, an End-Device sends a join-request including its unique, clear-text DevEUI. The
Network Server (NS) forwards this information to the Join Server (JS), which authenticates
the device. Then, the JS generates the JoinNonce and sends it to the NS. The NS produces
a new DevAddr and forwards it along with the nonce to the End-Device. At this point, both
End-Device and JS possess the same JoinNonce value and can derive various session keys:
40
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
41 CHAPTER 3. BACKGROUND
• The pre-shared AppKey derives the AppSKey, shared with the Application Server (AS),
and responsible for payload confidentiality via AES-128 CCM* encryption [65, sec. 6.2.5].
• The pre-shared NwkKey derives the NwkSEncKey, FNwkSIntKey, and SNwkSIntKey, all
shared with the NS.2 Network-specific information such as MAC commands and FOpts
are encrypted using the NwkSEncKey via AES-128 CCM*. FNwkSIntKey and SNwkS-
IntKey are leveraged to compute the MIC via AES-128 CMAC, guaranteeing integrity [65,
sec. 6.2.5].
Once every key is derived, the End-Device can start the actual communication, by inserting
its clear-text DevAddr in all following uplink messages. The same value is used by the NS in
downlink messages.
3.1.7 Constraints
LoRaWAN operates under several constraints, making it suitable primarily for low-bandwidth
applications, such as sensor networks, rather than high-throughput ones. As we propose a
new design in Chapter 7, it is crucial to consider these constraints.
First, End-Devices usually operate on battery power during long period of time, spanning
multiple years. Energy is a scarce resource and everything is made to optimize its usage.
For instance, high transmission costs encourage developers to reduce the payload to a mini-
mum [54].
Second, LoRaWAN devices transmit in unlicensed frequency bands, respecting the ETSI
EN 300 220-2 standard for channel access restrictions [110]. Along with low data rates (250
to 11000 bit/s [64, Sec. 2.1.3]) and limited payload sizes inherent to LoRaWAN, the standard
imposes strict duty cycles: an End-Device must not transmit more than 0.1% to 10% of the
maximum time depending on the band [110, Tab. B.1].
Third, End-Devices are generally low-cost sensors with limited computational and memory
capabilities, placing them between class 0 and class 1 of IETF classification of constrained de-
vices [49]. They are often equipped with CPUs operating at a few dozen megahertz (contrary
to the gigahertz of modern personal computers) [50]. Semtech, the leading manufacturer of
LoRaWAN chipsets, states that RAM can be as low as 8 kB. 3
Finally, contrary to low range wireless protocols such as Wi-Fi and BLE, LoRaWAN is
expected to work in large and dense environments (e.g. cities). In this context, the Packet
Loss Rate (PLR) is high; average values ranging from 1% to 40% have been reported in
real-world deployments [128, 136, 170].
Held-out
dataset
Machine learning is generally organized in sequential phases presented in Figure 3.4, in-
cluding [36]:
2 Here,
F corresponds to “Forwarding” and S to “Serving”, a terminology relative to roaming scenarios,
which are outside the scope of this thesis.
3 [Link]
ement/
41
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
42 CHAPTER 3. BACKGROUND
1. Data Collection: Gathering relevant data that will be used to train and test the machine
learning model. This data should include features (inputs) and corresponding labels
(outputs).
2. Data Preprocessing: Cleaning the data to make it suitable for modeling, and preparing
sub-datasets for training and testing.
3. Model Selection: Choosing the appropriate machine learning algorithm based on the
nature of the problem, the type of data, and the desired outcomes.
In real-world deployments, the initial steps involve selecting the most effective model to
be used in production (e.g. make predictions on incoming data). However, our research goal
is only to show that it is possible to build a robust model, not to actually exploit it in the
wild. Thus, some steps are left out, including as Hyperparameter Tuning (when possible, see
Section 8.6.3), model deployment, and continuously monitoring and maintaining the model.
42
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
43 CHAPTER 3. BACKGROUND
To balance the training dataset, multiple techniques exist. Under-sampling reduces the
number of majority class samples, while over-sampling increases the size of the minority
class [36]. Under-sampling may lead to loss of information from the majority class, while over-
sampling can result in overfitting (see Section [Link]). As loosing information can produce
subpar results, we prefer over-sampling, along with other methods to reduce overfitting. More
precisely, we select Synthetic Minority Oversampling Technique (SMOTE) [59, 129]. This
method generates synthetic samples for the minority class without directly duplicating original
data. It has been shown to correctly prevent overfitting and has been used in various network
security-oriented works [74, 201].
In any case, only the training dataset is balanced while testing is kept intact. First,
it avoids data leakage [115], since SMOTE selects certain data points during resampling,
potentially including them in testing (see Section [Link]). Second, maintaining the natural
imbalance in the testing dataset helps identify any biases present in the model by reflecting
the original dataset’s distribution [31].
43
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
44 CHAPTER 3. BACKGROUND
3.2.4 Summary
In this thesis, we leverage supervised machine learning based on labeled datasets extracted
from network traffic. Additionally, we take into account the inherent imbalance of such
datasets, by re-sampling the training dataset, using cross-validation, and measuring perfor-
mance via the balanced accuracy.
44
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Part II
Unveiling Vulnerabilities:
Privacy Threats in LoRaWAN
45
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 4
Linking LoRaWAN identifiers
Contents
4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
4.2 Threat model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
4.3 Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
4.4 Linking join requests and uplink messages . . . . . . . . . . . . 48
4.4.1 Intuitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
4.4.2 Features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
4.4.3 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
4.5 Experimental results . . . . . . . . . . . . . . . . . . . . . . . . . 52
4.5.1 Models comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
4.5.2 Features importance . . . . . . . . . . . . . . . . . . . . . . . . . . 52
4.6 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
This chapter explores how LoRaWAN separates the identity from the activity of a device
during the join process by using two identifiers in distinct, theoretically unlinkable messages.
By leveraging a large-scale real-world dataset of more than 27 millions messages, we analyze
various content, time, and radio-based domains features. We demonstrate that a machine
learning process can reliably link back the two relevant messages, reaching a balanced accuracy
of ∼0.999. Notably, we find that a single field of the header is particularly relevant during
classification, and should be obfuscated.
This chapter is partially based on our peer-reviewed and published work Device Re-
identification in LoRaWAN through Messages Linkage [166]. The present chapter
differs mainly by the following points:
• A larger dataset is leveraged and analyzed. Instead of roughly one year and a
half, experiments are conducted on three years of LoRaWAN traffic, making it
the most complete dataset to date in the literature [195, 196].
• We provide more details on the intuition behind the work and further justify the
feature selection process.
• We streamline our approach to meet current methodology standards and conduct
a thorough analysis of the FCnt’s impact on message linkage. Additionally, we
explore the relevance of each feature using a more robust solution.
46
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
47 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
4.1 Introduction
LoRaWAN utilizes two identifiers: the DevEUI, and the DevAddr. Looking back at the pri-
vacy definition for IoT devices given in Section 2.2.1, this mechanism allows LoRaWAN to
separate the identity (DevEUI) from the activity (DevAddr) of a device.1 As illustrated by
Figure 4.1, the DevAddr is sent to the device encrypted, preventing an eavesdropper from
trivially associating it to a DevEUI. In that sense, both identifiers are by design theoretically
unlikable.
End-Device Gateway
Join-request (DevEUI)
Join-accept (DevAddr)
Uplink (DevAddr)
Linking
Figure 4.1: Linking LoRaWAN identifiers during the join process.
Linking back the join-request with the following uplink message opens up the path
to two types of attacks. First, one can infer the identity corresponding to known activity
traces. For instance, in the case of a smart parking lot [74], an attacker could precisely
link network communications with a physical spot. Likewise, linking enriches the existing
traces with additional context: the DevEUI contains an OUI related to the End-Device’s
manufacturer, with some makers focusing only on a certain type of devices (e.g. smart home
sensors). Second, it enables tracking a known identity across its multiple communications if
a disconnection or re-join occurs, by linking a set of DevAddr to a single DevEUI. Hence, an
attacker could follow a specific mobile End-Device as it moves around between its (re)join
process.
In this section, we propose a solution to link back the DevEUI sent in the join-request
and the following DevAddr, contrary their initial design. First, we present the threat model
under which we operate, before detailing the dataset used to carry out experiments. Then,
we showcase the overall process and finish by discussing our experimental results.
47
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
48 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
ping on Wi-Fi or BLE, where proximity is required (tens of meters at most), an attacker
can easily deploy sniffing devices outside the immediate vicinity and thus stealthily cap-
ture traffic for longer periods of time.
• Low financial investment: an attacker can build a gateway using affordable hard-
ware, such as a Raspberry Pi and a LoRa concentrator, for less than 200€ [106].
• Multi-band access: an attacker can listen to all bands simultaneously and thus have
access to all uplink messages.
• Publicly available demodulation: while the LoRa radio modulation scheme is pro-
prietary to Semtech, Robyns et al. have published an open-source SDR implementa-
tion [178]. Thus, an attacker are not limited to official commercial gateways and can
build a listening station based on off-the-shelf hardware.
4.3 Dataset
To carry out our experiments, we utilize network traces captured via CampusIoT. This
experimental platform is managed by the University of Grenoble2 and focuses on teaching
and research projects. Thanks to its private LoRaWAN network and ∼50 gateways deployed
around the city of Grenoble, France, CampusIoT continuously listens to surrounding traffic
on the EU 868MHz band. More specifically, gateways listen to 8 channels: 867.1, 867.3, 867.5,
867.7,867.9,868.1,868.3,868.5 MHz. While other additional channels can be used by private
operators, they are not received by the current gateways and are irrelevant for our work.
We have access to these logs from the 23rd June 2020 to the 20th August 2023, representing
a total of 419,688,857 LoRaWAN messages. Table 4.1 provides an overview of messages
available to us. For more information on message types, please refer to Section 3.1.5.
For the current experiment, only join-request and uplink messages for which we know
the valid association are relevant. As this information is unavailable for third-party operators,
we focus solely on CampusIoT’s End-Devices, summing up to ∼1.2M join-request, and 26M
uplink messages. This represents respectively ∼1.9k DevEUI and 36k DevAddr.
We note such a dataset is based on experimental data, where End-Devices may be affected
by abnormal behaviors (e.g. high frequency of transmission). However, it is impossible to
produce a reliable ground truth based on more realistic third-party traces because of the
protocol’s design.
48
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
49 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
and uplink messages are similar, there is a high chance both originate from the same End-
Device. We begin by developing this intuition and further detail selected features, before
presenting our machine learning method.
4.4.1 Intuitions
We motivate our process with two intuitions: based on the protocol design, and on the
nature of End-Devices as well as their environment. First, the FCnt is initialized at zero
after a join process, which should be a reliable indicator of the relevant uplink following a
join-request. Additionally, LoRaWAN End-Devices should send their first uplink message
soon after joining the network. The End-Device is not obligated to immediately utilize the
newly acquired DevAddr upon reception. However, it is generally advantageous for a sensor to
either request its configuration or start data transmission promptly. Figure 4.2 displays the
distribution of times between a join-request (containing the DevEUI) with the first uplink
(with the DevAddr). Generally, both messages are sent close to each other, with a majority
following a ∼10 seconds delay.
15000
10000
Count
5000
00 10 20 30 40 50 60
Time difference (in seconds)
Figure 4.2: Time between a join-request and its associated first uplink message.
Second, End-Devices communication features should not vary significantly in between two
messages. If static (e.g. metering sensors), devices should exhibit stable radio signals [25], as
gateways distance and surroundings do not change. While mechanisms adapting transmission
power exist and could influence radio-based features, they are generally used after 20 uplink
messages, and not directly post join process [68]. If mobile (e.g. position trackers), other
header fields should still help us link back the DevEUI and DevAddr.
4.4.2 Features
We extract content, time, and radio-based features from on our real-world dataset. To deter-
mine which ones to select, we study their distribution for the actual uplink directly following
a join-request (valid pair), and for other unrelated uplink messages (invalid pair). This is
possible because we control the network and possess the ground truth of associations between
DevAddr and DevEUI. Some features correspond to a single value in the uplink, while others
are distances compared with the join-request of reference or previously known information.
A complete list is available below.
The feature is deemed important, i.e able to discriminate between valid and invalid pairs,
if the variation from one category to the other is significant. Figure 4.3 shows the distri-
bution discrepancy between normalized values from valid and invalid pairs, highlighting the
difference between them. For instance, the FCnt is equal to 0 for the first uplink following
49
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
50 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
a join-request, with some outliers when packet loss occurs. While this does not guarantee
relevance during the machine learning process, it hints at strong distinguishability.
0.6
0.4
0.2
0.0
RSSI D.
FCnt
FPort
GTD
PL
TD
DD
RGD
SF
DR
SNR D.
ESP D.
Feature
Figure 4.3: Distinguishing behavior of features based first uplink received after a
join-request, or other uplink messages.
50
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
51 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
• Gateway Time Distance (GTD): A message is sent once but can be received by multiple
gateways with a slight delay. The difference of time of arrival at gateways is computed
pair-wise and saved in vectors. A euclidean distance is computed between the vector
produced by the join-request, and following uplink messages. Assuming numerous
End-Devices are static, relative time differences and distances should be low.
• DevAddr Difference (DD): As seen in Section 3.1.6, the DevAddr is randomly generated
during the join process. Thus, we compute the last time such an identifier has been
seen, expecting the correct uplink to be its first occurrence.
• Receiving Gateways Distance (RGD): Static devices should be received by similar sets
of gateways for both the join-request and subsequent uplink messages. Hence, we
compute the Hamming distance using gateways as vector indexes, expecting low distance
values for valid links.
• Spreading Factor (SF) & DataRate (DR): Radio configuration utilizes default values for
the first uplink message(s) and may change for later ones. Transmission information
such as the data rate (number of bits transmitted per second, from 0 to 7) and SF
(number of bits encoded per symbol, from 7 to 12) are not directly linked to the uplink
order. However, they are determined dynamically based on the network’s Adaptive Data
Rate (ADR) mechanism, which re-evaluates the link quality after 20 uplink messages by
default [68]. This optimization adjusts SF and DR accordingly, with potential variations
expected between first and subsequent uplink messages.
• SNR: Defined as the ratio between received power signal and the noise floor power
level, the Signal-to-Noise Ratio (SNR) can be positive or negative, depending on the
transmission quality. In our dataset, the SNR varies from -26 to +30.5.
• ESP : Derived from both RSSI and SNR, the Estimated Signal Power (ESP) is used to
compare the channel quality in radio communications. It is an integer ranging from
0 to -160dBm in our dataset, providing more precise values than the RSSI when its
value drops around -120dBm. Experiments show that, contrary to the ESP, the RSSI
distribution saturates when signal is way below the noise floor [14]. It is computed via
the formula 4.1.
RSSI, SNR, and thus ESP, are available for each receiving gateways. Hence, we com-
pute the euclidean distance for each of these values between the join-request and uplink
messages, and expect low values for valid links (RSSI D., SNR D. ESP D.).
4.4.3 Method
In order to link a join-request with the following uplink messages, we consider a machine
learning-based method utilizing features detailed in Section 4.4.2. The problem can be de-
signed as a binary classification: either a given uplink corresponds to a join-request (valid
pair), or it does not (invalid pair). In this section, we present steps followed to prepare the
data, considered classifiers, performance metrics, and overall process.
51
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
52 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
52
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
53 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
Table 4.2: Performance comparison of various classifiers via 5-fold cross validation.
The process is repeated 15 times for each random seed to smooth values. Following
Section 4.4.2, we expect the FCnt to have significant influence on the results due to its
straightforward exploitation. Additionally, it can be easily obfuscated via encryption (see
Section [Link]) and we anticipate a future removal from available features.
Results are presented in Figure 4.4 (high value means high impact), from which we can
draw multiple insights. First, both FCnt and Time Difference are strongly relevant. This is
expected, as nearly all uplink messages have a FCnt equal to 0. Likewise, the time between
join-request and the following valid uplink is distinctively low (cf. Figure 4.2). Second,
when removing the FCnt, models tend to favor the Time Difference and Frame Port (as well
as the Payload Length more marginally), while other features remain low importance.
Decrease in balanced accuracy
0.20 Baseline
Without FCnt
0.15
0.10
0.05
0.00
RSSI D.
FCnt
FPort
GTD
PL
TD
DD
RGD
SF
DR
SNR D.
ESP D.
To complete this analysis by anticipating a possible obfuscation of the FCnt, we train and
compare models with all features (Baseline), with all features but no FCnt (Without FCnt ),
and with FCnt only.
Figure 4.5 shows that all machine learning models (with Naive Bayes (NB) as a notable
exception) follow a similar pattern: removing the FCnt slightly decreases performance, further
reduced when using FCnt only. For instance, Random Forest decreases from ∼0.999 to 0.997
when removing the FCnt, and as low as ∼0.986 with FCnt only.
First, this confirms that other features are able to effectively compensate a lack of FCnt:
the attack would work even if it was not available. Second, the FCnt is in itself highly relevant
for an eavesdropper, achieving high accuracy using only that piece of information.
4.6 Conclusion
We demonstrate the robust association of two theoretically unlinkable messages: the join-request
and the corresponding first uplink message. Through this linkage, we establish a connection
53
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
54 CHAPTER 4. LINKING LORAWAN IDENTIFIERS
1.0
Balanced accuracy 0.9
0.8
0.7
Baseline
0.6 Without FCnt
FCnt only
0.5 RF kNN LGBM DT LR AB NB
Machine learning model
Figure 4.5: Comparison of machine learning models performance.
between the identity (DevEUI) and activity (DevAddr) of a LoRaWAN End-Device. Our ma-
chine learning model matches both identifiers with a ∼0.999 balanced accuracy, revealing that
additional efforts on the protocol design are required. Notably, we underscore the significance
of the FCnt, which enables accurate linking of both messages, even when used alone. Our
analysis of each selected feature underscores the critical role of timing between join-request
and uplink, calling for additional countermeasures.
While linking identifiers poses a privacy risk, other potential attacks on communication
data warrant attention. Given our ability to reliably link identity and activity, in the next
chapter, we investigate privacy threats associated with the activity itself.
54
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 5
Fingerprinting LoRaWAN devices
Contents
5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
5.2 Motivation & threat model . . . . . . . . . . . . . . . . . . . . . 56
5.2.1 Fingerprinting as a privacy threat . . . . . . . . . . . . . . . . . . 56
5.2.2 Fingerprinting as security protection . . . . . . . . . . . . . . . . . 57
5.2.3 Updated threat and security model . . . . . . . . . . . . . . . . . . 57
5.3 Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
5.3.1 Selecting relevant data . . . . . . . . . . . . . . . . . . . . . . . . . 58
5.3.2 Characterizing uplink messages . . . . . . . . . . . . . . . . . . . 59
5.4 A new fingerprinting representation . . . . . . . . . . . . . . . . 59
5.5 Fingerprint-based linkage . . . . . . . . . . . . . . . . . . . . . . . 60
5.5.1 Content-based features . . . . . . . . . . . . . . . . . . . . . . . . . 61
5.5.2 Time-based features . . . . . . . . . . . . . . . . . . . . . . . . . . 61
5.5.3 Radio-based features . . . . . . . . . . . . . . . . . . . . . . . . . . 62
5.6 Evaluation methodology . . . . . . . . . . . . . . . . . . . . . . . 62
5.6.1 Dataset generation . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
5.6.2 Groundtruth: labeling sequences . . . . . . . . . . . . . . . . . . . 63
5.6.3 Linkage via machine learning . . . . . . . . . . . . . . . . . . . . . 63
5.7 Experimental results . . . . . . . . . . . . . . . . . . . . . . . . . 64
5.7.1 Fingerprint representations analysis . . . . . . . . . . . . . . . . . 64
5.7.2 Measuring the impact of sequence length and feature domains . . . 65
5.7.3 Fingerprinting mobile End-Devices . . . . . . . . . . . . . . . . . . 66
5.7.4 Impact of the number of controlled listening stations . . . . . . . . 67
5.8 Limitations and future works . . . . . . . . . . . . . . . . . . . . 68
5.9 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
This chapter studies how communication patterns of LoRaWAN devices can be leveraged
to accurately identify them. After analyzing relevant features and proposing a new, holistic
approach to fingerprinting, we utilize a large and realistic dataset consisting of 41 million
uplink messages from third-party professional operators. Our machine learning approach
successfully re-identifies the origin of message sequences with a ∼0.98 balanced accuracy.
Supported by various scenarios, such as identifying mobile devices and utilizing a limited
number of listening stations, we show that fingerprinting LoRaWAN devices is both largely
applicable and robust.
55
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
56 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
• We enhance the justifications for feature selection, offering more technical expla-
nations to support our choices.
• We outline the limitations inherent in our approach, acknowledging potential
constraints and challenges that may affect the interpretation and application of
our findings.
5.1 Introduction
Linking the join-request with the following uplink messages is only one of the steps in
tracking LoRaWAN End-Devices. The flow of messages is equally interesting for eavesdrop-
pers. It contains all the actual communication data and metadata, revealing information
about the identity, activity and location of a device.
End-Devices are distinguished by their DevAddr, which may change throughout their lifes-
pan, after each rejoin process or disconnection event. This process can be frequent (cf. Ta-
ble 4.1), and we find that around 50% of devices in our year-long dataset use an identifier
lasting less than a week. While such high numbers can be explained by short activity periods,
experiments and packet losses, other mechanisms including pseudonyms rotation, already uti-
lized in BLE [92, 143], may be further increase them in future versions of LoRaWAN if they
are deployed. In any case, a changing pseudonym breaks the flow of data for an eavesdropper,
leaving them with only part of the whole communication.
Fingerprinting is a generic tool for network traffic analysis, enabling tracking devices based
solely on their characteristics and bypassing pseudonym-based protections (see Section 2.3.5).
While already utilized across various protocols such as BLE [19], Wi-Fi [174, 213], and Web
browsing [121, 161, 189], it has seldom been studied in LoRaWAN networks.
In this chapter, we explore fingerprinting LoRaWAN End-Devices based on their distinct
communication patterns. We begin by outlining our motivations and the threat model under
which we operate. Then, we present the dataset leveraged during experiments and its relevant
characteristics. After proposing a new fingerprint representation, we illustrate its usage via
a method to link back sequences of uplink messages generated before and after a change of
DevAddr. We end by showcasing the experimental results of fingerprinting End-Devices, and
by exploring some limitations of our approach.
56
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
57 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
device, and 2) link the trace to an activity (e.g. occupying a parking space) [74]. As explored
in Section 2.4.5, some protocols support identifier rotations through pseudonyms (e.g. in
Bluetooth) [92]. While LoRaWAN does not support rotating addresses yet (see Chapter 7),
a disconnection or rejoin process induce the NS to generate a new DevAddr. An End-Device
can trigger this process independently of its application traffic by sending a “Rejoin-Request
message” [65, sec. 6.2.4]. Additionally, the server can initiate a rejoin via a “ForceRejoinReq”
command [65, sec. 5.13]. As seen in Table 4.1, this feature is already employed in real-world
deployments, thwarting tracking and profiling solutions.
Fingerprinting enables attackers to establish a consistent device identity by leveraging
diverse (meta)data. Contrary to others consumer-grade IoT devices such as smartwatches,
LoRaWAN devices are generally tied to a specific activity (e.g. water metering). Hence, iden-
tifying the devices is usually enough to gain information on the underlying activity. However,
fingerprinting can also serve as a tool for more general activity inference, by comparing the
targeted traffic with known, generic communication patterns. In this study, we focus on
device identification as a proof-of-concept.
57
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
58 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
DevAddr
update
Linking identifiers A A A A B B B B
Compare
Fingerprints Fingerprints
Change in
behavior
Anomaly detection A A A A A A
Time
5.3 Dataset
For the current work, we utilize a subset of the dataset previously employed for linking,
as discussed in Section 4.3. We focus on December 2021 to September 2022, because this
period shows regular activity, corresponding to 41 million uplink messages, a significant
enough number to draw insights from. Instead of relying on traffic generated directly by
CampusIoT to establish ground truth, we study the data of third-party operators. This
approach enables us to analyze realistic traces from professional deployments, aiming for
high accuracy in fingerprinting real-world traffic. The counterpart is a lack of knowledge
of the actual deployment nature: it is impossible to know precisely which End-Devices are
communicating, and from where. As shown in Section 5.7.3, we find alternative ways to
overcome these limitations.
58
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
59 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
As highlighted in Section 4.5.2, time-based features play a significant role in linking the
join-request with following uplink messages. To further study uplink messages, Figure 5.3
displays the distribution of time window between each message, named IAT. IAT measure-
ments are concentrated around recognizable values, such as 10, 15, 20 minutes. Albeit with
fewer occurrences, this periodicity is also observed for full hours, including 24 and 48 hours.
×106
1.0
0.8
Number of IATs
0.6
0.4
0.2
0.0
1
10
20
30
40
50
60
Minutes
Figure 5.3: Distribution of IAT occurrences (limited to 60 minutes for clarity).
F = (v, hb , v, σ 2 , σ, γ1 , Kurt, P )
59
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
60 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
The machine learning approach alleviates concerns regarding the high dimensionality of
the fingerprint. Through this method, models can effectively identify patterns and prioritize
the most pertinent features, thus minimizing the impact of the representation’s complexity.
In contrast to simpler formats like vectors of values, combining multiple representations
simultaneously, including Markov chains, requires computing distances between matrices. We
consider this feasible for two reasons: 1) fingerprinting occurs outside constrained nodes, and
2) LoRaWAN is designed as a low-throughput protocol.2 This ensures sufficient power and
time between each message to compute fingerprints, even dynamically.
Sequence of
messages A
Sequence of
messages B
Fingerprinting
Fingerprint Fingerprint
A B
Distance computation
Vector of distances
Machine learning
We select similar features as the ones presented in Section 4.4.2, focusing on both stable
and possibly stateful features. We remove low impact ones based on our previous results in
Chapter 4 and empirical study. Notably, we observe that parameters derived from the Adap-
tive Data Rate mechanism, such as Spreading Factor and DataRate, do not contribute signif-
icantly to the fingerprint. Additionally, the Receiving Gateways Distance is excluded because
similar information is already available via other radio-based features saved in gateway-based
vectors (RSSI, SNR, and ESP). A summary of the 8 selected raw features corresponding to
the same three domains (content, time, and radio) is available in Table 5.1.
2 We also note that, in our case, matrices are sparse due to the limited number of states an End-Device
occupies.
60
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
61 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
Considered stateful
Domain Raw feature
(Markov chains)
Content FPort YES
Payload length YES
Time Arrival time YES
(refined as: Hour of day, Day of week)
Inter-Arrival Time YES
Radio RSSI NO
SNR NO
ESP NO
• Arrival time: Defined as a Unix timestamp precise to the millisecond, it denotes the
moment a message is received by a listening station.
• IAT : The time interval between two consecutive uplink messages is computed based on
the first listening stations receiving them. Contrary to others features, the IAT requires
two messages.
Both features target End-Devices programmed to send recurring reports, for instance
communicating sensor data every hour with an occasional management frame or alert. They
are refined in two ways.
First, the arrival time serves as the basis for deriving additional time-based features,
including the hour of reception (in ranges of 10 minutes, e.g. 2:30 PM) and the day of the
week (e.g. Thursday).
Second, the millisecond precision measured by the listening station is too high to create
significant states, due to various latency and treatment delays, long propagation over the air,
and clock drift between the NS and End-Devices. For example, a message every 3600042 or
3599068 milliseconds roughly corresponds to 1 hour in both cases. To reduce the precision,
we bin values and generate multiple sizes of bins, letting the machine learning model pick the
most relevant one.3 Based on the apparent periodicity shown in Figure 5.3, we select peaks
corresponding to more than 1% of the total number of IAT values, translating them into bins
of 1, 5, 10, 15, 20, 30, 40 minutes, as well as 1, 2, and 4 hours.
Finally, as seen in Section 5.3, a majority of IAT values clusters around rounded increments
(e.g. a message every hour). Figure 5.5 illustrates how values theoretically sent at exactly
tj and tj+1 can be misclassified into the wrong bin due to various time perturbations. By
shifting the bins bi and bi+1 , we are able to group clustered values in the correct context. As
3 We empirically find that manually selecting only a subset of bins is detrimental to final results.
61
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
62 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
depicted in Figure 5.3, the time drift never exceeds 3 minutes around significant values. To
err on the side of caution, we shift bins by 5 minutes.
62
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
63 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
Uplink messages
corresponding to ...
a given DevAddr
Sequence 1 Sequence 2 Sequence 3
Figure 5.6: Splitting a set of uplink messages coming from the same End-Device into
sequences of length 5.
In practice, we select c ∈ {5, 10, 20, 50, 100, 200} based on the distribution of uplink
messages received by a DevAddr (see Figure 5.2). For instance, if a set contains 233 messages,
a total of 46, 23, 11, 4, 2, and 1 sequences of lengths 5, 10, 20, 50, 100, and 200 are respectively
created. End-Devices associated with a significant number of messages in the dataset generate
a large number of sequences, resulting in over-representation, particularly for short-length
sequences. To address this issue, we limit the maximum number of sequences per End-Device
for each length to 3.
Finally, we virtually match sequences of correct and incorrect origins to generate both
linked and not linked vectors of distances. With the DevAddr identifying the origin of each
sequence, we can precisely label the vectors, thereby establishing a reliable ground truth.
testing. However, it does not protect from End-Devices updating their DevAddr and randomly being present
in both datasets. Given the relative rarity of re-join processes (as indicated in table 4.1) and the inability to
detect such behavior third-party traffic, we believe this unlikely event does not significantly impact the overall
process.
63
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
64 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
0.8
0.6
0.4
0.2
0.0
Linked Not linked
Class
Figure 5.7: Normalized distances between linked and not linked fingerprints pairs w.r.t.
fingerprint representation.
Similar patterns are observed when analyzing individual features. As depicted in Fig-
ure 5.8, certain features demonstrate superior consistency and discriminatory power com-
pared to others. For instance, the length distribution exhibits discernible differences between
linked and non-linked pairs, contrary to the port variance which remains consistent across
both categories, hinting at lower relevance during classification.
While these findings suggest decent consistency and discriminatory power for fingerprints,
they do not reveal the actual performance variations for each representation during classifica-
tion. Therefore, we train models using only one fingerprint representation and compare their
balanced accuracy.
64
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
65 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
Figure 5.9 illustrates that the holistic fingerprint expectedly outperforms other represen-
tations from the literature. Remarkably, distributions and descriptive statistics yield similar
results for sequences longer than 100 uplink messages, indicating that these simpler represen-
tations remain relevant when sufficient data is available.
1.00
0.95
Balanced accuracy
0.90
0.85 Holistic fingerprint
0.80 Distributions only
Descriptive statistics only
0.75 Vectors only
Markov chains only
0.70 5 10 20 50 100 200
Number of uplink messages per sequence
Figure 5.9: Performance w.r.t. fingerprint representations.
Radio-based features are considered stateless and thus not represented as Markov chain.
This can explain why the Markov chains representation initially produces a comparatively
low balanced accuracy compared to other representations. However, upon removal of Markov
chains from the holistic fingerprint, we observe a slight reduction in balanced accuracy. There-
fore, we retain it as part of the holistic fingerprint representation despite its lower performance.
All other results presented in the following sections are derived from the holistic fingerprint
representation.
65
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
66 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
1.00
Balanced accuracy 0.95
0.90
0.85
0.80 All features
Content-based features only
0.75 Time-based features only
Radio-based features only
0.70 5 10 20 50 100 200
Number of uplink messages per sequence
Figure 5.10: Performance w.r.t. feature domains.
In the all features setting, the performance remains consistently high across all sequence
lengths, with a balanced accuracy ranging from ∼0.95 for sequences of 5 messages to 0.98
for sequences of 200 messages. In other words, even a few messages are sufficient to produce
reliable fingerprints. The improvement in performance with longer sequences is expected, as
more data points contribute to more accurate fingerprints.
Additionally, figure 5.10 illustrates the performance when using only one domain of fea-
tures. Content-based features alone achieve the highest performance, starting at a balanced
accuracy of 0.90 for sequence length of 5 and increasing to nearly match the performance of
the all features configuration for longer sequences. On the other hand, performances with
time-based features alone start lower at 0.76 and gradually increase but remain below the all
features setting.
When using only radio-based features, the balanced accuracy starts at 0.8532 for 5-message
sequences and is somewhat stable regardless of sequence length, ending at 0.8537 for 200-
message sequences. This could be attributed to the selection bias of End-Devices with longer
message sequences, where their radio features may be less reliable than other devices.
In defense-oriented scenarios such as network security monitoring, rapid fingerprinting and
linkage are imperative. Our sequences consist of a minimum of 5 uplink messages. Despite
this brief sequence length, LoRaWAN networks generally maintain a low transmission rate. In
our dataset, it takes approximately 6 hours, as per the median time, to collect 5 messages from
the same End-Device. While this duration might appear significant, the detection process
remains fast when considering the numerical aspect.
Moreover, attackers aiming to remain under the radar would likely avoid disrupting exist-
ing communication patterns and thus be limited by the current low-throughput constraints.
On the other hand, sending rapid successions of messages during the exploitation would result
in both a) an easily identifiable disruption of reference patterns, and b) a reduction of the
overall time of detection.
66
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
67 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
1.00
Balanced accuracy 0.95
0.90
0.85
0.80
0.75 With radio-based features
Without radio-based features
0.70 5 10 20 50 100 200
Number of uplink messages per sequence
Figure 5.11: Performance w.r.t. radio-based features.
67
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
68 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
1.00
Balanced accuracy 0.95
0.90
Baseline (50 receivers)
0.85 5 receivers
0.80 4 receivers
3 receivers
0.75 2 receivers
1 receiver
0.70 5 10 20 50 100 200
Number of uplink messages per sequence
Figure 5.12: Performance w.r.t. the number of controlled listening stations.
did not consider the positions of listening stations, it is conceivable that results could be
further enhanced with strategically deployed stations. In practical terms, even with only a
handful of listening stations, fingerprinting would remain effective in real-world scenarios
Dataset Imperfections: Despite our best efforts to curate the dataset, it is not flawless.
We may have failed to filter out experimental devices deployed by third-party operators
without external notice. However, it remains the most extensive and comprehensive dataset
available compared to related works [195, 196].
Holistic Fingerprint Applicability: While the holistic fingerprinting approach yields ex-
cellent results, its applicability to specific defensive use cases in other wireless protocols (ne-
cessitating rapid recognition) remains uncertain. This aspect warrants further investigation
beyond the scope of this thesis.
68
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
69 CHAPTER 5. FINGERPRINTING LORAWAN DEVICES
Device Similarities and Differentiation: Fingerprinting relies on the premise that dif-
ferent devices produce distinguishable data. However, End-Devices with identical software
and configurations may exhibit similarities in content and time-based features. While radio-
based features can aid in differentiation, distinguishing between End-Devices deployed in the
same location may pose challenges. Further experiments, beyond the scope of the current
dataset, are necessary to investigate this highly specific scenario.
5.9 Conclusion
In this study, we demonstrate the feasibility of reliably fingerprinting LoRaWAN End-Devices,
effectively linking sequences of messages even if they use different identifiers.
We first introduce the holistic fingerprint, a novel agnostic representation applied to Lo-
RaWAN End-Device communication. Through rigorous comparisons, we establish its superi-
ority over existing formats, demonstrating remarkable classification accuracy. Leveraging this
representation, we effectively fingerprint End-Devices using content, time, and radio-based
features.
Our experiments reveal the robustness of our approach, reliably identifying origins of
sequences as short as 5 messages with a balanced accuracy of 0.95, up to 0.98 based on
sequences of 200 messages. Additionally, we explore various simulated scenarios, including
mobile End-Devices and eavesdroppers with limited resources. Despite these challenges, our
fingerprinting technique remains effective across all scenarios, demonstrating its versatility
and resilience.
Beyond offensive applications, our work has broader implications for network security, of-
fering a reliable framework for device identification and classification. By studying LoRaWAN
fingerprinting, we pave the way for enhanced network monitoring and security measures.
69
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Part III
Safeguarding Privacy:
Countermeasures Against
Attacks on LoRaWAN Privacy
70
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 6
Feature-based countermeasures in
LoRaWAN
Contents
6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
6.2 Feature-based countermeasures . . . . . . . . . . . . . . . . . . . 72
6.2.1 Content-based features . . . . . . . . . . . . . . . . . . . . . . . . . 72
6.2.2 Time-based features . . . . . . . . . . . . . . . . . . . . . . . . . . 74
6.2.3 Radio-based features . . . . . . . . . . . . . . . . . . . . . . . . . . 75
6.3 Evaluation methodology . . . . . . . . . . . . . . . . . . . . . . . 76
6.4 Experimental results . . . . . . . . . . . . . . . . . . . . . . . . . 76
6.4.1 Applying countermeasures in isolation . . . . . . . . . . . . . . . . 76
6.4.2 Combining countermeasures . . . . . . . . . . . . . . . . . . . . . . 78
6.4.3 Countermeasures overhead and impact on communications . . . . 79
6.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
In this chapter, we explore ways to counter linkage and fingerprinting attacks (Chapters 4
and 5) based on machine learning models. By studying each feature and selecting or de-
signing specific mitigations w.r.t. LoRaWAN’s constraints, we apply countermeasures to the
dataset and observe the protection gained via the loss in attack efficiency. When combining
all mitigations at once, we reduce linking attacks by 12.5% and device fingerprinting by 7.32%
in realistic settings. We complete our analysis by estimating the inherent overhead of such
countermeasures, for instance the additional bytes transmitted due to padding.
71
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
72 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
6.1 Introduction
The privacy attacks discussed in previous chapters rely on robust machine learning models
that utilize diverse features, many of which overlap (e.g. radio-based features are all derived
from the radio power). Given the effectiveness of these attacks, there is a pressing need
for further investigation into potential countermeasures to safeguard End-Devices and their
users effectively. As explored in Section 2.4, numerous mitigations for wireless networks have
been proposed for various contexts, including dummy traffic [15, 27, 112, 131, 153, 229],
padding [22, 76, 145, 171, 186], random delay between messages [33, 139, 173, 197, 231], or
modifying the radio medium [20, 23, 198].
However, such countermeasures have seldom been studied in the specific context of Lo-
RaWAN, nor in a similar scale. In this chapter, our primary focus lies in formulating viable
mitigations for each of these features, and assessing their feasibility within the framework
of LoRaWAN. Through a methodical approach, we first study the potential deployment sce-
narios of these mitigations and their efficacy in thwarting attacks. Then, we evaluate their
accuracy in isolation, before examining their combined effects under diverse settings. Finally,
we measure how these mitigations impact End-Devices by inherently producing overhead,
both in computation and communication.
Supported by
Features Countermeasure Specified by
LoRaWAN standard
FCnt & FPort Encryption Standard No
Payload length Padding Application Yes
Time-based Delay Application Yes
Radio-based Transmit power modulation Application Yes
72
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
73 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
(see Section 5.5.1). In practice, the server relies on precise values from these fields, rendering
noise-based countermeasures ineffective. Therefore, we opt to encrypt both FCnt and FPort
using existing cryptographic primitives.
Encrypting the FCnt: The FCnt is encrypted using the NwkSEncKey, shared with the NS,
along other already obfuscated network-specific information including MAC commands and
FOpts.
Encrypting the FPort: The FPort is encrypted using the NwkSEncKey and decrypted by
the NS, before relaying the clear-text value to the AS. The FPort is leveraged by both the
NS and AS to correctly route traffic to relevant applications. For instance, when the FPort
is set to 0, it indicates that the payload contains encrypted MAC commands destined for the
NS, while values ranging below 223 are application-specific. This solution necessitates trust
in the NS, but it represents an improvement over the current situation.
With both fields encrypted, their corresponding information are now unavailable to a
passive eavesdropper. In practice, they are removed from datasets used in the machine
learning process.
Separator
Payload Padding
Payload Padding
Payload length
Second, the amount of padding is limited by the maximum payload length. In LoRaWAN,
it depends on the DataRate (DR) (from 51 to 222 bytes, see Table 3.1). Figure 6.2 presents
the observed payload sizes for uplink messages in our real-world dataset already presented
in Sections 4.3 and 5.3. We note that 99.82% of the 71 million uplink messages captured
from June 2020 to August 2023 contain a payload shorter than 51 bytes, allowing for padding
even in the most conservative scenarios where DR 0 to 2 are used (max. 51 bytes).
Given a message of payload size ℓ and a maximum size L, the padding amplitude pA is
calculated as L − ℓ. Based on these parameters, we explore two padding strategies, illustrated
by Figure 6.3.
Max-padding: We fill the payload to its maximum size, ensuring that all uplink messages
have a consistent length of L after padding. Messages that are already at the maximum size
remain unaffected by this process.
73
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
74 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
100
Percentage 80
60
40
20
0
0
50
0
10
15
20
Payload length (bytes)
Figure 6.2: Cumulative distribution function of payload lengths for uplink messages.
Max-padding
Uniform padding
Figure 6.3: Padding strategies, with original payload on the left and padding (hatched) on
the right.
remain unchanged. As discussed in Section [Link], applying this padding method provides
guarantees of Approximate Differential Privacy, a relaxed version of Differential privacy [75].
In practice, we arbitrarily select pA ∈ 5, 10, 40, 80, 200, corresponding to an Approximate
Differential Privacy of parameters (0, pLA ) [75].
74
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
75 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
Original Maximum
message delay
Messages sent
without delay
Messages sent
with delay
Random
delay
In practice, we draw random values from an interval [0..∆].1 To counter linking join-request
and following uplink messages, we select ∆ ∈ {30, 60, 300, 600, 1800, 3600} seconds (from 30
seconds to 1 hours). As the time between consecutive uplink messages is larger than between
join-request and following uplink messages, we increase the possible delay for fingerprint-
ing: ∆ ∈ {60, 300, 600, 900, 1200, 1800, 2400, 3600, 7200, 14400} seconds (from 1 minute to 4
hours).
1 We do not consider negative delay (i.e. sending a report before it should have been transmitted), as it
75
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
76 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
For the linkage attack (Chapter 4), we intentionally keep the Gateway Time Distance,
the DevAddr Difference, Receiving Gateways Distance, Spreading Factor, and DataRate, as
no countermeasure presented here influence them. When targeting the fingerprinting attack
(Chapter 5), we select a sequence length equal to 50 uplink messages. This offers a good
compromise between attack performance, realistic number of available sequences to an eaves-
dropper, and comparatively quick results generation. For combined countermeasures, we
study the impact evolution on all sequence lengths.
In both cases, evaluations are done using all listening stations, and are repeated 15 times
with different random seeds to smooth results, showcased using the Balanced Accuracy (BA).
76
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
77 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
the balanced accuracy of the attack from ∼0.97 to 0.92 (-4.7%) in all cases. Uniform-padding
provides less protection, producing a -2.92% decrease for a padding amplitude of 40 bytes. In
both cases, the decrease in balanced accuracy is ultimately limited, with a small progression
for uniform-padding when increasing the padding amplitude.
1.00
0.95
Balanced accuracy
0.90
0.85
0.80 Baseline (fingerprinting)
0.75 Uniform-padding
Max-padding
0.70 5 10 40 80 200
Maximum padding value (bytes)
Figure 6.6: Impact of the padding on fingerprinting performances.
As seen in figure 6.2, many uplink messages are 11 bytes long, with 78% of the dataset
transmitted using DR 0 - 2, allowing for 51-byte payloads at most. Such configuration can
explain the plateau of performance from 40 bytes and onwards, as all available bytes are
already padded.
Based on these results, padding alone does not significantly reduce the fingerprintability
of LoRaWAN devices, with max-padding showing a slightly higher impact.
1.0
0.9
Balanced accuracy
To summarize, adding a small delay effectively reduces the accuracy of the attack when
77
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
78 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
looking at time-based features only. However, it is useless when taking into account other
features.
Table 6.2: Impact of isolated countermeasures compared to the baseline of both attacks.
These small values can be attributed to models adjusting for the presence of counter-
measures by selecting alternative features. For example, as previously demonstrated in Sec-
tion 4.5.2, the linking attack heavily rely on the FCnt alone to achieve a high balanced
accuracy, which is still available in all but the “Encrypted FCnt ” case.
Additionally, as noise amplifies within certain features, machine learning models can more
easily discard outliers and concentrate on meaningful data. This pattern is observed in
fingerprinting, where increased perturbation of time-based features via random delay actually
improves the attack accuracy instead of decreasing it.
78
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
79 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
our approach, including the last time a DevAddr has been seen, or the list of gateways that
received both join-request and uplink messages.
1.00
0.95
Balanced accuracy
0.90
0.85
0.80 Baseline (fingerprinting)
0.75 Moderate
Aggressive
0.70 5 10 20 50 100 200
Number of uplink messages per sequence
Figure 6.8: Impact of combined countermeasures on fingerprinting performance.
Additionally, we train models using a limited number of listening stations with a similar
methodology as in Section 5.7.4. We observe that diminishing the number of listening stations
amplifies the effectiveness of countermeasures. For example, implementing transmit power
modulation across sequences of 50 messages results in a performance reduction of -0.03%
across all listening stations, yet it escalates to -2.50% when only three listening stations are
present under moderate settings. This suggests that countermeasures could potentially be
more impactful in real-world scenarios, where eavesdroppers are constrained to a limited
number of listening stations.
In summary, aggressive mitigations expectedly show greater impact on the attack’s per-
formance compared to moderate settings. However, all things considered, fingerprinting is
seldom affected, even when all countermeasures are combined. Moreover, their impact dimin-
ishes as the number of uplink messages per sequence grows: models are able to efficiently
classify End-Devices despite aggressive settings, thanks to data aggregations corresponding
to a single identity. Even more aggressive methods would be required to further thwart the
attack, which may not be acceptable from the point of view of LoRaWAN applications.
79
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
80 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
its exact impact without knowing the requirements of underlying applications. Hence, we
focus on encryption of FCnt and FPort, padding, and transmit power modulation.
100
Overhead evolution (% of TOA)
80
60
40
20 Uniform-padding
Max-padding
0 5 10 40 80 200
Maximum padding value (bytes)
Figure 6.9: Impact of padding on Time On Air.
80
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
81 CHAPTER 6. FEATURE-BASED COUNTERMEASURES IN LORAWAN
6.5 Conclusion
Our exploration into implementing feature-based countermeasures to mitigate privacy attacks
has revealed both opportunities and challenges. While it is feasible to introduce mitigations,
our findings highlight the importance of combined approaches. Merely applying isolated
countermeasures without considering the broader context proves ineffective, as demonstrated
by the adaptability of machine learning models. Nevertheless, our experiments demonstrate
that combining multiple mitigations can yield improved results. Despite this, even aggressive
countermeasures induce at most ∼ 15% of decrease in balanced accuracy. Moreover, the
adoption of these mitigations inherently introduces overheads, which poses challenges for
resource-constrained end-devices.
In light of these limitations, alternative strategies warrant investigation, such as reducing
the number of uplink messages associated with a single identity. Pseudonym-based technolo-
gies offer promising avenues for further research and development in enhancing the security
and privacy of LoRaWAN networks.
81
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 7
Privacy-preserving pseudonyms for
LoRaWAN
Contents
7.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
7.2 Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
7.3 Limits of existing pseudonym schemes . . . . . . . . . . . . . . . 84
7.3.1 Legacy schemes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
7.3.2 HASHA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
7.3.3 Pool-based pseudonyms . . . . . . . . . . . . . . . . . . . . . . . . 85
7.3.4 Resolvable pseudonyms . . . . . . . . . . . . . . . . . . . . . . . . 86
7.3.5 Encrypted link-layer pseudonyms . . . . . . . . . . . . . . . . . . . 87
7.4 Adapting pseudonym schemes to LoRaWAN . . . . . . . . . . . 87
7.4.1 Resolvable pseudonyms . . . . . . . . . . . . . . . . . . . . . . . . 87
7.4.2 Sequential pseudonyms . . . . . . . . . . . . . . . . . . . . . . . . 88
7.5 Renewal strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
7.6 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
7.6.1 P1 : Unlinkability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
7.6.2 P2 : Minimal communication overhead . . . . . . . . . . . . . . . . 91
7.6.3 P3 : Low computation/memory overhead . . . . . . . . . . . . . . . 91
7.6.4 P4 : Reliability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
7.6.5 P5 : Legacy support . . . . . . . . . . . . . . . . . . . . . . . . . . 95
7.6.6 Summary and artifacts . . . . . . . . . . . . . . . . . . . . . . . . . 95
7.7 Complementary security considerations . . . . . . . . . . . . . . 95
7.8 Applicability to other LPWAN protocols . . . . . . . . . . . . . 96
7.9 Limitations and future works . . . . . . . . . . . . . . . . . . . . 96
7.10 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
In this chapter, we design rotating pseudonyms tailored to the specific constraints of Lo-
RaWAN. We first analyze the limits of existing schemes, including alternatives already sup-
ported by the protocol’s standard and certificate-based approaches implemented in VANETs.
We then study how resolvable pseudonyms implemented in BLE, and sequential pseudonyms
designed for Wi-Fi, can be adapted to LoRaWAN, evaluating them based on various desired
properties. We reinforce our proposition with theoretical analyses and simulations on a large-
scale dataset, discussing limitations of each solution, and we find that sequential pseudonyms
match the requirements. Finally, we explore potential generalization to other LPWAN proto-
cols.
82
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
83 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
This chapter is partially based on our peer-reviewed and published work Privacy-
Preserving Pseudonyms for LoRaWAN [165]. The present chapter differs mainly by
the following points:
7.1 Introduction
Clear-text and stable identifiers allow eavesdroppers to track devices across time and space [89].
For instance, fingerprinting attacks discussed in Chapter 5, along with other activity infer-
ence threats [131], rely on aggregating numerous messages behind a single address. As seen
in Section 2.4.5, a viable countermeasure is to deploy temporary, unlinkable, and regularly
updated values called pseudonyms. For example, introducing such a mechanism would re-
duce the number of messages per sequence available as ground truth for the machine learning
model, effectively thwarting the fingerprinting attack presented in Chapter 5.
Existing solutions deployed in technologies such as BLE [92] or VANETs [138, 169] were
not specifically designed with strict energy usage optimization in mind. As outlined in
Section 3.1.7, LoRaWAN operates under distinct constraints, including the need for op-
timized energy consumption, low communication throughput, and limited computational
power. Given these differences, further research is needed to assess the feasibility of im-
plementing pseudonyms in the context of LoRaWAN. In this section, we explore how the
DevAddr, crucial for network access and communication, can be replaced by a rotating pseu-
donym to minimize the number of messages associated with a single identity, while preserving
its functionality.
According to Petit et al., a pseudonym [...] should be useable for authentication, but must
not contain any personal identifiable information that could link to the pseudonym holder’s
real identity [169]. First, we extend their definition by proposing multiple properties tailored
for the constraints of the LoRaWAN protocol. Then, we explore existing solutions in light
of those constraints, and select two contenders to evaluate, through both theoretical analysis
and simulations. Finally, we generalize our work to other LPWAN technologies and discuss
security-related issues.
7.2 Properties
We define several properties, stemming from our threat model based on a passive eavesdrop-
per lacking knowledge of the cryptographic keys shared between devices and servers (see
Section 2.3.1). More precisely, the attacker aims to compile sets of uplink messages cor-
responding to a specific device to track and/or fingerprint it. Currently, linking messages
is trivially achieved through the DevAddr. Our goal is to thwart such an approach while
respecting the constraints of the LoRaWAN protocol (see Section 3.1.7)
83
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
84 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
A suitable privacy-preserving pseudonym scheme for LoRaWAN should have the following
properties:
• P1 : Unlinkability: Two messages originating from the same device should not be
linked using the pseudonym (previously DevAddr) and other frame fields available in
clear (e.g. FCnt).
• P2 : Minimal communication overhead: Because of the strict duty cycle and energy
constraints of End-Devices, the communication overhead has to be kept minimal [54].
84
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
85 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
7.3.2 HASHA
Another potential avenue for pseudonyms, called HASHA, has been proposed by Zhang and
Zhang to protect the location privacy of wireless sensor networks [231]. Here, the pseudonym
is encrypted via a rolling key derived based on the hash of the previous payload. Unfortu-
nately, this approach has several shortcomings. First, an eavesdropper can circumvent address
randomization: 1) pseudonyms are generated based on original addresses and a fixed, publicly
known value, and 2) the original address is sent in clear-text via sniffable beacon frames.
Second, the unlinkability property hinges on the attacker missing a single data packet,
leading to a loss of synchronization with the targets. As long as the attacker is synchronized
with senders, they can predict following identifier values. One could argue that LoRaWAN’s
reported 40% packet loss in urban environments [136] would help in that regard, but it does
not offer a strong guarantee.
Third, devices must send acknowledgment frames to maintain synchronization. It would
require LoRaWAN End-Devices to receive and parse downlink for all of their uplink, incur-
ring higher energy consumption. Additionally, high numbers of downlink messages has been
shown to deteriorate the overall performance of the network [150].
Because of these various issues regarding properties P1 (Unlinkability) and P2 (Minimal
communication overhead), we exclude the HASHA scheme from further comparisons.
85
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
86 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
• Signature from the Sender (Vehicle): a mean for the receiver to verify the authenticity
of the sender. When the vehicle sends a message, it signs the message using its private
key, corresponding to the public key included in the certificate.
• Signature from the Trusted Authority: a mean for the receiver to trust the sender. The
vehicle’s certificate is signed by a trusted authority, ensuring authenticity and integrity.
The security of asymmetric cryptography depends particularly on key lengths. As of Jan-
uary 2019, the National Institute of Standards and Technology, USA (NIST) advise selecting
RSA keys of at least 2048 bits. When using Elliptic Curves, keys must be over 224 bits for
the same level of security [39]. These recommendations are relevant for years prior to 2030;
longer keys will be required later. Similar values are given by other state-funded organization,
such as the French National Agency for the Security of Information Systems (ANSSI) [26].
A 224-bit cryptographic key corresponds to 55% of the maximum LoRaWAN payload size,
in the most used DR case (see Section [Link]). Its usage is unrealistic in a protocol as con-
strained as LoRaWAN. Multiple projects propose lightweight alternatives to existing public
key cryptography schemes [202], both in key size and computations. However, they require
(at best) 80 bits for keys alone, not accounting for signatures and pseudonym.
Second, one vehicle stores multiple pseudonyms, valid for a specific period of time. Con-
strained LoRaWAN End-Devices may not be able to manage a pool of lengthy pseudonyms,
requiring additional costly communications to rotate them.
Thus, the implementation of pseudonyms in VANETs goes against properties P2 (Mini-
mal communication overhead), P3 (Low computation/memory overhead), and P5 (Legacy
support), and can not be transferred to the highly constrained LoRaWAN. We exclude
certificate-based pseudonyms from further comparisons.
86
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
87 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
Generation Resolution
=?
87
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
88 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
[Link] Challenges
The limited available space poses two significant challenges. First, 24-bit hashes in BLE
RPAs are highly likely to be resolved by a single key. In contrast, hashes in LoRaWAN can
be as short as 4 bits, leading to collisions, i.e. multiple keys generating the same hash, and
a pseudonym corresponding to more than one End-Device. To avoid identity conflicts during
the resolution, we propose to utilize the MIC (see Section 3.1.3), as illustrated by Figure 7.2.
In case of collisions, the NS can compare the received MIC with the re-computed value based
on the shared key, reliably identifying the actual sender.
Possible keys
...
Hash computation Match MIC computation Match
Device's
identity
Figure 7.2: Device identification via MIC computation when using resolvable pseudonyms.
Second, LoRaWAN networks are designed to host a higher number of devices than BLE,
leading to complexity issues when resolving pseudonyms. For instance, Android supports only
7 associated BLE devices simultaneously1 , in stark contrast with the thousands of concurrent
End-Devices in some LoRaWAN deployments. When the NS receives an uplink message, it
has to try all the keys of N active devices until it successfully resolves the received pseudonym
(1 + N 2−1 tries on average). This open challenge for the uplink message resolution leads us
to explore sequential pseudonyms in Section 7.4.2.
We note that this issue does not arise with downlink messages because an End-Device
simply needs to resolve the pseudonym using its own key. Additionally, End-Devices listen to
downlink channels during time slots closely following their uplink message, mostly receiving
messages intended for them, and therefore requiring only a few resolution attempts.
Finally, to protect the FCnt from being used for pseudonym linkage purposes (see Chap-
ter 4), we propose to encrypt this field like other network-specific information, as explored in
Section [Link], using the NwkSEncKey shared with the NS.
_target.h#1428
88
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
89 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
Pre-generated Corresponding
pseudonyms counters
aaa 1
Device 1 ddd 2
bbb 3
abb 15
Device 2 bba 16
abc 17
aaa 42
Device 3 ccc 43
fff 44
[Link] Challenges
As highlighted previously, LoRaWAN identifiers are tightly constrained in space: only 7 to
25 bits are available. Hence, the probability of collision, i.e. multiple devices using the same
89
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
90 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
sequential pseudonym, is high. For instance, both Device 1 and 2 in Figure 7.4 use the
pseudonym aaa, corresponding to different counters. Following propositions for resolvable
pseudonyms, we leverage the MIC to detect the actual origin when two or more devices
generate the same pseudonym.
The FCnt field can not be kept in clear to avoid linkability attacks (see Chapter 4). It
is redundant with sequential pseudonyms, already functioning as counters themselves. We
propose to extend the NwkAddr field by the 2 vacated bytes of the FCnt to reduce the number
of potential collisions. The sequential scheme can thus utilize pseudonyms of length ranging
from 23 (i.e. 16 + 7) to 41 (i.e. 16 + 25) bits, depending on the network type
7.6 Evaluation
We evaluate both resolvable and sequential pseudonym schemes w.r.t. the properties pre-
sented in Section 7.2. When possible, we complement theoretical analyses by conducting
simulations on a real-world dataset. Already presented in Sections 4.3 and 5.3, this dataset
corresponds to 71 million of uplink messages from third-parties captured from June 2020 to
August 2023.
7.6.1 P1 : Unlinkability
The unlinkability property is related to pseudonyms themselves as well as surrounding meta-
data. First, both resolvable and sequential schemes generate their pseudonyms for each
message via AES as a Cryptographically Secure PRNG (CSPRNG). This guarantees a high
entropy output, meaning the pseudonyms themselves are unlinkable.
Second, both the payload encryption and MIC computation require unique clear-text
FCnt values. Even if we encrypt the FCnt with resolvable pseudonyms or hide it in sequential
pseudonyms, its underlying value is still available to both the End-Device and NS. Thus,
for two payloads with equal inputs (e.g. temperature sensors reporting the same value), the
encrypted output is guaranteed to remain different [65, sec. 4.3.3].
2 In theory, the counter i of sequential pseudonyms could be incremented only every 10 messages.
90
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
91 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
Additionally, the rest of the header (FCtrl and FPort) contains flags and options amount-
ing for a total of 16 bits. In practice, only the FPort exhibits some contextual variation, with
∼3.5 bits of entropy. As outlined in Section [Link], it can be encrypted with limited overhead.
Hence, after encrypting the FCnt and FPort, we consider that both resolvable and sequential
pseudonyms are unlinkable.
Table 7.1: Computation and memory overhead of resolvable and pseudonym schemes.
We note values correspond to overhead, meaning they do not represent existing require-
ments of the current version of the standard.
Transmission: Sequential pseudonyms require only 1 AES encryption to encrypt the coun-
ter and generate the pseudonym on both End-Device and NS, and 1 to encrypt the FPort. On
the other hand, for resolvable pseudonyms, the End-Device has to do 5 encryption operations:
it generates a random value (1 operation), produces the hash via CMAC (2 operations) [194],
as well as encrypts of the FCnt (1 operation) and FPort (1 operation). Both FCnt and FPort
are also encrypted by the NS.
Reception: On the server side, the computation overhead depends on the number of colli-
sions c. By default, a collision is handled by computing the MIC, requiring 4 AES encryption
(2 CMAC operations), totaling to 4c AES encryption operations [65, sec. 4.4]. For resolvable
91
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
92 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
pseudonyms, the NS needs to resolve pseudonyms by computing the hash via CMAC using
the shared keys of all N active devices, so 2N AES encryption operations.3 Finally, it also
has to decrypt the FCnt and FPort, summing up to 4c + 2N + 2 encryption operations. For
sequential pseudonyms, the NS only has to maintain the list of pre-generated pseudonyms by
producing a new one once a message is received (1 operation) as well as decrypting the FPort
(1 operation), resulting in 4c + 2 encryption operations.
For End-Devices, resolvable pseudonyms require 4 encryption operations: 2 for hash com-
putation via CMAC to resolve the pseudonym, and 2 for FCnt and FPort decryption. Se-
quential pseudonyms only need 2 encryption operations, to produce a new pseudonym while
maintaining the pre-generated list length, and to decrypt the FPort.4
In summary, sequential pseudonyms show a significantely lower computation overhead
than resolvable pseudonyms, notably during reception server-side, as well as in both trans-
mission and reception device-side.
7.6.4 P4 : Reliability
We investigate two dimensions of pseudonym robustness. First, as outlined in Section 7.4,
both resolvable and sequential schemes can result in pseudonym collisions, necessitating addi-
tional operations. Second, packet loss might lead to desynchronization between End-Devices
and the NS, requiring a costly rejoin process upon detection.
tions.
4 We note an encryption operation is required when receiving a sequential pseudonym on the End-Device
because the NS do not answer with the last sequential pseudonym used in an uplink. Rather, it communicates
a pseudonym generated specifically for downlink communications, following the existing concept of direction-
based counters (FCntUp and FCntDown).
5 [Link]
ement/
92
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
93 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
Let T be the size of the address space; for instance, a 32-bit pseudonym has T = 232
available addresses. Then p = 1 − (1 − T1 )m . In conclusion, the probability law of X is the
same as 1 + Y where Y follows a binomial law of parameter N − 1, 1 − (1 − T1 )m .
Considering resolvable pseudonyms, half of the available bits are utilized to store a random
b
value. Thus, for a b-bit pseudonym, there are T = 2 2 available addresses. Using the expression
deduced above for Y , we notice that decreasing T increases the probability of collisions.
In conclusion, because resolvable pseudonyms have a smaller address space than sequential
pseudonyms, they have a higher probability of collisions.
Simulation: In addition to the analytical evaluation, we simulate the two pseudonym
schemes on the dataset detailed in Sections 4.3 and 5.3. To do so, we need to know which End-
Devices are active at a given time. Third-party data does not include such an information:
an End-Device could disconnect and its DevAddr be re-assigned later to a new End-Device
without external notice.
Hence, we chronologically analyze incoming uplink messages in sliding time windows of
one month, considering all observed DevAddr as unique active End-Devices during the whole
duration. We replicate the process for each month in the dataset, average the results, and
corresponding to a median of 4789 active End-Devices monthly. Selecting one month as a
time window insures enough End-Devices are considered active, simulating a well-populated
network.
For each received uplink message, we simulate both resolvable and sequential pseudonyms,
and enumerate the number of collisions. When a pseudonym matches n devices, we count
n−1 collisions (because one of them is the real receiver). We configure sequential pseudonyms
with m ∈ {5, 10, 15, 30} pre-generated pseudonyms, with and without utilizing the FCnt for
comparison purposes.
104
Median number of collisions
Resolvable
10 3 Sequential (30)
Sequential (15)
102 Sequential (10)
Sequential (5)
101
100
0
10 15 20 25 30 35 40
Number of bits available
Figure 7.5: Median number of collisions for an uplink message received by the NS (data
points for sequential pseudonyms overlapping for 0 collisions with b ≥ 20).
Figure 7.5 illustrates the median number of collisions for a single uplink message based on
the required number of bits per pseudonym. The sequential scheme generates fewer collisions
than the resolvable scheme across nearly all numbers of available bits, even in scenarios where
the FCnt is not utilized (less than 25 bits). As the number of collisions is computed for a
single incoming uplink message, resolvable pseudonyms producing a median of 1 collision
with 25-bits addresses could increase quickly as networks grow to thousands of active devices.
In summary, sequential pseudonyms outperform resolvable ones, demonstrating signifi-
cantly fewer collisions on the server-side, both analytically and through simulations.
93
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
94 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
by the receiving end, both potentially loosing data and necessitating an energy-consuming
rejoin process upon detection.
One option to detect desynchronization requires some sort of periodic acknowledgment in
downlink messages and has to be done on the End-Device side, as it is impossible for the
NS to directly contact a device. An End-Device would still send multiple uplink messages
before detecting the lack of periodic downlink acknowledgment and then proceed with a
rejoin process. Thus, it is significantly better to avoid desynchronization completely.
As seen in Section 3.1.7, the Packet Loss Rate (PLR) of LoRa networks highly depends on
the environment and deployment conditions, reaching as high as 40% PLR [136], suggesting
that desynchronization could be a recurring problem for sequential pseudonyms. To alleviate
such concerns, we analyze the probability of desynchronization both analytically and through
simulation.
Analytical evaluation: For a packet loss rate P LR, the average number of uplink messages
LR−m −1
sent before m consecutive losses is P 1−P LR [87], where m is the pre-generated list length.
1020
Average number of packets
before desynchronization
1015
PLR=1%
10 PLR=5%
10
PLR=10%
PLR=30%
105 PLR=40%
PLR=50%
100
5 10 15 20 25 30
Pre-generated list length
Figure 7.6: Average number of packets before desynchronization (cropped for clarity).
Table 7.2: Comparison between simulations and theoretically expected fractions of devices
desynchronizing at least once.
Table 7.2 compares the measurements based on the simulation versus the expected prob-
ability computed following the theoretical approach. 30.5% of devices experience desynchro-
nization with only 5 pre-generated pseudonyms, but this number falls to 0.0% with a longer
list of 30 pseudonyms, following expectation. For m ∈ {10, 15}, we measure a higher fraction
of desynchronized devices. This difference can be explained by End-Devices at the edge of
our coverage area, entering and leaving reception range. With smartly deployed gateways,
we believe effective desynchronizations would fall to numbers closer to theoretical values.
In conclusion, such results demonstrate that a long enough list of pre-generated pseudonyms
prevents desynchronization for most to all devices. Based on the theoretical framework and
94
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
95 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
simulations, at least 15 pseudonyms are enough to avoid this feared event, while leading to
acceptable memory overhead (see Section [Link]).
NS should have access to this information. Additionally, the timestamp has to be included into the MIC
computation to avoid tampering.
95
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
96 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
LPWAN protocols can be classified into two groups based on their identifier policy: a) a
lifetime static identifier (e.g. Sigfox), and b) a session pseudonym (e.g. LoRaWAN). In the
second case, the pseudonym is updated only if the device is disconnected from the network or
by explicitly sending a command (e.g. the ForceRejoinReq in LoRaWAN). Neither option
fully satisfies the unlinkability property P1 , demanding regular pseudonym rotations.
At a high level, pseudonyms technically require two building blocks that are already
available to LPWANs: enough bits to transmit them, and cryptographic primitives for their
generation and resolution. First, while LPWAN protocols leverage identifiers of limited size to
reduce time-on-air and energy consumption, we have demonstrated that this number is man-
ageable for sequential pseudonyms. Second, all LPWAN protocols presented here support the
cryptographic primitives required to produce encrypted counters for sequential pseudonyms.
Sigfox [134, Sec. 5.3] and Dash7 [224] operate using AES-128, and NB-IoT is based on the
LTE standard utilizing AES-128 for various tasks [12, Sec. [Link]].
Adapting sequential pseudonyms to other LPWAN protocols would require additional
research to correctly implement them while taking into account all the specificities. For
example, a similar update to the FCnt in LoRaWAN can be done in Sigfox to hide its 12-bits
Message Counter [134, sec. 3.7]. However, we believe regular pseudonym rotations is both
possible and highly beneficial for the privacy of LPWAN protocols.
Additional identification techniques: While we make sure to encrypt and/or hide meta-
data such as the FCnt and FPort, other information can be utilized to link uplink messages,
9 Other protocols are often cited, such as Ingenu [109] or Weightless-P. However, they did not achieve
widespread adoption. In any case, they show similar design trends, including identifiers of limited length,
respectively of 32 and 18 bits.
96
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
97 CHAPTER 7. PRIVACY-PRESERVING PSEUDONYMS FOR LORAWAN
notably the radio signal itself (see RSSI fingerprinting in Chapter 5). In this case, a lack of sta-
ble identifiers to group messages during training would warrant additional work in real-world
scenarios.
Major Version required: Multiple versions of the protocol can cohabit thanks to the
Major Version field. Hence, backward compatibility is guaranteed with introducing a new
major version. Previous implementations of NS would not be able to handle pseudonyms and
could dynamically discard them based on the protocol’s version number, while up-to-date
servers could support both solutions.
Limited to uplink messages: As we focus on the DevAddr, we did not explore possi-
ble pseudonyms replacing the 64-bit DevEUI used in linkage attacks (Chapter 4). Instead of
sending this unique identifier in clear to any eavesdropper, End-Devices could generate a tem-
porary and unlinkable value via their stored NwkKey, resolved by the NS akin to Resolvable
random Private Address in BLE. Then, linking the pseudonym to following uplink message
would be meaningless, as it would correctly separate the identity of a device from its activity.
Further research is required to study the feasibility of such a solution.
7.10 Conclusion
While pseudonyms have been widely studied in the literature and are already deployed in some
wireless technologies, we present the first solution tailored to LoRaWAN. After defining key
properties motivated by both privacy and resource consumption, we analyze various existing
approaches, from VANETs to LoRaWAN’s legacy features. Then, we propose two solutions
thwarting identifier-based tracking: resolvable and sequential pseudonyms. We evaluate them
both theoretically and via simulations on a real-world dataset, based on computation and
memory overheads, and probability of adverse events. We show that the sequential scheme
outperforms its resolvable counterpart, with a reduced memory overhead and limited risks of
collisions and desynchronization when configured to work with 15 pre-generated pseudonyms.
While our work focuses on LoRaWAN, based on our preliminary analyses, such a solution
could also be adapted to other LPWAN protocols, including Sigfox and NB-IoT. Moving
forward, this proposal still needs to be validated with an integration and evaluation in a real
LoRaWAN network.
10 [Link]
97
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Part IV
98
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 8
Identifying IoT devices via encrypted
DNS traffic
Contents
8.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
8.2 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
8.2.1 DNS in consumer-grade IoT . . . . . . . . . . . . . . . . . . . . . . 101
8.2.2 Encrypted DNS protocols . . . . . . . . . . . . . . . . . . . . . . . 101
8.2.3 DNS-over-HTTPS . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
8.3 Threat model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
8.3.1 Position in the network . . . . . . . . . . . . . . . . . . . . . . . . 103
8.3.2 Data access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
8.3.3 Alternative scenarios . . . . . . . . . . . . . . . . . . . . . . . . . . 103
8.4 IoT testbed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
8.4.1 IoT Devices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
8.4.2 DNS resolvers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105
8.5 Identifying IoT devices via DoH traffic . . . . . . . . . . . . . . 105
8.5.1 Features selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105
8.5.2 Features extraction . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
8.6 Experimental methodology . . . . . . . . . . . . . . . . . . . . . . 107
8.6.1 Dataset collection . . . . . . . . . . . . . . . . . . . . . . . . . . . 107
8.6.2 Handling dataset imbalance . . . . . . . . . . . . . . . . . . . . . . 107
8.6.3 Machine learning method selection . . . . . . . . . . . . . . . . . . 107
8.7 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
8.7.1 Comparison of machine learning methods . . . . . . . . . . . . . . 108
8.7.2 Device identification . . . . . . . . . . . . . . . . . . . . . . . . . . 109
8.7.3 Rapid identification . . . . . . . . . . . . . . . . . . . . . . . . . . 110
8.7.4 Importance of length vs IAT . . . . . . . . . . . . . . . . . . . . . 110
8.7.5 Model degradation over time . . . . . . . . . . . . . . . . . . . . . 111
8.7.6 Countermeasures . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
8.8 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113
8.8.1 Implementing padding mitigation . . . . . . . . . . . . . . . . . . . 113
8.8.2 Implications for privacy and broader scope . . . . . . . . . . . . . 113
8.8.3 Limitations and future work . . . . . . . . . . . . . . . . . . . . . . 114
8.8.4 Responsible disclosure . . . . . . . . . . . . . . . . . . . . . . . . . 114
8.9 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114
In order to broaden the scope of our research and adapt our methodology, we now study
on a cornerstone of the Internet (of Things): DNS. In order to protect end-users’ privacy,
multiple recent solutions, such as DNS-over-HTTPS (DoH), encrypt DNS.
99
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
100 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
In this chapter, we explore IoT device identification based solely on encrypted DNS traffic,
and assess if protections are sufficient to protect privacy. Focusing on the use-case of a
smart home, we leverage a extensive IoT dataset and show that consumer-grade devices can
be reliably identified, with up to 0.98 balanced accuracy. We study possible countermeasures
such as padding and show that they effectively thwart our attack. Finally, we discover that
half of the evaluated DNS resolvers do not respect the corresponding standard, compromising
users’ privacy.
This chapter is partially based on our work Does Your Smart Home Tell All? IoT
Device Identification through DNS-over-HTTPS (currently under review). The present
chapter differs mainly by the following points:
• We compare in greater details various machine learning methods with the state
of the art, both in terms of accuracy and prediction time. In addition, we further
analyze the cause of misbehaving DNS resolvers w.r.t. padding responses.
8.1 Introduction
Contrary to LoRaWAN devices studied in previous chapters, consumer-grade IoT devices
generally have access to greater amounts of computation power and battery. Equally ubiqui-
tous in everyday environment, there are often deployed in smart homes, including speakers,
baby monitors, and automation devices [203].
In this ecosystem, DNS plays a crucial role, allowing communication over the Internet be-
tween devices and their respective application servers. Efforts have been made in recent years
to protect this previously clear-text protocol: multiple alternatives encrypting its content are
now available and widely deployed, including DNS-over-HTTPS (DoH). Its ties with human
activity as well as large-scale deployment make it a prime candidate for studies to better
understand associated privacy implications. Indeed, as seen in Section [Link], determining
which device is currently deployed in a house represents a threat to users, allowing outsiders
to profile and identify them [15, 27, 66, 131, 186].
Inferring activity or identifying devices through their DNS traffic have been studied in
various contexts. Table 8.1 outlines the current state of research. While extensive work have
been conducted on clear-text DNS for both Web browsing and IoT devices, as well as DoH for
Web browsing, a notable gap remains in the analysis of IoT-based DoH traffic. Additionally,
it has been shown that IoT traffic exhibits distinct characteristics compared to Web browsing,
such as length of queried domain names and timing of communications [168, 220]. This
prompts for further research to confirm whether existing identification techniques remain
effective in encrypted DNS within IoT networks.
Protocol
Clear-text DNS DNS-over-HTTPS (DoH)
Traffic type
Web browsing [93, 122] [53, 69, 191]
IoT [18, 168, 203] Our work
In this chapter, we address this gap by analyzing a large dataset of DoH traffic generated
by 34 consumer-grade IoT devices representative of a smart home environment. First, we
summarize necessary background on DNS and its encrypted versions, before adapting the
threat model to a smart home setup. Then, we detail our testbed and our DoH-based IoT
100
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
101 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
device identification process. After presenting the evaluation methodology, we highlight key
results and discuss their implication for privacy as well as inherent limitations of our approach.
8.2 Background
In this section, we provide essential information on DNS in IoT devices and its encryption.
Application server
IoT device
Figure 8.1 illustrates a typical domestic IoT environment. As soon as devices are switched
on, they obtain their own IP address from a DHCP server within the local network, as well as
the IP address of a DNS server. A DNS resolver may be preconfigured within devices [220];
however, such occurrences are rare to prevent malfunctioning when the IP of the DNS resolver
is updated. Alternatively, DNS servers can be deployed locally, which does not significantly
change the setup nor the threat model and attack. Then, devices obtain IP addresses of
remote application servers by sending their corresponding domain names to the DNS resolver.
Following communications usually occur over TLS-encrypted HTTP sessions.1
While IoT devices continuously report to application servers [220], the majority of their
DNS requests are generated within the first few minutes of establishing a connection to fetch
updates and transmit status reports. Hence, analyzing communication patterns in the initial
traffic offers an effective and easily automated method for device identification [203] (see
Section [Link]).
101
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
102 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
We focus our study on DoH for two main reasons.3 First, it benefits from a large-scale
support in several major DNS resolvers, such as Cloudflare [2] and Google [4], and in Linux
systems via multiple solutions (e.g. BIND [1] or Unbound [9]). Second, routing DNS traffic
in port 443 along other existing HTTPS communications (e.g. Web browsing) [3] provides
another layer of privacy protection, as DNS needs to be successfully extracted, and then
analyzed. On the other hand, port 853 is utilized exclusively for DoT or DoQ (respectively
over TPC and UDP), facilitating the attack.
8.2.3 DNS-over-HTTPS
DoH is a simple encapsulation: a DNS request is sent as an HTTP POST or GET request to
a DNS server, encoded in the same binary format it would have directly over UDP:
Here, the dns URL query parameter contains a DNS request for [Link]’s A
record, in binary format (and base64 encoded to be used as URL parameter). Alternatively,
some resolvers such as Google and Cloudflare support a non-standardized JSON format for
easier human parsing and development [2, 4]. In this work, we focus on the official and widely
deployed standard using DNS in wire format.
From an eavesdropper’s perspective, there are two main distinctions between DoH and
traditional DNS traffic. First, DoH includes extra messages for TCP and TLS session initia-
tion and termination, in addition to packets carrying requests and responses, whereas DNS
traffic comprises only the latter messages. Second, the size of packets containing queries and
responses varies: while encapsulation in HTTPS adds a constant amount of data to each
message, the message size still depends on the content of the request or response.
All TLS-based DNS encryption (excluding DoDTLS because of possible fragmentation) should follow the same
trends.
102
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
103 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
B
A
IoT device Local router On path equipments DNS server
Figure 8.2: Threat model, with an attacker A in the local network, and an external attacker
B on path.
103
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
104 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
IoT devices
Smart plugs
Local server
Internet
Figure 8.3: Testbed overview, with smart plugs and IoT devices all connected via Wi-Fi to a
local orchestration server.
variety across several categories: Appliance (4), Baby Monitor (2), Camera (5), Doorbell (4),
Hub (2), Light (6), Pet (2), Plug (1), Medical (1), Sensor (2) and Speaker (5). While they do
not support DoH yet, we design a way to realistically simulate this protocol in Section [Link].
This selection is interesting for two reasons. First, some categories of devices are relevant
to eavesdropper. For instance, in case of burglary, knowing whether cameras are installed
inside a house or if a dog is present could be valuable information. Second, we want to test if
104
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
105 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
similar devices (same manufacturer) exhibit similar behaviors and can still be distinguished.
To study this question further, we also include two versions of the same device: Echodot4
and Echodot5.
1.0 0.030
Length
IAT
0.8 0.025
Inter-Arrival Time
Message length
0.020
0.6
0.015
0.4
0.010
0.2 0.005
0.0 0.000
ht
or
ra
er
ell
r
Pe
ito
nc
Hu
Plu
Lig
k
me
ns
orb
ea
on
Se
pli
Ca
Sp
Do
M
Ap
by
Ba
Device category
Figure 8.4: Discriminative behavior of length and IAT.
To detect possible discrepancies between devices, we normalize both message length and
IAT of DNS messages for all devices in our dataset, and examine their respective distributions.
Figure 8.4 illustrates these distributions by category for clarity, but later results are reported
per device. Analysis of the message size reveals significant variability across device categories,
hinting at possible length-based identification. Notably, plugs and hubs exhibit smaller mes-
sage length compared to sensors. IAT values remain consistently low across most device
categories, with pet-oriented devices and appliance showing higher values than speakers and
105
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
106 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
doorbells. Therefore, we opt to include both message length and IAT of DNS messages as
features.
Captured packets
Length
10
0 Time
IoT Device Name, 10, 5, 15, 0, ..., 0, 10, 16.67, 4.08, 0, -1.50
Contrary to our previous holistic approach on LoRaWAN fingerprinting (see Chapter 5),
we utilize less complex data representations, excluding distributions and Markov chains. First,
we empirically find that such complex implementations are not required to reach high ac-
curacy. Second, an alternative use-case of IoT device identification is security, with related
works focusing on fast identification on edge device (e.g. routers) [203]. In the case of Markov
chains, we observe a median of 5 DNS requests per power cycle. Based on previous chapters,
this is far fewer than the required number of messages to witness significant improvements
(>50, see Figure 5.7).
106
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
107 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
107
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
108 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
3) avoiding test snooping and demonstrating that the method correctly generalizes against
unseen data [31].
First, as presented in Chapter 4, we select various well-known multi-class machine learning
methods: Random Forest, K-Nearest Neighbors, Complement Naive Bayes, Logistic Regres-
sion, Support Vector Classification (linear, one-vs-one, one-vs-the-rest). We complete them
by re-implementing the solution deployed by Thompson et al. using a Neural Network [203],
with the same layers.4
Second, in order to evaluate each method to its full potential and correctly compare with
the state of the art where it is used, we select best performing hyperparameters via Halving
Random Search Cross-Validation [133].
Following best practices highlighted in Section 3.2, we split the dataset into training and
held-out data following a 80:20 ratio. Then, cross validation is used by Halving Random
Search, further splitting the training dataset into 5 subsets of 80:20 training:testing to avoid
overfitting on training data. After selecting the hyperparameters producing the highest bal-
anced accuracy for each machine learning method and comparing them, we pick the best
performing method itself. Finally, we train the model using relevant hyperparameters on the
original 80% of the dataset, and validate it against the held-out 20%, producing the final
results.
The machine learning pipeline is run 15 times, each initialized with a different seed for
the PRNGs. Unless otherwise specified, we aggregate model performances across all DNS
resolvers and a single day of replay. We ensure performance consistency across multiple days
through manual confirmation and designate a random day as an illustrative baseline.
8.7 Results
In this section, we present the performance analysis of identification over the 34 IoT devices
deployed in our testbed. We start by selecting the best performing machine learning method,
and then proceed with actual identification.
Table 8.4: Performance of machine learning methods during Halving Random Search
Cross-Validation.
For security purposes, rapid identification is crucial to promptly detect any unidentified
or malfunctioning devices. Contrary to LoRaWAN, DNS traffic is generally more frequent,
possibly generating multiple requests per seconds when an IoT device is switched on, and
requiring a fast responding model. To estimate the time required for a single prediction when
receiving a new message, we utilize the mean duration needed to compute balanced accuracy
during the cross-validation process.
4 We simply adapt the last layer to match the number of output classes (34).
108
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
109 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
In practice, the Random Forest method is faster than many other machine learning ap-
proaches, with a mean scoring time of ∼35 milliseconds.5 Given that this time frame cor-
responds to thousands of predictions for the balanced accuracy, we infer that the prediction
time is sufficiently brief for quick identification.
Petsafe Feeder
Aqara HubM2
Arlo Camera Pro4
Blink Mini Camera
Boifun Baby
Bose Speaker
Coffee Maker Lavazza
Cosori Air Fryer
Echodot4
Echodot5
Reolink Doorbell
Nanoleaf Triangles
While other results are presented as an average across all DNS resolvers, we also study the
identification performance for each of them in Table 8.5. There is minimal variation observed
among DNS resolvers, with stable values observed across all results. Said differently, changing
DNS resolver has no impact on the performance of the attack.
5 For comparison purposes with the state of the art, Neural Network require ∼520 milliseconds.
109
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
110 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
1.0
0.9
Balanced accuracy
0.8
0.7
0.6
0.5 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30
Number of messages
Figure 8.7: Evolution of the balanced accuracy based on the number DNS requests.
110
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
111 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
1.0
0.9
Balanced accuracy
0.8
0.7
0.6
0.5 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14
Days
Figure 8.8: Performance evolution over 15 days, with day 0 used as reference.
8.7.6 Countermeasures
Due to the limited number of raw features in the machine learning process (IAT and length),
options for mitigating our attack are restricted. In this section, we explore potential delays
and grouping DNS requests, as well as padded queries.
111
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
112 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
1.0
0.9
Balanced accuracy
0.8
0.7
No padding (length only)
0.6 Clostest to 128-bytes padding (length only)
Random-Block-Length Padding (length only)
Perfect protection (IAT only)
0.5 Cloudflare Google AdGuard Quad9 CleanBrowsing NextDNS
Resolvers
Figure 8.9: Impact of padding strategies on the performance, compared to no padding and
perfect protection.
Figure 8.9 shows that balanced accuracy is significantly reduced for Cloudflare, Google,
and AdGuard with both strategies. For instance, using 128-byte padding with Cloudflare
results in a ∼33% decrease. Random-Block-Length Padding yields similar results while gen-
erating considerably more overhead (respectively 125% and 67% additional bytes on the wire).
112
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
113 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
8.8 Discussion
In this section, we explore how padding could be deployed based on our threat model and
technical requirements. We analyze the impact of such an attack on privacy and discuss the
limitations of our approach.
igations directly at the IoT manufacturer level, anticipating new deployments, or upgrading existing devices.
8 We can not be certain of the DNS used by Google devices because our DHCP server advertises [Link]
113
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
114 CHAPTER 8. IDENTIFYING IOT DEVICES VIA ENCRYPTED DNS TRAFFIC
Dataset Completeness: Our dataset is limited to traffic originating from 34 devices, which
may appear to be a small sample size. While previous works [203, 220] operate using fewer
devices, hopping to reach internet scale in a lab is not feasible. However, we make sure to
match the diversity of IoT setups, considering similar devices: same usage, or same manufac-
turer and different version. In both cases, our model is able to reliably identify each device,
hinting at its robustness in broader deployments.
Traffic capture: We specifically focus on DNS traffic directly following a device power on,
as it generates a high number of requests [203]. This supposes either that a) an attacker
can monitor such an event, or b) the DNS traffic generated throughout the remainder of
a device’s lifespan is equally identifiable. While previous works have shown that long-term
monitoring via DNS queries is possible [168], its possible adaptation to DoH require additional
investigation.
Identifying DoH itself: Based on our threat model, we consider that DoH is already
extracted from other encrypted traffic on port 443. This assumption is supported by prior
research, which has demonstrated that distinguishing DoH from HTTPS is trivial via known
IP addresses (e.g. [Link] for Cloudflare, or [Link] for Google) [53, 157]. Additionally, it can
be achieved thanks to specific features, including shorter packet length and communication
timing [69, 157, 155, 214].
Static IP addresses: Some IoT devices may not be affected by our attack as they do
not rely on DNS and rather use hard-coded IP addresses. However, due to the maintenance
challenges associated with static IP addresses, very few IoT devices adopt this approach.9
8.9 Conclusion
While DoH brings significant privacy improvements in consumer-grade IoT devices, it does not
guarantee a perfect protection. In this study, we show that the remaining metadata, including
length and IAT, is enough to yield a 0.98 balanced accuracy when identifying 34 devices, across
all studied DNS resolvers. Likewise, devices distributed by the same manufacturer, or even
only differentiated by their version, are reliably classified.
Although other mitigations are too complex to implement, we find that padding is a
straightforward and appropriate countermeasure. However, we discover that some DNS re-
solvers do not respect the corresponding standard and threaten their user’s privacy.
More generally, we demonstrate that the trends observed in protocol operating under
completely different constraints, LoRaWAN, also apply to DoH. We are able to re-adapt
our methodology and successfully build an attack against privacy, while proposing similar
countermeasures.
9 We have yet to see a modern device behave like this in our dataset.
114
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Part V
Conclusion
115
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Chapter 9
Contributions and Future Directions
Contents
9.1 Contributions summary . . . . . . . . . . . . . . . . . . . . . . . . 116
9.2 Perspectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
9.2.1 Short-term . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118
9.2.2 Long-term . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119
This chapter concludes the thesis by first presenting a summary of our contributions, before
exploring both short and long term perspectives.
Q1 What are the privacy challenges of the LoRaWAN protocol? We discovered that
metadata inherently produced by the LoRaWAN protocol, including the size of encrypted
messages and various header fields, timing between messages, as well as radio characteristics,
poses significant privacy risks.
First, we showed that metadata generated during the join process allows an eavesdropper
to effectively link the identity of a device with its activity. By leveraging machine learning
techniques and various features from content, time, and radio domains, we were able to
associate two theoretically unlinkable identifiers, the DevEUI and DevAddr. We also identified
key features, such as the FCnt, that are crucial for reliable linking.
Second, we explored device identification through fingerprinting based on metadata and
traffic patterns observed in sequences of messages. We compared the performance of var-
ious fingerprint representations by formatting features from the content, time, and radio
domains, and demonstrated that combining all of them into a holistic representation yields
the best results. Additionally, we studied multiple scenarios, including mobile devices and
situations where the attacker has access to limited resources, such as a reduced number of
listening stations. In doing so, we showed the robustness of our approach and its reliability
in fingerprinting LoRaWAN devices. While this contribution is initially thought as a privacy
assessment, it can also be leveraged in defensive scenarios such as network monitoring.
116
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
117 CHAPTER 9. CONTRIBUTIONS AND FUTURE DIRECTIONS
such as increased transmitted bytes due to padding. Even with significant network impacts,
feature-based mitigations alone proved insufficient to effectively thwart the attacks.
Second, due to the fact that the attacks were not significantly impacted and the overhead
was considerable, we investigated the root cause. Using a stable identifier throughout the
entire session allows an eavesdropper to group sequences of messages. Replacing the sta-
ble DevAddr, we designed a privacy-preserving pseudonym scheme for LoRaWAN. We first
established desired properties for the scheme and excluded multiple existing solutions due
to the LoRaWAN’s resource constraints. We then further evaluated resolvable pseudonyms,
inspired by Resolvable random Private Addresses in BLE, and sequential pseudonyms, based
on Shroud by Greenstein et al. [90]. Our findings showed that sequential pseudonyms offer a
compelling solution, generating limited extra energy consumption, ensuring a reduced number
of collisions, and proving to be reliable. We concluded our study by expanding our analysis to
other LPWAN protocols, stating that privacy-preserving pseudonyms could be both valuable
and implementable.
Q3 Can this methodology be adapted to other IoT networks? We found that our
machine learning approach can be applied to different IoT contexts to assess privacy, with
similarly high accuracy.
To explore this, we focused on DNS, a protocol widely deployed in consumer-grade IoT
devices, specifically its encrypted version, DNS-over-HTTPS (DoH). Despite the differences in
resource consumption and communication patterns compared to LoRaWAN, we demonstrated
that devices can be identified based solely on packet lengths and timings. Our process shows
high reliability regardless of the DNS resolver used or whether the devices were from the same
manufacturer. Furthermore, we demonstrated rapid identification and model stability over
time. Finally, we explored countermeasures and found that padding has a significant impact
on mitigating our attack. Additionally, we discovered that some DNS resolvers do not respect
padding specifications, compromising users’ privacy.
To conclude, we showed that privacy attacks are possible against different IoT networks
such as LoRaWAN and encrypted DNS, leveraging the same machine learning approach.
While we explore some conceivable countermeasures, various protocol constraints, as well as
the inherent difficulty to update both the standard and devices, suggest that further efforts
are required for truly privacy-preserving communication in IoT networks.
As IoT devices become increasingly intertwined with human activities, privacy should
be a default consideration during protocol design. This assessment should go beyond basic
payload encryption and take metadata into account to fully protect the activity, location, and
identity of end-users.
117
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
118 CHAPTER 9. CONTRIBUTIONS AND FUTURE DIRECTIONS
9.2 Perspectives
We first presented how our contributions are built upon years of prior research in the state
of the art (see Chapter 2). In this section, we highlight how they could pave the way for new
explorations, whether in the near or distant future.
9.2.1 Short-term
While we outlined some limitations and associated future works in previous sections, we now
select five generic, short-term possible extensions of our work.
1 [Link]
2 [Link]
3 This would require serious discussions with the relevant ethics committee and possible safeguards to
118
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
119 CHAPTER 9. CONTRIBUTIONS AND FUTURE DIRECTIONS
9.2.2 Long-term
In addition to low-hanging fruits mainly targeting LoRaWAN, we propose four long-term
perspectives related to machine learning, protocol design, and energy consumption.
119
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
120 CHAPTER 9. CONTRIBUTIONS AND FUTURE DIRECTIONS
ones regrouping numerous major companies or more local alternatives such as SIDO5 ) or even
organize interventions with the public (e.g. Fête de la Science.6 ) To complete this bottom-up
approach, large-scale awareness campaigns could be organized in collaboration with official
structures, such as the CNIL7 , to inform citizens about privacy implications of the IoT and
latest measures taken to protect their privacy.
Finally, successful software such as Signal shows that there is a demand for privacy-
preserving solutions. Although private partnerships were seldom investigated during this
thesis, joint work with companies on IoT devices with privacy as a selling point could be an
interesting avenue to further popularize the concept.
5 [Link]
6 [Link]
7 The CNIL already works with the press: [Link]
120
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
Bibliography
121
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
122 BIBLIOGRAPHY
[16] Kuburat Oyeranti Adefemi Alimi, Khmaies Ouahada, Adnan M. Abu-Mahfouz, and
Suvendi Rimer. A Survey on the Security of Low Power Wide Area Networks: Threats,
Challenges, and Potential Solutions. Sensors, 20(20):5800, January 2020.
[17] Michiel Aernouts, Rafael Berkvens, Koen Van Vlaenderen, and Maarten Weyn. Sigfox
and LoRaWAN Datasets for Fingerprint Localization in Large Urban and Rural Areas.
Data, 3(2):13, June 2018.
[18] Ahmet Aksoy and Mehmet Hadi Gunes. Automated IoT Device Identification using
Network Traffic. In ICC 2019 - 2019 IEEE International Conference on Communica-
tions (ICC), pages 1–7, May 2019.
[19] Hidayet Aksu, A. Selcuk Uluagac, and Elizabeth S. Bentley. Identification of Wearable
Devices with Bluetooth. IEEE Transactions on Sustainable Computing, 6(2):221–230,
April 2021.
[20] Wahhab Albazrqaoe, Jun Huang, and Guoliang Xing. A practical Bluetooth traffic sniff-
ing system: Design, implementation, and countermeasure. IEEE/ACM Transactions
on Networking, 27(1):71–84, 2018.
[21] LoRa Alliance. RP2-1.0.3 LoRaWAN® Regional Parameters, 2021.
[22] Ahmed Alshehri, Jacob Granley, and Chuan Yue. Attacking and protecting tunneled
traffic of smart home devices. In Proceedings of the Tenth ACM Conference on Data
and Application Security and Privacy, pages 259–270, 2020.
[23] Joseph Henry Anajemba, Tang Yue, Celestine Iwendi, Pushpita Chatterjee, Desire
Ngabo, and Waleed S. Alnumay. A secure multiuser privacy technique for wireless
IoT networks using stochastic privacy optimization. IEEE Internet of Things Journal,
9(4):2566–2577, 2021.
[24] Lucas Ancian and Mathieu Cunche. Re-identifying addresses in LoRaWAN networks.
page 28, 2020.
[25] Mahnoor Anjum, Muhammad Abdullah Khan, Syed Ali Hassan, Aamir Mahmood, and
Mikael Gidlund. Analysis of RSSI Fingerprinting in LoRa Networks. page 7, 2019.
[26] ANSSI. Mécanismes cryptographiques | ANSSI.
[Link] 2021.
[27] Noah Apthorpe, Danny Yuxing Huang, Dillon Reisman, Arvind Narayanan, and Nick
Feamster. Keeping the smart home private with smart (er) iot traffic shaping. arXiv
preprint arXiv:1812.00955, 2018.
[28] Chrisil Arackaparambil, Sergey Bratus, Anna Shubina, and David Kotz. On the reli-
ability of wireless fingerprinting using clock skews. In Proceedings of the Third ACM
Conference on Wireless Network Security, pages 169–174, Hoboken New Jersey USA,
March 2010. ACM.
[29] Oscar Arana, Hector Benı́tez-Pérez, Javier Gomez, and Miguel Lopez-Guerrero. Never
Query Alone: A distributed strategy to protect Internet users from DNS fingerprinting
attacks. Computer Networks, 199:108445, November 2021.
[30] F. Armknecht, J. Girao, A. Matos, and R. L. Aguiar. Who Said That? Privacy at Link
Layer. In IEEE INFOCOM 2007 - 26th IEEE International Conference on Computer
Communications, pages 2521–2525, Anchorage, AK, USA, 2007. IEEE.
[31] Daniel Arp, Erwin Quiring, Feargus Pendlebury, Alexander Warnecke, Fabio Pierazzi,
Christian Wressnegger, Lorenzo Cavallaro, and Konrad Rieck. Dos and Don’ts of Ma-
chine Learning in Computer Security. In 31st USENIX Security Symposium (USENIX
Security 22), pages 3971–3988, 2022.
[32] Muhammad Rizwan Asghar, György Dán, Daniele Miorandi, and Imrich Chlamtac.
Smart meter data privacy: A survey. IEEE Communications Surveys & Tutorials,
19(4):2820–2835, 2017.
122
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
123 BIBLIOGRAPHY
[33] Tomer Ashur, Jeroen Delvaux, Sanghan Lee, Pieter Maene, Eduard Marin, Svetla
Nikova, Oscar Reparaz, Vladimir Rožić, Dave Singelée, Bohan Yang, and Bart Pre-
neel. A privacy-preserving device tracking system using a low-power wide-area network:
CANS 2017. Cryptology and Network Security - 16th International Conference, CANS
2017, Revised Selected Papers, pages 347–369, January 2018.
[34] Eyuel D. Ayele, Nirvana Meratnia, and Paul JM Havinga. Towards a new opportunistic
IoT network architecture for wildlife monitoring system. In 2018 9th IFIP International
Conference on New Technologies, Mobility and Security (NTMS), pages 1–5. IEEE,
2018.
[35] Leonardo Babun, Hidayet Aksu, Lucas Ryan, Kemal Akkaya, Elizabeth S. Bentley, and
A. Selcuk Uluagac. Z-IoT: Passive Device-class Fingerprinting of ZigBee and Z-Wave
IoT Devices. In ICC 2020 - 2020 IEEE International Conference on Communications
(ICC), pages 1–7, June 2020.
[36] Solveig Badillo, Balazs Banfai, Fabian Birzele, Iakov I. Davydov, Lucy Hutchin-
son, Tony Kam-Thong, Juliane Siebourg-Polster, Bernhard Steiert, and Jitao David
Zhang. An Introduction to Machine Learning. Clinical Pharmacology & Therapeutics,
107(4):871–885, April 2020.
[37] Miloud Bagaa, Tarik Taleb, Jorge Bernal Bernabe, and Antonio Skarmeta. A Machine
Learning Security Framework for Iot Systems. IEEE Access, 8:114066–114077, 2020.
[38] Tariq Al Balushi, Ayoub Al Hosni, Hashim Al Theeb Ba Omar, and Dawood Al Abri.
A LoRaWAN-based Camel Crossing Alert and Tracking System. In 2019 IEEE 17th
International Conference on Industrial Informatics (INDIN), volume 1, pages 1035–
1040, July 2019.
[39] Elaine Barker. Recommendation for Key Management: Part 1 – General. Technical
Report NIST Special Publication (SP) 800-57 Part 1 Rev. 5, National Institute of
Standards and Technology, May 2020.
[40] David Barrera, Glenn Wurster, and P C van Oorschot. Back to the Future: Revisiting
IPv6 Privacy Extensions. 2011.
[41] Arup Barua, Md Abdullah Al Alamin, Md Shohrab Hossain, and Ekram Hossain. Se-
curity and privacy threats for bluetooth low energy in iot and wearable devices: A
comprehensive survey. IEEE Open Journal of the Communications Society, 3:251–281,
2022.
[42] Gustavo EAPA Batista, Ana LC Bazzan, and Maria Carolina Monard. Balancing
training data for automated annotation of keywords: A case study. Wob, 3:10–8, 2003.
[43] E Bäumker, A Miguel Garcia, and P Woias. Minimizing power consumption of LoRa
®
and LoRaWAN for low-power wireless sensor nodes. Journal of Physics: Conference
Series, 1407(1):012092, November 2019.
[44] Johannes K Becker, David Li, and David Starobinski. Tracking Anonymized Bluetooth
Devices. Proceedings on Privacy Enhancing Technologies, 2019(3):50–65, July 2019.
[45] John Bellardo and Stefan Savage. 802.11 {Denial-of-Service} attacks: Real vulnerabil-
ities and practical solutions. In 12th USENIX Security Symposium (USENIX Security
03), 2003.
[46] Eleanor Birrell, Jay Rodolitz, Angel Ding, Jenna Lee, Emily McReynolds, Jevan Hut-
son, and Ada Lerner. SoK: Technical Implementation and Human Impact of Internet
Privacy Regulations. In 45th IEEE Symposium on Security and Privacy, San Francisco,
CA, USA, May 2024. IEEE.
[47] Bram Bonne, Arno Barzan, Peter Quax, and Wim Lamotte. WiFiPi: Involuntary
tracking of visitors at mass events. In 2013 IEEE 14th International Symposium on ”A
World of Wireless, Mobile and Multimedia Networks” (WoWMoM), pages 1–6, Madrid,
June 2013. IEEE.
123
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
124 BIBLIOGRAPHY
[48] Martin Bor and Utz Roedig. LoRa transmission parameter selection. In 2017 13th
International Conference on Distributed Computing in Sensor Systems (DCOSS), pages
27–34. IEEE, 2017.
[49] Carsten Bormann, Mehmet Ersue, and Ari Keränen. Terminology for Constrained-
Node Networks. Request for Comments RFC 7228, Internet Engineering Task Force,
May 2014.
[50] Taoufik Bouguera, Jean-François Diouris, Jean-Jacques Chaillout, Randa Jaouadi, and
Guillaume Andrieux. Energy consumption model for sensor nodes based on LoRa and
LoRaWAN. Sensors, 18(7):2104, 2018.
[51] Danah Boyd. It’s Complicated: The Social Lives of Networked Teens. Yale University
Press, 2014.
[52] Peter Brown, Collectif, Georges Duby, and Philippe Ariès. Histoire de la vie privée,
tome 1 : De L’Empire romain à l’an mil. Seuil, October 1999.
[53] Jonas Bushart and Christian Rossow. Padding ain’t enough: Assessing the privacy
guarantees of encrypted DNS. 2020.
[54] Lluı́s Casals, Bernat Mir, Rafael Vidal, and Carles Gomez. Modeling the energy per-
formance of LoRaWAN. Sensors, 17(10):2364, 2017.
[55] Guillaume Celosia and Mathieu Cunche. Saving private addresses: An analysis of
privacy issues in the bluetooth-low-energy advertising mechanism. In Proceedings of
the 16th EAI International Conference on Mobile and Ubiquitous Systems: Computing,
Networking and Services, pages 444–453, Houston Texas USA, November 2019. ACM.
[56] Guillaume Celosia and Mathieu Cunche. Discontinued privacy: Personal data leaks
in apple bluetooth-low-energy continuity protocols. Proceedings on Privacy Enhancing
Technologies, 2020:26–46, 2020.
[57] Guillaume Celosia and Mathieu Cunche. Valkyrie: A generic framework for verifying
privacy provisions in wireless networks. In Proceedings of the 13th ACM Conference on
Security and Privacy in Wireless and Mobile Networks, pages 278–283, Linz Austria,
July 2020. ACM.
[58] Bharat S. Chaudhari, Marco Zennaro, and Suresh Borkar. LPWAN Technologies:
Emerging Application Characteristics, Requirements, and Design Considerations. Fu-
ture Internet, 12(3):46, March 2020.
[59] Nitesh V. Chawla, Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer.
SMOTE: Synthetic minority over-sampling technique. Journal of artificial intelligence
research, 16:321–357, 2002.
[60] Roger Clarke. Introduction to dataveillance and information privacy and definitions of
terms. www. anu. edu. au/people/Roger. Clarke/DV/Intro. html, 1997.
[63] LAN/MAN Standards Committee. IEEE Recommended Practice for Privacy Consid-
erations for IEEE 802(R) Technologies, 2020.
[64] LoRa Alliance Technical Committee. LoRaWAN® Regional Parameters
v1.1rA. [Link]
parameters-v1-1ra, October 2017.
124
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
125 BIBLIOGRAPHY
[66] Bogdan Copos, Karl Levitt, Matt Bishop, and Jeff Rowe. Is Anybody Home? Infer-
ring Activity From Smart Home Network Traffic. In 2016 IEEE Security and Privacy
Workshops (SPW), pages 245–251, San Jose, CA, May 2016. IEEE.
[67] Semtech Corporation. Semtech LoRa Technology Overview | Semtech.
[Link]
[68] Semtech Corporation. LoRaWAN – simple rate adaptation recommended algorithm,
2016.
[69] Levente Csikor, Himanshu Singh, Min Suk Kang, and Dinil Mon Divakaran. Privacy
of DNS-over-HTTPS: Requiem for a Dream? In 2021 IEEE European Symposium on
Security and Privacy (EuroS&P), pages 252–271, Vienna, Austria, September 2021.
IEEE.
[70] Mathieu Cunche, Mohamed Ali Kaafar, and Roksana Boreli. I know who you will
meet this evening! Linking wireless devices using Wi-Fi probe requests. In 2012 IEEE
International Symposium on a World of Wireless, Mobile and Multimedia Networks
(WoWMoM), pages 1–9, San Francisco, CA, USA, June 2012. IEEE.
[71] Syed Muhammad Danish, Arfa Nasir, Hassaan Khaliq Qureshi, Ayesha Binte Ashfaq,
Shahid Mumtaz, and Jonathan Rodriguez. Network intrusion detection system for
jamming attack in LoRaWAN join procedure. In 2018 IEEE International Conference
on Communications (ICC), pages 1–6. IEEE, 2018.
[72] Aveek K. Das, Parth H. Pathak, Chen-Nee Chuah, and Prasant Mohapatra. Uncovering
Privacy Leakage in BLE Network Traffic of Wearable Fitness Trackers. In Proceedings
of the 17th International Workshop on Mobile Computing Systems and Applications,
pages 99–104, St. Augustine Florida USA, February 2016. ACM.
[73] Ralph Droms and Steve Alexander. DHCP Options and BOOTP Vendor Extensions.
Request for Comments RFC 2132, Internet Engineering Task Force, March 1997.
[74] Lea Dujić Rodić, Toni Perkovic, Maja Skiljo, and Petar Solic. Privacy Leakage of
Lorawan Smart Parking Occupancy Sensors, March 2022.
[75] Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni
Naor. Our data, ourselves: Privacy via distributed noise generation. In Advances in
Cryptology-EUROCRYPT 2006: 24th Annual International Conference on the Theory
and Applications of Cryptographic Techniques, St. Petersburg, Russia, May 28-June 1,
2006. Proceedings 25, pages 486–503. Springer, 2006.
[76] Kevin P. Dyer, Scott E. Coull, Thomas Ristenpart, and Thomas Shrimpton. Peek-a-
boo, i still see you: Why efficient traffic analysis countermeasures fail. In 2012 IEEE
Symposium on Security and Privacy, pages 332–346. IEEE, 2012.
[77] Aviv Engelberg and Avishai Wool. Classification of Encrypted IoT Traffic Despite
Padding and Shaping, October 2021.
[78] Sinem Coleri Ergen, Huseyin Serhat Tetikol, Mehmet Kontik, Raffi Sevlian, Ram Ra-
jagopal, and Pravin Varaiya. RSSI-fingerprinting-based mobile phone localization with
route constraints. IEEE Transactions on Vehicular Technology, 63(1):423–428, 2013.
[79] Bernat Carbonés Fargas and Martin Nordal Petersen. GPS-free geolocation using LoRa
in low-power WANs. In 2017 Global Internet of Things Summit (Giots), pages 1–6.
IEEE, 2017.
[80] Ellis Fenske, Dane Brown, Jeremy Martin, Travis Mayberry, Peter Ryan, and Erik Rye.
Three years later: A study of mac address randomization in mobile devices and when
it succeeds. Proceedings on Privacy Enhancing Technologies, 2021.
[81] Mohamed Amine Ferrag, Lei Shu, Xing Yang, Abdelouahid Derhab, and Leandros
Maglaras. Security and Privacy for Green IoT-Based Agriculture: Review, Blockchain
Solutions, and Challenges. IEEE Access, 8:32031–32053, 2020.
125
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
126 BIBLIOGRAPHY
[82] Rachel L. Finn, David Wright, and Michael Friedewald. Seven Types of Privacy. In
Serge Gutwirth, Ronald Leenes, Paul de Hert, and Yves Poullet, editors, European Data
Protection: Coming of Age, pages 3–32. Springer Netherlands, Dordrecht, 2013.
[83] Luciano Floridi. The Informational Nature of Personal Identity. Minds and Machines,
21(4):549–566, November 2011.
[84] Julien Freudiger, Mohammad Hossein Manshaei, Jean-Yves Le Boudec, and Jean-Pierre
Hubaux. On the Age of Pseudonyms in Mobile Ad Hoc Networks. In 2010 Proceedings
IEEE INFOCOM, pages 1–9, March 2010.
[85] Harald T. Friis. A note on a simple transmission formula. Proceedings of the IRE,
34(5):254–256, 1946.
[86] Flavio D. Garcia, David Oswald, Timo Kasper, and Pierre Pavlidès. Lock it and still
lose it—on the ({In) Security} of automotive remote keyless entry systems. In 25th
USENIX Security Symposium (USENIX Security 16), 2016.
[87] Paul Ginsparg. How many coin flips on average does it take to get n consecutive heads.
[Link] cit. cornell. edu/info2950 2012sp/mh. pdf, 2005.
[88] Google. MAC Randomization Behavior, Android documentation.
[Link]
2024.
[89] Ben Greenstein, Ramakrishna Gummadi, Jeffrey Pang, Mike Y Chen, Tadayoshi Kohno,
Srinivasan Seshan, and David Wetherall. Can Ferris Bueller Still Have His Day Off?
Protecting Privacy in the Wireless Era. page 6, 2007.
[90] Ben Greenstein, Damon McCoy, Jeffrey Pang, Tadayoshi Kohno, Srinivasan Seshan, and
David Wetherall. Improving wireless privacy with an identifier-free link layer protocol.
In Proceeding of the 6th International Conference on Mobile Systems, Applications, and
Services - MobiSys ’08, page 40, Breckenridge, CO, USA, 2008. ACM Press.
[91] Ulrich Greveler, Peter Glösekötterz, Benjamin Justusy, and Dennis Loehr. Multimedia
content identification through smart meter power usage profiles. In Proceedings of the
International Conference on Information and Knowledge Engineering (IKE), page 1.
The Steering Committee of The World Congress in Computer Science, Computer . . . ,
2012.
[92] Core Specification Working Group. Bluetooth® Core Specification 5.4.
[Link] February
2023.
[93] Saikat Guha and Paul Francis. Identity Trail: Covert Surveillance Using DNS. In
Nikita Borisov and Philippe Golle, editors, Privacy Enhancing Technologies, volume
4776, pages 153–166. Springer Berlin Heidelberg, Berlin, Heidelberg, 2007.
[94] Serge Gutwirth. Privacy and the Information Age. Rowman & Littlefield, 2002.
[95] Dalton A. Hahn, Arslan Munir, and Vahid Behzadan. Security and Privacy Issues in
Intelligent Transportation Systems: Classification and Challenges. IEEE Intell. Transp.
Syst. Mag., 13(1):181–196, 2021.
[96] Jetmir Haxhibeqiri, Eli De Poorter, Ingrid Moerman, and Jeroen Hoebeke. A Survey of
LoRaWAN for IoT: From Technology to Application. Sensors, 18(11):3995, November
2018.
[97] Inc Helium Systems. Buy An Organizationally Unique Identifier | Helium Doc-
umentation. [Link]
devaddrs, 2024.
[98] Noelia Hernández, Manuel Ocaña, Jose M. Alonso, and Euntai Kim. Continuous space
estimation: Increasing WiFi-based indoor localization resolution without increasing the
site-survey effort. Sensors, 17(1):147, 2017.
126
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
127 BIBLIOGRAPHY
[99] Frank Hessel, Lars Almon, and Flor Álvarez. ChirpOTLE: A Framework for Practical
LoRaWAN Security Evaluation. In Proceedings of the 13th ACM Conference on Security
and Privacy in Wireless and Mobile Networks, pages 306–316, July 2020.
[100] Frank Hessel, Lars Almon, and Matthias Hollick. LoRaWAN Security: An Evolvable
Survey on Vulnerabilities, Attacks and their Systematic Mitigation. ACM Transactions
on Sensor Networks, 18(4):1–55, November 2022.
[101] Paul E. Hoffman and Patrick McManus. DNS Queries over HTTPS (DoH). Request
for Comments RFC 8484, Internet Engineering Task Force, October 2018.
[102] Zi Hu, Liang Zhu, John Heidemann, Allison Mankin, Duane Wessels, and Paul E.
Hoffman. Specification for DNS over Transport Layer Security (TLS). Request for
Comments RFC 7858, Internet Engineering Task Force, May 2016.
[103] Anna Huang. Similarity measures for text document clustering. In Proceedings of the
Sixth New Zealand Computer Science Research Student Conference (NZCSRSC2008),
Christchurch, New Zealand, volume 4, pages 9–56, 2008.
[104] Leping Huang, Kanta Matsuura, Hiroshi Yamane, and Kaoru Sezaki. Enhancing wire-
less location privacy using silent period. In IEEE Wireless Communications and Net-
working Conference, 2005, volume 2, pages 1187–1192. IEEE, 2005.
[105] Christian Huitema, Sara Dickinson, and Allison Mankin. DNS over Dedicated QUIC
Connections. Request for Comments RFC 9250, Internet Engineering Task Force, May
2022.
[106] The Things Industries. Building a gateway with Raspberry Pi and IC880A.
[Link] 2022.
[107] The Things Industries. NetID and DevAddr Prefix Assignments.
[Link] 2024.
[108] Information Sciences Institue University of Southern California. Internet Protocol.
Request for Comments RFC 791, Internet Engineering Task Force, September 1981.
[109] Development Ingenu. An Idiot’s Guide to RPMA Development – Part 2, August 2016.
[110] European Telecommunications Standards Institute. ETSI EN 300 220-2, 2018.
[111] Taher Issoufaly and Pierre Ugo Tournoux. BLEB: Bluetooth Low Energy Botnet for
large scale individual tracking. In 2017 1st International Conference on Next Generation
Computing Applications (NextComp), pages 115–120, Mauritius, July 2017. IEEE.
[112] Arshad Jhumka, Matthew Leeke, and Sambid Shrestha. On the use of fake sources for
source location privacy: Trade-offs between energy and privacy. The Computer Journal,
54(6):860–874, 2011.
[113] Yu Jiang, Linning Peng, Aiqun Hu, Sheng Wang, Yi Huang, and Lu Zhang. Physical
layer identification of LoRa devices using constellation trace figure. EURASIP Journal
on Wireless Communications and Networking, 2019(1):223, December 2019.
[114] C Joshitha, P Kanakaraja, Mallela Divya Bhavani, Yerramsetti Naga Venkata Raman,
and Tadepalli Sravani. LoRaWAN based Cattle Monitoring Smart System. In 2021
7th International Conference on Electrical Energy Systems (ICEES), pages 548–552,
February 2021.
[115] Shachar Kaufman, Saharon Rosset, Claudia Perlich, and Ori Stitelman. Leakage in
data mining: Formulation, detection, and avoidance. ACM Transactions on Knowledge
Discovery from Data (TKDD), 6(4):1–21, 2012.
[116] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei
Ye, and Tie-Yan Liu. Lightgbm: A highly efficient gradient boosting decision tree.
Advances in neural information processing systems, 30:3146–3154, 2017.
127
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
128 BIBLIOGRAPHY
[118] Kevin H. Knuth. Optimal Data-Based Binning for Histograms, September 2013.
[119] Ron Kohavi. A study of cross-validation and bootstrap for accuracy estimation and
model selection. page 8, 1995.
[120] Roman Kolcun, Diana Andreea Popescu, Vadim Safronov, Poonam Yadav, Anna Maria
Mandalari, Richard Mortier, and Hamed Haddadi. Revisiting IoT Device Identification,
July 2021.
[121] Maciej Korczyński and Andrzej Duda. Markov chain fingerprinting to classify encrypted
traffic. In IEEE INFOCOM 2014 - IEEE Conference on Computer Communications,
pages 781–789, April 2014.
[122] Srinivas Krishnan and Fabian Monrose. DNS prefetching and its privacy implications:
When good things go bad. In Proceedings of the 3rd USENIX Conference on Large-scale
Exploits and Emergent Threats: Botnets, Spyware, Worms, and More, pages 10–10,
2010.
[123] Jacob Leon Kröger, Philip Raschke, and Towhidur Rahman Bhuiyan. Privacy implica-
tions of accelerometer data: A review of possible inferences. In Proceedings of the 3rd
International Conference on Cryptography, Security and Privacy - ICCSP ’19, pages
81–87, Kuala Lumpur, Malaysia, 2019. ACM Press.
[124] A. R. Kulaib, R. M. Shubair, M. A. Al-Qutayri, and Jason W. P. Ng. An overview of
localization techniques for Wireless Sensor Networks. In 2011 International Conference
on Innovations in Information Technology, pages 167–172, Abu Dhabi, United Arab
Emirates, April 2011. IEEE.
[125] Marc Langheinrich. Privacy by Design — Principles of Privacy-Aware Ubiquitous Sys-
tems. In Gerhard Goos, Juris Hartmanis, Jan Van Leeuwen, Gregory D. Abowd, Barry
Brumitt, and Steven Shafer, editors, Ubicomp 2001: Ubiquitous Computing, volume
2201, pages 273–291. Springer Berlin Heidelberg, Berlin, Heidelberg, 2001.
[126] Arash Habibi Lashkari, Mir Mohammad Seyed Danesh, and Behrang Samadi. A survey
on wireless security protocols (WEP, WPA and WPA2/802.11 i). In 2009 2nd IEEE
International Conference on Computer Science and Information Technology, pages 48–
52. IEEE, 2009.
[127] Alexandru Lavric and Valentin Popa. A LoRaWAN: Long range wide area networks
study. In 2017 International Conference on Electromechanical and Power Systems
(SIELMEN), pages 417–420, October 2017.
[128] Huang-Chen Lee and Kai-Hsiang Ke. Monitoring of Large-Area IoT Sensors Using a
LoRa Wireless Mesh Network System: Design and Evaluation. IEEE Transactions on
Instrumentation and Measurement, 67(9):2177–2187, September 2018.
[129] Guillaume Lemaı̂tre, Fernando Nogueira, and Christos K. Aridas. Imbalanced-learn: A
python toolbox to tackle the curse of imbalanced datasets in machine learning. Journal
of Machine Learning Research, 18(17):1–5, 2017.
[130] Martine S. Lenders, Christian Amsüss, Cenk Gündogan, Marcin Nawrocki, Thomas C.
Schmidt, and Matthias Wählisch. Securing Name Resolution in the IoT: DNS over
CoAP. Proceedings of the ACM on Networking, 1(CoNEXT2):1–25, September 2023.
[131] Patrick Leu, Ivan Puddu, Aanjhan Ranganathan, and Srdjan Čapkun. I Send, Therefore
I Leak: Information Leakage in Low-Power Wide Area Networks. In Proceedings of the
11th ACM Conference on Security & Privacy in Wireless and Mobile Networks, pages
23–33, Stockholm Sweden, June 2018. ACM.
128
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
129 BIBLIOGRAPHY
[132] Baoli Li and Liping Han. Distance Weighted Cosine Similarity Measure for Text Classifi-
cation. In David Hutchison, Takeo Kanade, Josef Kittler, Jon M. Kleinberg, Friedemann
Mattern, John C. Mitchell, Moni Naor, Oscar Nierstrasz, C. Pandu Rangan, Bernhard
Steffen, Madhu Sudan, Demetri Terzopoulos, Doug Tygar, Moshe Y. Vardi, Gerhard
Weikum, Hujun Yin, Ke Tang, Yang Gao, Frank Klawonn, Minho Lee, Thomas Weise,
Bin Li, and Xin Yao, editors, Intelligent Data Engineering and Automated Learning –
IDEAL 2013, volume 8206, pages 611–618. Springer Berlin Heidelberg, Berlin, Heidel-
berg, 2013.
[133] Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar.
Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization, June
2018.
[134] Alexandre Liaut. Sigfox Radio Specifications v1.7. 2023.
[135] Fen Liu, Jing Liu, Yuqing Yin, Wenhan Wang, Donghai Hu, Pengpeng Chen, and
Qiang Niu. Survey on WiFi-based indoor positioning techniques. IET Communications,
14(9):1372–1383, June 2020.
[136] Qian Liu, Yanyan Mu, Jin Zhao, Jingxia Feng, and Bin Wang. Characterizing packet
loss in city-scale LoRaWAN deployment: Analysis and implications. In 2020 IFIP
Networking Conference (Networking), pages 704–712. IEEE, 2020.
[137] Yan-Ting Liu, Bo-Yi Lin, Xiao-Feng Yue, Zong-Xuan Cai, Zi-Xian Yang, Wei-Hong Liu,
Song-Yi Huang, Jun-Lin Lu, Jing-Wen Peng, and Jen-Yeu Chen. A solar powered long
range real-time water quality monitoring system by LoRaWAN. In 2018 27th Wireless
and Optical Communication Conference (WOCC), pages 1–2, April 2018.
[138] Zishan Liu, Lin Zhang, Wei Ni, and Iain B. Collings. Uncoordinated Pseudonym
Changes for Privacy Preserving in Distributed Networks. IEEE Transactions on Mobile
Computing, 19(6):1465–1477, June 2020.
[139] Javier Lopez, Ruben Rios, Feng Bao, and Guilin Wang. Evolving privacy: From sensors
to the Internet of Things. Future Generation Computer Systems, 75:46–57, October
2017.
[140] Slim Loukil, Lamia Chaari Fourati, Anand Nayyar, and K.-W.-A. Chee. Analysis of
LoRaWAN 1.0 and 1.1 protocols security mechanisms. Sensors, 22(10):3717, 2022.
[141] Chaoyi Lu, Baojun Liu, Zhou Li, Shuang Hao, Haixin Duan, Mingming Zhang, Chun-
ying Leng, Ying Liu, Zaifeng Zhang, and Jianping Wu. An End-to-End, Large-Scale
Measurement of DNS-over-Encryption: How Far Have We Come? In Proceedings of
the Internet Measurement Conference, pages 22–35, Amsterdam Netherlands, October
2019. ACM.
[142] Gonzalo Marı́n, Pedro Casas, and Germán Capdehourat. RawPower: Deep Learning
based Anomaly Detection from Raw Network Traffic Measurements. In Proceedings of
the ACM SIGCOMM 2018 Conference on Posters and Demos, pages 75–77, Budapest
Hungary, August 2018. ACM.
[143] Jeremy Martin, Douglas Alpuche, Kristina Bodeman, Lamont Brown, Ellis Fenske,
Lucas Foppe, Travis Mayberry, Erik C. Rye, Brandon Sipes, and Sam Teplov. Handoff
All Your Privacy: A Review of Apple’s Bluetooth Low Energy Continuity Protocol.
arXiv preprint arXiv:1904.10600, 2019.
[144] Alexander Mayrhofer. The EDNS(0) Padding Option. Request for Comments RFC
7830, Internet Engineering Task Force, May 2016.
[145] Alexander Mayrhofer. Padding Policies for Extension Mechanisms for DNS (EDNS(0)).
Request for Comments RFC 8467, Internet Engineering Task Force, October 2018.
[146] Kerry McKay, Lawrence Bassham, Meltem Sönmez Turan, and Nicky Mouha. Report
on lightweight cryptography. Technical report, National Institute of Standards and
Technology, 2016.
129
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
130 BIBLIOGRAPHY
[147] Afef Mdhaffar, Tarak Chaari, Kaouthar Larbi, Mohamed Jmaiel, and Bernd Freisleben.
IoT-based health monitoring via LoRaWAN. In IEEE EUROCON 2017 -17th Interna-
tional Conference on Smart Technologies, pages 519–524, July 2017.
[148] Simon Meier, Benedikt Schmidt, Cas Cremers, and David Basin. The TAMARIN Prover
for the Symbolic Analysis of Security Protocols. In David Hutchison, Takeo Kanade,
Josef Kittler, Jon M. Kleinberg, Friedemann Mattern, John C. Mitchell, Moni Naor,
Oscar Nierstrasz, C. Pandu Rangan, Bernhard Steffen, Madhu Sudan, Demetri Ter-
zopoulos, Doug Tygar, Moshe Y. Vardi, Gerhard Weikum, Natasha Sharygina, and Hel-
mut Veith, editors, Computer Aided Verification, volume 8044, pages 696–701. Springer
Berlin Heidelberg, Berlin, Heidelberg, 2013.
[149] Markus Miettinen, Samuel Marchal, Ibbad Hafeez, N. Asokan, Ahmad-Reza Sadeghi,
and Sasu Tarkoma. IoT SENTINEL: Automated Device-Type Identification for Secu-
rity Enforcement in IoT. In 2017 IEEE 37th International Conference on Distributed
Computing Systems (ICDCS), pages 2177–2184, June 2017.
[150] Konstantin Mikhaylov, Juha Petäjäjärvi, and Ari Pouttu. Effect of downlink traffic on
performance of LoRaWAN LPWA networks: Empirical study. In 2018 IEEE 29th An-
nual International Symposium on Personal, Indoor and Mobile Radio Communications
(PIMRC), pages 1–6. IEEE, 2018.
[151] Abhishek Kumar Mishra, Aline Carneiro Viana, and Nadjib Achir. Introducing bench-
marks for evaluating user-privacy vulnerability in WiFi. In 2023 IEEE 97th Vehicular
Technology Conference (VTC2023-Spring), pages 1–7. IEEE, 2023.
[152] Satyajayant Misra and Guoliang Xue. Efficient anonymity schemes for clustered wireless
sensor networks. International Journal of Sensor Networks, 1(1/2):50, 2006.
[153] Frederik Möllers. Energy-Efficient Dummy Traffic Generation for Home Automation
Systems. Proceedings on Privacy Enhancing Technologies, 2020(4):376–393, October
2020.
[154] Mohammad Mezanur Rahman Monjur, Joseph Heacock, Rui Sun, and Qiaoyan Yu. An
Attack Analysis Framework for LoRaWAN applied Advanced Manufacturing. In 2021
IEEE International Symposium on Technologies for Homeland Security (HST), pages
1–7, Boston, MA, USA, November 2021. IEEE.
[155] Mohammadreza MontazeriShatoori, Logan Davidson, Gurdip Kaur, and Arash Habibi
Lashkari. Detection of doh tunnels using time-series classification of encrypted traffic.
In 2020 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf
on Pervasive Intelligence and Computing, Intl Conf on Cloud and Big Data Computing,
Intl Conf on Cyber Science and Technology Congress (DASC/PiCom/CBDCom/Cyber-
SciTech), pages 63–70. IEEE, 2020.
[156] Christoph Neumann, Olivier Heen, and Stéphane Onno. An empirical study of passive
802.11 Device Fingerprinting. arXiv:1404.6457 [cs], April 2014.
[157] Frank Nijeboer. Detection of HTTPS Encrypted DNS Traffic. 2020.
[158] Helen Nissenbaum. Privacy as contextual integrity. Wash. L. Rev., 79:119, 2004.
[159] Hassan Noura, Tarif Hatoum, Ola Salman, Jean-Paul Yaacoub, and Ali Chehab. Lo-
RaWAN security survey: Issues, threats and possible mitigation techniques. Internet
of Things, 12:100303, December 2020.
[160] Rebekah Overdorf, Mark Juarez, Gunes Acar, Rachel Greenstadt, and Claudia Diaz.
How Unique is Your .onion?: An Analysis of the Fingerprintability of Tor Onion Ser-
vices. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Com-
munications Security, pages 2021–2036, Dallas Texas USA, October 2017. ACM.
[161] Wubin Pan, Guang Cheng, and Yongning Tang. Wenc: Https encrypted traffic clas-
sification using weighted ensemble learning and markov chain. In 2017 IEEE Trust-
com/BigDataSE/ICESS, pages 50–57. IEEE, 2017.
130
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
131 BIBLIOGRAPHY
[162] Jeffrey Pang, Ben Greenstein, Ramakrishna Gummadi, Srinivasan Seshan, and David
Wetherall. 802.11 user fingerprinting. In Proceedings of the 13th Annual ACM Inter-
national Conference on Mobile Computing and Networking - MobiCom ’07, page 99,
Montr&233;al, Québec, Canada, 2007. ACM Press.
[163] UK Parliament. The Product Security and Telecommunications In-
frastructure (Security Requirements for Relevant Connectable Products).
[Link] 2023.
[164] Article 29 Working Party. Guidelines on Data Protection Impact Assessment (DPIA),
2017.
[165] Samuel Pélissier, Jan Aalmoes, Abhishek Kumar Mishra, Mathieu Cunche, Vincent
Roca, and Didier Donsez. Privacy-preserving pseudonyms for LoRaWAN. In 17th
ACM Conference on Security and Privacy in Wireless and Mobile Networks (WiSec
2024), 2024.
[166] Samuel Pélissier, Mathieu Cunche, Vincent Roca, and Didier Donsez. Device re-
identification in LoRaWAN through messages linkage. In Proceedings of the 15th ACM
Conference on Security and Privacy in Wireless and Mobile Networks, pages 98–103,
2022.
[167] Samuel Pélissier, Abhishek Kumar Mishra, Mathieu Cunche, Vincent Roca, and Didier
Donsez. Introducing Multi-Domain Fingerprints for Lorawan Device Linkage. Available
at SSRN 4728641.
[168] Roberto Perdisci, Thomas Papastergiou, Omar Alrawi, and Manos Antonakakis.
IoTFinder: Efficient Large-Scale Identification of IoT Devices via Passive DNS Traffic
Analysis. In 2020 IEEE European Symposium on Security and Privacy (EuroS&P),
pages 474–489, Genoa, Italy, September 2020. IEEE.
[169] Jonathan Petit, Florian Schaub, Michael Feiri, and Frank Kargl. Pseudonym schemes
in vehicular networks: A survey. IEEE communications surveys & tutorials, 17(1):228–
255, 2014.
[170] Tara Petric, Mathieu Goessens, Loutfi Nuaymi, Laurent Toutain, and Alexander Pelov.
Measurements, performance and analysis of LoRa FABIAN, a real-world implementa-
tion of LPWAN. In 2016 IEEE 27th Annual International Symposium on Personal,
Indoor, and Mobile Radio Communications (PIMRC), pages 1–7, Valencia, Spain,
September 2016. IEEE.
[171] Antônio J. Pinheiro, Paulo Freitas de Araujo-Filho, Jeandro de M. Bezerra, and Di-
vanilson R. Campelo. Adaptive packet padding approach for smart home networks: A
tradeoff between privacy and performance. IEEE Internet of Things Journal, 8(5):3930–
3938, 2020.
[172] Nico Podevijn, David Plets, Jens Trogh, Luc Martens, Pieter Suanet, Kim Hendrikse,
and Wout Joseph. TDoA-Based Outdoor Positioning with Tracking Algorithm in a
Public LoRa Network. Wireless Communications and Mobile Computing, 2018:1–9,
May 2018.
[173] Nelson Prates, Andressa Vergütz, Ricardo T. Macedo, Aldri Santos, and Michele
Nogueira. A defense mechanism for timing-based side-channel attacks on IoT traf-
fic. In GLOBECOM 2020-2020 IEEE Global Communications Conference, pages 1–6.
IEEE, 2020.
[174] Michael Quan, Eduardo Navarro, and Benjamin Peuker. Wi-fi localization using rssi
fingerprinting. 2010.
[175] S. R. Jino Ramson, S. Vishnu, A. Alfred Kirubaraj, Theodoros Anagnostopoulos, and
Adnan M. Abu-Mahfouz. A LoRaWAN IoT-Enabled Trash Bin Level Monitoring Sys-
tem. IEEE Transactions on Industrial Informatics, 18(2):786–795, February 2022.
131
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
132 BIBLIOGRAPHY
[176] Tirumaleswar Reddy.K, Dan Wing, and Prashanth Patil. DNS over Datagram Trans-
port Layer Security (DTLS). Request for Comments RFC 8094, Internet Engineering
Task Force, February 2017.
[177] Jingjing Ren, Daniel J Dubois, David Choffnes, Anna Maria Mandalari, Roman Kol-
cun, and Hamed Haddadi. Information exposure from consumer iot devices: A multi-
dimensional, network-informed measurement approach. In Proceedings of the Internet
Measurement Conference, 2019.
[178] Pieter Robyns, Peter Quax, Wim Lamotte, and William Thenaers. Gr-lora: An efficient
LoRa decoder for GNU Radio. Zenodo Ed, 10:5281, 2017.
[179] Ribana Roscher, Bastian Bohn, Marco F. Duarte, and Jochen Garcke. Explainable
machine learning for scientific insights and discoveries. Ieee Access, 8:42200–42216,
2020.
[180] Erik C. Rye, Robert Beverly, and kc claffy. Follow the Scent: Defeating IPv6 Prefix
Rotation Privacy. In Proceedings of the 21st ACM Internet Measurement Conference,
pages 739–752, November 2021.
[181] Said Jawad Saidi, Oliver Gasser, and Georgios Smaragdakis. One Bad Apple Can Spoil
Your IPv6 Privacy, March 2022.
[182] Yaman Sangar and Bhuvana Krishnaswamy. WiChronos: Energy-efficient modulation
for long-range, large-scale wireless networks. In Proceedings of the 26th Annual Inter-
national Conference on Mobile Computing and Networking, pages 1–14, London United
Kingdom, April 2020. ACM.
[183] Cristiana Santos, Nataliia Bielova, Vincent Roca, Mathieu Cunche, Gilles Mertens,
Karel Kubicek, and Hamed Haddadi. Feedback to the European Data Protection
Board’s Guidelines 2/2023 on Technical Scope of Art. 5(3) of ePrivacy Directive, Febru-
ary 2024.
[184] Amardeo C. Sarma and João Girão. Identities in the Future Internet of Things. Wireless
Personal Communications, 49(3):353–363, May 2009.
[185] David W. Scott. Sturges’ rule. WIREs Computational Statistics, 1(3):303–306, Novem-
ber 2009.
[186] Narmeen Shafqat, Daniel J. Dubois, David Choffnes, Aaron Schulman, Dinesh Bhara-
dia, and Aanjhan Ranganathan. ZLeaks: Passive Inference Attacks on Zigbee based
Smart Homes. arXiv:2107.10830 [cs], July 2021.
[187] Vishal Sharma, Ilsun You, Giovanni Pau, Mario Collotta, Jae Deok Lim, and Jeong Nyeo
Kim. LoRaWAN-Based Energy-Efficient Surveillance by Drones for Intelligent Trans-
portation Systems. Energies, 11(3):573, March 2018.
[188] Guanxiong Shen, Junqing Zhang, Alan Marshall, Linning Peng, and Xianbin Wang.
Radio frequency fingerprint identification for LoRa using spectrogram and CNN. In
IEEE INFOCOM 2021-IEEE Conference on Computer Communications, pages 1–10.
IEEE, 2021.
[189] Meng Shen, Mingwei Wei, Liehuang Zhu, and Mingzhong Wang. Classification of En-
crypted Traffic With Second-Order Markov Chains and Application Attribute Bigrams.
IEEE Transactions on Information Forensics and Security, 12(8):1830–1843, August
2017.
[190] Taeshik Shon and Jongsub Moon. A hybrid machine learning approach to network
anomaly detection. Information Sciences, 177(18):3799–3821, 2007.
[191] Sandra Siby, Marc Juarez, Claudia Diaz, Narseo Vallina-Rodriguez, and Carmela Tron-
coso. Encrypted DNS –> Privacy? A Traffic Analysis Perspective. arXiv:1906.09682
[cs], October 2019.
132
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
133 BIBLIOGRAPHY
[192] Ritesh Kumar Singh, Michiel Aernouts, Mats De Meyer, Maarten Weyn, and Rafael
Berkvens. Leveraging LoRaWAN technology for precision agriculture in greenhouses.
Sensors, 20(7):1827, 2020.
[193] Daniel J. Solove. The myth of the privacy paradox. Geo. Wash. L. Rev., 89:1, 2021.
[194] Junhyuk Song, Radha Poovendran, Jicheol Lee, and Tetsu Iwata. The aes-cmac algo-
rithm. Technical report, 2006.
[195] Pietro Spadaccino, Francesco Giuseppe Crinó, and Francesca Cuomo. LoRaWAN Be-
haviour Analysis through Dataset Traffic Investigation. Sensors, 22(7):2470, March
2022.
[196] Pietro Spadaccino, Domenico Garlisi, Francesca Cuomo, Giorgio Pillon, and Patrizio
Pisani. Discovery privacy threats via device de-anonymization in LoRaWAN. In 2021
19th Mediterranean Communication and Computer Networking Conference (MedCom-
Net), pages 1–8, June 2021.
[197] Vijay Srinivasan, John Stankovic, and Kamin Whitehouse. Protecting your daily in-
home activity information from a wireless snooping attack. In Proceedings of the 10th In-
ternational Conference on Ubiquitous Computing, pages 202–211, Seoul Korea, Septem-
ber 2008. ACM.
[198] Paul Staat, Simon Mulzer, Stefan Roth, Veelasha Moonsamy, Markus Heinrichs, Rainer
Kronberger, Aydin Sezgin, and Christof Paar. IRShield: A countermeasure against
adversarial physical-layer wireless sensing. In 2022 IEEE Symposium on Security and
Privacy (SP), pages 1705–1721. IEEE, 2022.
[199] Jothi Prasanna Shanmuga Sundaram, Wan Du, and Zhiwei Zhao. A Survey on LoRa
Networking: Research Problems, Current Solutions and Open Issues. arXiv:1908.10195
[cs, eess], August 2019.
[200] Latanya Sweeney. K-anonymity: A model for protecting privacy. International journal
of uncertainty, fuzziness and knowledge-based systems, 10(05):557–570, 2002.
[201] Xiaopeng Tan, Shaojing Su, Zhiping Huang, Xiaojun Guo, Zhen Zuo, Xiaoyong Sun,
and Longqing Li. Wireless sensor networks intrusion detection based on SMOTE and
the random forest algorithm. Sensors, 19(1):203, 2019.
[202] Vishal A. Thakor, Mohammad Abdur Razzaque, and Muhammad RA Khandaker.
Lightweight cryptography algorithms for resource-constrained IoT devices: A review,
comparison and research opportunities. IEEE Access, 9:28177–28193, 2021.
[203] Oliver Thompson, Anna Maria Mandalari, and Hamed Haddadi. Rapid IoT Device
Identification at the Edge. In Proceedings of the 2nd ACM International Workshop on
Distributed Machine Learning, pages 22–28, December 2021.
[204] Min Ye Thu, Wunna Htun, Yan Lin Aung, Pyone Ei Ei Shwe, and Nay Min Tun. Smart
Air Quality Monitoring System with LoRaWAN. In 2018 IEEE International Confer-
ence on Internet of Things and Intelligence System (IOTAIS), pages 10–15, November
2018.
[205] Stefano Tomasin, Simone Zulian, and Lorenzo Vangelista. Security analysis of lorawan
join procedure for internet of things networks. In 2017 IEEE Wireless Communications
and Networking Conference Workshops (WCNCW), pages 1–6. IEEE, 2017.
[206] Nuno Torres, Pedro Pinto, and Sérgio Ivan Lopes. Security Vulnerabilities in
LPWANs—An Attack Vector Analysis for the IoT Ecosystem. Applied Sciences,
11(7):3176, January 2021.
[207] Kun-Lin Tsai, Fang-Yie Leu, Ilsun You, Shuo-Wen Chang, Shiung-Jie Hu, and Hoony-
ong Park. Low-power AES data encryption architecture for a LoRaWAN. IEEE Access,
7:146348–146357, 2019.
133
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
134 BIBLIOGRAPHY
[208] A. Selcuk Uluagac, Sakthi V. Radhakrishnan, Cherita Corbett, Antony Baca, and Ra-
heem Beyah. A passive technique for fingerprinting wireless devices with Wired-side
Observations. In 2013 IEEE Conference on Communications and Network Security
(CNS), pages 305–313, October 2013.
[209] Publications Office of the European Union. Regulation (EU) 2016/679 of the European
Parliament and of the Council of 27 April 2016 on the protection of natural persons
with regard to the processing of personal data and on the free movement of such data,
and repealing Directive 95/46/EC (General Data Protection Regulation) (Text with
EEA relevance). [Link]
11bd-11e6-ba9a-01aa75ed71a1/language-en, April 2016.
[210] Muhammad Usama, Junaid Qadir, Aunn Raza, Hunain Arif, Kok-Lim Alvin Yau, Yehia
Elkhatib, Amir Hussain, and Ala Al-Fuqaha. Unsupervised machine learning for net-
working: Techniques, applications and research challenges. IEEE access, 7:65579–65615,
2019.
[211] Christian Vaas, Mohammad Khodaei, Panos Papadimitratos, and Ivan Martinovic.
Nowhere to hide? Mix-Zones for Private Pseudonym Change using Chaff Vehicles.
In 2018 IEEE Vehicular Networking Conference (VNC), pages 1–8, Taipei, Taiwan,
December 2018. IEEE.
[212] Eef Van Es, Harald Vranken, and Arjen Hommersom. Denial-of-Service Attacks on Lo-
RaWAN. In Proceedings of the 13th International Conference on Availability, Reliability
and Security, pages 1–6, Hamburg Germany, August 2018. ACM.
[213] Mathy Vanhoef, Célestin Matte, Mathieu Cunche, Leonardo S. Cardoso, and Frank
Piessens. Why MAC Address Randomization is not Enough: An Analysis of Wi-Fi
Network Discovery Mechanisms. In Proceedings of the 11th ACM on Asia Conference
on Computer and Communications Security, ASIA CCS ’16, pages 413–424, New York,
NY, USA, May 2016. Association for Computing Machinery.
[214] Dmitrii Vekshin, Karel Hynek, and Tomas Cejka. DoH Insight: Detecting DNS over
HTTPS by machine learning. In Proceedings of the 15th International Conference on
Availability, Reliability and Security, pages 1–8, Virtual Event Ireland, August 2020.
ACM.
[215] Paul A. Vixie. Extensions to DNS (EDNS). Internet Draft draft-ietf-dnsind-edns-03,
Internet Engineering Task Force, August 1998.
[216] Paul A. Vixie. Extensions to DNS (EDNS1). Internet Draft draft-ietf-dnsext-edns1-03,
Internet Engineering Task Force, August 2002.
[217] Tien Dang Vo-Huu, Triet Dang Vo-Huu, and Guevara Noubir. Fingerprinting Wi-Fi
Devices Using Software Defined Radios. In Proceedings of the 9th ACM Conference on
Security & Privacy in Wireless and Mobile Networks, pages 3–14, Darmstadt Germany,
July 2016. ACM.
[220] Yinxin Wan, Kuai Xu, Feng Wang, and Guoliang Xue. Characterizing and Mining
Traffic Patterns of IoT Devices in Edge Networks. IEEE Transactions on Network
Science and Engineering, 8(1):89–101, January 2021.
[221] Samuel Warren and Louis Brandeis. The Right to Privacy. In Tom Goldstein, editor,
Killing the Messenger, pages 1–21. Columbia University Press, 1890.
134
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
135 BIBLIOGRAPHY
[222] Stephan Wesemeyer, Ioana Boureanu, Zach Smith, and Helen Treharne. Extensive
security verification of the LoRaWAN key-establishment: Insecurities & patches. In
2020 IEEE European Symposium on Security and Privacy (EuroS&P), pages 425–444.
IEEE, 2020.
[223] Alan F Westin. Privacy And Freedom. 1968.
[224] Maarten Weyn, Glenn Ergeerts, Rafael Berkvens, Bartosz Wojciechowski, and Yordan
Tabakov. DASH7 alliance protocol 1.0: Low-power, mid-range sensor and actuator
communication. In 2015 IEEE Conference on Standards for Communications and Net-
working (CSCN), pages 54–59. IEEE, 2015.
[225] Maarten Weyn, Glenn Ergeerts, Luc Wante, Charles Vercauteren, and Peter Hellinckx.
Survey of the DASH7 Alliance Protocol for 433 MHz Wireless Sensor Communication.
International Journal of Distributed Sensor Networks, 9(12):870430, December 2013.
[226] Tim Wicinski. DNS Privacy Considerations. Request for Comments RFC 9076, Internet
Engineering Task Force, July 2021.
[227] Sanjay Yadav and Sanyam Shukla. Analysis of k-fold cross-validation over hold-out
validation on colossal datasets for quality classification. In 2016 IEEE 6th International
Conference on Advanced Computing (IACC), pages 78–83. IEEE, 2016.
[228] Yi Yang, Min Shao, Sencun Zhu, and Guohong Cao. Towards statistically strong source
anonymity for sensor networks. ACM Transactions on Sensor Networks, 9(3):1–23, May
2013.
[229] Lin Yao, Lin Kang, Pengfei Shang, and Guowei Wu. Protecting the sink location privacy
in wireless sensor networks. Personal and Ubiquitous Computing, 17(5):883–893, June
2013.
[230] Shui Yu, Guofeng Zhao, Wanchun Dou, and Simon James. Predicted Packet Padding
for Anonymous Web Browsing Against Traffic Analysis Attacks. IEEE Transactions on
Information Forensics and Security, 7(4):1381–1393, August 2012.
[231] Qiong Zhang and Kewang Zhang. Protecting Location Privacy in IoT Wireless Sen-
sor Networks through Addresses Anonymity. Security and Communication Networks,
2022:e2440313, April 2022.
[232] Yue Zhang and Zhiqiang Lin. When Good Becomes Evil: Tracking Bluetooth Low En-
ergy Devices via Allowlist-based Side Channel and Its Countermeasure. In Proceedings
of the 2022 ACM SIGSAC Conference on Computer and Communications Security,
pages 3181–3194, Los Angeles CA USA, November 2022. ACM.
135
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés
FOLIO ADMINISTRATIF
.THESE
.. . . . . . . . DE L’INSA
. . . . .. . . . . . . ..LYON,
. . . . . . ..MEMBRE DE
.. .. . .. ... . .. . . ..L’UNIVERSITE
.. .. .. . ... .. ... . ..DE LYON
.. . .. .. .. .
Prénoms : Samuel
Spécialité : Informatique
RÉSUMÉ :
Les dernières décennies ont été témoins de l’émergence et de la prolifération d’objets connectés, communément
appelés Internet des Objets (IdO). Cet écosystème divers correspond à une large gamme de dispositifs spécialisés,
allant de la caméra IP aux capteurs détectant les fuites d’eau, chacun conçu pour répondre à des objectifs et des
contraintes de consommation d’énergie, de puissance de calcul ou de coût. Le développement rapide de nombreuses
technologies et leur connexion en réseau s’accompagne de la génération d’un important volume de données, soule-
vant des préoccupations en matière de vie privée, en particulier dans des domaines sensibles tels que la santé ou les
maisons connectées.
Dans cette thèse, nous exploitons les techniques d’apprentissage automatique (machine learning) pour explorer les
problèmes liés à la vie privée des objets connectés via leurs protocoles réseau. Tout d’abord, nous étudions les
attaques possibles contre LoRaWAN, un protocole longue distance et à faible coût d’énergie. Nous explorons la
relation entre deux identifiants du protocole et montrons que leur séparation théorique peut être contrecarrée en
utilisant les métadonnées produites lors de la connexion au réseau. En nous appuyant sur une approche multi-
domaines (contenu, temps, radio), nous démontrons que ces métadonnées permettent à un attaquant d’identifier les
objets connectés de manière unique malgré le chiffrement du trafic, ouvrant la voie au traçage ou à la ré-identification.
Nous explorons ensuite les possibles contre-mesures, en analysant systématiquement les données utilisées lors de
ces attaques et en proposant des techniques pour les obfusquer ou réduire leur pertinence. Nous démontrons que
seule une approche combinée offre une réelle protection. Par ailleurs, nous proposons et évaluons diverses solutions
de pseudonymes temporaires adaptées aux contraintes de LoRaWAN, en particulier la consommation énergétique.
Enfin, nous adaptons notre méthodologie d’apprentissage automatique à DNS, un protocole largement déployé dans
l’IdO grand public. À nouveau basées sur les métadonnées, notre attaque permet d’identifier les objets connectés,
malgré le chiffrement du flux DNS-over-HTTPS. Explorant les contre-mesures potentielles, nous observons un non-
respect des standards liés au padding, entraînant la compromission partielle de la vie privée des utilisateurs.
MOTS-CLÉS : Vie privée ; Internet des Objets ; Objets connectés ; Sécurité ; Protocoles ; Réseaux sans fil ; Mé-
tadonnées ; Empreinte ; Re-identification ; Inférence d’activités ; LoRaWAN ; LPWAN ; DNS ; DNS-over-HTTPS.
Directeurs de thèse : Mathieu CUNCHE (Professeur des Universités, INSA Lyon) & Vincent ROCA (Chargé de
recherche, Inria)
Composition du Jury :
Thèse accessible à l'adresse : [Link] © [S. Pélissier], [2024], INSA Lyon, tous droits réservés