PAM 2023 Conference Proceedings
PAM 2023 Conference Proceedings
Marcel Flores
Marco Fiore (Eds.)
LNCS 13882
Marco Fiore
IMDEA Networks Institute
Madrid, Spain
© The Editor(s) (if applicable) and The Author(s), under exclusive license
to Springer Nature Switzerland AG 2023, corrected publication 2023
6 chapters are licensed under the terms of the Creative Commons Attribution 4.0 International License (http://
[Link]/licenses/by/4.0/). For further details see license information in the chapters.
This work is subject to copyright. All rights are reserved by the Publisher, whether the whole or part of
the material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation,
broadcasting, reproduction on microfilms or in any other physical way, and transmission or information
storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now
known or hereafter developed.
The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication
does not imply, even in the absence of a specific statement, that such names are exempt from the relevant
protective laws and regulations and therefore free for general use.
The publisher, the authors, and the editors are safe to assume that the advice and information in this book
are believed to be true and accurate at the date of publication. Neither the publisher nor the authors or the
editors give a warranty, expressed or implied, with respect to the material contained herein or for any errors
or omissions that may have been made. The publisher remains neutral with regard to jurisdictional claims in
published maps and institutional affiliations.
This Springer imprint is published by the registered company Springer Nature Switzerland AG
The registered company address is: Gewerbestrasse 11, 6330 Cham, Switzerland
Preface
We are excited to present the proceedings of the 24th Annual Passive and Active Mea-
surement PAM Conference. With this program, PAM continues its tradition as a venue
for thorough and compelling, but often early-stage and emerging, research on networks,
Internet measurement, and the emergent systems that they host. This year’s conference
took place on March 21–23, 2023. Based on learnings from recent years, this year’s PAM
was again virtual, both to accommodate the realities of modern travel, and to ensure the
accessibility of the conference to attendees who may not otherwise be able to travel long
distances.
This year we received 80 double-blind submissions from over 100 different insti-
tutions of which the Technical Program Committee (TPC) selected 27 for publication,
resulting in a similar sized program to previous years. As with last year, submissions
could be of either long or short form, and our ultimate program featured 18 long papers,
a notable increase over last year. This year also featured the option to submit papers to
an explicit replication track, in which submissions could explicitly explore past find-
ings with new experiments and conditions. We received 8 such submissions, 4 of which
ultimately were accepted. The 27 papers of the final program illustrate how network mea-
surements can provide important insights for different types of networks and networked
systems and cover topics such as applications, performance, network infrastructure and
topology, measurement tools, and security and privacy. It thus provides a comprehensive
view of current state-of-the-art and emerging ideas in this important domain.
As with last year, we conducted a call for TPC participation, and built a TPC that
included a mix of experience levels, backgrounds, and geographies, bringing both well
established and fresh perspectives to the committee. Each submission was assigned to
four reviewers, with each reviewer providing an average of just under 5 reviews. All but
8 papers received the assigned reviews, while the remaining 8 were evaluated based on
3 reviews, but had a clear consensus. Again following in the footsteps of previous years,
we established a Review Task Force (RTF) of experienced community members who
were able to guide much of the discussion amongst reviewers. Following in recent PAM
tradition, the TPC meeting was again held virtually and asynchronously through lively
and in depth discussions amongst the reviewers and the RTF. Finally, 19 of the accepted
papers were shepherded by members of the TPC who were reviewers of each paper. As
program chairs, we would like to extend a big thank you to our TPC and RTF members
for volunteering their time and expertise with such dedication and enthusiasm.
Special thanks to our hosting organization this year, IMDEA Networks Institute. Par-
ticular thanks to our web chair Orlando Martinez-Durive, the publications chair Aristide
Akem, and the virtual arrangements chair, Antonio Bazco Nogueras. Thanks to the PAM
steering committee for their guidance in putting together the conference. Further thanks
to last year’s TPC chairs, Cristel Pelsser and Oliver Hohlfeld, who offered consider-
able guidance and learnings from last year’s experience, and to Brandenburg Technical
University for hosting the submission site. Finally, thank you to the researchers in the
vi Preface
networking and measurement communities and beyond who submitted their work to
PAM and engaged in the process.
General Chair
Program Committee
TLS
Applications
Measurement Tools
Network Performance
Topology
DNS
Web
1 Introduction
The privacy of Internet users has become one of the most discussed issues in the
field of networking. New protocols and services are being developed with strong
privacy guarantees, while privacy-enhancing technologies are opening opportuni-
ties for new markets. iCloud Private Relay (PR) is a new service recently created
by Apple that is integrated into the company’s operating systems (i.e., MacOS,
iOS, iPadOS). Initially launched in 2021, it offers users the possibility of for-
warding traffic via a multi-party relay [19], offering a service that in many ways
resembles a VPN but differs in privacy guarantees. The architecture results in no
party (neither Apple nor their infrastructure partners) holding both user identity
and the contacted servers, whereas a VPN architecture simply shifts trust to the
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 3–17, 2023.
[Link]
4 M. Trevisan et al.
VPN which has access to both. The seamless integration of the service in the
Apple OSes, its low cost ($0.99 per month for the cheapest plan) and its low
entry barrier suggest that a large adoption of the service is very likely, with an
anticipated major impact on Internet traffic [16] moving forward.
iCloud Private Relay works with a multi-party relay architecture: The client
operating system connects to an ingress proxy (operated by Apple) using an
encrypted connection over QUIC [4]. The ingress proxy routes the client traffic
to an egress proxy (currently operated by one of Akamai, Cloudfare, and Fastly)
that forwards the traffic to the destination server requested by the user. With
this architecture, the ingress and egress proxies can only see the client’s or the
server’s IP address, respectively, but never both. Equally, eavesdroppers (e.g.,
ISPs) can observe the traffic of multiple users to/from ingress and egress proxies
and thus cannot easily profile individual users’ activity from the traffic [26].
The possibility of a major adoption of the service in the short term raises
questions about its impact on the internet. Similar privacy protection mecha-
nisms, such as VPNs, onion routing [24] and Tor [10] have been studied in terms
of both performance and privacy [2,13]. For example, the authors of [6] uncover
the websites a user is visiting when connected via Tor by relying on side channels
such as packet sizes and timing. Similarly, multiple authors [2,15] have studied
the impact of privacy-enhancing technologies on Internet performance. For PR,
however, we are aware of a single study focusing on the service [16], which focused
on describing the system architecture and its deployment footprint, neglecting
implications on performance and user-perceived quality of experience.
In this work, we focus on the impact of iCloud Private Relay on web perfor-
mance. We set up active experiments using Apple devices and design multiple
workloads to assess the effects of PR on different scenarios. We deploy our testbed
across multiple locations and gather several metrics associated with users’ Qual-
ity of Experience (QoE), such as page load time, throughput, and latency. Apple
notes [3] that iCloud Private Relay can negatively affect web speed tests as such
tests routinely use “several simultaneous connections to deliver the highest possi-
ble result”, but goes on to claim that “actual browsing experience remains fast.”
Therefore we design our experiments to assess these claims, including speed tests
and web browsing with and without PR in place.
Our results show that iCloud Private Relay does impact performance. We
confirm a significant reduction in the throughput measured with speed tests,
e.g., with up to 10-fold slower download throughput when using PR. We notice
a performance penalty in web browsing too, observing a 60% increase in page
load time in some cases. Performance impairments also occur in cases where a
single connection is used to download a large file, thus questioning the claim
that several simultaneous connections are the root cause of performance penal-
ties. Interestingly, the selection of the egress proxy operator appears to have
crucial implications on performance. We also observe that client traffic over PR
outperforms traffic over an unmodified connection in some cases, suggesting that
the system’s overlay routing can result in more optimal paths.
Overall, our study is a first step towards understanding the impact of large-
scale, well-provisioned, privacy-enhancing services such as iCloud Private Relay
Measuring the Performance of iCloud Private Relay 5
Akamai
Fastly
Proxy B
Fig. 1. Overview of the iCloud Private Relay architecture. Client traffic passes through
two proxies: i) Proxy A operated by Apple; and ii) Proxy B operated by one of Akamai,
Cloudflare, or Fastly.
on Internet performance. To increase the impact of our study and allow for
reproducible comparisons, we release our measurements and source codes [1].
Apple launched the PR service during its Apple Worldwide Developers Confer-
ence (WWDC) in 2021 [4]. The service employs a multi-hop proxy architecture,
also known as a Multi-Party Relay (MPR) [19]. The architecture provides privacy
benefits by decoupling the users’ network identity (i.e., the client IP address)
from their Internet usage (i.e., the destination servers). This is accomplished by
the client setting up two nested tunnels: the first to an ingress proxy (Proxy A
in Fig. 1), operated by Apple, which provides authentication and localization;
the second to egress proxy (Proxy B in Fig. 1) operated by one of Apple’s infras-
tructure partners (currently including Akamai, Cloudflare, and Fastly), which in
turn connects to the destination server(s) on the client’s behalf. Proxy A only
has visibility into the client’s IP address and cannot inspect the encrypted and
tunneled web traffic. Proxy B knows the servers that the clients connect to, but
cannot see the client’s IP address. Likewise, the destination server does not see
the client IP addresses as connections are initiated by Proxy B. PR is currently
limited to Apple-specific applications (i.e., Safari).
The PR architecture relies on well-known web protocols rather than custom
protocols. The connection to Proxy A uses QUIC by default, with a fallback to
HTTP/2 and TLS if the QUIC connection fails or is blocked. The connection
to Proxy B defaults to HTTP/3 and MASQUE [17,18], which allows building
efficient QUIC connections over a QUIC proxy. If HTTP/3 is not supported, the
connection to Proxy B falls back to the classical HTTP CONNECT over TLS.
The PR does not act as a classical VPN and handles the traffic coming
uniquely from the Safari web browser. Only HTTP(S) browser traffic goes
through the PR, while we notice that, in the case of audio/video calls, WebRTC
6 M. Trevisan et al.
traffic (RTP, DTLS, and STUN/TURN protocols) is not captured by the PR.
The traffic of other applications in the system uses classical routing too, includ-
ing other browsers and mail clients. Interestingly, the curl command-line facility
uses PR, but only for clear-text HTTP traffic. The fact that only some applica-
tions support PR is a problem, since PR may give users a false sense of privacy
while routing only a share of their traffic to the PR tunnel.
Moreover, the usage of such an architecture will impact the efficacy of existing
Internet services. For instance, services that rely on client IP address informa-
tion for localizing content (i.e., IP geolocation) no longer have access to clients’
actual IP addresses. Other services that require insight into user traffic, such as
middleboxes that provide content filtering (e.g., corporate networks or parental
control services) will be unable to access user content. Lastly, the additional
hops introduced by the service may hamper performance, as we investigate in
this paper.
command-line tool. When Private Relay is enabled, curl traffic uses it, allowing
us to easily test HTTP downloads in isolation. We use curl to download a 1 GB
file several times. We select a 1 GB test file made available on the Hetzner CDN,
a standard file used for evaluating content distribution speeds [11]. From each
location, we download the test file 200 times with and without Private Relay
and record the download time.
iCloud Private Relay is mainly designed to allow web browsing with stronger
privacy guarantees. Our goal is ultimately to study to what extent Private Relay
impacts the user’s perceived performance and, in turn, its implications for web
QoE. To this end, we instrument Safari to visit a set of web pages automatically
and collect statistics regarding page loading. We target the 100 most popular
websites in each country according to the public ranking provided by SimilarWeb
analytics [21].
We use the BrowserTime toolset to automate the visits to the websites and
the collection of the statistics [23]. For each website, we run five visits with and
without PR enabled. Out of each visit, we collect statistics about each HTTP
transaction carried out during the page loading. Essential to our analysis, we
collect the Page Load Time (also called onLoad time) that we use as a practical
proxy for measuring the web performance. Page Load Time represents the time
elapsed between the beginning of the visit and the instant when the last object
of the web page is retrieved. The Page Load Time has previously been shown to
be correlated with users’ QoE [9].
Finally, note that in our experimental campaign, we do not measure explicitly
the end-to-end RTT. Indeed, our measurement infrastructure cannot observe the
layer-4 RTT, as we rely on browser instrumentation. Measuring the RTT poses
some challenges in the case of tunneled traffic (such as PR), e.g., one could
instrument the SO kernel to monitor TCP statistics. This is by no means trivial,
in particular considering the proprietary software offering PR. We thus focus on
user-perceived quality, showing higher-level metrics such as Page Load Time or
Throughput, leaving these additional aspects for future work.
4 Results
We now present results across the three workloads. We observe that, in gen-
eral, PR negatively impacts performance, particularly for scenarios that require
long-lasting network flows, i.e., bulk download and speed test measurements.
Further, in these experiments, PR usage results in a higher level of variability
in performance, even for stable and fast Ethernet network connections. Inter-
estingly, these takeaways do not apply across all results: in one case, i.e., bulk
download in France, we observe that PR outperforms an unmodified connection.
Measuring the Performance of iCloud Private Relay 9
1.0
0.8
0.6
ECDF
0.4
PR
0.2
Native
0.0
0 250 500 750 1000 0 250 500 750 1000 0 250 500 750 1000
Throughput [Mbit/s] Throughput [Mbit/s] Throughput [Mbit/s]
1.0
0.8
0.6
ECDF
0.4
PR
0.2
Native
0.0
0 250 500 750 1000 0 250 500 750 1000 0 10 20 30 40 50
Throughput [Mbit/s] Throughput [Mbit/s] Throughput [Mbit/s]
Fig. 3. Upload throughput measured with speed test measurements. Note the different
x-scale for the US.
4.1 Throughput
1.0
CloudFlare
0.8 Akamai
ECDF 0.6
0.4
0.2
0.0
0 250 500 750 1000 0 250 500 750 1000
Throughput [Mbit/s] Throughput [Mbit/s]
path between the client and Proxy A. In sum, PR seemingly does not impact
performance when the client-side connections are the bottleneck.
Interestingly, we observe that speed tests performed when PR is enabled
result in a much higher performance variability. For example, we observe that in
multiple configurations, in particular for downstream experiments in France and
the US (Fig. 2a and 2c, respectively), experiments result in bimodal speed dis-
tributions, possibly caused by either ephemerally congested paths or congested
proxies that negatively impact performance in a subset of experiments.
We investigate this aspect further in Fig. 4, where we dissect throughput
distribution according to Proxy B selected as the egress node by PR. Sattler et
al. [16] found that Proxy B selection changes multiple times in a day. For France,
we observe that all speed tests achieving throughput below 200 Mbit/s are those
using a Cloudflare-owned Proxy B, while the faster ones are all using Akamai’s
Proxy B. In the US, we observe the opposite scenario, with CloudFlare Proxy B
leading to better performance compared with Akamai, even if the two distribu-
tions partially overlap. We do not report the figure for Italy as all experiments
for this case resulted in an Akamai Proxy B egress, leading to the performance
shown in Fig. 2. We also observe that when PR is in place, the Ookla’s measure-
ment server is often further from the user than without PR for both Italy and
France. With native connection, the speed test is served from a server within
120 km, while, with PR, the server is 200–300 km far away. We detail this in the
Appendix. In a nutshell, the choice of egress node has paramount implications
on the achieved throughput, and this choice is not under the user’s control.
Overall, these results appear to confirm Apple’s disclaimer that PR can neg-
atively impact speed test performance. Apple justifies this performance loss to
the normal behavior of speed test experiments. In particular, they state that
“Private Relay uses a single, secure connection to maintain privacy and perfor-
mance. This design may impact how throughput is reflected in network speed
tests that typically open several simultaneous connections to deliver the high-
est possible result.” To verify whether the performance loss experienced can be
solely linked to the use of multiple connections, in the next section, we replicate
Measuring the Performance of iCloud Private Relay 11
1.0
0.8
0.6
ECDF
0.4
PR
0.2
Native
0.0
0 50 100 150 0 100 200 300 400 0 50 100 150
Throughput [Mbit/s] Throughput [Mbit/s] Throughput [Mbit/s]
1.0
0.8
0.6
ECDF
0.4
PR
0.2
Native
0.0
0 1 2 3 4 5 0 1 2 3 4 5 0 5 10 15
Page Load Time [s] Page Load Time [s] Page Load Time [s]
5.2 Localization
iCloud Private Relay is designed to prevent destination servers from observing
client IP addresses. Clearly, this design negatively impacts the ability of IP
geolocation services to map clients to their geographical location. These services
are widely used by content providers to localize users and determine access rules
based on geographical constraints. The PR architecture aims to minimize this
issue by roughly localizing the client using Proxy A, and carefully selecting the
Proxy B egress based on the location that the client is purported to be. This
would preserve, at least roughly, the geographic location of the user from the
server’s point of view. To support IP geolocation services in mapping the users’
geographical location, Apple publishes Proxy B IP addresses along with the
location of the users aggregated through them [5].
In many cases, low-fidelity location information is sufficient to provide localized
content. Unfortunately, some services require very accurate location information
to serve content (e.g., live streaming of sporting events), which may not be possible
using services such as PR. Further study is required to study the tradeoff between
privacy and usability in terms of localization. Additionally, previous work [16] has
shown that the IP-to-location mappings offered by Apple’s partners are not always
14 M. Trevisan et al.
a direct representation of the physical location of the proxy holding the given IP.
This is done to overcome the lack of PR proxies in certain regions of the world.
This could impact performance for users who connect to the PR infrastructure
from locations that are not physically served by it. The network paths would be
extended beyond their geographical location, adding latency to communications
and crossing national borders.
5.3 Cost
While the PR design seems beneficial for privacy, the real benefits have been left
unquantified and largely unexplored. Future work is necessary to understand
the benefits offered by such as system. This is particularly true considering the
inherent tradeoffs that Multi-Party Relay architectures have on network traffic
and the capability required to process it: In PR, clients’ traffic passes through
multiple middleboxes in order to achieve the privacy guarantees associated with
decoupling network identity from behavior. This has implications on perfor-
mance, at the center of this paper, as well as on energy consumption (e.g., due
to the additional servers and the multiple layers of encryption they have to han-
dle). For example, by nesting encrypted channels as the PR architecture does,
Proxy A could be wasting significant computing resources “double encrypting”
traffic. To avoid this overhead, QUIC-Aware Proxying Using HTTP has been
proposed, where Proxy A simply moves the traffic along the path towards Proxy
B without double encryption [17,18]. Other similar optimizations are likely to be
introduced as the architecture becomes more mature and more widely adopted.
6 Conclusions
Apple’s iCloud Private Relay is one of the recent attempts at deploying Multi-
Party Relay architectures at scale. Given Apple devices’ pervasiveness and the
company’s push towards privacy, it is possible that this architecture will be
quickly adopted as the de facto standard for privacy-oriented network architec-
tures. In this work, we present a first study of the impacts that PR architecture
can have on users’ performance. Through experiments across three locations in
France, Italy, and the US, we find that PR not only impacts active through-
put measurements but also negatively affects page load time and file download,
indicating potential impacts on the users’ web QoE. We show for example that
PR substantially changes the paths taken by traffic (e.g., during speed tests),
impacting performance. Our paper sheds light on new problems and calls for fur-
ther research on how to avoid them when deploying privacy-preserving services.
This work opens up a number of potential future venues to explore Multi-
Party Relay architectures such iCloud Private Relay, not solely in terms of per-
formance, but also across multiple dimensions such as privacy-costs tradeoffs,
content access, and the impact on network routing at large. To engage the com-
munity to search for the answer to these questions, we release the source code
of the software used to perform the experiments presented in this paper.
Measuring the Performance of iCloud Private Relay 15
Appendix
In this Appendix, we break down the distance between the user and Ookla’s
speed test measurement servers, with and without PR. In the following two
tables, we show the measurement server chosen by Ookla, detailing its location
and distance from the testing location. We report data for the Italian and French
locations and separate the cases with and without PR. We omit the US location
as, in all cases, the measurement server is located at the same location, i.e.,
Hawaii.
When PR is in place, it is more likely that the speed test server is far away
from the client. For example, for the Italian location, without PR, speed tests
are served within 120 km, while with PR, servers are at 200 km or more from the
client.
Ookla obviously cannot identify the true location of the users, since its servers
observe only egress IP addresses. Indeed, hiding the users’ IP addresses is the
ultimate goal of PR and, as such, these differences are expected. We here show
that the servers selected by Ookla when PR is enabled deliver poorer throughput
figures, and our conjecture is that the root causes for such performance penalties
are in the path from clients to the selected servers.
The same situation may occur with other services relying on IP geolocation,
such as content providers and CDNs. Our measurements, while preliminary, show
that the introduction of the PR tunnels impact performance (see our discussion
on future work in Sect. 5) (Table 1).
Table 1. Share of Speed Tests served from servers in different locations. The distance
from the client is reported in brackets.
Ljubljana (70 km) Venice (120 km) Conegliano (120 km) Milan (200 km) Rome (400 km)
References
1. [Link]
2. Alsabah, M., Goldberg, I.: Performance and security improvements for tor: a sur-
vey. ACM Comput. Surv. 49(2), 1–36 (2016)
3. Apple. About iCloud Private Relay, December 2021. [Link]
en-us/HT212614
4. Apple. iCloud Private Relay Overview, December 2021. [Link]
privacy/docs/iCloud Private Relay Overview [Link]
5. Apple. Prepare Your Network or Web Server for iCloud Private Relay. https://
[Link]/support/prepare-your-network-for-icloud-private-relay/,
December 2021
6. Arp, D., Yamaguchi, F., Rieck, K.: Torben: a practical side-channel attack for
deanonymizing tor communication. In: Proceedings of the ASIA CCS, pp. 597–602
(2015)
7. Fast Company. How one second could cost amazon $1.6 billion in sales,
December 2021. [Link]
cost-amazon-16-billion-sales
8. Cui, H., Biersack, E.: On the relationship between QOS and QOE for web sessions.
EURECOM, Sophia Antipolis, France, Technical report, RR-12-263 (2012)
9. da Hora, D.N., Asrese, A.S., Christophides, V., Teixeira, R., Rossi, D.: Narrowing
the gap between QoS metrics and web QoE using above-the-fold metrics. In: Bev-
erly, R., Smaragdakis, G., Feldmann, A. (eds.) PAM 2018. LNCS, vol. 10771, pp.
31–43. Springer, Cham (2018). [Link] 3
10. Dingledine, R., Mathewson, N., Syverson, P.: Tor: the second-generation onion
router. In: Usenix Security Symposium (2004)
11. Hetzner. Test-files, December 2021. [Link]
12. Mandalari, A.M., et al.: Measuring roaming in Europe: infrastructure and impli-
cations on users’ QOE. IEEE Trans. Mob. Comput. 21(10), 3687–3699 (2021)
13. Mani, A., Wilson-Brown, T., Jansen, R., Johnson, A., Sherr, M.: Understanding tor
usage with privacy-preserving measurement. In: Proceedings of the Internet Mea-
surement Conference 2018, IMC 2018, New York, NY, USA, pp. 175–187 (2018).
Association for Computing Machinery
14. Ookla. Speedtest, December 2021. [Link]
15. Pudelko, M., Emmerich, P., Gallenmüller, S., Carle, G.: Performance analysis of
VPN gateways. In: 2020 IFIP Networking Conference (Networking), pp. 325–333
(2020)
16. Sattler, P., Aulbach, J., Zirngibl, J., Carle, G.: Towards a tectonic traffic shift?
Investigating Apple’s new relay network. In: Proceedings of the 22nd ACM Internet
Measurement Conference, IMC 2022 (2022)
17. Schinazi, D.: Proxying UDP in HTTP. RFC 9298, August 2022
18. Schinazi, D., Pardue, L.: HTTP Datagrams and the Capsule Protocol. RFC 9297,
August 2022
19. Schmitt, P., Iyengar, J., Wood, C., Raghavan, B.: The decoupling principle: a
practical privacy framework. In: ACM SIGCOMM Workshop on Hot Topics in
Networking (HotNets), November 2022
20. Selenium: Selenium automates browsers. that’s it!, December 2021 [Link]
[Link]
21. SimilarWeb. Effortlessly analyze your competitive landscape, December 2021.
[Link]
Measuring the Performance of iCloud Private Relay 17
22. Sirinam, P., Imani, M., Juarez, M., Wright, M.: Deep fingerprinting: undermining
website fingerprinting defenses with deep learning. In: Proceedings of the CCS, pp.
1928–1943 (2018)
23. [Link]. Documentation v16. [Link]
browsertime/, December 2021
24. Syverson, P.F., Goldschlag, D.M., Reed, M.G.: Anonymous connections and onion
routing. In: Proceedings. 1997 IEEE Symposium on Security and Privacy (Cat.
No. 97CB36097), pp. 44–54. IEEE (1997)
25. Trevisan, M., Drago, I., Mellia, M.: Impact of access speed on adaptive video
streaming quality: a passive perspective. In: Proceedings of the 2016 Workshop on
QoE-Based Analysis and Management of Data Communication Networks, Internet-
QoE 2016, New York, NY, USA, pp. 7–12 (2016)
26. Trevisan, M., Soro, F., Mellia, M., Drago, I., Morla, R.: Does domain name encryp-
tion increase users’ privacy? SIGCOMM Comput. Commun. Rev. 50(3), 16–22
(2020)
27. Wang, T., Cai, X., Nithyanand, R., Johnson, R., Goldberg, I.: Effective Attacks
and provable defenses for website fingerprinting. In: Proceedings of the USENIX
Security, pp. 143–157 (2014)
Characterizing the VPN Ecosystem
in the Wild
Abstract. With the increase of remote working during and after the
COVID-19 pandemic, the use of Virtual Private Networks (VPNs) around
the world has nearly doubled. Therefore, measuring the traffic and secu-
rity aspects of the VPN ecosystem is more important now than ever. VPN
users rely on the security of VPN solutions, to protect private and cor-
porate communication. Thus a good understanding of the security state
of VPN servers is crucial. Moreover, properly detecting and characteriz-
ing VPN traffic remains challenging, since some VPN protocols use the
same port number as web traffic and port-based traffic classification will
not help.
In this paper, we aim at detecting and characterizing VPN servers in
the wild, which facilitates detecting the VPN traffic. To this end, we per-
form Internet-wide active measurements to find VPN servers in the wild,
and analyze their cryptographic certificates, vulnerabilities, locations, and
fingerprints. We find 9.8M VPN servers distributed around the world using
OpenVPN, SSTP, PPTP, and IPsec, and analyze their vulnerability. We
find SSTP to be the most vulnerable protocol with more than 90% of
detected servers being vulnerable to TLS downgrade attacks. Out of all
the servers that respond to our VPN probes, 2% also respond to HTTP
probes and therefore are classified as Web servers. Finally, we use our list
of VPN servers to identify VPN traffic in a large European ISP and observe
that 2.6% of all traffic is related to these VPN servers.
1 Introduction
even a more dramatic increase of 20x has been reported [27], which shows a
prominent growth of remote work and e-learning. Additionally, several articles
find that remote work is here to stay [21,48]. According to recent statistics from
SurfShark [44], 31% of all Internet users use VPNs.
In order to facilitate network planning and traffic engineering, Internet Ser-
vice Providers (ISPs) have an interest in understanding the network applications
being used by their clients, and how these applications behave in terms of traffic
patterns and volume. Therefore, detecting and characterizing VPN traffic is an
important task for ISPs. Certain VPN protocols use known port numbers for
their operation, e.g. port number 4500 is used for IPsec, and port number 1723
is used for SSTP. Thus, the traffic using protocols over the known port numbers
can easily be detected as VPN traffic. However, some VPN protocols, e.g. SSTP,
and in some occasions, OpenVPN use port number 443 which is commonly used
for secure web applications. This makes it challenging to distinguish between
web and VPN traffic.
Moreover, VPN users might share sensitive private or corporate data over
VPN connections. As the number of cyber attacks has almost doubled after the
pandemic [9], it makes Internet users even more aware of their privacy and the
security of their VPN connections. Therefore, investigating the vulnerabilities of
the VPN protocols helps to highlight existing shortcomings in VPN security.
Previous studies focused on detecting VPN traffic using machine learning
[15,39], or DNS-based approaches [2,18]. Some studies have also analyzed the
commercial VPN ecosystem [29,47]. However, to the best of our knowledge, this
is the first work which conducts active measurements to detect and characterize
VPN servers in the wild.
In this paper, we aim to detect, characterize, and analyze the deployment of
VPN servers in the Internet using active measurements along with passive VPN
traffic analysis. Specifically, this work makes the following main contributions:
2 Background
VPNs establish cryptographically secured tunnels between different networks
and can be used to connect private networks over the public network. Thus, a
proper VPN connection should be encrypted in order to prevent eavesdropping
and tampering of VPN traffic. The tunneling mechanism of a VPN connection
also provides privacy since the traffic is encapsulated. Therefore, users remotely
accessing a private network appear to be directly connected.
While the exact tunneling process varies depending on the underlying VPN
protocol, it is quite common to categorize VPNs in two different groups:
The usage of VPNs has evolved over the past three decades. David Crawshaw
[13] gives a very comprehensive overview of how and why VPNs changed over
the years. While in the earlier days of the Internet, they were primarily used
by companies to connect their geographically distinct offices, VPNs nowadays
provide a variety of use cases for individuals as well and are used by millions of
end-users around the globe. Use cases include:
of institutional VPNs. Generally, they found that, while most VPN users are
concerned about their privacy, they are less concerned about data collection by
VPN companies.
Especially during the COVID-19 pandemic, VPNs increasingly gained sig-
nificance. The pandemic and the resulting lockdowns caused many employees
and students to work and study remotely from home. Feldmann et al. [18] ana-
lyzed the effect of the lockdowns on the Internet traffic. Their work included
the analysis of how VPN traffic shifted during the pandemic. They detected a
traffic increase of over 200% for VPN servers identified based on their domain
with increased traffic even after the first lockdowns. These findings highlight the
rising significance of VPNs. With progressing digitalization, VPN traffic can be
expected to increase even further.
Table 1. Overview of VPN protocols showing the transport protocol, port, (D)TLS
encryption, and possible detection.
3 Methodology
In this section, we introduce our methodology for our passive and active measure-
ments. We perform Internet-wide measurements in order to detect VPN servers
in the wild and create hit lists of identified VPN servers. Based on those results,
we conduct follow-up measurements to fingerprint the VPN servers and further
22 A. Maghsoudlou et al.
analyze them in terms of security. Finally, we look for the detected IP addresses
in the traffic from a large European ISP to find out the amount of VPN traffic.
issuer are both specified as localhost or [Link]. For the certificates signed
by a Certificate Authority (CA), we collect the most common issuing organi-
zations. We gather domain names corresponding to the responsive IP addresses
using reverse DNS (rDNS) look-ups, and collect certificates with and without the
Server Name Indication (SNI) extension using these domain names and compare
them against each other. SNI can be used by the client in the TLS handshake
in order to specify a hostname for which a connection should be established.
This might be necessary in cases where multiple domain names are hosted on
a single address. Finally, we test if the servers are susceptible to the Heartbleed
[50] vulnerability as well as a series of TLS downgrade attacks. The Heartbleed
attack is based on the Heartbeat Extension [49] of the OpenSSL library. In TLS
downgrade attacks, we try to force a server to establish a connection using an
outdated SSL/TLS version or using insecure cipher suites by suggesting those
outdated primitives in the TLS handshake. Table 8 summarizes all the vulner-
abilities and their requirements, i.e., what we have to test for or the version or
cipher suite to which we try to downgrade the TLS connection. For instance, in
order to check if a server is vulnerable to the FREAK attack, we suggest any
SSL/TLS version and only RSA EXPORT cipher suites in the TLS handshake.
3.3 Fingerprinting
We try to infer more information on the VPN servers based on our connection
initiation requests as well as from follow-up measurements in order to further
categorize them.
One aspect we examine is the server software deployment. For SSTP and
PPTP, we can extract information on the software vendor directly from the
responses to our initiation requests.
Furthermore, we perform OS detection measurements on a subset of 1000
VPN servers for each protocol using Nmap [34], a network scanner that can be
used for network discovery among other things. We use Nmap’s fast option and
target 100 instead of 1000 ports to decrease runtime and parse the results for
the most common open ports and OS guesses. With those results, we can learn
more about the VPN server infrastructure and potential other services running
on the same servers.
Then, we look for the IP addresses from our VPN hitlist on over a week
of network flow data from the ISP to find out the amount of traffic associated
with the VPN hitlist and compare the results with a port-based VPN traffic
detection, and also a state-of-the-art approach.
0.75
ECDF
0.50 IPv4
IPv6
0.25
0.00
1 10 100 1000 10000 1 10 100
In total, we find 9,817,450 responsive IPv4 addresses with our probes that we
can identify as VPN servers.
rDNS. We investigate the reverse DNS records corresponding to the responsive
IPv4 addresses. We aggregate results on the second-level domain and sort them
based on the number of responsive IPs that they correspond to. We find that all
the top 10 domain names belong to telecommunication companies (e.g., Open
Computer Network, a large Japanese ISP, and Telstra, an Australian telecom-
munications company). Next, we filter all rDNS records which contain vpn in
their second-level domain names in order to detect commercial VPN providers.
We find a single domain related to PacketHub which manages IP addresses for
several companies, including NordVPN, a major commercial VPN provider. This
domain name ranks 60th among all rDNS second level domains.
AS Analysis. Figure 1 shows the distribution of ASes to which our responsive
IP addresses belong. The responsive IP addresses are originated by 49625 and
334 ASes in total, while top 10 ASes contribute to 22% and 38% of the IP
addresses, for IPv4 and IPv6 respectively, as shown in Fig. 1. Top 10 ASes for
IPv4 responsive addresses are all telecommunication companies, while out of the
top 10 ASes for IPv6 responsive addresses, 8 are telecommunication companies
and 2 are academy-related ASes. Tables 2 and 3 further summarize the top 10
AS numbers as well as the AS names or organizations and the number of VPN
servers that are registered within the respective AS. As can be seen, most top
ASes are large ISP networks.
Moreover, we investigate the top ASes for commercial VPN providers. As
shown by Ramesh et al. [47] it is quite common for commercial VPN providers
to use shared infrastructure. 27 providers, including popular companies such as
NordVPN, Norton Secure VPN, or Mozilla VPN, use the same AS, namely AS
9009 operated by M247 Ltd. This AS is also visible in our measurements and it
ranks 14th with 74,894 identified VPN servers (0.76% of all addresses). Further-
more, Ramesh et al. [47] find that some IP blocks in AS 16509 (Amazon) are shared
across Norton Secure VPN and SurfEasy VPN. AS 16509 lands on rank 20 of our
list being shared by almost 60,000 VPN servers (0.6% of all addresses). Another AS
known to be used by VPN providers is AS 60068—again operated by M247 Ltd.—
which is used by NordVPN and CyberGhost VPN. It ranks on place 178 of our list
with 6,898 VPN servers (0.07% of all addresses). Overall, we find that although
the top ASes are dominated by large ISPs, a considerable number of VPN servers
are located in ASes used by commercial VPN providers.
Geolocation. We use Geolite Country Database [51] to determine the loca-
tion of the responsive IP addresses. Figure 2 shows a heatmap of the number of
responsive IPv4 addresses per country. We observe that responsive IP addresses
are scattered all over the world, in total over 241 and 52 countries for IPv4 and
IPv6, respectively. However, 64% and 86% of IP addresses belong to the top 10
countries for IPv4 and IPv6 respectively. Top 3 countries contributing to IPv4
responsive addresses are the United States, China, and UK, while top 3 countries
for IPv6 are the United States, Japan, and Germany.
26 A. Maghsoudlou et al.
# Responsive IPs
500000
1000000
1500000
2000000
Out of the around 1.4 million OpenVPN servers, 1,011,178 were detected over
UDP and 482,956 over TCP. Considering that the TCP version of OpenVPN is
generally rather considered as a fallback option, this disparity is to be expected.
Figure 3 visualizes the intersection of those two address sets in a Venn diagram.
We can see that the majority of the servers supports only a single transport
protocol.
Overlap Between Protocols. In the next step, we compare the IP address
sets for the four protocols to depict their intersections and to find out how many
of the servers support more than one VPN protocol. Figure 4 summarizes those
findings in an upset plot. The horizontal bars on the left visualize the sizes of
the four protocol sets. The vertical bars represent the different intersections and
the sets to be considered are indicated by the black dots below the vertical bars.
The first bar on the left, e.g., represents the number of VPN servers supporting
both PPTP and IPsec with roughly 550,000 servers making up for around 5.7%
of the whole detected VPN server ecosystem. The second bar on the right, on
the other hand, represents the number of servers supporting all four protocols,
which is close to zero with only around 2.8 thousand servers.
28 A. Maghsoudlou et al.
Fig. 4. VPN protocol summary: Number of detected VPN servers for each protocol
and the intersection between all protocols.
We can see that the majority of all VPN servers support only one of the four
protocols we consider in this work. Since commercial VPN providers usually
offer a variety of different VPN protocols to choose from, it is possible that a
large percentage of the servers supporting several protocols are commercial. This
might be the case especially for the ones supporting three or four protocols. We
investigate the rDNS records corresponding to the servers supporting all the
four protocols, and find that there are no commercial VPN provider in the top
10 s-level domains. All in all, we find that commercial VPN providers account
for only a fraction of the entire VPN server ecosystem considering the supported
protocols.
Different Protocol Versions. Some VPN protocols might include different
versions or configurations, like OpenVPN, for instance. We therefore try to trig-
ger VPN responses from OpenVPN servers suggesting the outdated key exchange
method key method 1. We also try to trigger responses with random HMAC sig-
natures.
We find that only 84 of the roughly 1.4 million servers accept our random
signature. Apart from that, none of the detected servers support the insecure
key exchange method. While most of the servers ignored our requests, we still
received around 6,500 responses specifying the default key exchange method key
method 2. We can therefore conclude that key method 1 is truly deprecated in
the OpenVPN ecosystem.
Characterizing the VPN Ecosystem in the Wild 29
TLS Certificate Analysis. We collect TLS certificates for the TLS-based VPN
servers which include SSTP and OpenVPN over TCP and consider only unique
certificates. For that, we compare the certificate fingerprints, i.e., the unique
identifier of the certificate, to make sure we do not consider the same certificate
more than once. Some certificates, however, do not include a fingerprint. There-
fore, the number of certificates that we analyze in the end might be higher than
the number of unique certificates. For OpenVPN, we find 129,143 unique certifi-
cates with a fingerprint for 312,095 servers. The most frequently occurring cer-
tificate is collected over 10,000 times and is issued for [Link].
For SSTP, there are 104,988 fingerprints for 184,047 servers. We detect a certifi-
cate issued for *.[Link] 2561 times and one for *.[Link] 1194 times.
These are commercial VPN providers that seem to use the same certificate for
all of their VPN servers. While we are able to collect certificates for nearly all
of the SSTP servers, we only receive TLS certificates for around 65% of the
detected OpenVPN servers. This is most likely caused by the fact that Open-
VPN performs a variation of the standard TLS handshake during connection
establishment. Therefore, some of the servers might not respond when trying to
initiate a regular TLS handshake.
Table 5. Expired, self-issued, and self-signed TLS certificates for OpenVPN and SSTP.
Table 5 summarizes the results of the certificate analysis and contains the
number of certificates that we analyzed after filtering out unique certificates
and certificates without fingerprints. We detect a large number of self-issued or
self-signed certificates for both protocols. Out of the self-issued certificates, we
characterize only around 4.7% as snake oil certificates for SSTP and close to
zero for OpenVPN with around 0.4%. However, 33% of the self-issued SSTP
certificates contain softether, an open-source and multi-protocol VPN software,
in the CN fields. 13% specify an IPv4 address in the CN sections. Upon looking
at the organization field, we find over 21,000 different organizations where almost
14,000 specify no organization at all. For the OpenVPN certificates, we find that
around 77% of the self-issued certificates include the Fireware web CA as CNs
specifying WatchGuard as organization. For the rest, we detect more than 21,000
different organizations.
30 A. Maghsoudlou et al.
Fig. 5. Distribution of expiry time (time between day of expiry and Aug. 15, 2022) for
expired certificates.
Table 8. Requirements for TLS vulnerabilities and number of vulnerable servers per
protocol.
The Effect of Not Using SNI. As we target only IP addresses in our follow-
up TLS measurements without the SNI extension, we want to investigate the
effect of not using SNI. Therefore, we first perform an rDNS resolution for our
IP addresses and find 259,910 domain names for about 480,000 OpenVPN TCP
servers and 86,630 domain names for roughly 180,000 SSTP servers. We now
collect certificates with the SNI extension and then re-run the TLS scans without
SNI for the respective addresses for whose domains we could gather certificates.
Table 9 shows the results of the comparison of those two types of certificates
and the number of certificates we could collect. Two certificates mismatch when
the fingerprints differ. We then compare different fields and summarize the mis-
match occurrences in the table. If those fields match and the certificate has only
been renewed, we do not count it as a mismatch.
While the results for both protocols are similar, relatively speaking, we find
more mismatches for SSTP. About 3% mismatch for OpenVPN, whereas for
SSTP 5.5% mismatch. To confirm that those mismatches are caused by using
SNI in the TLS handshakes, we perform a second measurement without SNI and
compare the certificates with the other non-SNI results. Without SNI, we find
less than half as many mismatches for SSTP and more than three times fewer
mismatches for OpenVPN.
32 A. Maghsoudlou et al.
OpenVPN SSTP
SNI Certificates 84,212 45,405
no SNI Certificates 81,379 45,026
Certificate Mismatches 2491 2515
Authority Key ID Mismatches 2051 1463
Subject Key ID Mismatches 2407 2379
Subject SANs Mismatches 2008 1677
Issuer CN Mismatches 1933 1476
Subject CN Mismatches 2021 1627
4.4 Fingerprinting
Server Software. For SSTP and PPTP, we can infer the server-side software
from the responses we receive to our initiation requests. For SSTP, we find that
around 80% of all detected servers use Microsoft HTTPAPI 2.0. Around 19% use
MikroTik-SSTP and less than 1% use something else or specify nothing at all.
However, the PPTP vendor software is a lot more heterogeneous compared
to SSTP. Table 10 shows the different software vendors we detect in the VPN
server responses. While there are four prominent vendors, over 15% of the PPTP
servers rely on 183 different types. This can have potential security implications
on the PPTP ecosystem. Assuming there was some kind of new vulnerability, the
rollout of a security update to counter this vulnerability would be significantly
slower compared to SSTP with fewer software vendors. A similar phenomenon
where vendor fragmentation leads to slower update rollout can also be observed
in the Android ecosystem. Thomas et al. [54] showed that almost 60% of all
devices ran insecure Android versions in July 2015. This share declines only
slowly after the discovery of a major vulnerability. They found out that the
bottleneck of this issue lies with the manufacturers and results in 87.7% of all
devices being exposed to at least 11 critical vulnerabilities. Jones et al. [26]
considered manufacturers between 2015 and 2019 and further showed that the
median latency of a security update is 24 days with an additional latency of 11
days before an end-user update.
Nmap OS Detection and Port Scans. In our Nmap OS detection measure-
ments, we first have a look at the most common ports for all four protocols.
Figure 6 summarizes the most frequently occurring open ports. As expected, the
default HTTP(S) ports 443 and 80 are the most common ports, with the excep-
tion of the PPTP servers for which the default PPTP port TCP/1723 obviously
Characterizing the VPN Ecosystem in the Wild 33
Vendor Percentage
Linux 32.3%
MikroTik 30.6%
Draytek 21.1%
Microsoft 6.9%
Cananian 2.0%
Fortinet PPTP 1.4%
Yamaha Corporation 1.4%
Cisco Systems, Inc. 1.2%
Others (162) 3.2%
is the most widely used port. As Ramesh et al. [47] pointed out, specific open
ports do not pose security risks by themselves, yet, they might still be abused
in order to identify and exploit particular services [23].
For the OS detection, we filter out the first guesses for every target and look
at the most common OSes and version ranges:
– IPsec: We receive 48 unique first guesses for 126 hosts out of 722 responsive
IPsec servers. Out of those, 40 guess the Linux Kernel ranging from version
2.6.32-3.10. In general, Linux is the most common OS with 67 guesses. How-
ever, Microsoft was barely guessed as an OS vendor with only nine guesses.
– PPTP: For 792 responsive hosts, Nmap was able to guess an OS for 216
addresses with 56 unique guesses. Linux was once again the primary occur-
rence. Out of those guesses, 88 specified Linux 2.6.32-3.10, where the majority
mewlie below version 3.2, however. As for IPsec, we have very few results for
Microsoft with only 15 guesses. For the PPTP servers, there were more hard-
ware guesses compared to the other protocols with 36 guesses specifying some
kind of hardware device.
– OpenVPN: The most frequent guesses are almost exclusively Linux again
in 33 unique guesses for 89 out of the 763 responsive hosts. 39 specify Linux
ranging from 3.2-4.11, i.e., the versions are not quite as outdated as for PPTP
and IPsec. We received only a single guess for Microsoft products.
– SSTP: The SSTP scans result in 44 different guesses for 178 out of 948
responsive hosts. This time, we have more results for Microsoft products with
a total of 49 guesses. The most prominent vendor is Linux again, however,
with 101 guesses where 53 range from Linux versions 2.6.32-3.10.
4.5 IPv6
VPN Server Detection. Targeting roughly 530 million IPv6 addresses in our
ZMapv6 port scans, we could detect 1,195,510 responsive hosts on port TCP/443
34 A. Maghsoudlou et al.
900
700
600
297 141 733 189
500
OpenVPN
400
541 166 129 120
300
200
SSTP
Fig. 6. Heatmap of most frequently detected open ports per VPN server.
which we target in our follow-up ZGrab2 scans for SSTP and OpenVPN over
TCP. We could not find any responsive addresses on port TCP/1723, the default
PPTP port. Since port TCP/1723 is used exclusively for PPTP and the protocol
is very outdated, it is not too surprising that there are no IPv6 servers supporting
PPTP. Apart from that, we do not get any responses on the UDP ports 500
(IPsec) and 1194 (OpenVPN over UDP).
Out of the roughly 1.2 million hits on port TCP/443, we could identify 2070
addresses as OpenVPN servers and 949 as SSTP servers with a total of 2221
VPN servers supporting IPv6. While those results seem very low, we have to
keep in mind that the rollout of IPv6 is still very slow in general. IPv6 is also
not yet supported by most commercial VPN providers.
As also observed in IPv4 results in Sect. 4.2, none of the OpenVPN servers
accepted our OpenVPN key method 1 requests with only 11 servers still respond-
ing with the secure key exchange method. Additionally, of the overall IPv6 VPN
servers we detect, around 36% support both protocols, i.e., compared to IPv4,
the overlap is higher.
Investigating the rDNS records corresponding to the responsive IPv6
addresses, we observe that the top 10 domains belong to hosting providers,
cloud providers, and research networks. Similar to the IPv4 results, we do not
find a domain name belonging to a commercial VPN provider among the top 10
domains. By filtering second-level domains to match *vpn* we find the commer-
cial VPN provider WhiteLabel VPN, ranking 25th among the top domains.
Therefore, we infer that most of the VPN servers that support IPv6 are, in
fact, not commercial VPN providers.
Characterizing the VPN Ecosystem in the Wild 35
TLS Certificate Analysis. The results of the TLS certificate analysis are
similar to IPv4. We could collect certificates for around 75% of the identified
OpenVPN servers with 816 unique fingerprints. Combined with the certificates
that do not contain a fingerprint, we analyze a total of 1882 certificates. We
collected certificates for every SSTP server resulting in 747 certificates after fil-
tering out 207 unique fingerprints. Less certificates are expired this time with
only 3.3% for OpenVPN and 2.1% for SSTP. This time, only 29% of the Open-
VPN certificates are self-signed. For SSTP, more certificates are self-signed for
the IPv6 servers with over 70% of all certificates. Out of those, we characterize
roughly 2% as snake oil certificates for both protocols. Furthermore, about two
thirds of the self-signed certificates for both protocols were issued by softether.
When examining the signing organizations for the CA-signed certificates, we
find that around 85% (709 certificates) of the OpenVPN certificates are signed by
Let’s Encrypt with a total of 43 organizations. For SSTP, around 73% are signed
by Let’s Encrypt (153 certificates). Here, we find a total of only 16 organizations.
TLS Vulnerability Analysis. The results of the TLS vulnerability analysis
are very similar to the IPv4 VPN servers. For both protocols, we are only able
to detect vulnerable servers for the same three prominent attacks as for the IPv4
analysis. Out of the 2070 OpenVPN servers, 31% are vulnerable to RC4 biases,
6% to Poodle and 74% to Robot. When analyzing the 949 SSTP servers, we find
that 67% are vulnerable to RC4 biases, 13% to the Poodle attack and roughly
98% to ROBOT. While the results are similar to our large-scale measurements,
we can conclude that the VPN servers supporting IPv6 are much more likely to
show any signs of vulnerability with the vast majority being vulnerable to the
ROBOT attack.
The Effect of Not Using SNI. The rDNS measurements for the IPv6 servers
resulted in 410 domain names for SSTP and 813 domain names for OpenVPN
over TCP. Again, we first collect TLS certificates using the SNI extension and
then try the same without SNI and compare the results. We find that only
around 3% of the certificates for both protocols mismatch in terms of fingerprints
and important certificate fields including authority and subject key IDs, subject
SANs, and CNs. When comparing those results by running a second TLS scan
without SNI, we find that only around 2.5% of the OpenVPN and less than 1%
of the SSTP certificates differ. Considering the overall number of certificates, the
effect of not using SNI is even less significant compared to IPv4 and is therefore
negligible.
VPN Server Software. Since we could not analyze the PPTP server software
ecosystem this time, we can only compare the results for SSTP. The results are
similar again with 91% of the SSTP servers specifying the Microsoft HTTP API
2.0. However, the rest did not specify any vendor, i.e., the IPv6 SSTP servers
seem to not use MikroTik-SSTP with Microsoft being the only vendor.
Nmap OS Detection and Port Scans. As for IPv4, we perform Nmap mea-
surements on the detected IPv6 VPN servers including 1000 random OpenVPN
36 A. Maghsoudlou et al.
TCP servers and all 949 SSTP servers. Out of those servers, 874 OpenVPN
servers and 852 SSTP servers are responsive.
The most commonly used open port is TCP/443 with 838 occurrences (96%)
for OpenVPN and 852 (97%) for SSTP. Compared to IPv4, the number of open
HTTPS ports is much higher for OpenVPN. Here, we have to keep in mind that
we can only consider OpenVPN servers over TCP. Thus, this disparity is to be
expected. The second most frequently open port for both protocols, in contrast
to IPv4, is TCP/22, the default SSH port. This port occurs 245 times (28%) for
OpenVPN and even 391 times (46%) for SSTP. Other common ports for both
protocols are ports TCP/8000 for OpenVPN (21%) and TCP/80 accounting for
around 17% of the open ports for both protocols.
We receive more OS guesses for the IPv6 servers compared to IPv4. As was
the case for IPv4, we filter out the first guesses for every target:
– OpenVPN: The measurement results in only four unique guesses for a total
of 481 hosts. 93% specify Linux with 416 guessing Linux 3.X and 33 guessing
version 2.6. Only 19 predictions include a Microsoft OS and only 13 an Apple
product.
– SSTP: For SSTP, there are five unique predictions for 406 addresses. The
majority specifies Linux again with 91%. Out of those, 333 guesses specify
Linux version 3.X and only 36 specify version 2.6. Microsoft OSes are pre-
dicted 36 times and only a single guess specifies a macOS.
addresses. To refine the reverse DNS results, we exclude any domain names con-
taining any order of the corresponding IP address bytes or octets in decimal or
hexadecimal format. Overall, we end up with the domain names corresponding
to 23.6% of the IP addresses from the VPN hitlist. Then, we apply the methodol-
ogy used by Feldmann et al. on the resulting domain names, i.e., we extract those
domain names that contain *vpn* on the left side of the public suffix [45], while
excluding any domain starting with www. to exclude web servers. We observe
that this methodology captures only 4.8% of our VPN hitlist. Therefore, our
approach can detect 4 times more VPN servers compared to the methodology
by Feldmann et al.
Finally, we look at a one-week snapshot of all the network flow traffic from
the large European ISP to find out the amount of traffic that can be attributed
to VPN.
To this end, we compare the amount of VPN traffic detected with three
methodologies:
1. VPN Hitlist: the methodology proposed in this paper, i.e. sending active
probes, including the responsive IP addresses in a hitlist, excluding those
IP addresses that answer to web requests, i.e. HTTP GET requests, then
measuring the traffic volume originated by or destined to these IP addresses.
2. Port-based : this methodology captures the traffic only based on port num-
bers, considering traffic with port numbers 500 (IPsec), 4500 (IPsec), 1194
(OpenVPN), 1701 (L2TP), 1723 (1723) both on UDP and TCP as VPN traf-
fic.
3. Domain-based : the methodology proposed by Feldmann et al., i.e. filtering
domain names based on certain keywords, then measuring the traffic volume
originated or destined to the IP addresses corresponding to these domain
names.
Figure 7 shows the traffic volume considered as VPN traffic by each of the
above-mentioned methodologies. The solid black line shows the total amount of
VPN traffic detected by either of the three approaches. The dashed line shows
the total traffic volume in the ISP. The left Y axis shows the VPN traffic volume
(including all the three approaches), and the right Y axis shows the total ISP
traffic volume. All the traffic values are normalized. While normalizing, we keep
the ratio between the VPN traffic and total traffic intact. Therefore, comparing
the left and right axis values shows that the total traffic is roughly 25 times as
much as all VPN traffic.
Compared to the Port-based approach, we detect twice as much traffic, and
compared to the Domain-based approach, we detect 8 times as much using the
VPN Hitlist.
The mean VPN traffic volume detected by all three approaches is 4.1% of
the mean total ISP traffic over the week, with VPN Hitlist contributing to 2.6%,
Port-based 1.3%, and Domain-based 0.3%.
Looking at the overlap between every two approaches, we find that only
2.7% of all the traffic detected by all three approaches is detected both by VPN
38 A. Maghsoudlou et al.
Fig. 7. Normalized VPN traffic volume for different traffic detection techniques.
Hitlist and Domain-based. We observe 1.2% overlap between the traffic detected
by VPN Hitlist and Port-based approaches.
We observe a diurnal pattern in the VPN traffic detected by all of the three
approaches. We find that VPN traffic pattern in the weekdays differs from that
of weekends. It peaks at noon in the weekdays, and at night in the weekend,
while the total ISP traffic always follows the same pattern, i.e. peaks at night.
It could indicate the fact that the VPN traffic is mostly work-related through
weekdays, while mostly entertainment-related throughout the weekend. In the
domain-based approach the amount of VPN traffic detected by the Domain-based
approach is much less in the weekends than in the weekdays. This could indicate
that the Domain-based approach detect mostly work-related VPN servers.
We investigate the domain names corresponding to the traffic we detect using
our approach and find that vpn., mail., www., and remote. are among the most
common prefixes left to the public suffix part of the domain names, with vpn.
being the most common prefix. The fact that we observe mail. and www. might
be either re-use of the same domain name for other purposes by the network
operators, or a mislabeling effect from our approach caused by not answering
our HTTP Get requests. Also, looking at the DNS records corresponding to the
IP addresses from our hitlist, using FlowDNS—a system to correlate DNS and
Netflow data at scale [36]—we find that 5 out of 10 top domains are related
to commercial VPN providers and the rest are CDN domains. We observe that
the most common source port/destination port combination is 4500/4500 which
belongs to IPsec, also port number 1194 which is registered for OpenVPN, and at
the same time 1193, which is practically used for VPN [42]. We also observe that
51820/51820 and 1337/1337 which belongs to WireGuard protocol are among
the top port number pairs observed in the traffic detected by our approach. Port
51820 also falls into the range of ephemeral ports numbers (49152 to 65535) which
Characterizing the VPN Ecosystem in the Wild 39
6 Discussion
In this work, we detect VPN servers in the wild by sending Internet-wide active
probes using different VPN protocols. We can distinguish between VPN servers
and Web servers by excluding those servers that respond to a Web request.
We compare the amount of traffic detected by our approach and two other
approaches over a week of traffic from a large European ISP and find out that the
approach proposed by this work detects much more VPN servers compared to
the state-of-the-art domain-based approach. In addition, our approach benefits
from detecting VPN servers that do not use any domain name, and can also
detect VPN traffic that is using unusual ports in case these servers answer VPN
probe on the usual VPN port numbers. Also, to be the best of our knowledge,
this is the first work to perform an Internet-wide active measurement of VPN
servers in the wild.
VPN Hitlist. We send active probes according to the specification of VPN pro-
tocols including SSTP, PPTP, OpenVPN, and IPsec to the whole IPv4 address
space and to an IPv6 hitlist. We make our list of detected VPN servers, namely
the VPN hitlist, publicly available at [Link]. This VPN hitlist
can be useful for network operators to find out about the amount and patterns
of VPN traffic in their networks. The VPN hitlist can also be used by fellow
researchers to investigate different behaviors of the VPN servers and VPN traf-
fic, e.g. investigating actual attacks to these servers.
Security. We also investigate the security of the OpenVPN and SSTP pro-
tocols in terms of different security aspects, including heartbleed attack, TLS
certificates security, and TLS downgrade attacks. We find that SSTP servers use
expired certificates 3x more than OpenVPN servers. We also find that 90% of
the SSTP servers are vulnerable to ROBOT attack. Therefore, we find SSTP
to be the most vulnerable protocol. This striking high percentage of vulnerable
servers for some of the protocols shows, that the VPN server ecosystem is not
as secure as some users believe it to be. Therefore, we hope that our analysis
can highlight these security risks with using each VPN protocol and also helps
network operators choose the right VPN protocols for their networks.
40 A. Maghsoudlou et al.
Limitations. Our approach builds upon receiving answers from the servers in
the wild and therefore, has its limitations. If there is a VPN protocol which uses
a pre-shared key in the first VPN request and does not respond otherwise, we are
unable to detect it. Examples of such VPN protocols are WireGuard and Cisco
AnyConnect. Therefore, we are unable to detect any VPN server which offers
only these two protocols. However, we observe that 8.6% of the detected traffic
is related to WireGuard which might be due to multiple protocols being served
by one VPN server. In addition, certain VPN servers might only work on non-
registered port numbers for better anonymization. Since in our work, we only
send probes to the port numbers registered for the VPN protocols by IANA [7],
we cannot detect VPN servers that work on unusual port numbers. Therefore,
our list of detected VPN servers is limited to those using the supported VPN
protocols and working on their registered port numbers.
Future work. In the future, our work can be complemented by including more
port numbers in the active scans. Results from previous studies on predicting
the services across all ports [25] can be used together with our approach to gain
more coverage. Despite the above-mentioned limitations, our proposed approach
detects much more VPN servers compared to the state-of-the-art domain-based
approach, and also, to the best of our knowledge, is the first work to perform an
Internet-wide active measurement of VPN servers in the wild.
Reproducibility. We make our analysis code and data [19], customized ZGrab2
modules [56], and our VPN hitlist publicly available1 for fellow researchers to
be able to reproduce our work and build upon it.
7 Related Work
1
[Link]
Characterizing the VPN Ecosystem in the Wild 41
have been previously applied for several intents including finding IPv6 respon-
sive addresses [20], responsive IPs to abnormal traffic [35], the usage of DNS
over encryption [33], and so on. However, to the best of our knowledge, this is
the first work applying active measurements to detect VPN servers in the wild
and detecting the traffic based on a VPN hitlist.
Investigating the security of the VPN servers is also an interesting research
problem which is already addressed by several studies. For example, Xue et al.
investigate the possibility and practicality of fingerprinting OpenVPN flows [1].
Tolley et al. investigate the vulnerability of known VPN servers to spoofed traffic
[55]. Crawshaw [13] addresses vulnerabilities that come with some of the proto-
cols themselves, such as outdated cryptographic cipher suites used in PPTP. In
his proposal for WireGuard [14], Donenfeld talks about disadvantages in current
popular VPN protocols. VPNalyzer requires a tool to be installed on the user’s
device to measure and collect data on the active VPN connections in terms differ-
ent security aspects including data leakage, open ports, and DNSSEC validation
[47]. Appelbaum et al. also identified vulnerabilities of commercial and public
online VPN servers [6].
A large body of literature also exists that empirically examines TLS vul-
nerabilities including self-signed root CA injection to intercept TLS connection
[24,46], and improper implementation of the protocol making version downgrade
attacks possible even with new TLS 1.3 [30].
We mainly focus on potential vulnerabilities that come with VPN proto-
cols which are built on top of SSL/TLS. Thus, we investigate SSL/TLS related
features of those protocols. For some identified OpenVPN servers, we can also
make assumptions on their security based on information we can infer about
their server configurations. All the previous works study the security of known
VPN servers, while in this paper, we measure the vulnerability of our detected
VPN server in the Internet.
8 Conclusion
twice as much as the trivial port-based approach. We publish our VPN hitlist,
our customized ZGrab2 modules for VPN scans, and the code to our analysis
for future researchers and network operators to use.
References
1. OpenVPN is open to VPN fingerprinting. In: 31st USENIX Security Symposium
(USENIX Security 22). USENIX Association, Boston, MA (2022). [Link]
[Link]/conference/usenixsecurity22/presentation/xue-diwen
2. ul Abideen, M.Z., Saleem, S., Ejaz, M.: VPN traffic detection in SSL-protected
channel. Secur. Commun. Netw. 2019, 1–17 (2019)
3. Adrian, D., et al.: Imperfect forward secrecy: how Diffie-Hellman fails in practice.
In: 22nd ACM Conference on Computer and Communications Security (2015)
4. Al-Fayoumi, M., Al-Fawa’reh, M., Nashwan, S.: VPN and Non-VPN network traffic
classification using time-related features. Comput. Mater. Continua 72, 3091–3111
(2022). [Link]
5. AlFardan, N., Bernstein, D.J., Paterson, K.G., Poettering, B., Schuldt, J.C.N.:
On the security of RC4 in TLS. In: 22nd USENIX Security Symposium (USENIX
Security 13), pp. 305–320. USENIX Association, Washington, D.C. (2013). https://
[Link]/conference/usenixsecurity13/technical-sessions/paper/alFardan
6. Appelbaum, J., Ray, M., Koscher, K., Finder, I.: vpwns: Virtual pwned networks.
In: 2nd USENIX Workshop on Free and Open Communications on the Internet.
USENIX Association (2012)
7. Authority, I.A.N.: Service name and transport protocol port number registry.
[Link] (2022). Accessed
25 Oct 2022
8. Aviram, N., et al.: DROWN: breaking TLS with SSLv2. In: 25th USENIX Security
Symposium (2016)
9. Bitaab, M., et al.: Scam pandemic: how attackers exploit public fear through phish-
ing. In: 2020 APWG Symposium on Electronic Crime Research (eCrime), pp. 1–10
(2020). [Link]
10. Böck, H., Somorovsky, J., Young, C.: Return of Bleichenbacher’s oracle threat
(ROBOT). In: 27th USENIX Security Symposium (USENIX Security 18), pp.
817–849. USENIX Association, Baltimore, MD (2018). [Link]
conference/usenixsecurity18/presentation/bock
11. Böttger, T., Ibrahim, G., Vallis, B.: How the internet reacted to COVID-19: a per-
spective from facebook’s edge network. In: Proceedings of the ACM Internet Mea-
surement Conference, pp. 34–41. IMC 2020, Association for Computing Machinery,
New York, NY, USA (2020). [Link]
12. Chair of Network Architectures and Services at TUM: ZMapv 6: internet scanner
with ipv6 capabilities, gitHub repository (2022). [Link]
zmap. Accessed 26 Oct 2022
13. Crawshaw, D.: Everything VPN is new again: the 24-year-old security model has
found a second wind. Queue 18(5), 54–66 (2020). [Link]
3439745
14. Donenfeld, J.: Wireguard: Next generation kernel network tunnel. Tech. Rep.
(2017). [Link]
Characterizing the VPN Ecosystem in the Wild 43
15. Draper-Gil, G., Lashkari, A.H., Mamun, M.S.I., Ghorbani, A.A.: Characterization
of encrypted and vpn traffic using time-related. In: Proceedings of the 2nd Inter-
national Conference on Information Systems Security and Privacy (ICISSP), pp.
407–414 (2016)
16. Durumeric, Z., Adrian, D., Mirian, A., Bailey, M., Halderman, J.A.: Tracking the
FREAK Attack. [Link] (2015). Accessed 19 Oct 2022
17. Dutkowska-Żuk, A., Hounsel, A., Morrill, A., Xiong, A., Chetty, M., Feam-
ster, N.: How and why people use virtual private networks. In: 31st USENIX
Security Symposium (USENIX Security 22), pp. 3451–3465. USENIX Associa-
tion, Boston, MA (2022). [Link]
presentation/dutkowska-zuk
18. Feldmann, A., et al.: The lockdown effect: implications of the COVID-19 pandemic
on internet traffic. In: Proceedings of the ACM Internet Measurement Conference,
pp. 1–18. IMC 2020, Association for Computing Machinery, New York, NY, USA
(2020). [Link]
19. Gasser, O.: Analysis scripts and raw data for VPN ecosystem measurements (2023).
[Link]
20. Gasser, O., et al.: Clusters in the expanse: understanding and unbiasing IPv6
hitlists. In: Proceedings of the 2018 Internet Measurement Conference. ACM, New
York, NY, USA (2018). [Link]
21. Haag, M.: Remote work is here to stay. Manhattan may never be the same. The
New York Times (2021). [Link]
[Link]
22. Hamzeh, K., Pall, G., Verthein, W., Taarud, J., Little, W., Zorn, G.: point-to-point
tunneling protocol (PPTP). RFC 2637 (Informational) (1999). [Link]
17487/RFC2637. [Link]
23. Horowitz, M.: TCP ports to test. [Link]
TCPports. Accessed 13 Oct 2022
24. Ikram, M., Vallina-Rodriguez, N., Seneviratne, S., Kaafar, M.A., Paxson, V.: An
analysis of the privacy and security risks of android VPN permission-enabled apps.
In: Proceedings of the 2016 Internet Measurement Conference, pp. 349–364 (2016)
25. Izhikevich, L., Teixeira, R., Durumeric, Z.: Predicting Ipv4 services across all ports.
In: Proceedings of the ACM SIGCOMM 2022 Conference, pp. 503–515. SIGCOMM
2022, Association for Computing Machinery, New York, NY, USA (2022). https://
[Link]/10.1145/3544216.3544249
26. Jones, K.R., Yen, T.F., Sundaramurthy, S.C., Bardas, A.G.: Deploying android
security updates: an extensive study involving manufacturers, carriers, and end
users. In: Proceedings of the 2020 ACM SIGSAC Conference on Computer and
Communications Security, pp. 551–567. CCS 2020, Association for Computing
Machinery, New York, NY, USA (2020). [Link]
27. Karamollahi, M., Williamson, C., Arlitt, M.: Zoomiversity: a case study of pan-
demic effects on post-secondary teaching and learning. In: Hohlfeld, O., Moura,
G., Pelsser, C. (eds.) PAM 2022. LNCS, vol. 13210, pp. 573–599. Springer, Cham
(2022). [Link] 26
28. Kenneally, E., Dittrich, D.: The Menlo report: ethical principles guiding informa-
tion and communication technology research. Available at SSRN 2445102 (2012)
29. Khan, M.T., DeBlasio, J., Voelker, G.M., Snoeren, A.C., Kanich, C., Vallina-
Rodriguez, N.: An empirical analysis of the commercial VPN ecosystem. In: Pro-
ceedings of the Internet Measurement Conference 2018, pp. 443–456. IMC 2018,
Association for Computing Machinery, New York, NY, USA (2018). [Link]
org/10.1145/3278532.3278570
44 A. Maghsoudlou et al.
30. Lee, S., Shin, Y., Hur, J.: Return of version downgrade attack in the era of TLS
1.3. In: Proceedings of the 16th International Conference on Emerging Networking
Experiments and Technologies, pp. 157–168 (2020)
31. Liu, S., Schmitt, P., Bronzino, F., Feamster, N.: Characterizing service provider
response to the COVID-19 pandemic in the united states. In: Hohlfeld, O., Lutu,
A., Levin, D. (eds.) PAM 2021. LNCS, vol. 12671, pp. 20–38. Springer, Cham
(2021). [Link] 2
32. Lotfollahi, M., Siavoshani, M.J., Zade, R.S.H., Saberian, M.: Deep packet: a novel
approach for encrypted traffic classification using deep learning. Soft Comput.
24(3), 1999–2012 (2020)
33. Lu, C., et al.: An end-to-end, large-scale measurement of DNS-over-encryption:
how far have we come? In: Proceedings of the Internet Measurement Conference,
pp. 22–35. IMC 2019, Association for Computing Machinery, New York, NY, USA
(2019). [Link]
34. Lyon, G.: Nmap. [Link] Accessed 26 Oct 2022
35. Maghsoudlou, A., Gasser, O., Feldmann, A.: Zeroing in on port 0 traffic in the
wild. In: Hohlfeld, O., Lutu, A., Levin, D. (eds.) PAM 2021. LNCS, vol. 12671, pp.
547–563. Springer, Cham (2021). [Link] 32
36. Maghsoudlou, A., Gasser, O., Poese, I., Feldmann, A.: FlowDNS: correlating Net-
flow and DNS streams at scale. In: Proceedings of the 18th International Conference
on Emerging Networking EXperiments and Technologies, pp. 187–195. CoNEXT
2022, Association for Computing Machinery, New York, NY, USA (2022). https://
[Link]/10.1145/3555050.3569135
37. Merget, R., Brinkmann, M., Aviram, N., Somorovsky, J., Mittmann, J.,
Schwenk, J.: Raccoon attack: finding and exploiting most-significant-bit-oracles
in TLS-DH(E). In: 30th USENIX Security Symposium (USENIX Security 21),
pp. 213–230. USENIX Association (2021). [Link]
usenixsecurity21/presentation/merget
38. Microsoft: microsoft security advisory 2743314. [Link]
en-us/security-updates/SecurityAdvisories/2012/2743314 (2012). Accessed 26 Oct
2022
39. Miller, S., Curran, K., Lunney, T.: Detection of virtual private network traffic using
machine learning. Int. J. Wirel. Netw. Broadband Technol. (IJWNBT) 9(2), 60–80
(2020)
40. Möller, B., Duong, T., Kotowicz, K.: This POODLE bites: exploiting the SSL 3.0
fallback. [Link] (2014). Accessed 19 Oct
2022
41. OpenVPN: deprecated options in OpenVPN. [Link]
net/openvpn/wiki/DeprecatedOptions#Option:-key-method. Accessed 26 Oct
2022
42. OpenVPN: typical network configuration. [Link]
server-manual/typical-network-configurations/. Accessed 28 Oct 2022
43. Partridge, C., Allman, M.: Ethical considerations in network measurement papers.
Commun. ACM 59(10), 58–64 (2016)
44. Jauniskis, P.: VPN statistics: users, markets, & legality. [Link]
com/blog/vpn-users (2022). Accessed 10 Oct 2022
45. PyPi: Public suffix PyPi. [Link] (2022).
Accessed 12 Oct 2022
46. Raman, R.S., Evdokimov, L., Wurstrow, E., Halderman, J.A., Ensafi, R.: Investi-
gating large scale https interception in Kazakhstan. In: Proceedings of the ACM
Internet Measurement Conference, pp. 125–132 (2020)
Characterizing the VPN Ecosystem in the Wild 45
47. Ramesh, R., Evdokimov, L., Xue, D., Ensafi, R.: VPNalyzer: systematic investi-
gation of the VPN ecosystem. In: Network and Distributed System Security. The
Internet Society (2022). [Link]
48. Robinson, B.: Remote work is here to stay and will increase into 2023, experts say.
Forbes (2022). [Link]
work-is-here-to-stay-and-will-increase-into-2023-experts-say/
49. Seggelmann, R., Tuexen, M., Williams, M.: Transport layer security (TLS) and
datagram transport layer security (DTLS) heartbeat extension. RFC 6520 (Pro-
posed Standard) (2012). [Link] [Link]
[Link]/rfc/[Link]. Updated by RFC 8447
50. Synopsis Inc: the heartbleed bug. [Link] (2020). Accessed
29 July 2022
51. The MaxMind company: Geolite2 free geolocation data. [Link]
[Link]/geoip/geolite2-free-geolocation-data (2022). Accessed 06 Oct 2022
52. The ZMap Team: Zgrab 2.0, gitHub repository. [Link]
zgrab2 (2022). Accessed 28 Sept 2022
53. The ZMap Team: Zmap: the internet scanner, gitHub repository. [Link]
[Link]/zmap/zmap (2022). Accessed 28 Sept 2022
54. Thomas, D.R., Beresford, A.R., Rice, A.: Security metrics for the android ecosys-
tem. In: Proceedings of the 5th Annual ACM CCS Workshop on Security and
Privacy in Smartphones and Mobile Devices, pp. 87–98. SPSM 2015, Association
for Computing Machinery, New York, NY, USA (2015). [Link]
2808117.2808118
55. Tolley, W.J., Kujath, B., Khan, M.T., Vallina-Rodriguez, N., Crandall, J.R.: Blind
In/On-Path attacks and applications to VPNs. In: 30th USENIX Security Sympo-
sium (USENIX Security 21), pp. 3129–3146. USENIX Association (2021). https://
[Link]/conference/usenixsecurity21/presentation/tolley
56. Vermeulen, L.: ZGrab2 VPN modules on GitHub. [Link]
vpnecosystem/zgrab2-vpn
57. Wang, W., Zhu, M., Wang, J., Zeng, X., Yang, Z.: End-to-end encrypted traffic
classification with one-dimensional convolution neural networks. In: 2017 IEEE
International Conference on Intelligence and Security Informatics (ISI), pp. 43–48
(2017). [Link]
58. Zirngibl, J., Steger, L., Sattler, P., Gasser, O., Carle, G.: Rusty Clusters? Dusting
an IPv6 Research Foundation. In: Proceedings of the 2022 Internet Measurement
Conference. ACM, New York, NY, USA (2022). [Link]
3561440
59. Zou, Z., Ge, J., Zheng, H., Wu, Y., Han, C., Yao, Z.: Encrypted traffic classifica-
tion with a convolutional long short-term memory neural network. In: 2018 IEEE
20th International Conference on High Performance Computing and Communica-
tions; IEEE 16th International Conference on Smart City; IEEE 4th International
Conference on Data Science and Systems (HPCC/SmartCity/DSS), pp. 329–334
(2018). [Link]
Stranger VPNs: Investigating the
Geo-Unblocking Capabilities of
Commercial VPN Providers
1 Introduction
Virtual Private Networks (VPN) allow us to act as part of a specific network
even from a physically remote location, with the added advantage of doing so
in a secure and private manner. As a consequence of this, not only are we able
to access our work environment while working from home, but users all over the
world benefit from better privacy or, for example, have the ability to circumvent
censorship [26]. Acting as part of a physically remote network, however, also
means that a user will appear to be in a different physical location than their
real one, thus allowing them access to content and services that are instead
typically bound to a specific geographic region. VPN providers have recently
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 46–68, 2023.
[Link]
Stranger VPNs 47
added one more service to their repertoire that allows them to capitalize on this,
namely a so-called streaming unblocking service.
VPN streaming services are born in the context of the arms race that seems
to exist between VPN providers and Video-on-Demand (VOD) providers. As a
consequence of non-technical constraints such as copyright negotiations, licensing
and business models, VOD providers typically use geolocation mechanisms to
identify if a user is allowed access to certain content. This phenomenon is known
as geo-blocking or geo-fencing, and it is often negatively perceived by end users.
Think about not being able to watch the new season of your favorite series
that is already streaming in the US, but that lengthy copyright negotiations are
holding back in other parts of the world. Since VPNs have the ability to let a
user’s request originate from a different geographical location, VOD providers
have developed methods to detect if a user connects to their service using a VPN,
and can block the request. Despite these countermeasures, VPN providers claim
to be able - and we can confirm that they succeed - to bypass VOD geo-blocking
and VPN detection mechanisms.
This paper investigates how VPN providers are able to bypass geo-blocking
and VPN detection. We refer to this as geo-unblocking. The challenge in this is
that VPN providers act as “black boxes”, hiding from the end-user how geo-
unblocking is achieved. This calls for an in-depth analysis of the VPN geo-
unblocking ecosystem.
The contributions of this paper are that:
– We empirically infer a model of the VPN geo-unblocking ecosystem and iden-
tify a methodology that gives us visibility into the population of proxy hosts
used to bypass geo-unblocking;
– We identify two distinct methods used by VPN providers to circumvent VOD
geo-blocking and VPN-detection mechanisms;
– We characterize these methods at the network-level, shedding light on how
VPN providers implement these methods and how their strategies adapt over
time and over different regions.
The remainder of this paper is organized as follows. In Sect. 2 we present
the relevant related work and background information. In Sect. 3 we describe
the VPN ecosystem in relation to geo-unblocking. We then describe how we
select relevant VPN providers in Sect. 4. In Sect. 5 we explain our methodology
to identify VPN exit nodes used for geo-unblocking, and we explain how we
collected our data set. Section 6 presents our results. We discuss ethical concerns
regarding our study in Sect. 7. Finally, we present our conclusions in Sect. 8.
the licensing rights to broadcast their content globally [11,19]. Therefore, they
have to ensure that only subscribers from regions for which the license has been
acquired may access the resource. In order to accurately do so, VOD providers
make use of geolocation databases to map users’ traffic to its geographical ori-
gin. Such geolocation databases are often maintained by commercial parties and
strive to have a high accuracy, but research has shown that this data is mostly
only accurate and consistent at a country level, compared to a more fine-grained
province or city level [24].
what is advertised and in one extreme case a provider claimed to have 190 dis-
tinct locations, but ultimately only 10 different data centers were responsible for
the hosting of the servers.
Similar research has been conducted by Weinberg et al., who tried to ver-
ify advertised proxy locations with the help of geolocation [32]. They conclude
that one third of all proxies are definitely not in the advertised location and
another third might not be. A different study by Winter et al. tried to geolo-
cate BGP prefixes, in order to better understand routing anomalies, outages
and more [33]. One of their data points showed that a /23 network geolocated
to 127 different countries (including Vatican City and North Korea). This is of
course highly unlikely and only after consulting WHOIS data, it became clear
that this was an IP range owned by a commercial VPN provider1 . The most
recent paper on the commercial VPN ecosystem from Ramesh et al. presents
measurement software called VPNalyzer which can run on end-user devices to
“collect 15 distinct measurements that test for aspects of service, security and
privacy essentials, misconfigurations, and leakages” [30]. Their system allows for
a systematic analysis of key security and privacy issues in VPN implementations
in the wild.
All previously mentioned studies highlight the usage of VPNs (or proxies) to
circumvent geo-blocking. Yet to the best of our knowledge, there has not been any
study so far that investigated the unblocking methodologies of these providers,
which we instead cover in this paper. There has been an assumption that geo-
unblocking can be facilitated through the sheer amount of servers operated2 .
Our contribution to this field presents new insights, by showing that this is not
necessarily the case.
3 Ecosystem
1
The provider in question is the same provider who claimed to operate 190 distinct
locations from the Khan et al. study [26].
2
Some commercial VPN providers claim to run between 2,000 and 4,000 servers [26].
50 E. Khan et al.
3
[Link]
[Link].
Stranger VPNs 51
Table 1. Manipulated DNS requests from NordVPN’s DNS servers for requests to
Netflix’s and Disney+’s homepage and CDN.
$ curl [Link] -v
* Trying [Link]:443...
* Connected to [Link] ([Link]) \
port 443 (#0)
...
* Server certificate:
* subject: C=US; ST=California; L=Los Gatos; \
O=Netflix, Inc.; CN=[Link]
* start date: Dec 14 00:00:00 2021 GMT
* expire date: Jan 14 23:59:59 2023 GMT
* subjectAltName: host "[Link]" \
matched cert’s "[Link]"
* issuer: C=US; O=DigiCert Inc; CN=DigiCert \
TLS RSA SHA256 2020 CA1
* SSL certificate verify ok.
...
VPN provider
DNS resolver
Specialized ISPs/
Hosting providers
VPN user
VPN provider TLS proxy
gateway
Residential
proxies
fully utilize the promised bandwidth without being throttled. The main dif-
ference though is the presence of a “streaming” criterion. This criterion rates
the compatibility of a commercial VPN provider to flawlessly work with VOD
providers. If we remember that VOD providers usually prohibit the use of VPNs,
this compatibility rating is an indication of whether VPN providers are capa-
ble of deploying geo-unblocking methods that successfully circumvent blocking
by VOD providers. Among the geo-unblocking VPN providers, we select six
commercial providers that occupy different roles in the market. In particular,
ExpressVPN [6] and NordVPN have been selected because they are established
providers with a large market share, as can be inferred by the fact that they are
able to allocate significant resources to marketing [7,21]. WeVPN, on the con-
trary, is a recent up-and-coming provider, which we expect to still have a limited
market share. Manual investigation of CyberGhost, PrivateVPN and Surfshark
places those instead as medium-sized providers. Finally, we also checked that all
the selected providers actually succeed in geo-unblocking content. To do so, we
followed the geo-unblocking instructions from each VPN provider as a regular
user would, and tried streaming content that would otherwise be not available
in our geographical area.
5 Methodology
In this section, we describe our measurement methodology, which focuses on
gaining visibility behind the TLS forwarding proxies we identified in Sect. 3.
The use of TLS forwarding proxies at most commercial VPN providers means
that network path information is not available for the entire route from client to
VOD CDN endpoint. Instead, we can only observe the path from our client to the
host on which the proxy is running, as demonstrated in Sect. 3. To understand
54 E. Khan et al.
Response
HTTP/1.1 403 Forbidden
Content-Type: application/octet-stream
...
Server: nginx
X-TCP-Info: addr=<Our Public IP>;port=58219;
Fig. 5. Shortened response of a GET request to Netflix’s video CDN, showing the
populated X-TCP-Info header.
5.2 Testbed
To automate the retrieval of the header and the extraction of the X-TCP-Info
field, we set up a testbed in the Netherlands which can support many con-
current VPN connections. A previous study [26] has used virtual machines for
this purpose to maintain isolation between the VPNs, but this is not feasible
for dozens of concurrent measurements. As a consequence we opted for a more
lightweight solution, by using containers with segregated network namespaces.
Within each container we establish an OpenVPN connection to a desired geo-
graphical unblocking region for every commercial VPN provider we consider.
Once the VPN connection is established, we send a single HTTPS request at an
interval of 30 s to Netflix’s video CDN. For each request we save the following
information: the time of the request, the status code of the request (i.e., OK or
timeout) and the IP address of the exit-node. We then enrich the collected IP
address with AS and (enhanced) geolocation information.
Most VPN providers limit the amount of concurrent connections per account. As
a result we cannot measure all geo-unblocking regions of a given VPN provider.
Instead, we focused on a limited number of regions making sure that these regions
are mostly supported by all chosen providers. We have selected the following four
regions: USA, Japan, Germany and the Netherlands. The rationale behind these
choices is as follows. We chose the USA because US-based VOD providers usually
offer a larger content library for their internal market (while rights need to be
negotiated for other countries), so we expect that this will draw the attention of
geo-unblocking services. Japan has instead been chosen because it is a content
creator for niche content, such as Anime, which could also be a reason to trigger
geo-unblocking requests. Finally, Germany and the Netherlands are chosen for
geographical diversity. In addition, the Netherlands has been chosen because it
is known to have a generally good Internet infrastructure, both for consumers as
well as for hosting. Table 2 shows that all chosen providers support these regions
except CyberGhost, who do not offer geo-unblocking in the Netherlands.
Most VPN providers support at least five concurrent connections, which is
why we chose to limit our vantage points to four, leaving one free connection as
buffer for timed-out sessions or for debugging purposes.
6 Results
In this section, we discuss the results of the two measurement campaigns we
ran. These measurement campaigns, both executed in 2022, are summarised in
Table 3. We start with a general overview of our measurements, and then dig
deeper into two particular mechanisms that VPN providers use to provide geo-
unblocking for their customers. We end the section with an analysis of potential
overlap in the backend infrastructures of VPN providers.
56 E. Khan et al.
ExpressVPN in Japan, where their unblocking strategy changes from one ASN
to many and then reverts back to one.
The difference in ASN usage is not the only noteworthy observation. We now
look at the IPv4/v6 usage and churn. Also in this case, the ecosystem shows
several different approaches. In the US, CyberGhost relies on many thousands
of distinct IPv4 addresses over the entire measurement period, many hundreds
of which also seem to be repeating between August and November. Differently,
PrivateVPN in Germany uses only 29 IPv4 addresses. These approaches are in
sharp contrast to Surfshark’s geo-unblocking solution. Instead of using IPv4,
their preferred method is IPv6. Additionally, for every HTTPS request we sent
we retrieved a unique IPv6 address, which is indicated by our churn-graph sitting
on top of the unique IP graph peaks.
Lastly, we sometimes seem to observe the VPN providers “experimenting”
with their settings. For example, Surfshark in the US has four distinct short-lived
periods in August, during which we recorded many different ASNs mixed in with
their “regulars”. The IP graph in the top plot also indicates the inclusion of
IPv4 addresses during these periods. Similarly, ExpressVPN in the Netherlands
displays several periods during which the amount of unique IPs spikes above
what is generally observed for that provider/region pairing.
These observations raise the question: Can the different patterns be explained
by VPN providers using different mechanisms to facilitate geo-unblocking? To
answer this question, we performed a detailed inspection of the IPs and ASNs
we observe for each provider in each region.
For each ASN with a substantial presence in our data set, we manually
obtained information on the type of service these ASNs provide, classifying these
into two groups: Specialized networks or hosting providers and (apparent) resi-
dential Internet service providers.
Specialized Networks/Hosting Providers—The first mechanism we identify
is the use of what we call “Specialized Networks” or “Hosting Providers”. This
category can be characterized as typically using a small number of ASNs (often,
but not always a single one) and a small number of IPs, and the ASNs are
characterized – at first glance, e.g., by inspecting their website – as residential
access networks. Deeper inspection of these networks, however, reveals that they
are in fact not residential access providers, but only masquerade as such. An
especially interesting case in this category is PrivateVPN, which is the only
provider in our set that does not use TLS proxies. Instead, it entirely relies on
the specialized ISP model. We discuss this category in more detail in Sect. 6.2.
Residential ISPs – The second mechanism we identify is the use of proxies
located at residential ISPs. VPN providers and regions that use this mechanism
can be characterized by the use of a large(r) number of ASNs and larger numbers
of unique IP addresses. The ASNs can all be categorized as legitimate ISPs that
provide residential and/or business services, and the IP addresses reflect this as
well (based on reverse DNS entries and geolocation). We discuss this category
in more detail in Sect. 6.3.
Stranger VPNs 59
to their customers, and streaming providers that want to stop viewers from cir-
cumventing geo-fencing of content [5].
In the remainder of our analysis, we will provide more details on the cat-
egories of “Specialized Networks/Hosting Providers” and “Residential ISPs”
in Sects. 6.2 and 6.3 respectively. The dynamics we observe suggest that the
VPN provider ecosystem w.r.t. geo-unblocking is in constant evolution, proba-
bly because providers constantly look for IPs that do not appear in blocklists.
This brings forward the question if providers also share these resources, or if
they differentiate their infrastructure. We will look into this in Sect. 6.4.
Key Takeaways: Analysis of regional Internet registry data for the IP addresses
involved in geo-unblocking indicates that there are two main approaches to geo-
unblocking: specialized networks and residential ISPs. Furthermore, we see strong
indications that VPN providers adapt their behavior. This may be driven by
attempts of the VOD providers to block them, or alternatively, the change in
behavior may also be a result of a need to have more bandwidth available to
satisfy the demand of their customers. As an external observer, however, we
cannot ascertain whether either of these two is the case or not.
Table 5. Matrix of PVDataNet AB and Telia Company AB’s ASes and IPs per vantage
point and the corresponding maintainer according to RIPE. IPs are represented by their
network prefix.
What differentiates this subclass from non top-tier hosting providers is that the
organisations in this category generally do not service any retail customers and
have optimized their business model to monetize IP addresses.
Key Takeaway: The use of specialised ISPs, especially those that are used by
multiple VPN providers, suggests that there may exist a specialised market that
caters to the needs of VPN providers for the combination of sufficient bandwidth
coupled with IP addresses that are not blocked by VOD providers.
Finally, we want to highlight one particular model that does not fit well into
either category. In particular we want to spotlight a company that appears to be
wholly owned and operated by PrivateVPN, for the sole purpose of appearing
to be a consumer ISP. This company, called Nordic Internet Service AB came to
light when we performed further investigation of the IP ranges we collected dur-
ing our measurement. In particular, this concerns IP ranges that belong to either
Telia Company AB (AS1299), Datacamp Limited (AS212238) or PVDataNet AB
(AS42201). Looking at the administrative information in the RIPE database,
the so-called maintainer object of these IP ranges shows that they are dele-
gated to either Nordic Internet Service AB, Privat Kommunikation Sverige AB
or PVDataNet AB as shown in Table 5.4
We consulted the Swedish companies registration office (Bolagsverket) for
additional information on the companies listed as RIPE maintainers, due to their
similarity in name. What we found is that the CEO for PrivateVPN Global AB
and Nordic Internet Service AB is the same person. Furthermore, this person
also acts as ordinary board member for PVDataNet AB. We therefore conclude
that PVDataNet AB, Nordic Internet Service AB and PrivateVPN Global AB,
are essentially the same entity.
4
Note that, at the time of writing, some IP ranges have already been re-allocated to
different providers and current RIR data might not reflect the data of this table.
62 E. Khan et al.
Table 6. Residential ISPs with the largest amount of unique IPs per core unblocking
region.
Country AS n
Germany Deutsche Telekom (AS3320) 2,412
Vodafone (AS3209) 1,143
Japan NTT (AS4713) 1,052
Softbank (AS17676) 724
Netherlands Vodafone Libertel (AS33915) 347
KPN (AS1136) 319
United States Comcast (AS7922) 48,288
AT&T (AS7018) 24,306
(∼ 0.1%) and no IPv6 addresses. These 273 IPv4 addresses in turn account for
only ∼ 5.8% of all observations (580,360 of 9,987,688). Even within this set of
IPv4 addresses, the distribution is not uniform. As Fig. 7 shows, roughly 10 IPs
make up half and about 40 IPs are responsible for 90% of our overlapping obser-
vations. In other words, based on these numbers we assume that it is highly
unlikely that the VPN providers share a common geo-unblocking platform.
Taking a step back though and looking at a coarser-grained set of data,
namely ASNs, we can see a different picture. In total, we observed 2,046 distinct
ASNs of which 464 (∼ 22.7%) have been found to overlap between the different
VPN providers. Table 8 shows how often ASNs that overlap occur in our dataset.
Generally speaking the amount of shared ASNs resembles a heavy-tailed dis-
tribution and can be seen in Fig. 8. Just two ASNs (AS212144 (45.3%) and
AS212238 (11.1%)) make up 56.4% of all overlapping observations. More impor-
tantly, though, they make up 48.8% of all observations in our data set. The first
ASN, AS212144, Trafficforce UAB (which we discussed previously in Sect. 6.2)
is shared among four of the VPN providers (NordVPN, Surfshark, CyberGhost,
ExpressVPN) and the latter, AS212238, Datacamp Ltd. (which announces IP
space for Trafficforce) is shared among five (PrivateVPN, CyberGhost, Surf-
shark, NordVPN, ExpressVPN). This can also be noticed in Fig. 6 by keeping
in mind that the same ASN is plotted in the same color throughout the picture.
Key Takeaway: The majority of overlap in infrastructure between VPN
providers lies with just a few distinct ASNs that seemingly belong to a class
of service provider that caters well to the need of VPN providers that want to
perform geo-unblocking. Nevertheless, we also still see evidence of overlap in
residential connections
Fig. 7. Number of overlapping IPv4 addresses and amount of occurrences in our data
set
Stranger VPNs 65
Fig. 8. Number of overlapping ASNs and amount of occurrences in our data set
7 Ethical Considerations
The work described in this paper has presented us with several different ethical
dilemmas that affect many of the different stakeholders in this context. In this
section we aim to provide an overview of the ethical challenges that we iden-
tified and explain how we handled them to minimize impact on the relevant
stakeholders.
7.1 Ecosystem
The ecosystem of VPN providers in the context of geo-unblocking and Video-
on-Demand (VOD) providers creates tensions. As described in Sect. 3 VOD
providers put measures in place to restrict material due to distribution right
restrictions. VPN providers in contrast advertise with geo-unblocking capabili-
ties to allow VPN users to circumvent the restrictions put in place by the VOD
providers. VOD providers in turn aim to detect the circumvention methods,
which causes VPN providers to come up with alternative ways to route traffic
and evade this detection.
This context presents ethical tensions:
– Users may use the VPN services to circumvent the geo-blocking measures of
VOD providers, which breaks the terms of service of most of these services,
and may even be illegal in some jurisdictions.
– VOD providers currently restrict content due to licensing agreements, and
have put detection capabilities in place to prevent geo-unblocking. Research
into this context provides them with additional information on the practices,
allowing (or perhaps even forcing) them to improve their detection capabili-
ties.
– VPN providers on the other hand advertise with the geo-unblocking capabil-
ity. Our research into this practice may hurt their business practices.
8 Conclusion
Both streaming and VPNs are multi-billion dollar industries [16,17]. The two are
constantly locked in an arms race where VPN providers are trying to offer geo-
unblocking to their customers and VOD providers are trying to enforce restric-
tions on the content delivery to certain regions to enforce licensing agreements.
In this paper we shed first light on how VPN providers circumvent geo-
blocking restrictions. Our main findings are three-fold. Firstly, VPN providers
use different mechanisms for bypassing geo-blocking, making use of special-
ized networks/hosting providers and residential ISPs. Secondly, VPN providers
use different mechanisms in different geographical regions, thus adapting their
behavior to what best ensures escaping VOD detection mechanisms. Finally
there are also temporal dynamics, i.e., the approaches change at different
moments in time, even within the same provider. These findings paint a pic-
ture of the geo-unblocking VPN providers’ ecosystem as highly dynamic and
adaptable. Given the value of the market we expect this arms race to continue
in a future where we might see even Stranger VPNs.
References
1. Tunnelr - maintenance mode - we’ll be right back! (2018). [Link]
com/
2. Digital element commemorates 20th anniversary (2019). [Link]
[Link]/digital element 20th anniversary/
Stranger VPNs 67
24. Huffaker, B., Fomenkov, M., claffy, k.: Geocompare: a comparison of public and
commercial geolocation databases. Tech. Rep. (2011)
25. Ikram, M., Vallina-Rodriguez, N., Seneviratne, S., Kaafar, M.A., Paxson, V.: An
analysis of the privacy and security risks of android VPN permission-enabled apps.
In: Proceedings of the 2016 Internet Measurement Conference, pp. 349–364. ACM,
Santa Monica California USA (2016). [Link]
26. Khan, M.T., DeBlasio, J., Voelker, G.M., Snoeren, A.C., Kanich, C., Vallina-
Rodriguez, N.: An empirical analysis of the commercial VPN ecosystem. In: Pro-
ceedings of the Internet Measurement Conference 2018, pp. 443–456. ACM, Boston
MA USA (2018). [Link]
27. McDonald, A., et al.: 403 forbidden: a global view of CDN geoblocking. In: Pro-
ceedings of the Internet Measurement Conference 2018, pp. 218–230. IMC 2018,
Association for Computing Machinery, New York, NY, USA (2018). [Link]
org/10.1145/3278532.3278552
28. Mi, X., et al.: Resident evil: understanding residential IP proxy as a dark service.
In: 2019 IEEE Symposium on Security and Privacy (SP), pp. 1185–1201. IEEE,
San Francisco, CA, USA (2019). [Link]
29. Perta, V.C., Barbera, M.V., Tyson, G., Haddadi, H., Mei, A.: A glance through the
VPN looking glass: IPv6 leakage and DNS Hijacking in commercial VPN clients.
Proceed. Priv. Enhan. Technol. 2015(1), 77–91 (2015). [Link]
popets-2015-0006
30. Ramesh, R., Evdokimov, L., Xue, D., Ensafi, R.: VPNalyzer: systematic investiga-
tion of the VPN ecosystem. In: Proceedings 2022 Network and Distributed System
Security Symposium. Internet Society, San Diego, CA, USA (2022). [Link]
org/10.14722/ndss.2022.24285
31. Tosun, A., De Donno, M., Dragoni, N., Fafoutis, X.: RESIP host detection: identi-
fication of malicious residential IP proxy flows. In: 2021 IEEE International Con-
ference on Consumer Electronics (ICCE). pp. 1–6. IEEE, Las Vegas, NV, USA
(2021). [Link]
32. Weinberg, Z., Cho, S., Christin, N., Sekar, V., Gill, P.: How to catch when proxies
lie: verifying the physical locations of network proxies with active geolocation. In:
Proceedings of the Internet Measurement Conference 2018, pp. 203–217. ACM,
Boston MA USA (2018). [Link]
33. Winter, P., Padmanabhan, R., King, A., Dainotti, A.: Geo-locating BGP prefixes.
In: 2019 Network Traffic Measurement and Analysis Conference (TMA), pp. 9–16.
IEEE, Paris, France (2019). [Link]
TLS
Exploring the Evolution of TLS
Certificates
1 Introduction
TLS has become the de-facto standard for securing the Internet; it is the under-
lying security procedure behind popular communication protocols like HTTPS
and SMTPS.
The widespread use of TLS has led to a lot of efforts from the community to make
the TLS certificate ecosystem more democratic, transparent and economically fea-
sible. Some efforts worth mentioning are the introduction of (1) ACME (specifically
Let’s Encrypt) [6] that allows valid certificates to be issued for free and removes
the need for human intervention for certificate issuance and (2) the CT standard,
which states that all compliant certificates must be published to append-only pub-
lic servers so that any mis-issuance is promptly discovered, thus can be revoked.
This work presents an audit on the evolution of the certificate ecosystem over
the last 8 years by using two sets of a large corpus of certificates; certificates
collected from full IPv4 scans [20] from 2013 to 2021 and certificates logged in
Google operated Certificate Transparency Logs till February 2021.
We make the following contributions. First, we explore how the overall valid-
ity of certificates has changed over time across most end-user applicable root
stores. We observe that while the percentage of valid certificates has improved,
a large portion of certificates are still invalid.
Second, we show that the use of template certificates to create new certificates
has led to certificates presenting invalid extension data. Some certificates using
template certificate fail to update key components in a certificates, which are
supposed to be unique such as subjectKeyIdentifier and ct precert scts
fields resulting in incorrect usage of these extensions.
Third, we show how the TLS certificate ecosystem has evolved over the past
8 years. We show that the overall security of certificates such as key strength
has improved, and the ecosystem has become more centralized over time with a
small number of CAs issuing a large percentage of total certificates.
2 Background
TLS Certificates: A TLS certificate binds a subject (domain) to a public key.
These certificates are usually issued and signed by Certificate Authorities (CAs)
once it successfully vets the subject. Thus, certificates usually have a certificate
chain rooted in a widely-trusted set of root certificates, which are self-signed.
X.509 [12] is the most commonly used certificate management standard.
X.509 certificates typically includes the subject (e.g., domain name), issuer (i.e.,
CA), public key, serial number (unique to a CA). It can also have additional infor-
mation such as CRL Distribution Points extension [12], which allows a client to
perform revocation check using URLs provided in the extension.
3 Related Works
Free and Automated CAs. While most measurement studies for PKI focus on
valid certificates, a previous study [11] showed that a majority of certificates
(88%) were invalid. The reason for these certificates being invalid was economical
and a vast majority of invalid certificates originated from IoT devices. Since then,
the introduction of Let’s Encrypt [15] and the concept of free, automated CAs
has made it increasingly easy and economically feasible to get valid certificates.
Other CAs (like Sectigo [3] and cPanel [25]) soon followed suit and added support
for automated certificate issuance.
Our dataset consists of certificates collected via full IPv4 scans in project Sonar
by Rapid7 [20] and the certificates logged to Certificate Transparency logs man-
aged by Google.
IPv4 Scans: These scans are conducted by Rapid7, are open for public use, and
are designed to find certificates from HTTPS endpoints. The timeline of this
dataset spans from September 2013 to December 2021 with a total of 358,575,204
unique certificates observed. Scans were conducted every week from September
2013 to June 2017, every two weeks from June 2017 to January 2019, and daily
afterwards. Additionally, from September 2013 to January 2018, only port 443
(the standard port for HTTPS) was scanned, while alternate HTTPS ports were
74 S. M. Farhan and T. Chung
also scanned afterwards. We use the number of unique certificates as the unit of
measurement for this dataset as a certificate is expected to be seen in multiple
scans.
Root Stores. When validating certificates, the first question that needs to be
answered is which root store should we use. Since our goal is to find out if end-
user applications will find these certificates to be valid, we use four different root
stores (Apple, Microsoft, Mozilla NSS, and Android root stores) to validate all
certificates collected from IPv4 scan. Previously, Zane et al. [24] showed that a
vast majority of root stores used in end user applications stem from a handful
of ‘root’ root stores.
Since our certificate transparency dataset only includes certificates logged to
Certificate Transparency logs managed by google, we want to use the root store
from google’s CT logs. Google states that the root stores for CT should be a
super set of all major root stores (Apple, Microsoft and Mozilla) to be inclusive
of all certificates that may be considered valid by these entities [4]. Looking at
the current snapshot of root certificates for all the CT logs in our dataset, we find
that the root stores for all active CT logs managed by google are the same. We
then backtrack through the entirety of the timeline that CTs have been active
to ensure we have all the historical root certificates in our root store.
5 Certificate Validity
5.1 IPv4 Scanning
Unless stated otherwise, we define a certificate as valid if it is valid for any end-
user root store in our dataset. After validating all certificates in our dataset,
we isolate 121,062,606 unique valid certificates (33.76% of all certificates). We
observe that the majority (66.23%) of certificates are invalid across all the root
stores we use in our test; more specifically, 55.41% (131,625,055) of the invalid
certificates are invalid because they are self-signed and 44.58% (105,770,334) are
invalid because they are signed by another invalid certificate. This accounts for
the vast majority of invalid certificates with only 1074 certificates invalid due to
some other reason.
Figure 1 shows the set relationship across different root stores. We find that
the CT root store contains all the root certificates from other root stores apart
from one certificate that is only present in the Apple root store. We also observe
that while a large number of certificates are shared between all end-application
root stores, a large number of certificates are unique to the Microsoft and Apple
root stores. However, there is little variation in the validity of certificates between
different root stores. We find that 120,200,897 certificates (99% of the certificates
valid in any root store) are valid across all end-application root stores in out
dataset. This echoes the work done by Pearet al. showing that a vast majority
of HTTPS servers use CAs that are trusted by all major trust stores [19].
76 S. M. Farhan and T. Chung
Fig. 2. The number of valid and invalid certificates via IPv4 scanning
Portion of certificates
1 Not issued to IP
0.8 Private (not routable) IP
Public IP
0.6 Other
0.4
0.2
0
2014 2015 2016 2017 2018 2019 2020 2021
The total number of certificates and percentage of valid certificates has been
on the rise throughout our scanning period as shown in Fig. 2. We surprisingly
saw a declining validity percentage after 2018, which we later found was an
artifact of more frequent scanning since a large percentage of invalid certificates
are seen in a single scan and are ephemeral. The corrected validity percentage
(where we sample our data after 2019 to be consistent with the prior data) shows
a constant increase over time. However, there remains a large percentage (45%)
of invalid certificates.
Why are there still invalid certificates? Almost half of the invalid certificates
are issued to IP addresses (i.e. the common name is an IP address) as highlighted
in Fig. 3 (the other half has a valid FQDN); we find that the percentage of invalid
certificates issued to private IPs has been on the rise for the past three years, while
the percentage of invalid certificates not issued to IP addresses has been on the
decline. Only a minute number of valid certificates are issued to IP addresses as
CAs will generally not issue certificates without a valid domain name.
Where are Invalid Certificates Hosted? We use the common name of the
certificate and the subject alternative names of the certificates to find all the
domains a certificate represents. Note that some certificates are issued to IP
addresses and are excluded from this analysis (and some certificates counted
multiple times in different domains) Fig. 4 shows the top 5 top level domains in
Exploring the Evolution of TLS Certificates 77
Fig. 4. Popular top level domains for valid and invalid certificates
our dataset. We observe that popular web domains like .com and .net hosting
publicly accessible web pages tend to present valid certificates. Throughout the
manual investigation, we find that invalid certificates for the .net domain are
largely issued by Kubernetes. The .box and .nas domains (over 99.9%) are almost
exclusively invalid and routers from AVM (fritz box) constitute a large majority
of these certificates; thus, we believe that such invalid certificates are used for
individual applications (such as using a network attached storage device). Since
the .local domain is reserved for use by Internet Engineering Task Force (IETF),
thus we exclusively observed them in invalid certificates. We could not identify
any patterns that might tell us their source.
6 Certificate Authorities
6.1 IPv4 Scanning
Figure 6 tracks how domains (CNs) have been migrated across different CAs. We
can infer that common names representing invalid certificates tend to be short-
lived, since the percentage of invalid certificates in Fig. 6 does not match the one
1
It is worth noting that these problematic certificates cannot be removed because the
CT log is a Merkle-tree based structure, which is append-only.
78 S. M. Farhan and T. Chung
400 1.6x109
350 Certificates 1.4x109
# of invalid certificates
Invalid certificates
(parts per million)
300 1.2x109
# of certificates
250 9
1x10
200 8x108
8
150 6x10
100 8
4x10
50 8
2x10
0 0
2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022
year
We find that Let’s Encrypt dominates the Certificate Transparency logs, con-
sistently issuing 80% of the certificates logged after 2016; Note that the volume
in CT was quite low prior to 2016. As shown in Fig. 7, cPanel and Sectigo also
consistently log a substantial portion of certificates. Since the vast majority of
the certificates from CT Logs are valid as shown in Fig. 5, we could not find any
distinctive pattern in terms of the population of invalid certificates acrossthe
CAs.
Exploring the Evolution of TLS Certificates 79
Fig. 6. Tracking yearly CA choice for all common names in the certificates from IPv4
scanning. The red flow represents invalid certificates while the blue flow represents
valid certificates. (Color figure online)
7 Host Networks
Now, we focus on where the certificates are hosted from by looking at the IP
address of hosts that serve certificates.
1
0.8
Proportion of Certificates
Invalid certificates Content
0.6 Enterprise
0.4 Transit/Access
Other
0.2
0
1
Valid certificates
0.8
0.6
0.4
0.2
0
2014 2015 2016 2017 2018 2019 2020 2021
Fig. 8. Distribution of AS types over time for certificates collected through IP scanning
for valid certificates while the proportion of valid certificates hosted via content
ASes is increasing.
8 Evolution
This section describes the major changes we observe in the IP scanned certifi-
cates.
Signature Algorithm: In early 2017, major browsers including Chrome, Fire-
fox and Safari officially depreciated the use of SHA1 as an encryption algorithm
for certificates [1] as there are significant collision attacks available for the algo-
rithm. Almost all valid certificates shifted to using SHA256 in favor of SHA1
by 2017. The same change happened much slower for invalid certificates with a
considerable portion of certificates still employing SHA1 as late as 2020.
Certificate Revocation. We rarely find a certificate revocation mechanism
defined for invalid certificates. For valid certificates, we find that CRLs were very
popular till 2015, after which we observe a constant decline in the percentage of
valid certificates supporting CRLs. OCSP was quite popular in 2013 with 90% of
Exploring the Evolution of TLS Certificates 81
Date
the valid certificates supporting it, but after 2015 practically all valid certificates
support OCSP.
CT Inclusion Proofs. Figure 10 shows the inclusion in CT for Rapid7 scanned
certificates over time. More than 99% of valid certificates after 2018 are included
in CT logs, while a small but increasing proportion of invalid certificates have
SCTs. While the rules governing Certificate Transparencies to filter certificates
are more lenient than those for validity, it is unexpected to find invalid certificates
with SCTs. We find that these certificates are sharing SCTs. SCTs should be
cryptographically generated by the CT log and thus valid SCTs must be unique.
Only 24% of invalid certificates have a unique SCT, in contrast we find no
valid certificate sharing an SCT with another valid certificate. We revalidate
these certificates (the ones that have unique SCTs) using the CT root store and
corresponding rules and find that a majority of these may be considered valid
by CT [17].
99% of the invalid certificates sharing SCTs are issued by Fortinet. In the
general case we find one valid certificate who’s SCT is shared by multiple invalid
certificates. Moreover, other (certificate specific) extensions, like Subject Key
Identifier are shared among these certificates even though they do not share a
public key. We believe that these are a result of using existing valid certificates
as a template to create new certificates. This explains why these certificates
present invalid SCT tags and share the Subject Key Identifiers. We reached out
to Fortinet for a comment but have not received a response.
9 Discussion
Free, Automated CAs like Let’s Encrypt and cPanel have been the biggest source
of change for the TLS certificate ecosystem. These services allow valid certificates
to be issued without any real investment from the domain owner (both in terms
82 S. M. Farhan and T. Chung
Proportion of certificates
1 Valid
Invalid
0.8
0.6
0.4
0.2
0
2014 2015 2016 2017 2018 2019 2020 2021
Fig. 10. Proportion of valid and invalid certificates with SCTs defined
of time and money). The improvement in validity percentage that we see is also
in large part due to these Certificate Authorities. We can also attribute the
decrease in validity period to these CAs.
On the other hand, these ACME-supporting CAs also put the overall ecosys-
tem at risk. The nature of the PKI ecosystem means that any domain can be
impersonated by compromising the least secure CA, and there are known attacks
against the domain validation employed by these CAs. We observe that central-
ity in issuers has increased considerably over the course of our scans, and this is
likely caused by the popularity of Let’s Encrypt and cPanel. Only 10 keys are
responsible for directly signing 80% of the valid certificates, which means that
in case these keys are compromised, a mast majority of valid certificates should
become invalid as the certificates with these keys are revoked, or are removed
from trust stores. These keys are likely to be compromised when compared to
keys used for root certificates, as the root certificate keys are rarely held in
memory to sign other certificates, while these are continuously held in memory.
10 Conclusion
This work presents a bird’s eye view of how the web’s PKI ecosystem has evolved
over the past 8 years. The validity of certificates has improved consistently, but a
large proportion of certificates are still invalid. Over time, most indicators show
that the ecosystem is moving towards better security practices. However, there
are a few alarming trends including the incorrect use of template certificates
causing invalid extensions and increasing centrality in issuers.
A Ethics
This paper does not pose any ethical issues as all the data we use for analysis is
collected by third parties and is available for public use.
Exploring the Evolution of TLS Certificates 83
References
1. Browser security icon updates And sha-1 deprecation — digicert. Com. https://
[Link]/blog/browser-security-icon-updates-sha-1-deprecation
2. incident response: november 2015 google ‘pilot’ And ‘aviator’ Logged 3 Certs With
Invalid Signatures
3. Sectigo Adds Acme Protocol Support In Certificate Manager Platform To Auto-
mate Ssl Lifecycle Management. 2019. [Link]
go-adds-acme-protocol-support-in-certificate-manager-platform-to-automate-ssl-
lifecycle-management
4. Certificate transparency Website. Google, March 2022. [Link]
le/certificate-transparency-community-site/blob/4ae90f78cdd821b9cb3a68848851
86d71b529c87/docs/google/[Link]
5. Laurie, et al. Certificate Transparency. RFC 6962 (Proposed Standard), IETF,
June 2013
6. Barnes, R., Hoffman-Andrews, J., McCarney, D., Kasten, J.: Automatic Certificate
Management Environment (acme). IETF, March 2019
7. Sloot, V., et al.: Towards a complete view of the certificate ecosystem. In: Pro-
ceedings of the 2016 Internet Measurement Conference, New York, NY, USA, pp.
543–549, Association for Computing Machinery (2016)
8. Caida asclassifications Dataset. [Link]
9. Cangialosi, F.,et al.: Measurement and analysis of private key sharing in the
https ecosystem. In: ACM Conference on Computer and Communications Security
(CCS), Vienna, Austria, October 2016
10. Certificate Transparency In Chrome (2019). [Link]
policy/blob/master/ct [Link]
11. Chung, T., et al.: Measuring and applying invalid SSL certificates: the silent major-
ity. In: ACM Internet Measurement Conference (IMC), Santa Monica, California,
USA, November 2016
12. Cooper, D., Santesson, S., Farrell, S., Boeyen, S., Housley, R., Polk, W.: Internet
X.509 public key infrastructure certificate and certificate revocation list (CRL)
profile. RFC 5280, IETF, May 2008. [Link]
13. Dukhovni, V., Hardaker, W.: The DNS-based authentication of named entities
(DANE) protocol: updates and operational guidance. RFC 7671, IETF, October
2015
14. Gasser, O., Hof, B., Helm, M., Korczynski, M., Holz, R., Carle, G.: In log we trust:
revealing poor security practices with certificate transparency logs and internet
measurements. In: Passive and Active Measurement Conference (PAM) (2018)
15. Aas, J., et al.: Let’s encrypt: an automated certificate authority to encrypt the
entire web. In: Proceedings of the 2019 ACM SIGSAC Conference on Computer
and Communications Security, New York, NY, USA, pp. 2473–2487, Association
for Computing Machinery (2019)
16. Korzhitskii, N., Carlsson, N.: Characterizing the root landscape of certificate
transparency logs. In: 2020 IFIP Networking Conference, Networking 2020, Paris,
France, 22–26 June 2020, pp. 190–198. IEEE (2020)
17. Laurie, B., Langley, A., Kasper, E.: Certificate Transparency. RFC 6962, IETF,
June 2013. [Link]
18. Li, B., et al.: Certificate transparency in the wild: exploring the reliability of mon-
itors. In: ACM Conference on Computer and Communications Security (CCS)
(2019)
84 S. M. Farhan and T. Chung
19. Perl, H., Fahl, S., Smith, M.: You won’t be needing these any more: on removing
unused certificates from trust stores. Financial Cryptography and Data Security
(FC), Christ Church, Barbados, March 2014
20. Project sonar. [Link]
21. Scheitle, Q., et al.: The rise of certificate transparency and its implications on the
internet ecosystem. In: Proceedings of the Internet Measurement Conference 2018,
New York, NY, USA, pp. 343–349. Association for Computing Machinery (2018)
22. Stark, E., et al.: Does certificate transparency break the web? Measuring adoption
and error rate. In: IEEE Symposium on Security and Privacy (IEEE S&P) (2019)
23. Working Together To Detect Maliciously Or Mistakenly Issued Certificates.
[Link]
24. Ma, Z., Austgen, J., Mason, J., Durumeric, Z., Bailey, M.: Tracing your roots:
exploring the TLS trust anchor ecosystem. In: Proceedings of the 21st ACM Inter-
net Measurement Conference, New York, NY, USA, pp. 179–194, Association for
Computing Machinery (2021)
25. Cpanel Auto SSL. [Link]
ssl/#autossl
Analysis of TLS Prefiltering for IDS
Acceleration
1 Introduction
2 Related Work
2.1 Software-only Solutions
The problem of meeting the throughput of current networks is the leading issue
for intrusion detection systems. There are multiple approaches to how network
administrators can deploy such systems. The first and most common one is a
software-only-based IDS. There are multiple options to choose from, notably,
IDS systems like Snort [8], Suricata [10], and Zeek (formerly Bro) [11] are the
most popular. The software-only approach leads to the easiest deployment but
is also tied to the most limited performance. As an example described in [19],
reaching a network throughput of 100Gbps with Zeek requires a complicated
deployment scheme involving multiple switches and servers, where each runs an
instance of Zeek.
On the other hand, some research published in papers, e.g., Pigasus [22], Snort
Offloader [18], propose a heavily-oriented hardware solution. The aim is to put
the whole/most of the IDS into the FPGA to accelerate the overall processing
performance. Snort offloader puts an entire IDS/IPS solution to the FPGA,
however, it lacks certain capabilities (e.g. TCP reassembly) to be deployed in
the production. Pigasus implementation is more complete than [18]. It puts
most parts of the IDS into the FPGA while the CPU is used for exact pattern
matching. However, due to resource constraints of FPGAs, Pigasus limits the
number of rules to 10 000 and the size of the flow table to 100 000 flows. Enabled
detection rules are statically inserted into the FPGA so they cannot be changed
dynamically.
independent parts - acceleration logic and IDS. For our experiments, we chose
Suricata IDS for its multi-threaded, high-performance architecture and, gener-
ally, good extensibility for further additions.
Suricata IDS supports multiple capture modules for packet sniffing with cap-
ture module AF PACKET being the most prevalent in the Suricata deployments.
As mentioned previously, AF PACKET can be extended with eXpress Data Path
(XDP) filtration program running in kernel space to improve the performance
of the capture module.
Suricata also contains a capture module based on Data Plane Development
Kit (DPDK) [6]. Unlike other packet capture interfaces, Suricata constantly polls
packets of the NIC to improve the throughput (performance) of the system. The
results of this approach have been evaluated in [23] where the results of DPDK
were compared to the AF PACKET non-bypassed capture interface. We decided
to have the DPDK capture module as a foundation for our work and compare
our results to baseline measurements of both AF PACKET (with XDP) and
DPDK capture modules.
The paper contains an analysis of real-world network traces, a comparison of
various TLS bypass methods coupled with the proposed acceleration solution,
and an elaborate analysis of detection results. The analysis of the detection
results proves the correctness of the results of the proposed solution.
We have classified the traffic into two groups, where one group represents traffic
with TLS protocol and the other group represents the remaining traffic.
Academic Academic
networks TLS networks
TLS
Commercial Other Commercial Other
networks networks
% of flows
Fig. 1. TLS analysis based on the number of bytes, packets, and flows
the client and the server. Based on the type of the message, a sniffing applica-
tion (IDS) can identify whether the payload of the TLS record is encrypted or
not. TLS handshake between communicating parties follows right after the TCP
handshake and is still unencrypted. During the TLS handshake, several mes-
sages are exchanged primarily for acknowledgment and verification of both sides
and for agreeing on the used cryptographic algorithms and session keys for the
subsequent encrypted communication. Once communication becomes encrypted,
it does not downgrade back to unencrypted. The graph in Fig. 2 represents an
occurrence of various TLS records within the TLS traffic of the datasets. TLS
record type ApplicationData serves as the main message for encrypted commu-
nication. This type of message is not valuable for IDS as it only contains the
length of the content and the encrypted message content itself. IDS can process
the remaining message types to match specific fields, e.g. server name identifica-
tion (SNI), or analyze JA3/JA3s hashes of the handshake messages as described
in [5].
This hints that only approximately 4% of TLS traffic is actually useful for
IDS and the remaining 96% can be bypassed. Taking an example of a 100 Gbps
network and given the facts mentioned previously, on average, only 30% of the
total network traffic (non-TLS) and 4% of TLS traffic is relevant for IDS. This
reduces data throughput to only 33 Gbps of valuable traffic from the overall 100
Gbps. The results also imply that IDS can be potentially sped up by up to 3
times in our network. Obviously, the overall acceleration is dependent on the
network traffic mix.
The main concept behind TLS Prefilter is to look for encrypted TLS records
and then make a decision about the bypass. The header of TLS records can be
found in the packet after the TCP header. From the TLS record, it is possible
to deduct whether the TLS payload is encrypted or not. Once TLS Prefilter
detects an encrypted TLS payload then it can safely assume that the flow remains
encrypted until the connection is terminated.
During our experiments and prototypes, we have come up with two solutions
to TLS prefiltering that are displayed in Fig. 5. The first solution works in a
stateless mode (Fig. 5a) where TLS Prefilter is analyzing packets and upon the
first encrypted TLS record, TLS Prefilter creates a bypass record in its internal
bypass flow table. The solution greatly helps Suricata to increase its performance.
Even though the traffic analysis of the current NREN and commercial networks
shows a majority of the traffic is encrypted, unencrypted services can still exist.
The advantage of plain traffic is that it can be completely monitored by IDS and
its pattern-matching engine. However, the naive approach of bypassing the flow
on the first encrypted TLS record can possibly open a security vulnerability in
these network-monitored environments. For example, monitoring of the unen-
crypted HTTP server can be circumvented by sending, among valid requests,
also a fabricated request containing a TLS-encrypted record. After this packet,
the stateless bypass-triggering logic of TLS Prefilter would bypass the flow, and
Suricata (IDS) would be completely blind to the given flow.
To mitigate this issue we have converted TLS Prefilter to be stateful (Fig. 5b).
That means the bypass table stores the states of each tracked flow. In this case,
TLS Prefilter does not issue flow bypass based on one packet. After detecting
an encrypted TLS record, TLS Prefilter creates a new record in the flow table
but does not issue a bypass yet. The flow remains inspected by the IDS until
TLS Prefilter receives the second encrypted TLS record but from the opposite
direction. This means that even if the attacker sends a forged packet with the
encrypted TLS record to the server’s direction, TLS Prefilter still passes all
packets of the flow to the IDS. TLS Prefilter enables the flow bypass only if the
server also replies with an encrypted TLS record.
Figure 6 presents the general flow control of the analyzed packets. Receiving
a packet from the top, it first parses the packet to an internal representation.
If the parsing fails (e.g. due to non-TCP or non-IPv4/6 packets), the packets
are passed to the IDS. The TLS Prefilter is alternatively able to unpack packet
encapsulation, e.g., VLAN protocol. Internal packet representation is formed by
a 5-tuple. Flow 5-tuple is defined as pairs of source and destination addresses and
Analysis of TLS Prefiltering for IDS Acceleration 95
ports and transport layer protocol (TCP/UDP). Even though TLS is based only
on the TCP protocol and therefore the information about the transport layer
protocol can be considered redundant, Prefilter follows the convention and uses
the standard 5-tuple flow identification. For efficient lookups, n-tuples are unified
for both directions of the communication by placing the lower IP address as the
first one in the pair. The same applies to the port pairs. This allows storing only
1 entry per-flow while also performing only 1 lookup for the subsequent packets
coming from both directions. After successful parsing, TLS Prefilter does a table
lookup. If the packet should not be bypassed, TLS Prefilter inspects the packet
for encrypted TLS records.
In case the packet contains some, TLS Prefilter checks if a TLS flow has been
registered in the flow table. Registering a flow means detecting the first encrypted
TLS record. After the flow entry is created, TLS Prefilter also notes the direction
(based on, e.g., IP addresses and ports) from which the packet came. At this
point packets of the flow are still passed to the IDS as the received packet can
still be forged by an attacker. If a flow entry is already present in the flow table,
96 L. Sismis and J. Korenek
then TLS Prefilter has already encountered at least one TLS-encrypted packet
of the given flow. The received packet is further inspected by TLS Prefilter
to determine whether it comes from the opposite direction as from which the
first encrypted TLS traffic has arrived. Receiving a packet from the reversed
direction means both sides of the connection use an encrypted TLS protocol to
communicate and therefore the flow can be bypassed.
Bypass Removal
Coming back to the beginning of the control flow diagram displayed in Fig. 6, if
the packet’s flow tuple is contained by the bypass table, TLS Prefilter analyzes
TCP header flags of the matched (bypassed) packets for the closing/opening
flow signs (e.g. RST/SYN/FIN+ACK flags). Flows of respective packets that
Analysis of TLS Prefiltering for IDS Acceleration 97
contain these flags remove bypass entries from the bypass flow table. Packets
not containing these flags are bypassed. Bypassed packets are then discarded by
TLS Prefilter.
TLS Prefilter can delete a bypass flow entry if a packet has its SYN flag set
as this is usually a sign of a new flow. If the communication continues (e.g. SYN
bit set but the sequence number outside of the expected window as described in
RFC 5961), TLS Prefilter removes bypass entry after receiving a such packet.
This way Suricata can receive more traffic than anticipated. However, if TLS
Prefilter encounters encrypted traffic from both directions of the flow again, the
bypass entry would be reinserted.
Similar behavior can happen with other closing TCP flags as well. To com-
pletely eliminate the issue, TLS Prefilter could wait with the bypass entry
removal until it is also confirmed with the other side of the communication.
For instance, this can be detecting a packet with FIN+ACK flags that is part
of the bypassed flow and only deleting the bypass entry from the table after
encountering a packet from the opposite side of the communication with the
same closing flags.
requires yet another separate CPU core(s) (as compared to running a bare IDS
without the TLS Prefilter), it was crucial to amortize this cost by distributing
packets in 1:N relationship, meaning from a single Prefilter core the packets are
scattered to multiple IDS cores/workers. The distribution operation had to take
into account the previously mentioned workers’ requirement (bi-directional flow
capture by one worker). The distribution problem has been solved by Receive
side scaling (RSS) [1] on the NIC card and symmetric RSS [21]. The number of
NIC queues is determined by the number of subscribed IDS workers.
Avoiding resource sharing was an important objective in the TLS Prefilter
design process. As can be seen in Fig. 7, each core of Prefilter operates on the
designated NIC queues, has separate rings for the assigned Suricata workers, and
has a unique bypass flow table. The bypass flow table is realized as a fast hash
table with a flow key and flow data as a key-value pair.
We started to experiment with different types of hash tables and finally
selected the Least Recently Used (LRU) hash table. The LRU hash table on
the inability to insert an entry into the hash table bucket (i.e. on the full hash
table bucket), replaces the least recently used entry with the inserting one. This
provides a simple mechanism to evict stale flows with new ones. LRU hash table
does not guarantee that all flows will be bypassed all the time. In case an active
flow is evicted, then it is bypassed with the pair of the following encrypted TLS
records. This attribute of the hash table can be manually adjusted by the size
of the flow bucket. However, the lightweight operation of the LRU hash table
proved to provide the best throughput while also keeping the detection results
consistent.
5 Results
– AF PACKET
– AF PACKET + XDP TLS bypass
– DPDK
– DPDK + internal TLS bypass
– DPDK + internal TLS + TLS Prefilter rules bypass - TLS Prefilter
rules bypass flows after encountering TLS-encrypted data received from both
directions of the communication
– Stateless DPDK TLS Prefilter - bypassing flows after detecting a single
TLS-encrypted packet
– Stateful DPDK TLS Prefilter - bypassing flows after encountering TLS-
encrypted data received from both directions
Analysis of TLS Prefiltering for IDS Acceleration 101
All tested variants were manually tuned to improve the performance from
the default configuration. However, sections of the configuration file not related
to the packet capture modules (e.g. detection settings) were shared among all
experiments.
Before any results are presented, it is important to note that the performance
of Suricata/IDS, in general, varies on many factors. Of the main ones, we can
mention the size of the ruleset, complexity of the individual rules, memory and
NUMA placement, CPU speed, or traffic composition. For this reason, we present
mainly relative measures rather than absolute values.
The first experiment was focused on measuring the Suricata throughput
under increasing load. The experiment was based on 4 Suricata workers run-
ning on separate CPU cores. Each worker executed detection on the full rule-
set. Figure 9 presents the results where a horizontal axis describes the speed at
which packets were transmitted against Suricata and a vertical axis describes the
ratio of processed packets. The ratio was calculated as the number of processed
packets by Suricata in proportion to the number of transmitted packets. In the
measurements that include bypass functionality, the dividend in the ratio of pro-
cessed packets was calculated as the number of processed packets by Suricata
plus the number of bypassed packets. Since the DPDK TLS Prefilter runs on
1 core and distributes packets to 4 Suricata workers, this measurement used 5
cores in total. From the graph, it is possible to notice that AF PACKET with 4
workers and the complete ruleset starts to discard packets at the bandwidth of
around 2 Gbps. AF Packet in combination with the XDP filter (AF PACKET +
XDP TLS bypass) increases the throughput to 2.6 Gbps. Performance-oriented
DPDK packet capture interface outperforms the default AF PACKET capture
module with a 2.4 Gbps throughput. When bypass functionality is enabled in the
same way as it was used in AF PACKET + XDP TLS bypass measurement, the
102 L. Sismis and J. Korenek
Table 1. Traffic that reached Suricata after applying individual bypass methods where
100% is all transmitted traffic
DPDK capture interface (DPDK + internal TLS bypass) reaches an even higher
throughput of 2.8 Gbps per 4 Suricata workers. On top of that, the next variant
(DPDK + internal TLS + TLS Prefilter rules bypass) with a set of bypassing
rules that emulates the TLS Prefilter in Suricata pushes the performance to 3.8
Gbps.
The performance of Suricata paired with DPDK TLS Prefilter in the next
two measurements increases to up to 7 Gbps. This means that Suricata can
handle over three times more traffic compared to the commonly deployed Suri-
cata setup running on the AF PACKET packet capture interface. The stateless
mode of TLS Prefilter almost doubles the performance of standalone Suricata
running with DPDK packet capture interface, has Suricata-induced TLS bypass
and bypass rules triggered after detecting encrypted TLS records from both sides
of the communication (DPDK + internal TLS + TLS Prefilter rules bypass vari-
ant). Suricata with TLS Prefilter running in stateful mode reach a performance of
6 Gbps. This again triples the performance of the AF PACKET capture interface
and increases the performance by almost 60% compared to the most performant
standalone Suricata setup running on DPDK (DPDK + internal TLS + TLS
Prefilter rules bypass). Stateful processing took some toll on the performance
but, in our opinion, is required to prevent evasion of network monitoring by the
IDS. Bad actors are therefore not able to easily trigger the Suricata bypass by
forging a TLS encrypted packet as it is possible in the stateless mode. Addition-
ally, we see a huge opportunity in transferring the stateful mode to the hardware
platform.
With regards to the results of the measurement displayed in Fig. 9 we have
also examined how many packets Suricata avoids thanks to individual bypass
methods. For this measurement, we were only interested in the variants with
enabled bypass and those are:
Table 1 summarizes how much of the transmitted traffic the IDS is required
to process. From the relationship between the amount of bypassed packets and
bytes, it is possible to observe that encrypted TLS packets are larger than aver-
age. It can be also noted that Suricata with the internally-controlled TLS bypass
(AFP + XDP bypass variant) has the lowest bypass rate of all examined vari-
ants. This is most likely caused by strict requirements for triggering bypass from
the IDS-side as it requires processing a complete TLS handshake and only issues
bypass if IDS considers it as a valid handshake. Other variants bypass around
the same amount of traffic.
To better interpret the data presented in Fig. 9, we transformed them to the
performance per one CPU core. This is shown in Fig. 10 where individual cap-
ture interfaces are sorted in ascending order. The results for the DPDK TLS
Prefilter also account for 1 extra core. From Fig. 10 it is possible to observe
the differences between the performance of individual experiments. Comparing
the achieved results of the stateful DPDK TLS Prefilter with the other variants
we can notice a more than doubled increase in performance from the default
AF PACKET capture interface. Compared with the AF PACKET or DPDK
capture modules with the enabled bypass functionality, the one core of the state-
ful DPDK TLS Prefilter can also handle about double the amount of the traffic
bandwidth. In other words, this means the DPDK TLS Prefilter requires only
half the number of cores compared to the existing capture interfaces. Compared
to the standalone Suricata running with DPDK capture interface, TLS internal
bypass, and bypass rules (DPDK + internal TLS + TLS Prefilter rules bypass),
stateful TLS Prefilter boosts the performance by over 26%.
Table 2 helps to visualize how many CPU cores would Suricata need in a
regular network to reach common network speeds of 10, 40, and, 100 Gbps. The
results are always rounded up. The table shows that Suricata with a DPDK
TLS Prefilter can reach at least 40 Gbps using a commodity NIC and a regular
36-core Xeon CPU. Without the DPDK TLS Prefilter, it would be required to
use a load balancer and a set of Suricata servers. Even though the table shows
a theoretical estimate of cores required to analyze the specified speeds, we see
a great acceleration opportunity in having the TLS Prefilter directly embedded
inside the NIC’s hardware.
Based on results presented in Figs. 9 and 10 and Table 1 we have presented
the main differences between the current commonly used Suricata bypass solu-
tion (AF PACKET + XDP TLS bypass) and our newly proposed stateful TLS
Prefiltering method. Both the XDP program and TLS Prefilter receive packets
prior to Suricata. This proves to be a suitable place to discard unwanted packets.
However, the two solutions are distinct not only in the bypass-decision process
but also in the architecture.
XDP program is running in kernel space and has a hash-based flow table
shared with Suricata that runs in the user space. To secure thread consistency,
the flow table needs to contain locks. The solution leads to the lock contention.
Since the XDP program relies on Suricata-induced bypass, Suricata is required
to perform flow record insertion to bypass new flows, lookup to evaluate flow
activity, and deletion to remove stale flows. XDP program then performs lookup
and update operations based on the incoming packets.
During the analysis of the results, we have also looked at a possible explana-
tion for worse results of the AF PACKET + XDP TLS bypass variant mentioned
in the [2]. XDP program runs asynchronously to Suricata operation and passes
packets to Suricata through receive buffers. At the same time, the bypass pro-
cess in the XDP program is dependent on Suricata-induced bypass. As a result,
Suricata bypass decisions are reflected in the XDP bypassing process with a
certain delay. It is therefore possible that by the time the bypass decision is
propagated to the XDP bypass table, the receive buffers already contain packets
of the bypassed flows. This problem is potentially significant with short flows
where the whole flow is inserted into Suricata’s receive buffer earlier than the
bypass decision is reflected in the XDP program. The problem seemed relevant
as the analysis in Fig. 3a shows that 90% of flows are shorter than 100 pack-
ets. However, further analysis showed that while this problem was present, it
has not proved to be the main cause of the worse performance results of the
AF PACKET + XDP TLS bypass variant. The rate of bypassed packets that
were delivered to Suricata (and instead should be bypassed) was around 1%
of the total number of bypassed packets. These packets were then bypassed by
Suricata internal bypass as mentioned in Subsect. 2.4.
As the TLS Prefilter bypass decision does not depend on Suricata, subsequent
packets of the flow are bypassed as soon as the first packet triggers the bypass
decision. As a result, this completely avoids the aforementioned issue. With
regard to data consistency, individual cores of the TLS Prefilter contain separate
Analysis of TLS Prefiltering for IDS Acceleration 105
Fig. 11. Reported alerts (N > 5000) after 50 million packets (evaluated after a single
run)
bypass tables. The bypass tables are solely managed by the TLS Prefilter cores
and as a result, do not need to contain any synchronization mechanism.
This can stem from the fact that during flow termination (as the flow starts to
be encrypted), the DPDK TLS Prefilter uses the ACK/SEQ numbers to gen-
erate the RST packet. In case the packet has an incorrect SEQ/ACK number,
the DPDK TLS Prefilter unaware of this fact, sends the crafted packet to Suri-
cata. Sending the crafted RST packet is a TLS Prefilter optimization to evict
encrypted flow from Suricata sooner than Suricata would bypass the flow based
on flow inactivity.
The presented analysis of the alerts showed only general Suricata engine-
related events. To perform an analysis of the TLS Prefilter impact on Suri-
cata and Emerging Threats, it was necessary to use a different dataset as the
dataset used in the previous measurements did not produce any ET alerts. For
this reason, we have also evaluated TLS Prefilter against malware traffic which
can trigger rules from the Emerging Threats ruleset. The next analysis includes
AF PACKET + XDP TLS bypass and TLS Prefilter and was done on a publicly
available traffic capture of Dridex malware [3]. Dridex malware communicates
over TLS. The presence of Dridex on the network can be detected with ET rules
focusing on JA3 hash and SSL certificate. These are parts of the TLS protocol
which can be obtained or derived from an unencrypted TLS handshake.
The analyzed Dridex traffic capture contains 4294 packets. Table 3 presents
the results of individual variants after analyzing the traffic capture. Variant ”Full
inspection, no bypass” was reading packets from the PCAP directly and exe-
cuted full detection analysis to provide baseline results for the following bypass
variants. Suricata with AF PACKET + XDP TLS bypass as a capture module
bypassed 2889 packets and generated 14 alerts. Suricata in combination with
TLS alerts bypassed 2934 packets and also generated 14 alerts.
The generated alerts were the same in both measurements. The same pair of
alerts was hit multiple times during the analysis of the Dridex traffic capture. The
signature IDs of the generated alerts were 2028765 and 2023476. They detected
the Dridex malware by inspecting the SSL certificate and JA3 hash.
Analysis of TLS Prefiltering for IDS Acceleration 107
Based on these findings we can conclude that DPKD TLS passes all necessary
traffic to Suricata and does not impact Suricata’s visibility/correctness in a
negative way. From the earlier analysis of Suricata engine alerts it is possible
to observe that DPDK TLS Prefitler can even save Suricata from generating
invalid alerts by cutting off the encrypted traffic.
6 Conclusion
References
1. Scalable networking: eliminating the receive processing bottleneck-introducing RSS
(2004)
2. Suricata and xdp (2019). [Link]
suricata-and-xdp
3. Dridex malware traffic capture (2020). [Link]
2020/06/03/[Link]
4. Introduction to eBPF and XDP support in suricata (2021). [Link]
[Link]/hubfs/Library/Documents%20(PDFs)/StamusNetworks-WP-eBF-
[Link]
5. TLS fingerprinting by JA3(s) method (2021). [Link]
ja3
6. DPDK support in suricata (2022). [Link]
configuration/[Link]#data-plane-development-kit-dpdk
7. Emerging threats open ruleset official webpage (2022). [Link]
[Link]/open/suricata/rules/
8. Snort official webpage (2022). [Link]
9. SSL dynamic preprocessor (SSLPP) (2022). [Link]
ssl
10. Suricata official webpage (2022). [Link]
11. Zeek official webpage (2022). [Link]
12. Baker, Z.K., Prasanna, V.K.: High-throughput linked-pattern matching for intru-
sion detection systems. In: 2005 Symposium on Architectures for Networking and
Communications Systems (ANCS), pp. 193–202 (2005). [Link]
1095890.1095918
13. Ceška, M., et al.: Deep packet inspection in FPGAs via approximate nonde-
terministic automata. In: 2019 IEEE 27th Annual International Symposium on
Field-Programmable Custom Computing Machines (FCCM), pp. 109–117 (2019).
[Link]
14. González, J., Paxson, V., Weaver, N.: Shunting: A hardware/software architecture
for flexible, high-performance network intrusion prevention, pp. 139–149 (2007).
[Link]
15. Jamshed, M.A., et al.: Kargus: a highly-scalable software-based intrusion detection
system. In: Proceedings of the 2012 ACM Conference on Computer and Commu-
nications Security. p. 317–328. CCS 2012, Association for Computing Machinery,
New York, NY, USA (2012). [Link]
16. Kučera, J., Kekely, L., Piecek, A., Kořenek, J.: General ids acceleration for high-
speed networks. In: 2018 IEEE 36th International Conference on Computer Design
(ICCD), pp. 366–373 (2018). [Link]
Analysis of TLS Prefiltering for IDS Acceleration 109
17. Mitra, A., Najjar, W., Bhuyan, L.: Compiling PCRE to FPGA for accelerating
snort IDS. In: Proceedings of the 3rd ACM/IEEE Symposium on Architecture for
Networking and Communications Systems, pp. 127–136. ANCS 2007, Association
for Computing Machinery, New York, NY, USA (2007). [Link]
1323548.1323571
18. Song, H., Sproull, T., Attig, M., Lockwood, J.: Snort Offloader: a reconfigurable
hardware NIDS filter. In: International Conference on Field Programmable Logic
and Applications, pp. 493–498 (2005). [Link]
19. Stoffer, V., Sharma, A., Krous, J.: 100G Intrusion Detection, August 2015. https://
[Link]/wp-content/uploads/2016/09/Berkeley-100GIntrusionDetection.
pdf. Accessed 27 May 2022
20. Weaver, N., Paxson, V., Gonzalez, J.M.: The shunt: an FPGA-based accelera-
tor for network intrusion prevention. In: Proceedings of the 2007 ACM/SIGDA
15th International Symposium on Field Programmable Gate Arrays, pp. 199–206.
FPGA 2007, Association for Computing Machinery, New York, NY, USA (2007).
[Link]
21. Woo, S., Park, K.: Scalable TCP session monitoring with symmetric receive-side
scaling (2012)
22. Zhao, Z., Sadok, H., Atre, N., Hoe, J.C., Sekar, V., Sherry, J.: Achieving
100Gbps intrusion prevention on a single server. In: 14th USENIX Sympo-
sium on Operating Systems Design and Implementation (OSDI 20), pp. 1083–
1100. USENIX Association (2020). [Link]
presentation/zhao-zhipeng
23. Šišmiš, L.: Optimization of the Suricata IDS/IPS, Master’s thesis, Faculty of Infor-
mation Technology (FIT), Brno, Czech Republic (2021)
DissecTLS: A Scalable Active Scanner
for TLS Server Configurations,
Capabilities, and TLS Fingerprinting
1 Introduction
Transport Layer Security (TLS) is currently the de facto standard for encrypted
communication on the Internet [18]; thus, providing a good common base to
analyze, compare, and relate servers. The protocol is influenced by libraries,
hardware capabilities, custom configurations, and the application build on top,
resulting in an a server specific TLS configuration. A large amount of meta-
data from this configuration can be collected because in the initial TLS hand-
shake clients and servers must exchange their capabilities such that a mutual
c The Author(s) 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 110–126, 2023.
[Link]
DissecTLS 111
cryptographic base can be found. There are at least two possibilities to collect
this metadata: on the one hand, TLS server debugging tools like [Link] [33]
or SSLyze [10] perform resource intensive scans that dynamically adapt to the
server and can reconstruct a human-readable representation. On the other hand,
active TLS fingerprinting approaches like JARM [4] or Active TLS Stack Fin-
gerprinting (ATSF) [31] use a small set of fixed requests that are designed to be
good in differentiating TLS server configurations. Their light-weight approaches
enable them to be used for Internet-wide scans; e.g., [Link] already provides
JARM fingerprints [8].
Related works have shown that collecting and analyzing TLS configurations
from a large amount of servers enables further use cases, e.g., monitoring a fleet
of application servers [4] or detecting malicious Command and Control (C&C)
servers [4,31]. To be able to collect this data, the respective scanning approach
needs to be efficient, to both reduce the time it takes to collect the data and the
impact the scan has on third parties.
However, using a fixed set of probes will always leave open the possibility for
redundant data to be collected and for useful information to be overlooked; there-
fore, the performance of subsequent applications (e.g., detecting C&C servers)
might not reach their full potential. An alternative is to exhaustively scan a
server until the full TLS configuration can be reconstructed. However, current
tools are not efficient enough to be used on a large scale.
This work investigates whether a dynamically adapting scan can be imple-
mented efficient enough to be used on a large scale and if this provides a benefit
over existing work and tools. We propose DissecTLS as an efficient tool to collect
TLS server configurations and provide the following contributions:
(i) a model of the TLS stack on a server that explains its behavior towards
different requests and that can be used to craft TLS Client Hellos (CHs)
on a per-server level to reconstruct its underlying configuration;
(ii) a comparison of five popular TLS scanners regarding their capabilities to
detect different configurations and their scanning costs performed both in
a controlled testbed environment and on toplist servers;
(iii) a measurement study of one top- and two blocklists over nine weeks com-
paring a C&C server detection using fingerprinting tools and this work,
complimented with an overview of common TLS parameters; and
(iv) published measurement data [29], scanner, and comparison scripts [30].
2 Methodology
During the initial handshake of the TLS protocol, clients and servers share sev-
eral pieces of information related to their capabilities to negotiate a mutual
encryption base. Part of this can be configured by the user (e.g., ciphers the
server is allowed to select), only limited by the actual capabilities of the soft-
ware and hardware. However, TLS servers only react to clients; therefore, reveal
only a portion of their internal configuration with every response (e.g., the server
112 M. Sosnowski et al.
Property Representation
selects only a single cipher from the list of proposed ciphers). This means, mul-
tiple requests (i.e., CHs) must be sent to collect the full amount of information
hidden in the TLS stack. It is not feasible (regarding time and resources) to send
every possible CH to a server. Thus, every active TLS scanner uses a strategy
to select CHs depending on the information it wants to collect. With DissecTLS
we aim to reconstruct the configuration that cause the observed TLS behavior
in a scalable manner that can be used even for Internet-wide scans. Therefore,
we need to reduce the number of requests as far as possible. This is achieved
by defining a general model of the TLS configuration on a server and use the
minimum number of requests necessary to learn the parameters of the model.
Additionally, we defined the output such that it can be used for fingerprinting;
i.e., exclude session, timing, and instance related data. Depending on the previ-
ous responses from a TLS server we use the model to craft the most promising
CH that should reveal new information about the server.
The following sections will explain our model of the TLS stack, how we
represent its features, and how our scanner is implemented on an abstract level.
options and the server selects one according to its internal preferences. Itera-
tively removing each parameter from new requests that was previously selected
by the server, the full list of length n can be scanned with n + 1 requests. This
is the optimal approach using the “lowest number of connections necessary [. . . ]
for one host”, explained by Mayer et al. [20]. However, if the server prefers client
preferences, only a set of supported parameters can be acquired instead of a
priority list. Clients can inform the server about their own priorities through the
order of parameters in the CH. We tested whether servers respect this priority
as follows: after learning at least two parameters, we also learned which one
the server selected first. Then, we send a new CH where the order of the two
is reversed; we know a server prefers its own preferences if this had no influ-
ence on the selection. We scan cipher suites, supported groups, and Application
Layer Protocol Negotiations (ALPNs) with the currently 350, 64, and 27 possi-
ble values listed by IANA [17], respectively. Some servers provide the full list of
supported groups [27] directly as extension, in these cases we do not explicitly
scan them. However, the presence of a pre-computed key share can influence
the priorities of the supported groups; hence, we collect the preference without
a pre-computed key share and afterwards test whether the presence influenced
the decision. Support of most TLS extensions is indicated by their presence or
absence and does not need a particular logic, they just need to be triggered
in the CH with their presence. Others need specific logic because they modify
the encryption (Encrypt Then Mac and Extended Master Secret), are mutually
exclusive with other extensions (Record Size Limit and Max Fragment Size), or
multiple values can be send (Heartbeat). Sometimes, the content of extensions
is of interest because it reveals information about the server capabilities, and in
these cases we store the raw byte content. The man-in-the-middle inappropriate
fallback protection needs special logic because it only makes sense to send the
signaling cipher [21] if multiple TLS versions are detected. Lastly, servers can
respond differently in cases of problems, some report an error on the Transmis-
sion Control Protocol (TCP) layer, some send TLS alerts, and others just ignore
the problematic part of the handshake (e.g., using a default value). An example
is shown in Appendix A.
In summary, this model is an abstract and human-readable representation of
the TLS stack on a server that can explain its behavior in TLS handshakes.
Fig. 1. Example for merging multiple observations of TLS extensions into a single
format. If the graph contains cycles after merging, the extension order is inconsistent.
internal order on the server as close as possible. We created a DAG for each
observation, merged these graphs and removed duplicate and transitive edges.
If the graph contains cycles after merging; this means, the observations were
inconsistent and the extension order cannot be reconstructed. An example for
this process is illustrated in Fig. 1.
In conclusion, a DAG allows to represent multiple observations of extensions
in a compact format that is as close as possible to the internal server order.
Table 2. Detected number of Nginx configurations for each test case. An ideal scanner
detects every alteration made on the server and finds “Goal” number of configurations.
Test Case DissecTLS DissecTLS (lim.) ATSF JARM SSLyze [Link] Goal
TLS Versions 15 15 13 11 15 14 15
Cipher Suites 1 956 359 115 11 63 1 956 1 956
ALPNs 2 2 2 2 1 2 2
Preferences 2 2 2 2 1 2 2
Session Tickets 2 2 2 2 2 2 2
Used CHs per Server
Minimum 8.0 8.0 9.0 9.0 423.0 9.0
Average 14.3 10.0 10.0 10.0 450.1 132.7
Maximum 42.0 15.0 12.0 10.0 455.0 224.0
ent TLS server configurations. Without analyzing the scanner output we argue
that whenever a scanner is able to differentiate two different TLS configurations,
the scanner has detected the relevant piece of information. The more configu-
rations it can differentiate, the more valuable is its output. However, we also
measured the costs of the scanner by counting the amount of requests it needed
to perform the scan. The lower the costs are, the more servers can be scanned
in the same time and the lower the impact is on individual servers. An ideal
scalable approach collects a high amount of information with low costs.
We compared [Link] [33], SSLyze [10], JARM [4], ATSF [31], and this
work. We selected them because from our knowledge they are the relevant rep-
resentatives that either fingerprint or reconstruct the TLS configuration. We
configured our approach in two versions, one tries to fully reconstruct the TLS
configuration (DissecTLS), the other completes using 10 handshakes (DissecTLS
lim.). We interpret the textual output of each scanner as its representation of
the server. If two outputs are equal, they detected no difference in the configu-
ration. We were able to directly use the output of JARM, ATSF, and this work.
We had to remove information regarding timing (e.g., scan time), sessions (e.g.,
cryptographic keys), and server instances (e.g., the domain name) from the out-
put of [Link] and SSLyze to get stable results for the same TLS configuration.
Additionally, we disabled the vulnerability detection of these tools.
In our local testbed we compared the TLS scanners based on a ground truth. We
challenged them in different scenarios where we systematically made alterations
to a server and checked whether the scanners were able to detect it.
The experiment was designed as follows: we selected a parameter we could
configure on the TLS server (Test Case), launched an Nginx 1.23 docker con-
tainer for each configuration we could generated for this parameter, and scanned
the containers with every scanner. We used tcpdump [32] to measure the num-
ber of CHs the scanners were using. The results can be seen in Table 2. An ideal
116 M. Sosnowski et al.
Table 4. Overview of the collected data for the Top- and Blocklist study.
does not mean a scanner collected only a super-set from another, as discussed
in Sect. 6. This work uses just a sixth of the requests compared to [Link],
with 24 CHs on average. JARM used less than 10 requests on average because
sometimes the TCP connection failed and no CH was sent. In contrast to the
last section, the limited version of DissecTLS performed a bit worse than ATSF.
Apparently, our approach only detects the finer details that help to differentiate
TLS configurations when it completes the scan.
This sections showed that the dynamic scanning approach from [Link],
DissecTLS, and SSLyze is superior to the fixed selection of CH regarding col-
lected data. However, this comes with increased scanning costs. We argue that
only JARM, ATSF, and DissecTLS are resource efficient enough to be used for
large-scale measurements. Additionally, in the following we refrain from limiting
the number of requests of DissecTLS. While roughly doubling the scanning costs
it provides a more complete; hence, a more useful view on the TLS stack.
This section transfers the findings from the previous section to a larger scale
where we collected more than 15 Million data samples with each scanner (data
available under Ref. [29]). However, we only used DissecTLS, ATSF, and JARM
for this study because [Link] and SSLyze did not scale well enough for this
use case.
We scanned servers from the complete Tranco [19] toplist and two C&C
server blocklists: the [Link] Feodo Tracker [1] and the [Link] SSLBL [2].
We collected nine weekly snapshots starting from July 01, 2022. Table 4 presents
an overview of these measurements. We resolved domains from the toplist and
118 M. Sosnowski et al.
Fig. 2. Precision and Recall for classifying C&C servers each week based on the data
collected in previous weeks. Using the fingerprints from the respective scanner as input.
scanned each combination of IPv4 and IPv6 address together with the domain as
Server Name Indication. We call each IP and domain name combination a target.
The two blocklists only list IP addresses; hence, the number of targets is equal
to the number of entries on these lists. We count a “success” if the respective
scanner produced an output. For DissecTLS and ATSF this is the case if a
TCP connection could be established. JARM additionally needed the server to
respond at least once with a Server Hello. DissecTLS and JARM implement a
retry mechanism on failed TCP handshakes, together with the different success
definition, this can explain the variations in the success rates.
Until this section this work analyzed TLS configurations as a single unit; how-
ever, DissecTLS produces an output (see example in Appendix A) that can be
used to understand how a server is configured. This can help to explain why fin-
gerprinting was possible. In Appendix B we present statistics from the top- and
blocklist servers that can deepen our understanding of TLS parameter usages on
the Internet. We have analyzed the support for different TLS versions; computed
a popularity ranking of cipher suites, supported groups, and ALPNs; analyzed
whether servers prefer client preferences or not; and looked how many servers
supported deprecated cipher categories.
In conclusion, an exhaustive TLS scanning approach can be used for finger-
printing but additionally provides valuable insights into the TLS ecosystem.
5 Related Work
Fingerprinting TLS clients in passive network traces is a well established disci-
pline, shown by multiple related works [3,5–7,16]. This concept has been adapted
by Althouse et al. [4] and Sosnowski et al. [31] through active scanning to be
able to fingerprint servers. Both approaches use a fixed set of 10 requests that
“have been specially crafted to pull out unique responses in TLS servers” [4] and
“empirically optimized to provide as much information as possible” [31], respec-
tively. They capture variations of the TLS configuration in their fingerprints;
however, they do not actively search for them; additionally, the explainability
of their output, or fingerprint, is low and it is difficult to understand what has
caused the specific fingerprint. Both works show that they can find malicious
C&C servers on the Internet. A fundamentally different approach is proposed by
Rasoamanana et al. [26], they define a State Machine describing TLS handshakes
and argue that the transitions between states can be used to fingerprint specific
120 M. Sosnowski et al.
implementations; especially, if these transitions are not conform to the TLS spec-
ification and, sometimes, even pose a security vulnerability. Their focus on the
behavior of the library in the context of erroneous input does not consider the
parameters that are the cause of the non-erroneous behavior. Dynamically scan-
ning TLS servers is a common practice in the context of analyzing and debugging
servers with tools like [Link] [33] or SSLyze [10]. Both make assumptions how
the TLS on the server works and adapt their scanning to this model. However,
they focus on the configurable part of the server, do not export every finger-
printable information, and are not optimized for Internet-wide usage (e.g., use
more than 100 requests to scan a single server). Mayer et al. [20] showed that
cipher suite scanning can be optimized to use 6% of the connections compared
to related works. However, they ignore the rest of the TLS configuration.
6 Discussion
This work proposes an exhaustive but optimized TLS scanning approach that
can be used for large-scale Internet measurements and for TLS fingerprinting.
The following paragraphs discuss several aspects we found worth mentioning.
C&C TLS Configurations. In general, configurations we could relate to C&C
servers had just slight alterations in their parameters compared to common con-
figurations (e.g., the position of a single cipher). However, we collected interesting
results (see Appendix A) from several servers labeled as Trickbot, according to
the Feodo Tracker [1]. These servers supported TLS 1.0 and downgraded higher
versions, which is already a rare behavior. In contrast to the low TLS version,
the ciphers were strong and some used a modern key agreement, i.a., X25519
(standardized 2016 [14] - 8 years after TLS 1.2 [28]). This led us to the conclusion
that this was a modern server where some modern features were disabled.
Completeness of the Testbed. Every TLS scanner from Sect. 3.1 was capable to
detect more configurations than the ones we have tested, e.g., TLS versions prior
to TLS 1.0 or other cipher suites. We selected the tested values because they
were configurable on the Nginx server. Some features, e.g., the extension order,
cannot be configured. Our choice of the six ciphers was arbitrary and it is possible
that there are combinations of ciphers where the performance of the scanners is
different. However, our tool sends 350 different ciphers and the analysis shows
that it can effectively identify permutations of those on the server.
Completeness of the TLS Server Model. Sections. 3.2 and 3.1 showed that Dis-
secTLS and [Link] were able to detect the most TLS configurations. How-
ever, looking into their output, no scanner provided a super-set of the other;
hence, our proposed model cannot be complete. We manually investigated cases
where [Link] was able to differentiate configurations while DissecTLS was
not, and vice versa. Both scanners rely on consistent server responses; however,
Sosnowski et al. [31] reported inconsistent behaviors for 1% of their fingerprinted
targets. If servers behave inconsistently, both scanners might have collected an
incomplete view of the TLS stack and reported different configurations on each
DissecTLS 121
7 Conclusion
This work proposes a scalable active scanning approach to reconstruct the TLS
configuration on servers. The approach is compared with four active TLS scan-
ners and fingerprinting tools. While we are able to collect a comparable amount
of information to single server TLS debugging tools, we also keep up with the per-
formance of scalable active TLS fingerprinting tools using around twice the num-
ber of requests. Our approach collects more data than the fingerprinting tools
and produces human-readable representations of a TLS configuration, improv-
ing the explainability of the approach. We performed a nine week measurement
study of top- and blocklists, analyzed common TLS parameter usages, and fin-
gerprinted potentially malicious C&C servers. Similar to related work, the fin-
gerprinting achieved a precision of more than 99% for the most conservative
detection threshold of 100%; however, at the same time DissecTLS achieved a
recall 2.8 times higher than the related ATSF [31]. This was achieved by a scan
that dynamically adapts based on a TLS stack model and previously learned
information. The model was used to explain server responses and to craft new
122 M. Sosnowski et al.
requests that should reveal new data. This paper shows that an exhaustive TLS
parameter scanner can be implemented efficiently enough to be used on a large
scale. Moreover, it can replace existing active TLS fingerprinting approaches
because it provides a similar fingerprinting performance but additionally pro-
duces a valuable dataset. In the future, it can help to acquire a global view on
the TLS parameter usage to deepen our understanding of the TLS ecosystem.
Property Value
supported TLS versions TLS 1.0 support
TLS 1.1 downgrade
TLS 1.2 downgrade
TLS 1.3 downgrade
cipher suites (priority list) TLS_ECDHE_RSA_WITH_AES_256_CBC_SHA
TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA
TLS_RSA_WITH_AES_256_CBC_SHA
TLS_RSA_WITH_AES_128_CBC_SHA
TLS_RSA_WITH_CAMELLIA_256_CBC_SHA
TLS_RSA_WITH_CAMELLIA_128_CBC_SHA
cipher Preference server
supported groups (priority list) x25519
secp256r1
group preference server
with key share client
ALPN (set) http/1.1
ALPN preference unknown
extension data EC Point Format → [uncompressed
ansiX962_compressed_prime,
ansiX962_compressed_char2]
order of TLS extensions (DAG) renegotiation_info → max_fragment_length →
ec_point_formats → session_ticket → ALPN →
encrypt_then_mac → extended_master_secret
version error behavior Ignore
cipher error behavior TLS Alert
groups error behavior Ignore
ALPN arror behavior Ignore
DissecTLS 123
Table 6. Support for different TLS versions from successfully scanned targets.
Support for the TLS versions can be seen in Table 6. Although TLS 1.0
and 1.1 is deprecated since 2021 [22], we saw a high amount servers supporting
it. Some servers even downgraded the handshake by responding with a lower
version than the one we requested. This was expected for TLS 1.3 because the
TLS 1.3 CHs is basically a TLS 1.2 CH with special extensions. A server that
does not understand these extensions should continue with a TLS 1.2 handshake.
However, we rarely observed this behavior also for other versions.
We collected cipher suites, supported groups, and ALPNs as priority lists.
This enables combining them to get the overall most popular values as shown
in Table 7 (full list available under [30]). This problem is similar to a voting
problem where multiple individuals can vote with a list of descending preference
and can be solved with scoring rules as discussed by Fraenkel et al. [13]. We
decided to use the Dowdall rule, which favors parameters with top preferences.
This way parameters of a low priority, usually only kept for backward compli-
ance, are given a low score. The ranking worked as follows: from each priority
list the parameters [p1 , . . . , pn ] are scored with [1, 12 , . . . , n1 ], the scores for each
parameter are summed up, and ranks based on the highest scores are computed.
We analyze the parameters independent of the TLS version; hence, the TLS
1.3 ciphers are ranked above the others because, in general, higher versions are
preferred over ciphers.
Some servers selected the cipher suites, supported groups or ALPNs based
on the preference of the client. This leaves security decisions open to the client
but can be beneficial to the user if the client has limited hardware capabilities.
However, in our measurements we saw this was rarely the case and most servers
preferred their own priorities as shown in Table 8. This is different if a client
already pre-computed a TLS 1.3 key share for one of the supported groups;
then, 29% of the servers used the key share to avoid an additional round trip.
An important security feature on servers is the support against version down-
grade attacks. If this is not given, even when security issues are fixed in a newer
TLS version, a downgrade can reopen these attack vectors. Such a downgrade
could be achieved, e.g., by a man-in-the-middle attacker blocking connections
124 M. Sosnowski et al.
Table 7. Most preferred TLS parameters (IANA names [17]) ranked separately with
the Dowdall rule and total distinct values. Each scanned target was used as vote.
for a higher TLS version expecting the client will attempt to reconnect with a
lower version. Table 8 shows most servers were protected.
Several servers still support categories of deprecated ciphers [9,25] as shown
in Table 9. These ciphers are known to be insecure; however, they are not per-se
a security vulnerability because an attacker would still need to force a client and
server to agree on them.
Table 9. Servers supporting at least one deprecated cipher suite per category. Per-
centages are in relation to the successfully scanned targets.
References
1. [Link]: Feodo Tracker. [Link] Accessed 28 Oct 28
(2022)
2. [Link]: SSL Certificate Blacklist. [Link] Accessed 28 Oct 2022
3. Althouse, J., Atkinson, J., Atkins, J.: TLS Fingerprinting with JA3 and
JA3S (2019). [Link]
ja3s-247362855967
4. Althouse, J., Smart, A., Nunnally Jr., R., Brady, M.: Easily identify malicious
servers on the internet with JARM (2020). [Link]
easily-identify-malicious-servers-on-the-internet-with-jarm-e095edac525a
5. Anderson, B., McGrew, D.: OS fingerprinting: new techniques and a study of infor-
mation gain and obfuscation. In: 2017 IEEE Conference on Communications and
Network Security (CNS) (2017). [Link]
6. Anderson, B., McGrew, D., Kendler, A.: Classifying Encrypted Traffic With TLS-
Aware Telemetry. FloCon (2016)
7. Anderson, B., McGrew, D.A.: Accurate TLS fingerprinting using destination con-
text and knowledge bases. CoRR (2020). [Link]
01939
8. Censys: JARM in Censys Search 2.0 (2022). [Link]
articles/4409122252692-JARM-in-Censys-Search-2-0. Accessed 14 Oct 2022
9. Dierks, T., Rescorla, E.: The Transport Layer Security (TLS) Protocol Version 1.1.
RFC 4346 (2006). [Link]
10. Diquet, A.: SSLyze. [Link] Accessed 13 Oct 2022
11. Dittrich, D., Kenneally, E., et al.: The Menlo Report: Ethical principles guiding
information and communication technology research. US Department of Homeland
Security (2012)
12. Durumeric, Z., Wustrow, E., Halderman, J.A.: ZMap: fast internet-wide scanning
and its security applications. In: Proceedings of the USENIX Security Symposium
(2013)
13. Fraenkel, J., Grofman, B.: The Borda Count and its real-world alternatives: com-
paring scoring rules in Nauru and Slovenia. Aust. J. Pol. Sci. (2014). [Link]
org/10.1080/10361146.2014.900530
14. Friedl, S., Popov, A., Langley, A., Emile, S.: Transport Layer Security (TLS)
Application-Layer Protocol Negotiation Extension. RFC 7301 (2014). [Link]
org/10.17487/RFC7301
15. Gasser, O., Sosnowski, M., Sattler, P., Zirngibl, J.: Goscanner (2022). https://
[Link]/tumi8/goscanner
16. Husák, M., Cermák, M., Jirsík, T., Celeda, P.: Network-based HTTPS client iden-
tification using SSL/TLS fingerprinting. In: 2015 10th International Conference on
Availability, Reliability and Security (2015). [Link]
35
17. IANA: Transport Layer Security (TLS) Parameters. [Link]
assignments/tls-parameters/[Link]. Accessed 13 Oct 2022
18. Labovitz, C.: Internet traffic 2009–2019. In: Proceedings of the Asia Pacific
Regional Internet Conference on Operational Technologies (2019)
19. Le Pochat, V., Van Goethem, T., Tajalizadehkhoob, S., Korczyński, M., Joosen,
W.: Tranco: a research-oriented top sites ranking hardened against manipulation.
In: Proceedings of the 26th Annual Network and Distributed System Security Sym-
posium (2019). [Link]
126 M. Sosnowski et al.
20. Mayer, W., Schmiedecker, M.: Turning active TLS scanning to eleven. In: De Cap-
itani di Vimercati, S., Martinelli, F. (eds.) SEC 2017. IAICT, vol. 502, pp. 3–16.
Springer, Cham (2017). [Link]
21. Moeller, B., Langley, A.: TLS Fallback Signaling Cipher Suite Value (SCSV) for
Preventing Protocol Downgrade Attacks. RFC 7507 (2015). [Link]
17487/RFC7507
22. Moriarty, K., Farrell, S.: Deprecating TLS 1.0 and TLS 1.1. RFC 8996 (2021).
[Link]
23. Mozilla: SSL configuration generator (2022). [Link]
Accessed 13 Oct 2022
24. Partridge, C., Allman, M.: Addressing ethical considerations in network measure-
ment papers. In: Proceedings of the 2015 ACM SIGCOMM Workshop on Ethics
in Networked Systems Research. Association for Computing Machinery (2016).
[Link]
25. Popov, A.: Prohibiting RC4 Cipher Suite. RFC 7507 (2015). [Link]
17487/RFC7465
26. Rasoamanana, A.T., Levillain, O., Debar, H.: Towards a systematic and auto-
matic use of state machine inference to uncover security flaws and fingerprint TLS
stacks. In: Computer Security - ESORICS (2022). [Link]
031-17143-7_31
27. Rescorla, E.: The Transport Layer Security (TLS) Protocol Version 1.3. RFC 8446
(2018). [Link]
28. Rescorla, E., Dierks, T.: The Transport Layer Security (TLS) Protocol Version 1.2.
RFC 5246 (2008). [Link]
29. Sosnowski, M., Zirngibl, J., Sattler, P., Carle, G.: DissecTLS Measurement Data.
[Link]
30. Sosnowski, M., Zirngibl, J., Sattler, P., Carle, G.: DissecTLS: Additional Material
(2023). [Link]
31. Sosnowski, M., et al.: Active TLS stack fingerprinting: characterizing TLS server
deployments at scale. In: Proceedings of the Network Traffic Measurement and
Analysis Conference (TMA) (2022)
32. The Tcpdump Group: tcpdump. [Link] Accessed 27 Oct 2022
33. Wetter, D.: Testing TLS/SSL encryption. [Link] Accessed 27 Oct 2022
Open Access This chapter is licensed under the terms of the Creative Commons
Attribution 4.0 International License ([Link]
which permits use, sharing, adaptation, distribution and reproduction in any medium
or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons license and indicate if changes were
made.
The images or other third party material in this chapter are included in the
chapter’s Creative Commons license, unless indicated otherwise in a credit line to the
material. If material is not included in the chapter’s Creative Commons license and
your intended use is not permitted by statutory regulation or exceeds the permitted
use, you will need to obtain permission directly from the copyright holder.
Applications
A Measurement-Derived Functional
Model for the Interaction Between
Congestion Control and QoE in Video
Conferencing
1 Introduction
Fig. 1. High-level architecture of a video conference with N clients. Video traffic orig-
inating at Client 1 shown: (i) peer-to-peer connection between two clients (ii) in SFU
mode, the SFU will forward the video from Client 1 to all other clients.
videos from each client to other clients without decoding [44]. SFUs can manip-
ulate videos prior to forwarding using techniques such as subsampling and layer
selection which can be applied to the encoded video. A high-level overview of a
video conferencing session is given in Fig. 1.
VCAs can be bandwidth intensive because of their extensive use of video.
Commercial systems carry the video on top of UDP, with congestion control
implemented in the application. Recognizing the important role of congestion
control, there is a long history of effort dedicated to building application layer
congestion control on top of UDP (e.g., [4,11,15]) with more recent efforts geared
towards specific use in VCAs [13]. Congestion control mechanisms present in
commercial VCAs are typically not shared publicly, though some open-source,
production-quality VCAs do exist [3,25].
The main goal of congestion control is to determine an application sending
rate that does not congest the shared network path used by the application.
After congestion control functions determine an acceptable target sending rate,
the application must enforce this target rate through video rate control. The
application can adjust the rate of the video being sent using video encoding at
the sending client or through layer selection and/or subsampling at the SFU1 .
Congestion control in VCAs, therefore, has a direct impact on video quality
and, in turn, the VCA user’s quality of experience (QoE). Understanding the
user-perceived QoE is important for network operators as it helps with efficient
bandwidth provisioning, though in this paper we show that congestion control
and video rate control can differ between VCAs, potentially complicating the
matter of estimating QoE. Researchers have reported on the throughput and
video metric performance of specific VCAs using structured experiments. Yet
prior work rarely examines the interaction between congestion control mecha-
nisms and rate adjustment techniques that produces the observed throughput
and QoE metrics. Understanding this interaction at a functional level paves the
way to explain observed performance, to pinpoint commonalities and key func-
tional differences across different VCAs, and to contemplate opportunities for
innovation. This is the aim of this paper.
1
In principle, the total client sending rate includes video, audio, and control data
(e.g., RTCP packets). In this paper, we focus only on video data as it is the most
bandwidth intensive.
Measurement-Derived Functional Model for Video Conferencing 131
surement results, along with data from publicly available documents and source
code. Our models reveal additional details and complexity in these systems and
demonstrate how, despite some uniformity in function deployment, there is sig-
nificant variability among the VCAs in the implementation of these functions.
We believe our more detailed models can better serve the research community
as we continue to investigate the design and performance of VCAs.
The experiments presented in this work were highly controlled and did not
involve any real users. There was no personal or private information sent or
received as part of any experiment. This work raises no ethical concerns.
2 Related Work
The surge in video conferencing use in recent years allow us to split prior work
into two groups; those from before this surge [1,3,16,24,42,43], and those from
after it [9,27,28,30,34]. In general, previous studies cover VCAs that were popu-
lar at the time of publication. For example, Xu et al. [42] and Yu et al. [43] from
2012 and 2014 cover Skype and FaceTime, whereas more recent work studies
VCAs like Zoom [9,27,28,30,34], Microsoft Teams [27,30], and WebEx [9,28].
“Legacy” VCAs, such as Skype, are typically not subject to examination in
more recent work, reflecting their decline in use in favor of newer products such
as Zoom and Microsoft Teams. Notably, interest in WebRTC remains consistent
in both groups of prior work [1,3,24,28,30].
VCA measurements typically involve subjecting various VCAs to different
network conditions and recording the results. These results are often presented
as measurements of throughput and video metrics, typically the video resolu-
tion and frame rate. A recent study aimed to infer models for congestion con-
trol [28] but without exploring the fully functional dependence between conges-
tion control and video QoE. In typical studies, measurements are often taken
in a highly controlled laboratory environment, though Fund et al. [16] perform
their measurement campaign in two different outdoors settings, one urban and
one suburban/rural, with the user devices connected to WiMAX base stations
in the vicinity. Varvello et al. [39] evaluate the performance of Zoom, Webex and
Google Meet using clients distributed around the world.
The VCAs studied and measurements taken in this paper result are most
directly related to work by MacMillan et al. [30]. In that work, the authors
study Zoom and WebRTC-based VCAs and consider measurements of the video
metrics without considering the underlying architectures. However, in our work
we focus on measurement for the purpose of building a generalized understand-
ing of critical parts of the VCA, rather than to measure performance of specific
VCAs. Sander et al. [34] focus on improving Zoom’s performance in the pres-
ence of competing flows by implementing different queuing policies at the bot-
tleneck, and discuss Zoom’s insensitivity towards packet loss and queuing delay.
We observe similar results for packet loss, but show that Zoom will respond to
Measurement-Derived Functional Model for Video Conferencing 133
3 Measurement Design
In this section we describe the measurement design. We first provide a breakdown
of the possible test conditions and describe those relevant to our goal of building
the generalized models. Second, we describe the experimental testbed that will
enable measurement under the defined set of test conditions.
Table 1. Summary of the considered test conditions and the values selected for each.
of the congestion or video rate control, we collect packet traces and video metrics
available in the VCA statistics or in conference recordings.
Two laptop clients are used for the majority of measurements, as this is
sufficient to understand the congestion control and video rate control on the
client side. To fully understand the SFU, we take measurements with additional
clients to examine how the SFU works when multiple videos are to be forwarded
to a single client. The background traffic on each client is kept to a minimum to
avoid competition with other flows. Measurements with competing TCP flows
are presented in Appendix A, but they do not help directly inform the functional
models. Lastly, as it is known that device type can impact video conferencing
performance and QoE [9,39,40], we use phone and tablet clients alongside the
laptops in SFU mode to understand their effect.
3.2 Testbed
The testbed used for measurements reflects the high-level architecture in Fig. 1
and must support two clients, with support for more in select experiments.
Furthermore, the testbed supports video conferencing in both peer-to-peer and
SFU-mediated modes, as the two are important for fully understanding the VCA
client and SFU, respectively. For measurements taken in peer-to-peer mode, only
Clients 1 and 2 and the Signalling Server, typically located within the network
owned by the VCA developer [31], are involved. The signalling server is only
used to establish the peer-to-peer video call and is uninvolved after this. The
SFU measurements make use of Clients 1 and 2 as well as the VCA SFU for
forwarding videos. Some measurements will also include additional clients.
While we desire minimal background traffic on the devices themselves, this is
not necessarily a realistic environment in practice. Therefore, to provide a semi-
realistic environment that we can model the VCAs within, the measurements
were taken with the two laptops connected to an uncontrolled WiFi access point
(AP). This access point was found to deliver a minimum of 10 Mbps upload and
download to each device at all times, which is sufficient for all of the considered
VCAs to send and receive video at maximum bitrate. Furthermore, packet loss
Measurement-Derived Functional Model for Video Conferencing 135
at the AP is minimal, and the latency is stable even during peak periods, such as
evenings and weekends. These properties ensure that all effects under the various
applied network conditions can be reliably observed, which is further guaranteed
using repeated measurements.
We are not able to control the traffic condition at the Zoom or BlueJeans
SFU, though we assume in our measurements that the SFU experiences minimal
congestion. We believe this to be a fair assumption, given the large cloud-based
networks that Zoom and BlueJeans control. We find that the clients typically
connect to Zoom SFUs located in or around New York City, as reported by
the Zoom application. The precise location of the BlueJeans SFU could not be
ascertained, though we suspect it is the eastern USA from traceroutes to the
SFU IP address. The Jitsi SFU is set up on a Google Cloud VM located in the
us-east1-b datacenter region in South Carolina. As the Jitsi SFU runs on a
VM which we control, we can ensure background traffic is kept to a minimum.
Fig. 3. Architecture of Client 1. Client 2 only uses the VCA and virtual camera input.
Both Clients 1 and 2 run Mac OS with native Zoom and BlueJeans applica-
tions (version 5.10.4 and 2.35.0, respectively), and Chrome version 100, which
is used for Jitsi and WebRTC. We note that VCAs are subject to frequent soft-
ware updates, which may bring slightly different behavior to the VCAs over time.
Given the relatively short time period over which measurements were taken, we
do not anticipate such effects to impact the accuracy of the measurements or
derived models. Different operating systems can also influence the behavior of
native VCA apps [30,39]. In addition, we observed over the course of our mea-
surements that the administrator-controlled settings for the enterprise VCAs,
particularly Zoom, could impact how the VCAs can be used. For example, over
Summer 2022, we noticed that peer-to-peer mode calls could not be established
on the Zoom accounts provided by our institution, but accounts from a different
institution would permit peer-to-peer calls under the same conditions.
All measurements are taken on Client 1, which has additional software run-
ning as shown in Fig. 3 to support automated measurement and data processing.
We apply the network conditions on Client 1; we use Network Link Conditioner
(NLC), available as part of the Xcode developer tools on Mac OS. NLC allows
the bandwidth, packet loss rate, and latency to be changed on upload and down-
load independently. Both clients use a virtual camera to display a talking head
136 J. He et al.
video, a still from which is shown in Fig. 3. This video was chosen for its typical-
ity and realistic resolution of 1280 × 720. The VCAs are operated in full-screen
mode during the two-client measurements. The measurement process was fully
automated using the Selenium library [32] and the Chrome WebDriver [10] for
WebRTC/Jitsi, and using the PyAutoGUI library [38] for Zoom and BlueJeans.
The automation code is made open source on GitHub [20].
Client 1 performs two types of data logging. The first is collecting packet
traces using Tshark. We note that packets are collected after NLC operates in
the uplink (UL) and the downlink (DL); this operating principle is important for
putting the results in later sections into context. Secondly, video statistics data
is logged. This logging varies depending on the VCA; for WebRTC and Jitsi, the
data is collected from the Chrome instance running the VCA using the developer
tools at chrome://webrtc-internals. Zoom and BlueJeans provide statistics
within the application itself; these are extracted via recording the screen area
containing the statistics and processing this screen recording with the Tesseract
text recognition software [17], with appropriate error-checking and correction to
ensure accurate reconstruction of results. Zoom provides statistics on the frame
rate, resolution, and bitrate; BlueJeans only provides the resolution.
Zoom’s built-in local recording feature produces videos that are high-fidelity
representations of the exact video that was sent or received, and thus can be used
to analyze the quantization parameter (QP) as well. The recording is processed
using an FFMpeg-based tool [33] to extract the QP for each frame; the tool also
outputs whether the frame is an I-frame or a P-frame. While BlueJeans supports
cloud-based recordings, they are post-processed to add black borders when the
video degrades from maximum resolution, and thus it is impossible to determine
if the QP is faithful to the encoder’s configuration.
4.1 Measurements
Figure 4 shows throughput plots for Zoom and WebRTC under the bandwidth,
latency, and packet loss conditions. To conserve space, nine out of the 12 network
conditions are shown.
Fig. 4. Response of Zoom and WebRTC peer-to-peer clients to different network condi-
tions. (a)-(c) upload bandwidth limits, (d)-(f) additional upload latency, (g)-(i) upload
loss rates.
We note that there are differences in the sending rate for each VCA even in
the absence of induced network conditions. Zoom’s sending rate varies consider-
ably over time with peaks between 3.5 and 4 Mb/s, while WebRTC has a much
more stable sending rate that rarely exceeds 2.8 Mb/s.
138 J. He et al.
Bandwidth. Removal of the 512 kb/s bandwidth limit in Fig. 4(b) shows that
Zoom takes over 160 s to return to its original sending rate of around 3.7 Mb/s;
WebRTC returns to its original peak sending rate in only around 40 s. Zoom also
exhibits a step-like sending rate increase.
Latency. Unlike WebRTC’s GCC which responds to delay variation, Zoom
responds to the actual delay value, and only does so above a certain thresh-
old, at most 500 ms. The response is severe, with the sending rate dropping to
less than 300 kb/s when 500 ms of latency is added.
Packet Loss. While WebRTC’s GCC uses a loss rate threshold of 10% before
reducing the sending rate, we find a small response at 5% loss rate. This may
be due to NLC’s probabilistic packet drop mechanism causing packet loss rates
above 10% at some points. Figures 4(g) and 4(h) show how GCC’s sending rate
decreases further when the packet loss rate is higher.
We note that the observed drop in Zoom’s sending rate is equal to the loss
rate. Given that the packet trace is recorded after NLC as described in Sect. 3.2
and shown in Fig. 3, we are observing only the packet loss enforced by NLC,
rather than Zoom lowering its sending rate. The lack of the step-like increase
upon removal of the packet loss is further evidence that Zoom’s congestion con-
trol algorithm is insensitive to packet loss, even at a loss rate of 50%.
The ability of Zoom to function in the presence of such extreme loss rates
may be due to forward error correction (FEC) data that is included with the
video stream [29]. Observation of the received video shows intermittent freezing,
but the picture quality and resolution remain high.
Overall, the measurements taken for WebRTC match well with the GCC
algorithm as described in the literature. We learn that Zoom in peer-to-peer
mode exhibits three notable behaviors: (i) a slow, step-like sending rate increase
function with a step multiplicative factor of 1.2, (ii) a severe response to latency
above 500 ms, and (iii) no response to extremely high packet loss rates. We note
that properties (i) and (iii) are highly similar to the TCP BBR algorithm [7],
which has been gaining popularity on the Internet [6]. Though the time between
each bandwidth probe event for Zoom is much longer than for BBR, Zoom’s
probing gain value of 1.2 is very similar to BBR’s 1.25. It seems likely that
Zoom uses a BBR-like congestion control algorithm, which contradicts studies
that find a fit between Zoom’s congestion control and GCC [28].
Zoom’s lack of response to packet loss means it behaves essentially opposite
to loss-based TCP congestion control. Therefore, in the presence of competing
TCP flows, Zoom is likely to starve the TCP flows, as long as the latency does
not increase beyond the threshold at which Zoom lowers its sending rate. This
effect has been noted in previous work [9,28], and we confirmed this behavior
with measurement shown in Appendix Sect. A. Appendix Fig. 15 shows how the
Zoom network flow is unhindered by up to ten concurrent TCP flows started
at the same time, instead managing to increase its sending rate while the TCP
flows are active.
Measurement-Derived Functional Model for Video Conferencing 139
Lastly, Zoom takes a considerably longer time to recover its sending rate
compared to WebRTC, as can be seen most clearly in Fig. 4(a). At face value,
this means the user needs to wait almost 90 s longer for Zoom to reach the same
sending rate as WebRTC after recovery from a drop in network bandwidth.
Fig. 5. Response of Zoom, Jitsi, and BlueJeans clients to network conditions while
in SFU-mediated mode. (a)-(c) upload bandwidth limits, (d)-(f) additional upload
latency, (g)-(i) upload loss rates.
140 J. He et al.
Latency. Zoom has a similar response to the peer-to-peer case when 500 ms
latency is added, including how the sending rate will begin to recover before this
additional latency is removed. This particular effect is further investigated from
the perspective of the video metrics in Sect. 5.2. This may occur as Zoom realizes
the latency has not decreased as a result of reducing the sending rate. Jitsi has
a similar transient response to latency changes as WebRTC has in peer-to-peer
mode. BlueJeans is notable as it has no observable response to latency changes,
except for small transients when the latency is added and removed.
Packet Loss. When the packet loss rate increases, there is an increase in Zoom’s
sending rate, as opposed to remaining unchanged in the peer-to-peer case. This
is possibly because Zoom is adding a larger amount of FEC code for resilience to
packet loss. The ability of Zoom to add FEC was noted previously in [30] and is
explained further in patents held by Zoom [29]; these results show that the Zoom
client itself is also capable of adding FEC. Jitsi starts to show a response to loss
above 5%, which becomes more severe as the loss rate increases. BlueJeans does
not have a significant response to packet loss except for the 50% case; before
this point we are measuring the reduction in effective sending rate due to NLC’s
packet drops, as explained in Sects. 3.2 and 4.2 for Zoom. At 50% packet loss in
Fig. 5(i), there is a pronounced reduction in BlueJeans’ sending rate. This figure
also shows that after packet loss, BlueJeans takes over 30 s longer than Jitsi to
start recovering sending rate.
BlueJeans stands out among the three VCAs in that it has a purely loss-
based congestion control. Zoom and Jitsi have responses to both loss and delay
when the SFU is being used.
SFU Behavior. The SFU also performs congestion control on forwarded videos,
leveraging the SVC or simulcast mechanisms described in Appendix B. By lim-
iting the download bandwidth of Client 1, we can measure the response of the
SFU’s congestion control.
Figure 6 shows the results of measurements taken similarly to the ones in
Sects. 4.2 and 4.3. We can see that the Jitsi SFU uses a similar multiplicative
rate increase as the client uses, and also takes around 40 s to complete the rate
increase. However, with higher download limits, we observe some variance in
the measurement results. This is likely caused by the SFU’s video selection algo-
rithm, which is choosing a different resolution simulcast stream or different frame
rate subsampling for each of the five measurement runs.
Measurement-Derived Functional Model for Video Conferencing 141
The Zoom SFU behaves almost identically to the client, whereas the Blue-
Jeans SFU does not keep a steady sending rate, and takes longer to recover from
a bandwidth constraint than the client. The SFU takes around one minute to
recover to the original sending rate; the client takes around 20 s.
Remarks. The first notable difference is how Zoom in SFU mode uses a signif-
icantly lower sending rate when no network conditions are applied. The reason
for this is likely to reduce congestion at the SFU, though we do observe that
this is an institutional configuration setting, as accounts from a different institu-
tion were capable of the same sending rate in both SFU and peer-to-peer mode.
We report the results for the reduced sending rate case to best illustrate this
markedly different mode of operation. BlueJeans uses a maximum resolution of
1280 × 720 even though it also operates in SFU mode like Zoom. Consequently,
the bandwidth usage is almost double that of Zoom.
We see that whether the video conference is held in peer-to-peer mode or via
the SFU, the Zoom and WebRTC/Jitsi clients have mostly the same congestion
control, aside from the lower sending rate for Zoom in SFU mode. Instead, we
see significant differences between the VCAs; BlueJeans stands out as having
no response to latency increase, and Zoom stands out having little response to
packet loss, with the client increasing its sending rate when in SFU mode.
5.1 Measurements
Unlike the measurements described in Sect. 4.1, we keep a constant network con-
dition applied for the duration of the measurements. By adjusting the available
bandwidth or the latency of this constant condition, we are able to study how
the three video quality metrics are affected: (i) frame rate, (ii) resolution, and
(iii) quantization parameter (QP). These will give insight into how the video
rate control is responding to different target encoding bitrates provided to it by
the congestion control.
Repeated measurements were taken for a range of upload bandwidth limits
between 96 kb/s and 2 Mb/s on Client 1. We also took measurements with upload
latency values between 200 and 1,200 ms for Zoom in particular, as the results
in Sects. 4.2 and 4.3 show that Zoom responds to high, long-term latency. Each
measurement consists of five minutes with the network condition held in steady
state; we take five repeats for each test condition, leading to each plotted point
representing 25 min of data. Each measurement repeat is given two minutes to
142 J. He et al.
stabilize before data is collected. The average and standard deviation for each
video metric is computed over the entire 25 min of data.
We observe that the receiver’s window size impacts the quality of the received
video for Zoom. Therefore, in a call with Client 1 and Client 2, if Client 2 changes
its window size, it can cause Client 1 to send a lower maximum video quality.
This occurs in both peer-to-peer and SFU mode. We ensure the Zoom window
on both clients is of sufficient size to allow for the maximum video quality.
Here, we present the results for measurements with limited upload bandwidth
and different upload latencies.
Fig. 7. Client sent video quality metrics in peer-to-peer mode, when adjusting the
available upload bandwidth.
Fig. 8. Video quality metrics over time with upload bandwidth limited to 256 kb/s.
Measurement-Derived Functional Model for Video Conferencing 143
QP. Note that Zoom and WebRTC use different encoders with different QP
scaling. As a consequence of WebRTC’s preference of maintaining 30 FPS, the
QP can reach values above 90, leading to a significant loss of picture quality.
Zoom reaches a plateau around 27 for I-frames and 25 for P-frames, achieved by
reducing the frame rate as seen in Fig. 7(a).
Figure 8 illustrates how the frame rate, resolution, and QP change over time
with an upload bandwidth of 256 kb/s. The frame rate and QP can change by
small amounts over time, most clearly seen in Zoom’s frame rate and WebRTC’s
QP. However, these small changes will occur around a clear target value. For
example, Fig. 8(a) shows two distinct frame rate levels for Zoom.
Conversely, the resolution for Zoom and WebRTC takes discrete values, and
typically changes on a much longer time scale compared to the frame rate and
QP. This suggests that, given a specific target bitrate, the encoder is able to
minutely adjust the frame rate and QP to meet it as best as possible, but will
settle on one of a small set of available resolutions.
Remarks. We note that Zoom and WebRTC have taken different approaches
on how the encoder responds to low target bitrates. The video metrics measure-
ments in Fig. 7 show that WebRTC considers frame rate to be more important
than QP at low bandwidths, whereas Zoom considers the opposite. Therefore,
these two VCAs behave very differently in low bandwidth regimes, with WebRTC
offering smooth frame rate but extremely low picture quality, and Zoom offering
low frame rate and moderate picture quality.
Figure 8 shows how Zoom’s frame rate and resolution have a periodic charac-
ter at the 256 kb/s upload bandwidth limit. This is also shown by the relatively
high variance at lower sending rates in Fig. 7. Furthermore, the periods of higher
frame rate appear to coincide with the periods of lower resolution, and vice versa.
Periodic changes in quality are not ideal, and this behavior may be caused by a
combination of the low bandwidth limit and a limited set of target resolutions
and frame rates that Zoom chooses from. Specifically, the upload bandwidth of
256 kb/s seems to lie between the minimum bandwidths which support 640 × 360,
26 frame/s and 800 × 450, 20 frame/s targets.
Lastly, results for Zoom under different latency conditions are presented in
Appendix Sect. C. The discrete levels of average sending rate as a function of
latency shown in Fig. 16(d) suggest that the Zoom congestion control has four
modes of sending rate reduction in response to latency: (i) no reduction, (ii) a
limited, temporary reduction, (iii) an increased but still temporary reduction,
and (iv) a severe and long-term reduction. Each of these operation modes has
a clear threshold in terms of the additional upload latency. Furthermore, within
operation mode (iii), the time taken for the sending rate to recover varies as a
function of the additional latency.
Client Behavior. Figure 9 shows the metrics for video sent by the VCA clients
as a function of the available upload bandwidth. Zoom’s SVC stream is shown,
along with Jitsi’s three simulcast streams. Chrome’s built-in WebRTC statistics
page provides metrics for each of Jitsi’s simulcast streams. The BlueJeans appli-
cation provides information regarding the resolution of the current video; it is
likely this represents the resolution of the highest bitrate simulcast stream. The
individual simulcast streams cannot be measured for BlueJeans.
Figures 9(a) and 9(b) show that the Zoom client will send video with a max-
imum resolution of 640 × 360 and frame rate of 26 frames/s when the upload
bandwidth is 512 kb/s or greater. In the peer-to-peer case, the sent video reaches
1280 × 720 and 30 frames/s. Figure 9(c) shows how Zoom uses additional band-
width above 512 kb/s to reduce the QP.
The three simulcast streams for Jitsi each turn on at certain upload band-
widths. The 320 × 180 stream is always on, while the 640 × 360 stream is enabled
at 512 kb/s and the 1280 × 720 stream is enabled at 2 Mb/s. Interestingly,
Fig. 9(a) shows that the video streams require additional bandwidth to reach
30 frames/s beyond the amount needed to enable them. This is a different result
compared to the WebRTC peer-to-peer case, where 30 frames/s was achieved in
all but the lowest bandwidth cases. The QP for the 640 × 360 and 1280 × 720
streams also does not exceed 50, showing that the bandwidth thresholds for
enabling the simulcast streams are chosen to achieve some minimum picture qual-
ity rather than achieve the maximum frame rate. We found that the 320 × 180
and 640 × 360 streams have a maximum sending rate of around 200 kb/s and
700 kb/s respectively, and each successive stream will not turn on until the lower
quality stream is sent at its maximum rate. The resolution of the highest bitrate
BlueJeans stream in Fig. 9(a) shows a step increase, demonstrating how succes-
sive simulcast streams are enabled as bandwidth increases.
Measurement-Derived Functional Model for Video Conferencing 145
SFU Behavior. To understand how the SFU decides what video quality to for-
ward, we take similar measurements to Sect. 5.3 with the download rate limited
on Client 1. Figure 10 shows the video metrics as a function of the applied down-
load limit. Overall, the quality metrics for the forwarded video from Zoom’s SFU
are similar to those for the encoded video from the Zoom client. The main differ-
ence is a slight reduction in frame rate and resolution at lower bandwidths. This
suggests that the Zoom SFU has almost the same flexibility via subsampling as
the client does via encoder parameter selection.
Jitsi achieves a stable 30 frames/s at 512 kb/s download bandwidth as seen
in Fig. 10(a), which is higher than the frame rate achieved by Zoom. However,
at 128 kb/s, the frame rate is almost zero on average, while Zoom will still
be able to forward video. This is likely a consequence of the limited frame rate
subsampling options available to Jitsi’s SFU [26]. The zero frame rate is also what
causes the close-to-zero measured QP in Fig. 10(c). Figure 10(b) shows that the
average resolution does eventually exceed Zoom’s maximum of 640 × 360, which
is expected as Jitsi has a 1280 × 720 simulcast stream. The high variance suggests
that the received video resolution changes frequently over time, which implies
that the Jitsi SFU is switching between simulcast streams very often.
BlueJeans’ resolution is also prone to switching as shown by the variance
in Fig. 10(b), however it is not as prevalent as in Jitsi. The BlueJeans SFU is
able to forward the maximum resolution stream consistently once the download
bandwidth reaches 3 Mb/s.
Remarks. We find Zoom’s SFU behaves very similarly to the Zoom client in
the presence of bandwidth limits, as seen in Figs. 9 and 10. This shows that the
SVC mechanism used by Zoom allows for stream selection at the SFU which
is almost as flexible as the client’s encoder. The same does not hold for the
simulcast systems Jitsi and BlueJeans; for example, the resolution sent by the
BlueJeans client at 1.5 Mb/s is lower than the resolution that can be received at
1.5 Mb/s, likely due to the overhead from sending multiple simulcast streams.
The received video quality metrics measurements in Fig. 10 show that the
Jitsi SFU is unstable even at 3 Mb/s client download rate. The client should be
able to download the maximum bitrate simulcast stream at this point, but the
Jitsi SFU still switches it with a lower bitrate stream. This unintended behavior
manifests as a received video that changes quality often, possibly degrading QoE.
146 J. He et al.
All SFU will be prone to such a problem, especially if the network bandwidth is
close to the threshold between two available bitrate streams.
The experiments in Sects. 4 and 5 involve only two users in the setup described in
Fig. 1. In this section, we consider video conferencing sessions in SFU-mediated
mode with more than two users. These experiments demonstrate certain behav-
iors of the SFU, including how it considers different device types, and how it
chooses which videos to forward to a client with constrained download band-
width. These insights cannot be provided by the two-user experiments.
6.2 Observations
We consider two effects in our experiments. The first is the impact of device type
on the recorded statistics. Secondly, we try to understand how the SFU decides
which videos to send, and at what quality to send those videos.
2
The Zoom packet structure determined by Michel et al. [31] could aid this process.
Measurement-Derived Functional Model for Video Conferencing 147
Fig. 11. Received frame rate and resolution with different device types on Zoom.
Impact of Device Type. Device type may have an impact on VCA perfor-
mance as described in [9,40]. Therefore, we can evaluate whether there are any
differences between the three device types: laptop, tablet, and phone.
Zoom. Figure 11 shows the received frame rate and video resolution for the
three additional devices on a Zoom call with Clients 1 and 2. Client 1 was kept
in gallery mode for this experiment. The video resolution received from each
device type generally follows the same trend, though at the lowest download
bandwidths, it appears that the laptop gets some priority in bandwidth allo-
cation. However, for frame rate, there is a significant difference; even at higher
bandwidths, the phone and tablet will only be received at 15 frames/s maximum.
In single-speaker mode, we found that all devices had the same behavior.
This shows that the phone and tablet are capable of sending at the maximum
26 frames/s that was found in the two-client video call for Zoom. However, the
SFU intervenes to limit the forwarded video from the phone and tablet to 15
frames/s if the client receiving the forwarded video is in gallery mode.
BlueJeans. While the reported video statistic for BlueJeans is difficult to con-
trol, we observed that a 320 × 180 resolution video was received from the phone
client in gallery mode, while all others were 640 × 360. The frame rate, while not
reported in the BlueJeans application, appeared to be the same for all devices.
Jitsi. There was no observed impact of device type for Jitsi.
SFU Decisionmaking. Section 5.3 describes how the SFU adjusts the video
quality metrics of the forwarded video streams in response to download band-
width constraints at the receiving client. If multiple client videos are available
for forwarding, the SFU must also make decisions on which videos to forward
to a receiving client. Combined with the adjustment of per-stream video quality
metrics, this represents the full congestion control at the SFU. Known decision
making processes such as Last-N [19] are based on the measured download band-
width for the receiving client, as well as a priority queue based on when each
client last spoke (produced sound through their microphone).
In all three VCAs, we found that the SFU will only decide which videos to
forward when the receiving client is in gallery mode, displaying multiple other
client videos at once. In single-speaker mode, all SFUs will forward the cur-
rently focused video at maximum possible quality. Typically, they will forward
only thumbnail videos at the lowest possible resolution for the remaining video
participants, which may be displayed as small insets in the VCA interface.
148 J. He et al.
Fig. 12. Jitsi received video quality metrics as a function of download bandwidth when
receiving four client videos (C1–C4) simultaneously.
Zoom. The limited information available within the application and packet
traces mean an observational approach must be taken. Client 1’s download band-
width is limited between 0.5 and 7 Mbps with the other five available devices
connected with video on. Notably, Zoom does not turn off any received videos; at
lower bandwidth, some videos may freeze for an extended time but will remain
visible. At moderate bandwidths, we observed that Zoom will allocate the max-
imum possible quality to the currently focused video, and then share remaining
bandwidth fairly between all others.
BlueJeans. As BlueJeans also provides limited information in the application
and packet trace, the same approach as Zoom is taken. Like Zoom, BlueJeans
does not turn off received videos even at very low bandwidths. However, unlike
Zoom, BlueJeans has a more clear prioritization of videos; the currently focused
video will receive the most bandwidth possible, but then successive videos will
receive a bandwidth share in order of when they were last active in audio. So
the remaining bandwidth after the currently focused video is not shared equally,
it is prioritized similarly to Jitsi’s Last-N algorithm.
Jitsi. Jitsi’s Last-N algorithm is well described [19]. Furthermore, the Jitsi source
code describes how bandwidth is allocated to videos according to the Last-
N prioritization [25,26]. Altogether, the decisionmaking process is very similar
to BlueJeans. However, the chrome://webrtc-internals statistics allow us to
evaluate how Last-N impacts the received video quality metrics.
Figure 12 shows how the video metrics for each of the four received video
streams change as a function of download bandwidth. The Last-N algorithm
is clear in the plots for resolution, QP, and video bitrate; each video stream
increases in quality in turn. However, for the frame rate, all three videos are
being received at the maximum 30 frames/s as soon as 1.5 Mbps download rate
is reached. This matches the general behavior of WebRTC and Jitsi as described
in Sects. 5.3 and 5.3 where the frame rate is maximised with greater priority
than the resolution and QP.
Altogether, the three studied SFU-mediated VCAs have similar behavior
when deciding which videos to forward and at what quality. All use a prioritiza-
tion based on which client was last active in audio; this client becomes the cur-
rently focused video in gallery mode. Zoom will share the remaining bandwidth
fairly among other clients, whereas BlueJeans and Jitsi will allocate bandwidth
to clients prioritized according to audio activity.
Measurement-Derived Functional Model for Video Conferencing 149
Fig. 14. Functional model for the VCA SFU and client.
congestion control as a unified block in the functional model, which uses RTCP
network feedback to make adjustments to the target video encoding bitrate.
Video Rate Control. The video rate control is split into two parts, the encoder
parameter adjuster, and the encoder itself. The encoder parameter adjuster
accepts a target frame rate and resolution from the FPS and resolution selection
and makes adjustments to the targets as well as generates the QP. The objective
of this is to most closely match the target bitrate that is provided by the con-
gestion control. As seen in Fig. 8, the QP and frame rate can be finely adjusted
over time to match the target bitrate as best as possible, but the resolution takes
a set of discrete values and does not change on such a small timescale. There-
fore, as illustrated in Fig. 13, the resolution emerges from the encoder parameter
adjuster unchanged, whereas the final frame rate may be different from the tar-
get, indicated by the asterisk notation.
The encoder uses the resolution, frame rate, and QP provided by the encoder
parameter adjuster to produce the video that will be sent over the network. The
encoder also uses the target bitrate produced from the congestion control to
compute the utilization, which it feeds back to the encoder parameter adjuster.
This forms the first of three closed feedback loops that the encoder drives. The
second feedback loop involves the encoder’s CPU utilization, and the third feed-
back loop involves the frame drop rate and the QP. These two feedback loops
connect back to the system monitors.
Section 5 shows how Zoom and WebRTC/Jitsi share a similar control for the
video resolution as a function of the available bandwidth, but opposite behavior
for frame rate and QP. At low bandwidths, Zoom will drop the frame rate to
maintain a reasonable QP, while WebRTC/Jitsi will maximize the QP to main-
tain a high frame rate. Therefore, while each VCA may employ a different policy
or algorithm to control the video metrics for the video rate control, they all make
use of the general functionality described in Fig. 13.
differences specific to operation in SFU mode. We develop the SFU side of the
model to be as analogous to the client side as possible, and so it is split into the
congestion control and video rate control components, with the associated boxes
colored and bordered as in Fig. 13.
Changes to the Client Model. The client now includes M encoders and
encoder parameter adjusters to support simulcast streams at M different resolu-
tions. Zoom, with its single SVC stream, has M = 1. In our measurements and
examination of the source code, we observe M = 3 for Jitsi. We are not able to
determine specific parameters for BlueJeans. Additionally, the client may also
receive feedback from the SFU providing additional constraints for what resolu-
tion and frame rate to send. These constraints are based on the SFU’s knowl-
edge of how the client’s video is being displayed by other clients. As observed
for Zoom, if all other clients are displaying a particular client’s video in a small
viewport, then that client has no need to send high quality video and the SFU
will provide parameters to constrain the quality of the video being sent.
We note that the measurement results in Sects. 4.3 and 5.3 show that the
congestion control and video rate control mechanisms on the client side are
generally identical between peer-to-peer and SFU modes. Therefore, we retain
the same congestion control and video rate control components as in the peer-
to-peer model.
SFU Congestion Control. Section 4.3 shows that the SFU also performs con-
gestion control when forwarding video to a client. As on the client side, the SFU
must implement congestion control by adjusting the video metrics for the videos
it forwards. Furthermore, the SFU can choose to not forward all videos if there
is insufficient bandwidth to do so. Therefore, the foremost job of the SFU con-
gestion control is to generate an available bandwidth for the video rate control
to use for decision making. This is equivalent to the target encoding bitrate
generated by the congestion control on the client side.
The specific behaviors of the SFU congestion control seen in Sect. 4.3 typically
mirror those of the client, and so we are left with the same conclusion that
generalizing the SFU congestion control is infeasible. With the observation that
each VCA’s SFU will implement its own specific congestion control algorithm,
typically similar to the client’s, we use the same unified block as on the client.
SFU Video Rate Control. The video rate control on the SFU side of the
model can be viewed as a step-by-step procedure. First, the M (N − 1) received
client videos are processed by the prioritizer, which orders the videos according
to a specific criteria. For Jitsi, the Last-N algorithm which prioritizes by most
recent speakers is used [19], and we observed similar behavior for BlueJeans as
described in Sect. 6.2. Zoom uses a similar algorithm, but instead attempts to
share the remaining bandwidth fairly after allocating the most recent/pinned
speaker as much bandwidth as possible. We note the prioritization process can
Measurement-Derived Functional Model for Video Conferencing 153
be done in a centralized manner by the SFU and applied to the video rate control
for all receiving clients.
After video prioritization, the SFU will run resolution selection for each of the
N − 1 client videos that are to be forwarded using the available bandwidth esti-
mated by the congestion control. For VCAs which use simulcast, this will involve
the selection of one out of the M simulcast streams that were received. For VCAs
that use SVC-based encoding, this will involve a subsampling procedure. In any
case, the number of videos which emerge from the resolution selection will be
N − 1. The SFU will then run FPS selection for each of the N − 1 video streams,
again taking the available bandwidth into account. FPS selection can typically
be performed via subsampling on all VCAs, as it is supported by H.264, VP8,
and VP9, the most common encoders used.
The available options for resolution and frame rate are determined by several
factors, including the devices used by the N − 1 sending clients and receiving
client, and the receiving client’s viewing mode. All will be known by the SFU
as a result of signaling operations. The effect of different device types for Zoom
is shown in Fig. 11, and the video metrics for each of four received streams at
a bandwidth-limited client are shown in Fig. 12. This figure clearly shows the
prioritization method, with four distinct traces for each client. We note that the
SFU must perform all of its video rate control on the encoded video streams, as
the decoding would put an unscalable computational load on the SFU.
8 Concluding Remarks
VCAs deploy congestion control and video rate control functionality at both the
client and SFU. In both instances, target rates are determined by congestion
control functions and then used to influence the rate of video transmitted by the
clients and the SFU. The adjustment of the video rate has direct consequences
on the video quality metrics such as frame rate and resolution. Given this base-
line level of understanding, we constructed more detailed functional models for
the VCA client and SFU which are based predominantly off of a directed mea-
surement campaign using a subset of commonly used VCAs. We believe these
models, along with the accompanying measurement results, provide a level of
understanding which was previously unavailable in related literature.
We expect the functional models will serve to inform further research into
video conferencing. We will use the functional models in our future work to relate
the congestion control mechanisms employed by different VCAs to end user
quality of experience, which is an important functionality for network operators
in order to best provision network services to maximize end user QoE.
Fig. 15. Competition with different types of TCP flows. (a-c) Jitsi, (d-f) Zoom in SFU
mode.
As Zoom has been reported to take a majority share of the bandwidth when
competing with TCP [28,30,34], we took measurements with different types
of TCP flows to compare and explain results using the functional models and
measurements in Sect. 4.
VCA Simulcast SVC Subsample FPS Subsample Res. Codec Alternate Codecs
Zoom No Yes Yes Yes H.264 None
Jitsi Yes No Yes No VP8 VP9, H.264
BlueJeans Yes No Yes No VP8 None
The testbed was used in two-person SFU mode with Zoom and Jitsi as in
Sect. 4.3. In all cases, a total bandwidth limit of 4 Mb/s was applied to the
upload of Client 1. After 30 s of measurement, ten TCP flows begin in various
patterns: (i) started at the same time with nine minute duration, (ii) started at
30 s intervals, all finishing at the nine minute mark, and (iii) started and stopped
at the same time every 20 s.
Figure 15 shows the results of these measurements. Figures 15(a), 15(b), and
15(c) all show that Jitsi immediately gives up practically all of its bandwidth in
the presence of TCP. In particular, Fig. 15(b) shows that this occurs with only a
single TCP flow sharing the link at the 30 s mark. This behavior is likely caused
by GCC’s sensitivity to changes in delay; once GCC gives up some bandwidth,
TCP takes it, compounding the response. Conversely, Figs. 15(d), 15(e), and
Measurement-Derived Functional Model for Video Conferencing 155
15(f) show that Zoom sending rate actually increases in TCP’s presence. This
is likely a consequence of how it decides to add FEC as described in Sect. 4.3,
the general insensitivity that Zoom has to packet loss, and the high threshold
for responding to latency.
As mentioned in Sect. 1, the SFU is able to adjust the quality of forwarded video
streams via subsampling. The exact method used for this differs between the
considered VCAs, as described below, and summarized in Table 2.
Zoom uses a custom implementation of H.264 Scalable Video Coding (SVC) [2,
23,36] to encode base and enhancement video layers. The SFU can then add/drop
layers before forwarding to clients. Because H.264 SVC is used, subsampling of
both the resolution and the frame rate is available to the SFU in Zoom as an
orthogonal technique to adjust video bit rate [21,37,44,45].
Jitsi and BlueJeans use video simulcast, in which the VCA client sends multiple
independent video streams at different resolutions. Jitsi uses simulcast when the
VP8 encoder is used. In particular, the Jitsi clients make use of three simulcast
streams, with resolution scaling factors of 1, 2, and 4. At 1280 × 720 native video
resolution, this means the two other streams will be 640 × 360 and 320 × 180. Blue-
Jeans also uses the VP8 encoder, implying use of simulcast, though there are no
means of measuring the individual video streams. In a VCA with simulcast, the
SFU chooses which one of the received video streams to forward, which determines
the resolution. Additionally, it may choose to adjust the frame rate without re-
encoding. Note that resolution changes without re-encoding are not possible with
the VP8 codec [5,41], demanding the use of simulcast.
Fig. 16. Client sent video quality metrics in peer-to-peer mode, when adjusting the
upload latency. The colors indicate the distinct modes of operation.
156 J. He et al.
Figure 16 presents the video metrics for Zoom as a function of the additional
latency on the upload. Overall, the sent video metrics begin to see an impact at
400 ms additional upload latency, and beyond 600 ms additional latency, there
are no further impacts.
Frame Rate and Resolution. We group the consideration of both of these
metrics as they share a very similar response to the additional upload latency.
The most notable feature is the evolution of high variance in the measurement
results for added latencies above 500 ms. We observe that much of the measured
variance is not due to fluctuation in the frame rate or resolution; instead it is
caused by Zoom returning to the maximum resolution and frame rate some time
after the experiment starts. The time at which this occurs is a function of the
latency; the higher the latency, the longer Zoom takes to begin recovering its
sending rate. At higher latencies beyond 600 ms, the recorded variance is low.
This is either because the time taken for Zoom to begin recovering is longer
than the measurement duration, or because Zoom keeps the low sending rate
indefinitely if the latency is beyond this value.
Measured Video Sending Rate. The measured video sending rate corre-
sponds well to the frame rate and resolution trends. Specifically, there appear
to be two intermediate levels of video sending rate; the first averaging around
2.25 Mb/s between 400 and 500 ms added latency, and the second averaging
around 1.5 Mb/s between 500 and 600 ms. Beyond 600 ms, a flat 100 kb/s send-
ing rate is used, which corresponds to what is observed in Fig. 4(f).
QP. The method for obtaining Zoom’s QP outlined in Sect. 3.2 is incompatible
with latency measurements. This is because starting a screen recording in the
Zoom application prompts a sudden, very large spike in latency. Once this spike
dissipates, the Zoom application begins sending at the maximum rate, even when
the steady-state latency is above 1,000 ms. We are uncertain if this is intended
behavior to maintain maximum quality for recording purposes, or if this is a bug.
In any case, we are unable to ascertain the Zoom QP for these measurements.
References
1. Amirante, A., Castaldi, T., Miniero, L., Romano, S.P.: Performance analysis of the
Janus WebRTC gateway. In: Proceedings of ACM Workshop on All-Web Real-Time
Systems (2015)
2. Amon, P., Li, H., Hutter, A., Renzi, D., Battista, S.: Scalable video coding and
transcoding. In: Proceedings of IEEE International Conference on Automation,
Quality and Testing, Robotics (2008)
3. André, E., Le Breton, N., Lemesle, A., Roux, L., Gouaillard, A.: Comparative study
of WebRTC open source SFUs for video conferencing. In: Principles, Systems and
Applications of IP Telecommunications (IPTComm) (2018)
Measurement-Derived Functional Model for Video Conferencing 157
4. Arif, A.S.M., Hassan, S., Ghazali, O., Nor, S.A.: The relationship of TFRC con-
gestion control to video rate control optimization. In: Proceedings of IEEE Inter-
national Conference on Network Applications, Protocols and Services (2010)
5. BlueJeans: Technical Specifications for Services (2021). [Link]
com/s/article/BlueJeans-Technical-Specifications
6. Cao, Y., Jain, A., Sharma, K., Balasubramanian, A., Gandhi, A.: When to use and
when not to use BBR: an empirical analysis and evaluation study. In: Proceedings
of ACM Internet Measurement Conference (2019)
7. Cardwell, N., Cheng, Y., Gunn, C.S., Yeganeh, S.H., Jacobson, V.: BBR:
congestion-based congestion control. ACM Queue 14, 20–53 (2016)
8. Carlucci, G., de Cicco, L., Holmer, S., Mascolo, S.: Analysis and design of the google
congestion control for web real-time communication (WebRTC). In: Proceedings
of ACM International Conference on Multimedia Systems (2016)
9. Chang, H., Varvello, M., Hao, F., Mukherjee, S.: Can you see me now? A mea-
surement study of Zoom, Webex, and Meet. In: Proceedings of ACM Internet
Measurement Conference (2021)
10. Chromium: ChromeDriver - WebDriver for Chrome (2022). [Link]
[Link]/home
11. Chundong, S., Chaojun, L., Shaohua, L.: Research on congestion control algorithms
for real-time audio and video stream. In: Proceedings of IEEE International Con-
ference on Computer and Communications (2018)
12. Cisco: Cisco Ultra Traffic Optimization (CUTO) Data Sheet (2021). https://
[Link]/c/en/us/products/collateral/wireless/ultra-traffic-optimization/
[Link]
13. De Cicco, L., Carlucci, G., Mascolo, S.: Congestion control for WebRTC: standard-
ization status and open issues. IEEE Commun. Stand. Mag. 1(2), 22–27 (2017)
14. Federal Communications Commission: FCC Proposes Higher Speed Goals for Small
Rural Broadband Providers (2022). [Link]
higher-speed-goals-small-rural-broadband-providers-0
15. Floyd, S., Handley, M., Padhye, J., Widmer, J.: Equation-based congestion con-
trol for unicast applications. In: Proceedings of ACM Conference on Applications,
Technologies, Architectures, and Protocols for Computer Communication (2000)
16. Fund, F., Wang, C., Liu, Y., Korakis, T., Zink, M., Panwar, S.S.: Performance
of DASH and WebRTC video services for mobile users. In: Proceedings of IEEE
Packet Video Workshop, pp. 1–8 (2013).[Link]
17. Google: Tesseract Open Source OCR Engine (2022). [Link]
io/
18. Google Git: WebRTC Native Code Package (2022). [Link]
com/src/
19. Grozev, B., Marinov, L., Singh, V., Ivov, E.: Last N: relevance-based selectivity for
forwarding video in multimedia conferences. In: Proceedings of ACM Workshop on
Network and Operating Systems Support for Digital Audio and Video (2015)
20. He, J., Ammar, M., Zegura, E.: Automation Code (2023). [Link]
jh4001/PAM2023 Automation
21. High Scalability: A Short On How Zoom Works (2020). [Link]
com/blog/2020/5/14/[Link]
22. Holmer, S., Lundin, H., Carlucci, G., de Cicco, L., Mascolo, S.: A Google Conges-
tion Control Algorithm for Real-Time Communication (2017). [Link]
[Link]/doc/html/draft-ietf-rmcat-gcc-02
158 J. He et al.
23. International Telecommunications Union: H.264: Advanced video coding for generic
audiovisual services (2016). [Link] [Link]?lang=e&
id=T-REC-H.264-201602-S!!PDF-E&type=items
24. Jansen, B., Goodwin, T., Gupta, V., Kuipers, F., Zussman, G.: Performance eval-
uation of WebRTC-based video conferencing. ACM SIGMETRICS Perform. Eval.
Rev. 45(3), 56–68 (2018)
25. Jitsi: Github repository (2022). [Link]
26. Jitsi: Selective Forwarding Unit implementation of the Jitsi Videobridge (2022).
[Link]
estimations
27. Kumar, R., Nagpal, D., Naik, V., Chakraborty, D.: Comparison of popular video
conferencing apps using client-side measurements on different backhaul networks.
In: Proceedings of ACM Symposium on Theory, Algorithmic Foundations, and
Protocol Design for Mobile Networks and Mobile Computing (2022)
28. Lee, I., Lee, J., Lee, K., Grunwald, D., Ha, S.: Demystifying commercial video
conferencing applications. In: Proceedings of ACM International Conference on
Multimedia (2021)
29. Liu, Q., Jia, Z., Jin, K., Wu, J., Zhang, H.: Error resilience for interactive real-time
multimedia application. U.S. Patent #10348454 (2019)
30. MacMillan, K., Mangla, T., Saxon, J., Feamster, N.: Measuring the performance
and network utilization of popular video conferencing applications. In: Proceedings
of ACM Internet Measurement Conference (2021)
31. Michel, O., Sengupta, S., Kim, H., Netravali, R., Rexford, J.: Enabling passive
measurement of zoom performance in production networks. In: Proceedings of
ACM Internet Measurement Conference (2022)
32. Muthukadan, B.: Selenium with Python (2022). [Link]
[Link]/
33. Robitza, W., Goring, S., Lebreton, P., Trevivian, N.: ffmpeg debug qp (2022).
[Link]
34. Sander, C., Kunze, I., Wehrle, K., Rüth, J.: Video conferencing and flow-rate fair-
ness: a first look at Zoom and the impact of flow-queuing AQM. In: Proceedings
of Passive and Active Measurement, pp. 3–19 (2021)
35. Sandvine: 2022 Global Internet Phenomena Report (2022)
36. Schwarz, H., Marpe, D., Wiegand, T.: Overview of the scalable video coding exten-
sion of the H.264/AVC standard. IEEE Trans. Circuits Syst. Video Technol. 17(9),
1103–1120 (2007)
37. Streaming Media: Q&A: Zoom CTO Brendan Ittelson (2021). [Link]
[Link]/Articles/Editorial/Featured-Articles/QA-Zoom-CTO-
[Link]
38. Sweigart, A.: PyAutoGUI - GitHub (2022). [Link]
pyautogui
39. Varvello, M., Chang, H., Zaki, Y.: Performance characterization of videoconferenc-
ing in the wild. In: Proceedings of ACM Internet Measurement Conference (2022)
40. Vucic, D., Skorin-Kapov, L.: QoE Assessment of mobile multiparty audiovisual
telemeetings. IEEE Access 8, 107669–107684 (2020)
41. WebRTC Glossary: Temporal Scalability (2022). [Link]
temporal-scalability/
42. Xu, Y., Yu, C., Li, J., Liu, Y.: Video telephony for end-consumers: measurement
study of Google+, IChat, and Skype. In: Proceedings of ACM Internet Measure-
ment Conference (2012)
Measurement-Derived Functional Model for Video Conferencing 159
43. Yu, C., Xu, Y., Liu, B., Liu, Y.: Can you SEE me now? A measurement study of
mobile video calls. In: Proceedings of IEEE Conference on Computer Communi-
cations (2014)
44. Zoom: Zoom: Architected for Reliability (2019). [Link]
doc/Zoom Global [Link]
45. Zoom: Here’s How Zoom Provides Industry-Leading Video Capacity (2022).
[Link]
capacity/
Effects of Political Bias and Reliability
on Temporal User Engagement with News
Articles Shared on Facebook
1 Introduction
Despite 74% of all Americans believing that the propagation of online misinfor-
mation is a big problem [9], a very large fraction of users today obtains their
news via social media [16]. In this environment, news articles are often prop-
agated based on other users’ interactions with the news (e.g., through likes,
comments, and sharing of posts linked to various news articles). Indeed, users’
interactions (and their engagement) with different news are becoming the big
driver for which news are most likely to be viewed by others, and hence also
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 160–187, 2023.
[Link]
Temporal Effects of Political Bias and Reliability on User Engagement 161
which news are given the best chance to impact other users’ views of the world,
including their opinions and thoughts on various current issues.
With increasing (political) polarization [11] and news articles often having
vastly different reliability levels, it is therefore important to measure and under-
stand whether there are fundamental disparities in the users’ interaction dynam-
ics with news articles that have different levels of reliability and political bias.
In this paper, we provide a rigorous temporal analysis in which we identify cases
of statistically significant disparity in the user interaction dynamics with differ-
ent classes of news articles. Our findings provide insights into how and when to
better protect against and/or slow down the spread of misinformation.
use a limited news article classification (e.g., binary), focuses on a limited set of
user interactions, and ignores the users’ engagement dynamics over time.
Main Contribution: This paper addresses the above shortcomings of the cur-
rent literature by presenting the first temporal analysis of the user interaction
dynamics with news articles of varying degrees of (political) bias and reliabil-
ity. We consider a spectrum of user interactions and study the impact of bias
and reliability in combination. In contrast to prior works studying the interaction
dynamics as part of the political conversations (in online social networks) during
elections and other events [7,8,18], our focus here is instead on the roles that the
bias and reliability play in the dynamics. Another novel aspect of our temporal
analysis is that we compare the temporal dynamics seen using different classes of
interactions with the news, including likes and shares of posts linking the news
articles. Only a few works have considered all types of user interactions (e.g.,
Edelson et al. [4]) but none of them consider the relative dynamics or the impact
of bias and reliability on the dynamics. Finally, we examine the predictability of
the total amount of user engagement that news articles of different classes may
receive based on the interactions it has received thus far.
RQ1 How do the bias and reliability of a post affect the temporal dynamics
of a user’s engagement with it?
RQ2 Using its intermediate interactions as a predictive criterion, how does the
bias and reliability of a post affect the prediction of the total engagement
it will receive?
RQ3 How do the temporal dynamics of user engagement differ across different
interaction types, and how does this variation relate to the bias and the
reliability of the post?
dynamics of the bias or reliability classes they belong to. In terms of interac-
tion changes over time, the “most reliable” posts and the “most unreliable” posts
exhibit opposite trends. Here, the “most reliable” news is experiencing a faster
decrease (than average) in the interaction rates, whereas the “most unreliable”
news experience a faster than average increase in the interaction rates, as seen
over time.
We find that when examining just the number of likes that a post receives
within the first hour of publication, the reliability of the post is positively associ-
ated with the normalized (over the total number of interactions) number of likes
received. In other words, during this period of time, the posts that are considered
“most reliable” receive the highest number of likes. Finally, when considering the
outlet-specific analysis, we find that despite Fox News and the New York Times
having different political biases, in both cases, relatively unbiased posts receive
greater interaction rates during the initial stages compared to their biased posts.
Bias class Far left Skews left Balanced Bias Skews right Far right
Bias range [–42, –18] (–18, –6] (–6, 6) [6, 18) [18, 42]
There exist several independent evaluation efforts to asses the bias and/or relia-
bility of individual news articles and/or news sources. Examples include Media
Bias Fact Check2 , Ad Fontes Media3 , AllSides4 , and NewsGuard5 . Of these, we
selected to use data from Ad Fontes Media for the following primary reasons: (1)
they evaluate individual news articles, (2) each evaluated article is scored with
regard to both bias and reliability, (3) the dataset contains over 30K articles cov-
ering over more than 1,500 sources, and finally (4) they provide a transparent
strategy, published and explained in a white paper [14].
For each news source, Ad Fontes Media selects sample articles that are promi-
nently featured on each source’s website over multiple news cycles. To prioritize
popular news sources, they rank the news sources and organize them into tiers
that are given different sample frequencies. Specifically, they label approximately
15 articles per month for the top-15 sources, seven articles per month are labeled
for the next 15 sources, the rest of the top-200 sources are assigned approximately
five new labeled articles per quarter, and the following 200 articles (ranks 201–
400) are updated approximately five times per six months. As mentioned in their
white paper [14], they attempt to strike a balance between rating new sources
and updating current ones with more recent samples. As a result, the dataset
consists of news articles spanning both a broad range of news sources and cap-
turing many samples from popular new sources seen over time.
Each article in the Ad Fontes Media dataset is evaluated with regard to both
bias and reliability by at least three human analysts with a balance of right, left,
and center self-reported political perspectives. The bias scores reported by Ad
Fontes Media range from –42 to +42, with greater negative values indicating a
2
[Link]
3
[Link]
4
[Link]
5
[Link]
Temporal Effects of Political Bias and Reliability on User Engagement 165
more leftward bias and positive values leaning toward the right party. For the
reliability scores, they use grades from 0 to 64, with 64 being the most reliable
news. Note that 42 and 64 (not usual numbers used for scales) are arbitrarily
selected by Ad Fontes Media, as described in [15]. For the analysis presented
here, we binned the bias scores into belonging to one of five bias ranges and
we binned the reliability scores into four reliability ranges. Ranges and assigned
labels are provided in Tables 1 and 2, respectively. Due to the smaller sample
size of extremely biased news articles (both to the left and right) we used larger
bin sizes for articles labeled as “Far left” [–42, –18] or “Far right” [18, 42].
We next used the CrowdTangle API to collect (1) all their Facebook posts includ-
ing one of the news article URLs, as well as (2) the temporal data of users’
interaction with these posts. The CrowdTangle platform [6], which is owned by
Facebook, indexes the posts and engagement data for around 7 million pages,
including “more than 50K likes pages, all public Facebook groups with 95k+
members, all US-based public groups with 2k+ members, and all verified pro-
files” [6], as well as any pages added to a CrowdTangle list by those with access
to it. For collecting the Facebook posts, we opted to use the “/Links” endpoint of
the CrowdTangle API. This ensured us that all shortened versions of the URLs
were also collected. To collect the maximum number of posts related to each
article, we passed the canonical form of the URL to this end point. In addition,
we strived to account for instances in which query strings were included in the
URLs’ canonical form.
The data collection was done on or after Sept. 1 (2022) for all posts pub-
lished before Sept. 1. By including only articles published before Aug. 2, our
methodology ensures that at least 4 weeks had passed since the publication date
of any articles included in our dataset. Since most posts sharing news articles
occur soon after an article is posted, the 4-week gap (between the collection
of articles and posts) allows us to collect (the 21 days) temporal interaction
data for all posts associated with the studied news articles. Similarly, the 4-week
threshold also ensures that we can catch most of the posts linking an article. In
this study, we removed any articles that did not have any published posts. After
this filtering, the dataset included 21,872 labeled articles for which we extract
the temporal interaction data.
Using CrowdTangle, we compile temporal user interaction data for the num-
ber of likes, shares, comments, and emoji-based interactions such as Like(s),
Wow(s), Sad(s), Angry(s), Love(s), and Haha(s). For each of the above metrics,
as well as for the total interactions (across all actions allowed by users), Crowd-
Tangle breaks the first (approximately) 21 days after the post is published into
74 roughly exponentially increasing time steps and provides the number of user
interactions for each of the user interactions at each of these time steps. The
increasingly sparse sample rate used by CrowdTangle is most likely motivated
by most posts being short-lived and the interaction rates quickly reducing over
time. We illustrate this in Fig. 1, where we show the cumulative fraction of all
interactions that have taken place after some time since the posting time of each
studied post in our dataset (with time on log scale).
For most posts, we have temporal data for the full 21-day period (the max-
imum age at the final data point for any posts observed in our dataset was 23
days). In addition to this temporal data, we also extract other post-related data
from CrowdTangle, including the date that the post was published.
Here, it should be noted that a user sharing a post essentially pushes the
post to the timeline of their friends and followers, and their statistics do not
include the shares of a post (on the original shares of a post). For comments
Temporal Effects of Political Bias and Reliability on User Engagement 167
statistics, the API counts all comments on the post and all first-level replies to
those comments.
For studying the temporal dynamics of the posts, we break up the 21-day time
period into smaller time buckets and then study the dynamics of user interactions
over each of these time-bucket sequences. For the analysis presented here, we used
four time buckets and selected the time thresholds used to define the bucket sizes
so that each bucket had roughly the same total number of interactions. More
specifically, we selected the time thresholds so that they represent the points
where 25%, 50%, and 75% of the overall interactions (sum of over all interaction
types) have been observed by CrowdTangle (and apply linear interpolation when
thresholds fall between sample points). The determined threshold values are
shown and highlighted (using red lines) in Fig. 1. As expected, the decreasing
interaction rates, result in increasing time bucket sizes.
While we observe approximately straight-line behavior for part of the param-
eter range, we note that the above selection process does not require any assump-
tions about the actual probability distribution. This selection also helps provide
fair head-to-head comparisons (using statistical tests) between the interaction
differences observed during the four different stages, effectively maximizing the
information gains from comparing the interaction dynamics of the users across
the four phases.
We next use the bias-reliability labels of the articles associated with each post
to compute statistics for each bias-reliability pair and time bucket. For most of
our analysis presented here, we report the mean values observed for each time
bucket and interaction type, as well as perform statistical tests on the relative
mean values.
number of posts that shared them, the eleven articles in the most unreliable-left
and the reliable-right classes were the most successful, as shown in this figure.
Moreover, we observe that, among all classes of reliability, the two extreme classes
are shared the most.
We also provide summary statistics for the total number of interactions (irre-
spective of interaction type), calculated as the sum of all interactions.7 As shown
in Fig. 2d, 81.9 million interactions have been recorded for the posts included
in our final dataset. As expected, the interactions are correlated with the num-
ber of posts. To determine objectively which class performs better in terms of
interactions per post, we present the normalized number of interactions (over
the number of posts) in Fig. 2e. As is noticeable, the right party (both “far right”
and “right”) receives more interactions. Regarding reliability, however, it is shown
that the “most unreliable” news are the least engaging for users. We next present
the results and analysis of the temporal sequences.
3 Results
3.1 High-level Analysis of the All Interactions Dynamics
Let us first consider the cases when all interactions are aggregated into one
interaction metric, calculated as the sum over all interaction types. Figure 3
shows the temporal interaction dynamics of this metric in terms of TICR. Here,
we again show the five categories of the bias and an “All” category (that combines
all observations regardless of bias) as rows in each sub-plot and show the four
categories of the reliability plus an aggregate “All” category (that combines all
observations regardless of reliability) as columns. The four sub-plots, going from
left to right, show the results for the time buckets containing all sample points
(as described in Sects. 2.3 and 2.5) associated with the following time buckets:
(1) 0 to 1 h and 17 min, (2) 1 h and 17 min to 5 h and 16 min, (3) 5 h and 16 min
to 17 h and 28 min, and (4) 17 h and 28 min until the end of the timeline of
each post we study (typically 21 days). We use a timeline with green markers to
illustrate this bucketization. As expected from the definition of TICR (Sect. 2.5)
and our selection of time bucket thresholds (Sect. 2.4), the TICR value for the
“all news” case (i.e., the right-top-most cell) of each bucket is 25%.
In each bucket, the mean TICR value for all posts belonging to the respective
class and bucket is depicted (using heatmap colors). To capture the variances
of each class and thereby quantify the reliability of the mean reported for each
group, the coefficient of variation of the mean (cvmean )(i.e., standard error of
7
For example, if a post receives 6 likes, 2 comments, 3 shares, and no other interac-
tions, the value of the total interactions for this post is 11.
Temporal Effects of Political Bias and Reliability on User Engagement 171
Fig. 3. Temporal dynamics of the total interactions (: coefficient of variation of the
mean is smaller than 4%, and : has deviation from the previous time bucket with
p-value < 0.05).
the class divided by its mean) of that class in percentage is computed. Then,
we mark the class with an asterisk () if cvmean is smaller than a threshold.
In the following, we decided on the value of 4% as the threshold. Accordingly,
the classes for which the cvmean is higher than 4% (due to not having enough
samples or having high variances) do not receive the asterisk.
Regarding selecting 4% as the cvmean threshold used for the above statistical
tests, we first note that this value is small. For example, for the general popula-
tion, which has a mean of 25, this threshold is equal to a standard error of 1.0.
The use of such a small threshold allows the comparison of all classes to be made
in a more reliable way. Furthermore, we have found that with this selection, any
two classes with “asterisks” within the same time bucket whose TICR values are
at least 0.2 units apart (from each other) always have statistically significantly
different means at the 90% confidence level. This finding has been validated for
all category pairs and time buckets using t-tests for comparing the means of
these classes, and the p-values are always less than 0.18 .
To come to the above thresholds, we performed pairwise comparisons between
all the classes for each bucket using different example thresholds. For each case,
this corresponds to calculating a 30 × 30 table of pairwise tests, in which each cell
include the p-value (capturing the statistical significance of the pairwise mean
comparisons). Clearly, showing this table – even for a single bucket (and example
threshold) – takes a lot of space. For this reason, we instead simply report the
determined thresholds (in our case 4% and 0.2 point difference) and mark the
classes that satisfied the 4% criteria with an “asterisks”. As an example, if we
turn our attention to the first bucket, we note that the “most reliable” group’s
8
Here, we use independent samples t-test when the classes are independent and depen-
dent samples t-test when the classes are not independent. Examples when the depen-
dent test is used, include cases when a class is compared to its parent bias or parent
reliability class (that it belongs to); e.g., comparison between the “right-unreliable”
class and the “right” (over all biases).
172 A. Mohammadinodooshan and N. Carlsson
TICR mean is higher than that of the “unreliable” class (with more than a 0.2
difference) and that both classes are marked with an “asterisks”. Therefore, we
can say that these two classes have statistically different means.
More than comparing the interaction levels of classes within the same time
bucket, it is also interesting to capture the changes in the interaction level of
one class between the time intervals. To cover this aspect, we annotated the
cells of the figure with an arrow for any class in a bucket for which the difference
between its mean in this bucket and its mean in the previous bucket is statisti-
cally significant at the 95% confidence level (i.e., the p-value of the paired t-test
is smaller than 0.05). Here, the direction of the arrow indicates whether this
variation is increasing () or decreasing ( ). As an example, we note that the
“most reliable” class receives these temporal significance indicators between the
first two buckets. This class (which was outperforming the other classes in terms
of receiving user interactions in the first bucket) hence performs more similar
to the other classes in the second bucket. For this group, the decrease pattern
between the second and the third bucket is also significant, although the change
is not as high. This is mainly due to this class having many samples (23,416
posts) and therefore more easily passing the t-test. The decreasing pattern of
this class also continues in the last bucket, but with a sharper slope.
We make several other observations from Fig. 3. As an example, there is a pos-
itive correlation between the reliability level of news and the level of interaction
they receive in the first bucket. Here, more reliable posts receive interactions at
a higher rate during the first hour after posting. In contrast, for the bias param-
eter, the two extremely biased classes (i.e., “far right” and “far left”) receive less
interaction rate than the unbiased (balanced) class in this bucket. In the final
bucket, the pattern is reversed, suggesting that unbiased postings are more suc-
cessful in the early stages of their lifespan compared to strongly biased posts.
Notably, even in the fourth bucket, unbiased postings receive higher interaction
volume because their total interaction (the denominator of the TICR values) is
significantly more than those of the other two extremely biased classes (37M vs.
7M and 3M interactions). This is one of the reasons why we chose to provide
the temporal dynamics of the TICR values as opposed to the actual interaction
values, as the TICR values capture these dynamics more precisely.
Another important observation we want to highlight is that for some classes,
identifying the class that a sample trend belongs to is easier to profile when we
look at the bias-reliability class not the reliability or bias class independently.
As an example, consider the “most unreliable - left” class which receives statis-
tical significance, and “asterisk” in the third bucket. Both of the means of the
bias and reliability classes it belongs to is statistically different than this class.
This observation suggests that the interaction level is best captured when bias
and reliability are evaluated jointly. Another observation worth mentioning is
regarding the temporal changes of the most unreliable news. This class has an
increasing pattern in terms of the rate of change they experience (for all buck-
ets, it is statistically significant). We have seen the exact opposite trend when it
comes to the “most reliable” news sources.
Temporal Effects of Political Bias and Reliability on User Engagement 173
Key Observations: The “most reliable” posts and the “most unreliable”
posts experience opposite trends in the interaction changes over time. In
the first hour following the publication of the post, there is a positive
correlation between the reliability of the post and the level of interaction
it receives.
We now turn our attention to each individual interaction types, including shares,
likes, and comments. In order to capture how much of the total interactions is
covered by each of the interactions in each time bucket, we use the same denom-
inator as total interactions for each of these interaction types. As an example, if
a post receives a maximum of 600 total interactions during our timeline of study
and receives 90 likes during the first time bucket, we say that the TICR of likes
for this period is 15%.
Shares: With sharing having perhaps the most direct effect on what news
people may be exposed to, we start our analysis with shares. The TICR scores
for the number of shares (of posts linking articles) are shown in Fig. 4. First,
note that the average TICR value for all posts (the top-right cell in each table)
has decreased from 25% to approximately 4% in each bucket (due to normalizing
over the total interactions). Comparing the top-right cells in Figs. 4-6, we note
that the fraction of shares is almost the same as the number of comments but
smaller than the number of likes.
Second, while it should not be expected that the top-right cell of all buckets
to have equal values when considering individual interaction classes, we note
that they are almost the same (in the range of 4-4.5%). Since we picked the
bucket thresholds to have roughly the same number of total interactions over all
posts to each bucket (but not necessarily the same volume for each interaction
type), this suggests that shares as an aggregate (over all classes of news articles)
represents a relatively stable fraction of the total number of interactions.
Third, across all buckets, the “most unreliable” class is the clear winner.
When compared to the other classes of reliability (first row) and even all classes
of bias (last column), they receive a greater proportion of the shares during all
buckets. Although this pattern could not occur for the “total interactions”, it is
feasible here since TICR values here are normalized over the total interactions.
Referring back to the “total interactions”, for which this class saw the lowest ratio
of interactions in the first two buckets(when compared with the other reliability
classes), we, therefore, expect the number of likes (Fig. 5) and comments (Fig. 6)
to be comparatively less (than for the other reliability classes) for these two
time buckets. This shows that the “most unreliable” news often is relatively
more shared early, despite not seeing as many likes and comments, but that this
evens out over time.
174 A. Mohammadinodooshan and N. Carlsson
Fig. 4. Temporal dynamics of the total interactions covered for shares (: coefficient
of variation of the mean is smaller than 4%, and : has deviation from the previous
time bucket with p-value < 0.05).
Fig. 5. Temporal dynamics of the total interactions covered for likes (: coefficient of
variation of the mean is smaller than 4%, and : has deviation from the previous
time bucket with p-value < 0.05).
Fig. 6. Temporal dynamics of the total interactions covered for comments (: coefficient
of variation of the mean is smaller than 4%, and : has deviation from the previous
time bucket with p-value < 0.05).
Temporal Effects of Political Bias and Reliability on User Engagement 175
Fourth, some classes consistently (throughout the four time buckets) see a
larger relative sharing fraction than the other classes. For example, the two
extreme bias classes (“far right” and “far left”) perform better than the remaining
bias classes in all buckets. Moreover, for both of these extreme bias classes as
well as for the “most unreliable” group, the trend of shares rate is increasing
as time goes on. As a result of this trend, the last bucket exhibits a negative
correlation between the total number of interactions covered for shares and the
reliability of news, with the “most unreliable” news seeing relatively more late
sharing.
Fifth, we observe several classes with relatively different temporal dynamics
than the bias and reliability classes they belong to. As an example, we can clearly
observe that for the third bucket, the “unreliable-right” class has a markedly
different pattern than both the corresponding bias and reliability classes that it
belongs to. This suggests that it is important to consider both these parameters
in combination when predicting the sharing of the news in a bucket.
Key Observations: Among all the reliability classes, the “most unreli-
able” posts experience the greatest gains in terms of share rates. During
the late stages of the posts’ lifetime (17 hours after publishing), there is a
negative correlation between reliability levels and share rates. The most
reliable postings receive the least normalized number of shares.
Likes: A more passive way to (indirectly) impact how visible posts on Facebook
is to like various posts. One reason for this is that posts with many likes are more
likely to occur higher up in the timelines of friends. A like also represents a user’s
(in most cases positive) interaction with the news. Figure 5 shows the temporal
dynamics of the total interactions covered for the number of likes. First, again
it is evident that considering bias and reliability simultaneously will yield more
reliable results. As an example, in the second bucket, the “unreliable-balanced”
class deviates from the bias and reliability classes it belongs to.
Second, in the first bucket, a positive correlation is observed between the reli-
ability level and the rate of likes in the early stages of the posts’ lifetime (initial
hour). In other words, during the initial time period, people more frequently like
reliable news. This is in contrast to the share rates (Fig. 4), which happens more
for unreliable posts during the very first stages of the posts’ lifetime.
These observations may suggest that the sharing patterns and like patterns
are substantially different and depend on the reliability and bias of the news.
Yet, some similarities between their patterns can also be observed. For example,
if we consider “all” posts, both metrics observe an increase in the third bucket.
176 A. Mohammadinodooshan and N. Carlsson
Other Interaction Types: While Facebook also allows other interaction types,
these typically see smaller interaction volumes and, therefore may have a less
clear impact on the dissemination patterns of news. We include results for some
of the other used interactions in Appendix A.2.
Reuters). For the classification of the outlets, we used Ad Fontes Media outlet-
based ratings.9 Table 4 lists these sources and high-level statistics extracted from
our dataset, including their bias class (from Ad Fontes ratings), the number of
articles in our data, the number of posts sharing these articles, the number of
interactions related to these posts, the number of posts per article, the number
of interactions per post, and their popularity in terms of their monthly visits.
Table 4. Statistics of the outlets (†: M stands for million, ‡: Website monthly visits
reported by [Link] (Oct. 2022)).
We next present temporal analysis results for the total interactions of arti-
cles published by The New York Times (left-biased), Fox News (right-biased),
and NPR (unbiased). Results for the other three outlets are found in Appendix
A.3. Furthermore, using the code we publish, interested researchers can conduct
similar analyses for the remaining media outlets we studied, although not all of
the results will be statistically significant.
The New York Times: Figure 7 shows the results for The New York Times.
We note that the white boxes represent categories of news for which we did not
have data. As perhaps expected, for The New York Times, we did not have data
for any of the right-biased categories (irrespective of reliability).
First, note that we are reporting the TICR statistics for the total interactions.
As discussed previously, given the selection of bucket thresholds, we anticipate
around 25% of TICR of the total interactions for all buckets when considering the
overall population (reported in the top-right cell of each bucket). However, when
considering individual publishers, this is not necessarily the case. For example,
as seen in Fig. 7, the user engagement with posts linking news articles by The
New York Times that are older than 17 h is lower than average. Instead, posts
linking their news appear most successful during the third bucket (5–17 h after
posts first appear). Second, a definite association between interactions with The
New York Times news related posts receive and their reliability can also be seen
when we focus on the early stages of postings (first two buckets) and late stages
9
Ad Fontes Media provides evaluations of both publishers and individual articles.
178 A. Mohammadinodooshan and N. Carlsson
Fig. 7. Temporal dynamics of total interactions for The New York Times (white boxes:
no data available, : coefficient of variation of the mean is smaller than 4%, and : has
a deviation from the previous time bucket with p-value < 0.05).
Fig. 8. Temporal dynamics of the total interactions covered for Fox News (white boxes:
no data available, : coefficient of variation of the mean is smaller than 4%, and : has
a deviation from the previous time bucket with p-value < 0.05).
Fig. 9. Temporal dynamics of the total interactions covered for NPR (white boxes: no
data available, : coefficient of variation of the mean is smaller than 4%, and : has
a deviation from the previous time bucket with p-value < 0.05).
Temporal Effects of Political Bias and Reliability on User Engagement 179
(17 h onward) and exclude the most unreliable news (which has only six articles
in our dataset). Here, the first stages’ correlation is positive, whereas for the late
stage this correlation is highly negative. While aiming to receive early engage-
ments, it is clear that the reliability of the news plays a significant factor in the
actual interaction levels achieved.
When considering bias, it is clear that less biased news receives higher inter-
actions in the early time slots (although The New York Times belongs to the left
party). Again, the trend of deviation of a class from the reliability and bias class
that it belongs to can be seen in different buckets. As an example during hours 1
to 5 (after publishing a post), a typical post belonging to the “most reliable-left”
class does not follow the pattern of the “most reliable” nor the “left group”.
Fox News: Figure 8 shows the temporal results for Fox News. For the first and
the last bucket, similar to The New York Times, we see a correlation between
reliability engagement, when again discarding the non-significant results of the
“most unreliable” class. After around five hours, the “most reliable” class loses
its first-place ranking to the “unreliable” group. The large increase in unreliable
news after 5 h is statistically supported. Again, we can see a big, normalized
decline in the most reliable news 17 h after posting. Similarly to what we have
observed for The New York Times, we may observe that in the earliest phases
of a post’s lifetime, balanced news is more engaging than biased ones, although
Fox News itself is a right-biased biased news outlet. Here, statistical evidence
supports the divergence we see for the bias class from the average population,
until 5 h after posting.
Key Observation: In spite of Fox News and the New York Times being
biased publishers, for both, related unbiased posts receive a higher inter-
action rate than biased ones in the first hour following posting.
NPR: Finally, we used NPR as an example of an outlet with very limited bias.
As seen in Fig. 9, again balanced news receives higher interaction rates than the
unbiased ones in the very first bucket, and the trend changes in the last bucket.
The biased news published by this outlet tends to receive the most interaction
during the late stages of the posts’ lifetime. Moreover, the statistically significant
decreasing pattern of the interaction rate with the “most reliable” news is worth
noting.
Table 5. Minimum time required for reaching high correlations between the current
and ultimate interactions (m: minutes, h: hours, and d: days).
Reliability
Most unreliable Unreliable Reliable Most reliable All
r2 > .6 r2 > .8 r2 > .6 r2 > .8 r2 > .6 r2 > .8 r2 > .6 r2 > .8 r2 > .6 r2 > .8
Bias Far left 15 m 21 m 25 m 1 h, 15 m 15 m 31 m 31 m 15 m 31 m
51 m
Left 15 m 15 m 25 m 2 h, 31 m 2 h, 15 m 37 m 31 m 1 h,
13 m 13 m 17 m
Balanced 9 h, 35 m 1d, 10 h 15 m 13 h, 6 h, 16 h, 21 m 1 h, 1 h, 17 m 11 h,
48 m 39 m 33 m 51 m 30 m
Right 15 m 15 m 21 m 1 h, 25 m 2 h, 18 m 18 m 21 m 2 h,
51 m 40 m 13 m
Far right 15 m 15 m 15 m 15 m 15 m 18 m 15 m 37 m 15 m 15 m
All 15 m 31 m 15 m 53 m 37 m 11 h, 21 m 1 h, 31 m 6 h,
30 m 4m 39 m
First, we divide the time axis into exponentially increasing time buckets. For
the first bucket, we use a size of 15 min, and then we use a factor of 1.2 to increase
the bucket sizes. Then, in each bucket and for each group, we compute the
coefficient of determination (r2 ) as the squared value of the Pearson correlation
coefficient between the current interaction values of the posts of the class and
the total interaction they receive in the future. Finally, we recorded the moments
in which the (r2 ) reached .6 and .8, respectively. Table 5 summarize the results.
While more advanced prediction models might be used in practice, not limiting
the discussion to a particular predictive model provides quantifiable insights
into the extent to which we can rely on predictive models to estimate the total
number of interactions from the current value of the interaction a post received
(even with simple models). We next share some of our key observations.
First, note that in most classes a Pearson correlation coefficient of 0.8 (r2 of
0.6) is achieved within one hour of posting, suggesting that the total number of
interactions is relatively well predicted very early. Second, if all posts are taken
into consideration, this can be accomplished within 30 min of posting. Third,
considering all reliability classes (last row), we can see that we can achieve this
level of predictability within around 40 min of posting. Fourth, as we examine
all bias classes in the last column, we can see that the biased classes are able to
reach this level earlier than the unbiased classes.
Fifth, note that reaching the high value r2 level of 0.8 for the general pop-
ulation (last row and last column of the table) is feasible within 7 h after the
posting. With regard to our definition of 4-bucket thresholds, we can say that
for all the classes except for 3 we can reach the 0.6 level of r2 in the first bucket.
In the second bucket, it is also feasible to achieve an r2 level of 0.8, except for
the six classes. Finally, we note that for all classes except one, we can reach the
r2 level of 0.8 before the fourth bucket, allowing us to apply patterns observed
in this bucket more broadly.
Temporal Effects of Political Bias and Reliability on User Engagement 181
5 Related Work
This paper relates to the works modeling and understanding the behavior of
users, their interactions with various kinds of news and contents, and the factors
that play roles in this context. For example, Aldous et al. [1] focus on the topic
and emotional factors and analyze their effects on posting on five social media
platforms (Facebook, Instagram, Twitter, YouTube, and Reddit) to demonstrate
that user engagement is strongly influenced by the content’s topic, with certain
topics being more engaging on a particular platform. Their work shows that the
engagement level is impacted differently on various platforms and by different
topics. They also demonstrate that post emotion is indeed a significant factor.
Karami et al. [10] demonstrate how social engagement may be used as a distin-
guishing characteristic between false and true news spreaders. However, they do
not consider the temporal patterns of different user interactions in their study.
The most comparable work to ours is the recent work by Edelson et al. [4].
Their large-scale study explores how consumers engage with news inside the
Facebook news ecosystem, as well as with specific pieces of news from unre-
liable suppliers and also between the suppliers and their audiences. However,
their methodology is distinct from ours in that they base their study on pub-
lisher ratings rather than independent bits of news, they use binary classes for
reliability, and they do not account for the temporal dynamics of the user inter-
actions. Galen et al. [21] carried out a similar investigation as Edelson et al. on
Reddit rather than Facebook. They also employ publisher-based rankings and
demonstrate that low-factual content receives 20% fewer upvotes and 30% fewer
cross-posting exposures than neutral or more factual information.
In another line of research, Allcott et al. [2] examine how users engage with
fake news information and websites. Their findings indicate that through the
end of 2016, user interactions with fraudulent information increased consis-
tently on both Facebook and Twitter. Since then, engagements on Facebook
have decreased significantly while continuing to increase on Twitter. Another
group of studies related to our work are the ones which examine the temporal
dynamics of user interactions but in different contexts. For example, Vassio et
al. [19] examine how influencer-generated material draws interactions over time.
Their findings indicate that while the growth rate of interactions naturally decays
with time, the decay rate differs substantially between posts and social media
platforms. As another related work and with a different methodology from the
above works, in [12] the authors use NLP techniques to analyze over 2,5 million
social media comments. The results show that Social media misinformation is
largely disregarded by users.
6 Limitations
Our study has four main limitations that the researchers should consider when
generalizing the findings. First, we dropped the posts with less than 10 total
interactions from our study. While these types of postings constitute a significant
182 A. Mohammadinodooshan and N. Carlsson
portion of the total number of posts on Facebook, they make up a very small
fraction of the total interactions (less than 5% in our dataset) and typically are
of little interest to both Facebook content moderators (wanting to ban large
interactions with misinformation) and also content publishers.
Second, similar to some other works (e.g., [4], we limited the study to news
postings and interactions on Facebook public forums (the most popular social
media platform [17]). Therefore, interactions with news articles on other social
media platforms and on the publisher’s website were not considered. We consider
a combined analysis that also takes into account these aspects as an interesting
future work. It should also be noted that our study is based on the CrowdTangle
dataset and does not consider every public page on Facebook. Yet, CrowdTangle
covers many pages from the whole public pages distribution. As as example, they
index more than 99% of the pages with more than 25K followers [5].
Third, despite the t-test results indicating that the results are significant for
several classes, the significance of the results may differ between different classes.
To help interpret the significance of individual results the interested reader can
consider also the number of articles in our dataset for each specific class. To
help the interested reader to reproduce the results and more easily consider such
additional dimensions, we will share our code. Here it should also be noted that
we utilized the TICR distributions of the posts, not the aggregated results across
the articles. One reason for this is that the number of posts for the flagged classes
was sufficient for the findings to frequently have p-values less than 0.05.
Finally, The study focuses on the impact of bias and reliability on user
engagement but does not account for other potential factors such as the rel-
evance, timeliness, or credibility of the news source, as well as the user’s indi-
vidual preferences and views. Further research can consider these factors and
their impact on user engagement, as well as investigate the effects of alternative
labeling methods or different time frames compared to those used in the current
study.
7 Ethical Considerations
All data was collected via public APIs while adhering to the rate limits of the
companies hosting the data. The study is done at the aggregate level and no
specific individuals are revealed. The likelihood of a substantial portion of the
analyzed posts having been removed from Facebook is low due to the 28-day
temporal separation between the publication date of the article and the date of
data collection.
8 Conclusions
This paper presented a large-scale investigation of the temporal dynamics of
various user interactions with Facebook posts belonging to different classes of
bias and reliability. Using a carefully designed methodology, our investigation
has answered and provided statistically supported insights into the research
Temporal Effects of Political Bias and Reliability on User Engagement 183
A Appendix
A.1 Procedure of Computing the Canonical Form of an Article Url
The following procedure is taken to transform URLs to canonical form. We
begin by converting all text to lowercase. We then delete the protocol schema
(e.g. ‘[Link] and remove any prefix instances of the strings ‘www.’ that
may be present. Next, we remove any # signs from the URL except for the
domains that it could not be removed from the canonical form (e.g., some
of edsource or npr domains URLs). Then, we remove all URL query param-
eters except for the domains for which this was part of their canonical form
184 A. Mohammadinodooshan and N. Carlsson
Fig. 10. Temporal dynamics of the total interactions covered for angry counts (:
coefficient of variation of the mean is smaller than 4%, and : has deviation from
the previous time bucket with p-value < 0.05).
Temporal Effects of Political Bias and Reliability on User Engagement 185
The temporal dynamic results for CNN (as our second left-based example) and
New York Post (as our second right-based outlet) and Reuters (as the second
least biased publisher) are presented in Figs. 12, 13 and 14. In contrast to the
other biased example outlets, we observed both right-biased and left-biased arti-
cles published by New York Post.
Fig. 11. Temporal dynamics of the total interactions covered for haha counts (: coef-
ficient of variation of the mean is smaller than 4%, and : has deviation from the
previous time bucket with p-value < 0.05).
Fig. 12. Temporal dynamics of the total interactions covered for CNN (: coefficient
of variation of the mean is smaller than 4%, and : has deviation from the previous
time bucket with p-value < 0.05).
186 A. Mohammadinodooshan and N. Carlsson
Fig. 13. Temporal dynamics of the total interactions covered for The New York Post
(: coefficient of variation of the mean is smaller than 4%, and : has deviation from
the previous time bucket with p-value < 0.05).
Fig. 14. Temporal dynamics of the total interactions covered for Reuters (: coefficient
of variation of the mean is smaller than 4%, and : has deviation from the previous
time bucket with p-value < 0.05).
References
1. Aldous, K.K., An, J., Jansen, B.J.: What really matters?: characterising and pre-
dicting user engagement of news postings using multiple platforms, sentiments and
topics. Behaviour & Information Technology, pp. 1–24 (2022)
2. Allcott, H., Gentzkow, M., Yu, C.: Trends in the diffusion of misinformation on
social media. Res. Polit. 6(2), 205316801984855 (2019). [Link]
2053168019848554
3. Barfar, A.: Cognitive and affective responses to political disinformation in Face-
book. Comput. Human Behav. 101, 173–179 (2019). [Link]
chb.2019.07.026
4. Edelson, L., Nguyen, M.K., Goldstein, I., Goga, O., McCoy, D., Lauinger, T.:
Understanding engagement with U.S. (mis)information news sources on Facebook.
In: Proceedings of the ACM SIGCOMM Internet Measurement Conference, IMC,
pp. 444–463 (2021). [Link]
Temporal Effects of Political Bias and Reliability on User Engagement 187
1 Introduction
That network latency is an important factor of network performance has long
been known [8]. Various studies have shown that users’ Quality of Experience
(QoE) for many different applications, such as web searches [3], live video [32]
and video games [31], is strongly related to end-to-end latency, where network
latency can be a major component. For highly interactive applications envisioned
for the Tactile Internet or Augmented and Virtual Reality (AR/VR), reliable low
latency will be even more crucial [24]. It is therefore of great interest to Internet
Service Providers (ISPs) to be able to monitor their customers’ network latency
at large. Furthermore, network latency monitoring has a wide range of other use
cases like: verifying Service Level Agreements (SLAs), finding and troubleshoot-
ing network issues such as bufferbloat [28], making routing decisions [34], IP
geolocation [12] and detecting IP spoofing [18] and BGP routing attacks [4].
c The Author(s) 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 191–208, 2023.
[Link]
192 S. Sundberg et al.
There exists many tools for actively measuring network latency by sending
out network probes, such as ping [16], IRTT [11], and RIPE Atlas [22]. While
active monitoring is useful for measuring connectivity and idle network latency
in a controlled manner, it is unable to directly infer the latency application traf-
fic experience. The network probes may be treated differently from application
traffic by the network, due to for example active queue management and load
balancing, and therefore their latency may also differ. Furthermore, many active
monitoring tools require agents to be deployed directly on the monitored target,
which is not feasible for an ISP wishing to monitor the latency of its customers.
Passive monitoring techniques avoid these issues by observing existing appli-
cation traffic instead of probing the network. Additionally, passive monitoring
can often run on any device on the path that sees the traffic, not limited to end
hosts. Several tools for passively inferring TCP round trip times (RTTs) already
exist: Tcptrace [25] can compute TCP RTTs from packet traces, but is unable
to operate on live traffic. Wireshark and the related tshark [9] can operate
on live traffic, but are unsuitable for continuous monitoring over longer periods
of time, due to keeping a record of all packets in memory. On the other hand,
PPing [21] uses a streaming algorithm, which allows for continuous monitoring
of live traffic. However, like most other software based passive network moni-
toring solutions, PPing relies on traditional packet capturing techniques such as
libpcap. Packet capturing imposes a high overhead and is unable to keep up
with the high packet rates encountered on modern network links [17].
To enable passive network monitoring at higher packet rates, several recent
works [7,10,23,35] propose solutions based on P4 [6]. While these P4-based solu-
tions can achieve high performance, they require hardware support for P4, com-
monly found in Tofino switches. It could be possible to modify such P4 programs
to compile with Data Plane Development Kit (DPDK), however, this would
compromise on the guaranteed performance provided by the hardware. Beyond
DPDK and P4, there are many more Linux devices relying on kernel network
stacks that could still benefit from monitoring network latency. Examples include
commodity web servers, routers, traffic shapers and Network Intrusion Detection
Systems (NIDS), which use the Linux network stack for their normal operation.
In recent years, the introduction of eBPF [29] in the Linux kernel added the
ability to attach small programs to various hooks that run in the kernel. This
makes it possible to inspect and modify kernel behavior in a safe and performant
manner, without having to recompile a custom kernel. eBPF is in general well
suited for monitoring processes in the kernel, and the BPF Compiler Collection
(BCC) repository already contains two tools to passively monitor TCP RTT:
tcpconnlat and tcprtt. While these tools expose RTT metrics in an efficient
manner, they rely on the RTT estimations from the kernel’s own TCP stack,
and can therefore only run on end hosts.
While retrieving statistics from the kernel certainly has its uses, Linux Traffic
Controller (tc) BPF and eXpress Data Path (XDP) [13] hooks go a step further
and essentially enable a programmable data plane in the Linux kernel [30]. eBPF
programs attached to tc and XDP hooks can process and take actions on each
Efficient Continuous Latency Monitoring with eBPF 193
packet early in the Linux network stack, without the overhead from cloning
the packet and exposing it to a user space process like packet capturing does.
XDP and tc-BPF have been used to implement for example efficient flow mon-
itoring [1], load balancers [19] and a Kubernetes Container Network Interface
(CNI) [14]. Of particular relevance for this work, [33] proposes an in-band net-
work telemetry approach for measuring one-way latency. It uses eBPF to add
timestamps to a fraction of the packet headers. However, this approach requires
full control over the part of the network that should be monitored as well as
synchronized clocks between source and sink nodes.
In this paper we instead propose using eBPF to efficiently inspect packets
and use a streaming algorithm, such as the one used by PPing, to calculate the
RTT for the packets as they traverse the kernel. Such a solution can continu-
ously monitor network latency from any Linux-based device that is able to see
the traffic in both directions of a flow. Also, our proposal does not require the
control of any other device in the network or end hosts. Furthermore, it avoids
the overhead of packet capturing, and it does not require any modifications to
the Linux kernel or special hardware support. To show the feasibility of this
approach, we make the following contributions:
– We implement an evolved Passive Ping (ePPing), inspired by PPing, but using
eBPF instead of traditional packet capturing.
– We evaluate the accuracy and overhead of ePPing, demonstrating that it pro-
vides accurate RTTs and can operate at high packet rates with considerably
lower overhead than PPing, being able to process upwards of 16x as many
packets at a third of the CPU overhead.
– We identify that reporting a large number of RTT values makes up a signifi-
cant part of the overhead of ePPing, and implement simple in-kernel sampling
and aggregation to mitigate it.
The design and implementation of ePPing is covered in Sect. 2, while the accu-
racy and performance of ePPing is evaluated in Sect. 3. Finally, we summarize
our conclusions in Sect. 4.
Ethical Considerations This work does not raise any ethical issues as all
experiments have been performed in a controlled testbed with no real user traffic.
However, the presented ePPing tool reports IP addresses and ports, which in
other contexts may contain sensitive information. Like any tool that can collect
and report IP addresses, great care should therefore be taken to ensure that such
information does not leak to unauthorized parties before deploying ePPing in a
public network.
The principle behind ePPing and most other passive latency monitoring tools is
to match replies to previously observed packets and to calculate the RTT as the
194 S. Sundberg et al.
time difference between these. How ePPing performs this task is illustrated in
Fig. 1. First, each incoming or outgoing packet is parsed for a packet-identifier
that can be used to match the packet against a future reply . 1 If such an
identifier is found, the current time is saved in a hash map using a combination of
the flow tuple and the identifier as a key to uniquely identify the packet .
2 Then
the program checks if the packet contains a suitable reply identifier, which it can
use to match with a previously seen packet in the reverse direction, and queries
the hash map . 3 If a match is found, the RTT is calculated by subtracting the
stored timestamp from the current time . 4 Finally, the RTT report is pushed
to user space ,5 which prints it out . 6 Additionally, ePPing also keeps track
of some state for each flow, e.g., number of packets sent and minimum RTT
observed.
Both ePPing and PPing use the TCP timestamp option [5] as identifiers.
With TCP timestamps, each TCP header will contain two timestamps: TSval
and TSecr. The TSval field will contain a timestamp from the sender, and the
receiver will then echo that timestamp back in the TSecr field. One can thus
use the TSval value as an identifier for an observed packet and later match it
against the TSecr value in a reply. It should be noted that TCP timestamps are
updated at a limited frequency, typically once every millisecond. Thus, multiple
consecutive packets may share the same TSval, which is therefore, especially
at high rates, not a reliable unique identifier. To avoid mismatching replies to
packets and getting underestimated RTTs, we only timestamp the first packet
for each unique TSval in a flow and match it against the first TSecr echoing it. By
only using the edge when TCP timestamps shift, the frequency rather than the
accuracy of the RTT samples is limited to the update rate of TCP timestamps.
Note that matching the first instance of a TSval against the first matching TSecr,
combined with the algorithm for how the receiver sets the TSecr, means that
the calculated RTT will always include a delay component of delayed ACKs [5].
We further discuss the implications of using TCP timestamps as identifiers to
passively monitor the RTT in Appendix A.
Efficient Continuous Latency Monitoring with eBPF 195
3 Results
traffic between the end hosts. In all experiments, the (partial) RTT between the
middlebox and receiver end host is passively monitored from the interface on
the middlebox facing the receiver, unless otherwise specified.
The network offloads Generic Receive Offload (GRO), Generic Segmentation
Offload (GSO) and TCP Segmentation Offload (TSO) are disabled on the mid-
dlebox, but left enabled on the end hosts. With this, we force the middlebox
to process every packet. This is not necessary for PPing or ePPing, however, it
provides a more accurate view of how packets traverse the wire. Furthermore,
disabling the offloads makes it easier to fairly compare performance across a
varying amount of concurrent flows, as the offloads tend to become less effective
as the rate per flow decreases. With the offloads left enabled, the middlebox
would have inherently performed much better for a few flows with very high
packet rates compared to if the same packet rate is distributed across many
flows, even without passive monitoring.
Section 3.1 focuses on the accuracy of the RTTs reported by ePPing by com-
paring them to the RTTs reported by PPing, which also relies on TCP times-
tamps, and tshark, which instead calculates the RTTs from the sequence and
acknowledgement numbers. Section 3.2 covers the overhead ePPing incurs on the
system compared to PPing, thereby evaluating if implementing a similar algo-
rithm in eBPF programs instead of relying on packet capturing is a feasible way
to extend passive latency monitoring to higher packet rates.
(a) RTTs reported over the duration of (b) The distribution of RTTs after sub-
the test. tracting configured delay.
Fig. 3. RTT values reported by tshark, PPing and ePPing for a single TCP flow with
0 to 100 ms of additional latency added in 10 ms steps.
To evaluate the accuracy of the RTT values ePPing reports, we use iperf3
to send data at a paced rate of 100 Mbps over a single flow from the sender to
the receiver end host. To test that ePPing is able to accurately track changes in
RTT, we apply a fixed netem delay, which is increased in 10 ms steps every 10 s,
Efficient Continuous Latency Monitoring with eBPF 197
going from 0 to 100 ms, see Fig. 3a. In addition to running ePPing at the capture
point, we capture the headers of all packets by running tcpdump on the same
interface. PPing, tshark and tcptrace calculate the TCP RTT values from the
capture file, but tcptrace is omitted from the results as it yields identical RTT
values as tshark. To avoid small latency variations from the CPU aggressively
entering different sleep states, we use the tuned-adm profile latency-performance
on the middlebox during these tests.
Figure 3a shows a timeseries of the RTT values calculated by each tool. All
tools provide RTT values closely following the configured netem delay. Figure 3b
instead shows the distribution of how much higher the reported RTT values
are compared to the configured netem delay, to avoid the scale of RTT values
to dwarf the variation. However, in both Figs. 3a and 3b the magnitude of the
RTT values and their variation are much larger than the differences between the
tools. Therefore, Fig. 4 shows the pairwise difference between each RTT value for
ePPing compared to PPing and tshark, respectively. Note that tshark reports
an RTT value for every ACK, whereas PPing and ePPing only produce an RTT
for ACKs with a new TSecr value, thus providing 13 % fewer RTT samples than
tshark in this experiment (see the count field in Fig. 3b). Therefore, Fig. 4b only
includes the RTT values from tshark that correspond to those from PPing and
ePPing, i.e. the ones from the first ACK with each TSecr value. Furthermore,
differences below 1 µs may be due to rounding as the RTT values from tshark
and PPing have microsecond resolution.
(a) Difference between ePPing and PPing. (b) Difference between ePPing and
tshark.
Fig. 4. Pairwise difference between RTT values reported by ePPing compared to other
tools.
Overall, ePPing reports slightly lower RTT values than PPing. This is
expected as the XDP hook used by ePPing for ingress traffic is triggered before
the packet enters the rest of the Linux network stack, and can be captured by
tcpdump. On the other hand, ePPing provides RTT values that are around 1 to
5 µs higher than those from tshark, which is explained by tshark calculating the
RTT in a different way. Both PPing and ePPing use TCP timestamps, and will
198 S. Sundberg et al.
therefore always include the additional latency caused by delayed ACKs. Mean-
while, tshark instead matches sequence and acknowledgement numbers, which
will often exclude this delay component. We have verified that the differences
between ePPing and tshark correspond to the additional latency component
from delayed ACKs. While the difference in how delayed ACKs are handled
result in very small differences in Fig. 4b, it can create larger differences for
some particular traffic patterns. In Appendix A we further discuss how relying
on TCP timestamps affect the calculated RTT values.
(a) Throughput (b) CPU (average across all (c) Packets processed
6 cores)
Fig. 5. Forwarding performance without monitoring (baseline), with PPing and with
ePPing for 10 concurrent iperf3 flows, when middlebox uses all CPU cores.
than ePPing, PPing is actually only processing just over 6 % of the packets. This
is due to the packet capturing being unable to keep up with the high packet rate,
and therefore missing the majority of the packets. In contrast, ePPing runs in line
with the rest of the network stack, and sees every packet, meaning it processes
roughly 16 times as many packets. While not apparent from Fig. 5, also note
that PPing is implemented as a single-threaded user space application, and is
therefore limited to how fast a single core can process all the logic. While the user
space component reporting the RTT values in ePPing is also single-threaded, the
eBPF programs that contain the logic for calculating the RTT values run on the
cores that the kernel assigns to process each packet, thus distributing the load
across multiple cores in the same manner as the normal network stack processing.
Table 1. Average packets per second processed on single core at capture point when
only forwarding (baseline scenario).
Although the results in Fig. 5 are promising, the end hosts are usually the
bottleneck here, especially as we increase the number of flows. These experiments
are consequently unable to push the middlebox and ePPing to their limits. We
therefore constrain the middlebox to using a single CPU core in the remaining
experiments, moving the bottleneck to the middlebox CPU. This means that the
middlebox is already using all of its CPU capacity just forwarding the traffic,
and any additional overhead from the passive monitoring results in decreased
throughput. Furthermore, we emphasize the total packet rate (sum of trans-
mitted and received packets) rather than the throughput. Packet rate is more
relevant for the performance of PPing and ePPing, as their logic has to run per
packet, and also stays more consistent across a varying number of flows as Table 1
shows. As the number of flows increases, the number of ACKs sent back by the
receiver increases (seen by the increase in received packets at the middlebox).
This results in less capacity to forward data packets by the middlebox (decrease
in transmitted packets), and thereby a lower throughput, while the total packet
rate handled remains similar.
Figure 6a summarizes the impact PPing and ePPing have on the forward-
ing performance of the middlebox when it is constrained to a single core. Both
ePPing and PPing now have a considerable impact on the forwarding perfor-
mance, but ePPing clearly sustains a higher packet rate than PPing, at least at
a limited number of flows. As the number of flows increases, the packet rate with
200 S. Sundberg et al.
ePPing drops from 1.53 to 1.13 Mpps. The reason for this drop in performance
as the number of flows increases is that, due to the limited update rate of TCP
timestamps, the number of potential RTT samples that ePPing has to process
increases with number of flows. This is evident in Fig. 6b, which shows that while
ePPing reports the expected 1000 RTT values per second for a single flow, at
1000 flows this increases to roughly 125,000 values per second.
(a) Average packet rate (b) Average RTT report rate (log scale)
Fig. 6. Middlebox performance when just forwarding (baseline), with PPing or ePPing
on a single CPU core. PPing misses most packets and thus processes (PPing-proc)
packets at a much lower rate than they are forwarded (PPing-fw).
Meanwhile, the forwarded packet rate with PPing actually appears to increase
slightly with number of flows (0.97 Mpps at one flow, 1.18 Mpps at 1000 flows),
but this is merely due to the packet capturing missing a larger fraction of packets.
The packet rate actually handled by PPing drops from approximately 170 kpps at
1 and 10 flows, to just 60 kpps at 100 and 1000 flows, meaning ePPing processes
packets at approximately an 18 times higher rate than PPing at 1000 flows.
PPing missing the majority of packets results in it also missing many RTT
samples, which can be seen by the much fewer RTT values reported by PPing in
Fig. 6b. Furthermore, the algorithm for matching packets to replies that PPing
and ePPing uses, relies on matching the first instance of each TSval to the first
matching TSecr. As PPing does not see every packet, it cannot guarantee this,
and it may therefore introduce small errors in its RTT values.
However, limiting the load by sampling, as PPing in practice does by miss-
ing packets, can be a valid approach. The high rate of RTT values reported by
ePPing may not be necessary, or even desirable, for many use cases. We there-
fore implement sampling for ePPing, and evaluate if it can be an effective way to
reduce the overhead. While it would be possible to only process a random subset
of the packets, similar to PPing, such an approach has several drawbacks. As
already mentioned, missing packets may interfere with the algorithm for match-
ing packets and replies, thus resulting in less accurate RTT values. Furthermore,
a random subset of packets is likely to mainly yield RTT samples from elephant
flows, and largely miss sparse flows. However, sparse flows often carry control
Efficient Continuous Latency Monitoring with eBPF 201
traffic and other latency sensitive data, and being able to monitor their RTT
may therefore be at least as important as the RTT of the elephant flows. Instead
of eliminating ePPing’s advantage of being guaranteed to see every packet, we
opt to implement a simple per-flow sample rate limit. With the sample rate limit,
ePPing must wait a time period t after saving a timestamp entry for a packet
before it can timestamp another packet from the same flow. This t may either
be set to a static value, or it can be dynamically adjusted to the RTT of each
flow, so that flows with shorter RTTs get more frequent samples than flows with
longer RTTs.
We repeat the experiments from Fig. 6, setting the sample limit t to 0, 10,
100 and 1000 ms, in practice corresponding to at most 1000, 100, 10 and 1
RTT values per flow and second, respectively. Figure 7a summarizes the results,
and clearly shows that less frequent sampling greatly reduces the overhead of
ePPing. Already at a sample limit of 10 ms we see great improvements. When
limiting it to a single sample every 1000 ms per flow (the default rate of ping),
ePPing is able to sustain a packet rate of 1.54 Mpps for 1000 flows, compared to
1.14 Mpps without sampling. The drop in forwarding performance when going
from 1 to 1000 flows thus decreases from 27 % without sampling, to just 1.6 %
at t = 1000 ms. The drawback of such coarse sampling is that the granularity of
the monitoring is reduced, and one might miss important RTT variations.
Fig. 7. Impact of different levels of per-flow sample limiting and aggregation. Note that
the Y-axis does not start at 0.
high sample limit, and consequently few RTT values to aggregate, the aggrega-
tion yields a very modest improvement. However, for smaller sample limits the
aggregation becomes more beneficial. In the scenario without any sampling, the
aggregation increases the packet rate at 1000 flows from 1.14 to 1.45 Mpps. By
combining sampling and aggregation, we therefore expect ePPing to be able to
maintain a high level of performance, while still providing useful RTT metrics,
at significantly more than 1000 concurrent flows. Due to limitations with the
current testbed, we are however unable to validate performance beyond 1000
flows.
4 Conclusion
In this paper we propose using eBPF to passively monitor network latency,
and demonstrate the feasibility of this by implementing evolved Passive Ping
(ePPing). By using eBPF, ePPing is able to efficiently observe packets as they
pass through the Linux network stack without the overhead associated with
packet capturing. It does not require any modifications to the kernel or replacing
the network stack with DPDK, nor any special hardware support. Our evaluation
shows that ePPing delivers accurate RTTs and has much lower overhead than
PPing, being able to handle over 1 Mpps on a single core, corresponding to more
than 10 Gbps of throughput. We also demonstrate that sampling and aggregation
of RTT values in the kernel can be used to further reduce the overhead from
handling a large amount of RTT samples.
While ePPing overall performs well in our experiments, our evaluation is
heavily based on bulk TCP flows generated by iperf3. In future work we intend
to evaluate how ePPing fares with a more realistic workload by using traffic from
an ISP vantage point. Another important aspect to consider is what impact the
passive monitoring has on end-to-end latency. We are currently working on bet-
ter understanding ePPing’s impact on end-to-end latency. Preliminary findings
indicate that while ePPing only adds a couple of hundred nanoseconds of pro-
cessing latency to each packet (99th percentile of approximately 350 ns), it may
under certain scenarios increase end-to-end latency by hundreds of microseconds.
Furthermore, our current implementation of ePPing has some limita-
tions. Limitations inherent to using TCP timestamps are further discussed in
Appendix A, with one of the primary ones being the lack of ability to monitor
flows where TCP timestamps are not enabled. Some of the these limitations could
be avoided by using sequence and acknowledgement numbers instead, although
that has its own set of limitations. We are also considering adding support for
other protocols, such as DNS and QUIC. Additionally, the sampling and aggre-
gation methods we employ in this work are relatively simple, and we are working
on more sophisticated ways to sample, filter and aggregate RTTs in-kernel to
provide enhanced RTT metrics while maintaining low overhead.
Efficient Continuous Latency Monitoring with eBPF 203
The decision to use TCP timestamps for ePPing was mainly based on having
a simple algorithm that avoids the TCP transmission ambiguity. As illustrated
in Fig. 8a, retransmissions can cause approaches that match sequence and ACK
numbers to greatly overestimate or underestimate the RTT, unless they also
detect retransmissions to filter out such spurious RTT samples. However, for
TCP timestamps, the retransmission will typically have a newer TSval, and,
therefore, no additional precautions are needed to calculate a correct RTT.
Due to how TSecr is updated, TCP timestamps also handle delayed ACKs
a bit differently compared to sequence and ACK number matching: The echoed
TSecr value is not necessarily the latest TSval. Rather, RFC 7323 [5] speci-
fies that TSecr should be set to a recent Tsval, which is updated according
to Algorithm 1, and essentially results in TSecr being set to the TSval from
the oldest in-order unacknowledged segment. The effect of this is that RTTs
based on TCP timestamps will, by design, always include the additional latency
from delayed ACKs. On the other hand, matching sequence and ACK numbers
will only include the delayed ACK if it is triggered by a timeout, as shown in
Fig. 8b. Consequently, matching sequence and ACK numbers will usually result
in RTTs that are a bit closer to the underlying network latency, whereas using
TCP timestamps will result in RTTs more similar to those experienced by the
TCP stack. Both methods are, however, prone to include RTT spikes caused by
delayed ACKs timing out.
There are also two noteworthy drawbacks with relying on TCP timestamps:
Firstly, TCP timestamps are optional, and ePPing can therefore only monitor
TCP traffic with TCP timestamps enabled. A recent study [2] found that out
of the most common operating systems (Android, iOS, Windows, MacOS and
Linux), Windows was the only one not supporting TCP timestamps by default.
A lot of traffic these days goes through mobile devices running Android and
iOS, but Windows is still the dominant desktop OS, making this a noteworthy
limitation. Secondly, the TCP timestamp update rate limits how frequently we
can collect RTT samples within a flow. The study in [2] found that among servers
for popular websites, the most common update rate was once per millisecond,
which is what Linux uses since v4.13, but some updated at a slower rate of every
4 ms or every 10 ms. For most applications we deem that 1000–100 RTT samples
per second per flow is plenty, but for very fine-grained analysis requiring an RTT
sample for every ACK this could be problematic.
Furthermore, there are two edge cases in which matching TCP timestamps
may result in slightly overestimating the RTT beyond the delayed ACK com-
ponent: The first case is when a retransmission happens fast enough that the
TSval is not updated from the original transmission. For example, consider if
the retransmission in Fig. 8a would still use T Sval = 1. This can only occur if
the retransmission occurs faster than the TCP timestamp update rate, and may
at most overestimate the RTT with the TCP timestamp update period. With
TCP timestamps typically being updated every millisecond, this should be very
rare in most environments outside of for example data center networks. The sec-
ond case is when the TSval is updated during a delayed ACK and persists into
packets being acknowledged by the next ACK. For example, consider if the third
packet sent by A in Fig. 8b would still have T Sval = 2. In that case, the RTT for
the second ACK sent by B (ACK = 400) would incorrectly be calculated from
the second packet sent by A (Seq = 200) instead of from the third packet sent
by A (Seq = 300). This error can occur in the presence of delayed ACKs, and if
multiple packets within a flow have the same TSval. Thus, this is also bounded
to at most overestimate the RTT with one TCP timestamp period. This edge
Efficient Continuous Latency Monitoring with eBPF 205
(a) Difference from including and exclud- (b) Additional overestimation of RTT
ing the additional latency component of from TCP timestamps due to the second
delayed ACKs. edge case.
In summary, using TCP timestamps may result in slightly higher RTT values
than matching sequence and ACK numbers, mainly due to different handling of
delayed ACKs. While ePPing could be modified to instead operate on sequence
and ACK numbers, it would then risk missing valid RTT samples, especially on
lossy links, and would still capture the largest RTT spikes from delayed ACKs.
206 S. Sundberg et al.
References
1. Abranches, M., Michel, O., Keller, E., Schmid, S.: Efficient network monitoring
applications in the Kernel with eBPF and XDP. In: IEEE NFV-SDN 2021 (2021).
[Link]
2. Barbette, T., Wu, E., Kostić, D., Maguire, G.Q., Papadimitratos, P., Chiesa,
M.: Cheetah: A high-speed programmable load-balancer framework with guaran-
teed per-connection-consistency. IEEE/ACM Trans. Netw. 30(1), 354–367 (2022).
[Link]
3. Barreda-Ángeles, M., Arapakis, I., Bai, X., Cambazoglu, B.B., Pereda-Baños,
A.: Unconscious physiological effects of search latency on users and their click
behaviour. In: SIGIR 2015 (2015). [Link]
4. Birge-Lee, H., Wang, L., Rexford, J., Mittal, P.: SICO: surgical interception attacks
by manipulating BGP communities. In: CCS 2019 (2019). [Link]
3319535.3363197
5. Borman, D., Braden, R.T., Jacobson, V., Scheffenegger, R.: TCP Extensions for
High Performance. Technical report. RFC 7323, Section 3, Internet Engineering
Task Force (2014). [Link]
6. Bosshart, P., et al.: P4: programming protocol-independent packet processors. SIG-
COMM Comput. Commun. Rev. 44(3), 87–95 (2014). [Link]
2656877.2656890
7. Chen, X., Kim, H., Aman, J.M., Chang, W., Lee, M., Rexford, J.: Measuring TCP
round-trip time in the data plane. In: SPIN 2020 (2020). [Link]
3405669.3405823
8. Cheshire, S.: It’s the Latency, Stupid (2001). [Link]
rants/[Link]. Accessed 07 May 2022
9. Combs, G.: Tshark (2022). [Link]
html. Accessed 17 May 2022
10. Ghasemi, M., Benson, T., Rexford, J.: Dapper: data plane performance diagnosis
of TCP. In: SOSR 2017 (2017). [Link]
11. Heist, P.: IRTT (Isochronous Round-Trip Tester) (2021). [Link]
heistp/irtt. Accessed 31 Oct 2022
12. Hillmann, P., Stiemert, L., Rodosek, G.D., Rose, O.: Dragoon: advanced modelling
of IP geolocation by use of latency measurements. In: ICITST 2015 (2015). https://
[Link]/10.1109/ICITST.2015.7412138
13. Høiland-Jørgensen, T., et al.: The eXpress data path: Fast programmable packet
processing in the operating system kernel. In: CoNEXT 2018 (2018). [Link]
org/10.1145/3281411.3281443
14. Isovalent: Cilium - Linux Native, API-Aware Networking and Security for Con-
tainers (nd). [Link] Accessed 21 Oct 2022
15. Iyengar, J., Thomson, M.: QUIC: A UDP-Based Multiplexed and Secure Transport.
Technical report, RFC 9000, Section 17.4, Internet Engineering Task Force (2021).
[Link]
16. Kuznetsov, A., Yoshifuji, H.: Iputils (2022). [Link]
Accessed 03 May 2022
17. Li, J., Wu, C., Ye, J., Ding, J., Fu, Q., Huang, J.: The comparison and
verification of some efficient packet capture and processing technologies. In:
DASC/PiCom/CBDCom/CyberSciTech 2019 (2019). [Link]
DASC/PiCom/CBDCom/CyberSciTech.2019.00177
Efficient Continuous Latency Monitoring with eBPF 207
18. Maheshwari, R., Krishna, C.R., Brahma, M.S.: Defending network system against
IP spoofing based distributed DoS attacks using DPHCF-RTT packet filtering
technique. In: ICICT 2014 (2014). [Link]
19. Meta: Katran: A high performance layer 4 load balancer (2022). [Link]
com/facebookincubator/katran. Accessed 21 Oct 2022
20. Mockapetris, P.: Domain names - implementation and specification. Technical
report, RFC 1035, Section 4.1.1, Internet Engineering Task Force (1987). https://
[Link]/10.17487/RFC1035
21. Nichols, K.: PPing: Passive ping network monitoring utility (2018). [Link]
com/pollere/pping. Accessed 21 Sep 2021
22. RIPE NCC: Home—RIPE Atlas (nd). [Link] Accessed 20 Oct
2022
23. Sengupta, S., Kim, H., Rexford, J.: Continuous in-network round-trip time moni-
toring. In: SIGCOMM 2022 (2022). [Link]
24. Sharma, S., Woungang, I., Anpalagan, A., Chatzinotas, S.: Toward tactile internet
in beyond 5G era: recent advances, current issues, and future directions. IEEE
Access. 8, 56948–56991 (2020). [Link]
25. Shawn Ostermann: Tcptrace (2013). [Link] Accessed
03 Apr 2022
26. Sundberg, S., Brunstrom, A., Ferlin-Reiter, S., Høiland-Jørgensen, T., Brouer,
J.D.: Efficient continuous latency monitoring with eBPF - Resources (2023).
[Link]
27. Sundberg, S., Høiland-Jørgensen, T.: BPF-examples: PPing using XDP and TC-
BPF (2022). [Link]
Accessed 26 Jan 2023
28. The Bufferbloat community: Buff[Link] (nd). [Link]
projects/. Accessed 05 May 2022
29. The Linux Foundation: eBPF - Introduction, Tutorials & Community Resources
(2021). [Link] Accessed 03 May 2022
30. Vieira, M.A.M., et al.: Fast packet processing with eBPF and XDP: concepts, code,
challenges, and applications. ACM Comput. Surv. 53(1), 1–36 (2020). [Link]
org/10.1145/3371038
31. Vlahovic, S., Suznjevic, M., Skorin-Kapov, L.: The impact of network latency on
gaming QoE for an FPS VR game. In: QoMEX 2019 (2019). [Link]
1109/QoMEX.2019.8743193
32. Wang, H., Zhang, X., Chen, H., Xu, Y., Ma, Z.: Inferring end-to-end latency in live
videos. IEEE Trans. Broadcast. 68(2), 517–529 (2022). [Link]
TBC.2021.3071060
33. Xhonneux, M., Duchene, F., Bonaventure, O.: Leveraging eBPF for programmable
network functions with IPv6 segment routing. In: CoNEXT 2018 (2018). https://
[Link]/10.1145/3281411.3281426
34. Zhao, Z., Gao, S., Dong, P.: Flexible routing strategy for low-latency transmis-
sion in software defined network. In: ICCBN 2021 (2021). [Link]
3456415.3456444
35. Zheng, Y., Chen, X., Braverman, M., Rexford, J.: Unbiased delay measurement
in the data plane. In: APOCS 2022 (2022). [Link]
77059.2
208 S. Sundberg et al.
Open Access This chapter is licensed under the terms of the Creative Commons
Attribution 4.0 International License ([Link]
which permits use, sharing, adaptation, distribution and reproduction in any medium
or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons license and indicate if changes were
made.
The images or other third party material in this chapter are included in the
chapter’s Creative Commons license, unless indicated otherwise in a credit line to the
material. If material is not included in the chapter’s Creative Commons license and
your intended use is not permitted by statutory regulation or exceeds the permitted
use, you will need to obtain permission directly from the copyright holder.
Back-to-the-Future Whois: An IP Address
Attribution Service for Working
with Historic Datasets
1 Introduction
Structure: First, we introduce the datasets we use and our methodology for
BTTF whois in Sect. 2. Next, we evaluate BTTF whois against Team Cumry’s
bulk whois in a sample case. Finally, we first discuss our results and limitations
in Sect. 4, before concluding in Sect. 5.
the period from April 2004 up until today, with a quarterly resolution. However,
this reduced resolution will lead to a reduced reliability of the AS2ORG map-
pings, meaning that changes of ownership/authority over an AS may be reflected
up to three months too late, while temporary changes of a duration less than
three months may remain completely unnoticed, see Sect. 4.2.
Team Cymru Whois Data. As a base-line, we requested bulk whois data from
Team Cymru’s bulk whois service for all unique addresses in January 2023. We
used the Team Cymru whois to resolve all 14M unique IP addresses in the
university dataset. For each IP address the bulk whois service of Team Cumry
returns the currently associated AS number, the requested address, and the AS
Name and location of the corresponding AS.
2.2 Methodology
In this section, we describe how we organized the CAIDA AS2ORG and
AS2Prefix datasets in our service daemon to enable quick queries for individ-
ual addresses against the dataset. The major challenge–preventing a traditional
RDBMS from being used–is that these datasets contain whole prefixes, instead of
individual IP addresses, and relations between objects are complex. This would
212 F. Streibelt et al.
lead to, for example in SQL, a nested JOIN structure which limits performance
of an RDBMS. To prevent this bottleneck, our implementation uses a completely
in-memory prefix trie, i.e., pytricia [1].
Prefix Tree (Trie). Next, we iterate through the list of available files by date,
and add the prefixes we find to an IP trie [1]. In that trie, each added prefix
holds a list at date ranges when it was observed. For each prefix in our input
files, we check if the prefix exists in the trie. Here, we have to handle four cases:
– Prefix is not in the trie: We add the prefix to the trie, setting the ’first
seen’ field to the date of the collection date of the currently processing file.
– Prefix is in the trie:
• No gap to last-seen date: If the last-seen date of the prefix is the date
of the day before the collection time of the currently processing file, we
update the last-seen date of the most recent date-range to the date of the
currently processing file.
• Gap to last-seen date: If the last-seen date of the prefix is not the date
of the day before the collection time of the currently processing file, we
add a new date-range to the list of date-ranges, and set the first seen date
to the date of the currently processing file.
A Historic IP Attribution Service for Network Measurement 213
In all cases, the prefix is attributed to the ASes we observe as announcing the
prefix. There, we also have to handle several special cases:
– Prefix originated by exactly one AS: If a prefix is originated by exactly
one AS, we add this AS as the authoritative AS.
– MOAS prefix: If a prefix is announced by multiple ASes at the same time,
commonly known as a MOAS (Multi Origin AS) prefix, we add all these ASes
to the announcement state, see the section on handling requests for details
on the presentation.
– ASSET aggregate: ASes may aggregate prefixes received from downstream
ASes. Fore example, if AS65536 announces [Link]/25 to AS65538, and
AS65537 announces [Link]/25 to AS65538, AS65538 can aggregate
these announcements to [Link]/24, only announcing that to its peers,
while also aggregating AS65536 and AS65537 to { AS65536, AS65537 } in the
AS path of that announcement. The information whether [Link]/25 was
originated by AS65536 or AS65537 is lost in this process. As this is suggested
to occur only on provider aggregatable IP space [6], we attribute the whole
/24 to the aggregating AS, i.e., AS65538 in this case.
2
[Link]
[Link].
3
[Link]
[Link].
214 F. Streibelt et al.
originated by private and reserved AS numbers, i.e., 0 [18,19], 23456 [34], 64496–
64511 [17], 64512-65534 [14,24], 65535 [13], 65536-65551 [17,34], 65552-131071
(IANA Reserved), 4200000000-4294967294 [24], and 4294967295 [13].
Lookups. The implementation of the historic whois service allows lookups with
daily granularity. When an IP address or prefix is looked up, we first identify
the most specific match. Next, we check if the prefix has been announced at the
given date, i.e., if it has a date-range covering the requested date. If it does not
have a corresponding date range, we traverse the tree until we either find a less
specific prefix with a covering date-range or arrive at the root of the address
tree. If we reach the root, we return that the prefix was not found at that date.
For the most specific prefix with a covering date range, we return the
requested IP address or prefix, the requested date, and the result set. The result
set contains the dates when the prefix was first and last observed for the date-
range covering the requested date, with the last-seen date being null if the prefix
was still being observed in the newest file imported into the daemon. Additionally
we return the identified prefix and the list of ASes associated with the prefix. For
each AS we also return an AS2ORG mapping, listing the ASN, the ASNAME
and RIR where the ASN has been registered. Furthermore, we return all orga-
nizations associated with the AS at the time of the request, which includes the
country code registered for the organization, the RIR the organization object
has been obtained from, and the name of the organization.
3 Results
In this section, we describe how we evaluate the efficacy of BTTF whois using
the work of Fiebig et al. [10] as a case-study. We first introduce the results Fiebig
et al. obtained by using Team Cymru’s bulk whois service. We then compare the
attribution of address ownership between Team Cymru’s bulk whois service and
BTTF whois. Finally, we revisit the results of Fiebig et al., and describe how
using BTTF whois influences them.
A Historic IP Attribution Service for Network Measurement 215
Fig. 1. Cloud use attribution for universities in the U.K. and the U.S. (January 2015–
May 2022) based on Team Cymru bulk-whois data.
As outlined in Sect. 2.1, Fiebig et al. use the Farsight SIE dataset to identify
IP addresses to which names under universities domains point with a monthly
granularity. Using Team Cymru’s bulk-whois service, they then attribute these
IPs to AS numbers. For their final analysis, they then calculate the share of
universities in a country under whose domains at least one name ultimately
points to an IP address announced by one of Amazon’s, Google’s, or Microsoft’s
ASes. Naturally, a university may have multiple names under its domain that
point to addresses announced by different cloud providers. Figure 1 depicts their
results from January 2015 to October 2022 for 115 U.K. universities and 260
U.S. universities, with each bar in the bar-plots representing the distribution
observed during a single month.
For both, the U.S. and the U.K., they find an overall high prevalence of at
least one service or site being run on Amazon, Google, or Microsoft systems.
Notably, the U.S. already shows an over 90% saturation in cloud use, with the
main development being that the prevalence of universities having infrastructure
located at all three major cloud providers continuously rises over time. For the
U.K., still around 75% of universities have at least one service in the major three
clouds in January 2015, followed by a gradual increase across all platforms. Still,
even in October 2022, the use of Google systems in the U.K. is lower comparison
to the observed U.S. usage.
216 F. Streibelt et al.
Fig. 2. Percentage of prefixes in the dataset on which the historic whois service we
implemented and the data from Team Cymru’s bulk whois service disagree. The shaded
background indicates the distribution of disagreement over AS tuple, i.e., the tuple of
the ASes to which Team Cymru attributes a prefix and the ASes BTTF attributes a
prefix to. Note that, as common with centralization, only a minor fraction of AS tuples
is responsible for the bulk of disagreement.
Fig. 3. Difference in cloud use attribution for universities in the U.K. and the U.S.
(January 2015–May 2022) between Team Cymru bulk-whois data and historic bulk-
whois data as absolute percent values as relative change considering Team Cymru as
the base-line, i.e., positive values mean more based on Team Cymru whois data, while
negative values mean more based on historic bulk whois data.
al. presented [10]. To this end, we compared the final cloud hosting verdict for
several countries between an analysis where our historic whois service has been
used and one where Team Cymru’s whois has been used (see Fig. 3). Over all
countries in our analysis, we only observe a significant impact in the U.K. and
the U.S.. For the U.K. and the U.S., we find that, overall, the number of universi-
ties attributed to Amazon (i.e., Amazon, Amazon+Google, Amazon+Microsoft,
Amazon+Google+Microsoft) are estimated higher by data from Team Cymru’s
whois until May 2016 by around 12.5%. Additionally, we find a minor (≤5%)
underestimation for Google use in 2015, and a high overestimation of Microsoft
use in January 2015 only.
Focusing on the Amazon case, we were able to attribute it to [Link]/8, the
IPv4 address block formerly allocated to the Massachusetts Institute of Tech-
nology (MIT). In 2017, MIT announced its intent to sell large parts (87.5%) of
this address block to Amazon [30]. The transfer of addresses was finalized in
2019, with the creation of associated route objects [2], but the networks to be
sold were cleared ahead of time. As several Universities in the U.S. and U.K.
had names under their domain pointing to IP addresses from MIT, and – based
on currently accurate attribution information – these now belong to Amazon,
these addresses were wrongly attributed to Amazon.
To better understand the significance of this attribution error, we compare
the cloud usage graphs generated when using whois data sourced via the Team
218 F. Streibelt et al.
Fig. 4. Cloud use attribution for universities in the U.K. and the U.S. (January 2015–
May 2022) based on historic bulk-whois data.
Cymru whois service (see Fig. 1) with the updated version relying on our historic
whois (see Fig. 4). We find that for both countries, the U.S. and the U.K., using
the historic whois service reveals an initially lower usage of Amazon based host-
ing, followed by a more rapid increase. For example, in the U.S., we find that
the initial share of universities also using Amazon hosted services now hovers
around 60% instead of the 75% initially observed.
In the U.K. the effect has been more pronounced. Instead of the gradual
increase initially assumed based on Team Cymru’s whois data, we now a lower
share of Amazon service usage for the U.K. in 2015 (around 12.5% instead of
25%), increasing during 2016. Hence, by initially not using a historic whois ser-
vice, Fiebig et al. missed an important growth effect in the data.
4 Discussion
In this section, we discuss lessons learned for research on historic datasets, discuss
the limitations of our approach, and outline further work.
4.2 Limitations
5 Conclusion
In this paper, we introduce and evaluate BTTF whois as a public community ser-
vice. This historic whois service allows more accurate estimations of IP address
ownership, especially when the concerned IP address has been observed in the
past. Based on a case-study, we demonstrate how the use of an accurate historic
whois service allows deeper insights into datasets, and reveal developments that
would remain shrouded when only relying on current whois information.
Nevertheless, several challenges exist, which should be resolved in further
iterations of the development of our service. This includes aggregating the Route-
Views dataset ourselves – especially as older data-sets are available than aggre-
gated by CAIDA – and continuously collecting RIR provided data for generating
AS2ORG maps ourselves, including addressing the issue of organizational fami-
lies more reliably. Furthermore, future implementations should include IRR and
RPKI data to make the implementation more robust against data noise due to
prefix hijacks and the announcement of prefixes by ASes not belonging to the
prefix-holder’s organization.
Service Availability: You can use a publicly available instance of BTTF whois
at [Link] port tcp/43. See Appendix A for usage details and
[Link] for further information.
from 2022.2, via 2022.4, to 2023.2 for seeding the idea to implement the BTTF whois
service and their continuous encouragement to pursue this work. Finally, Sebastian
Lohff’s input on the implementation and performance tuning were invaluable to real-
ize the service in a production-ready manner. This work was partially funded by the
German Federal Ministry of Education and Research under the project 6G-RIC, grant
16KISK027. Any opinions, findings, and conclusions or recommendations expressed in
this material are those of the authors and do not necessarily reflect the views of Farsight
Security, Inc., DomainTools, the German Federal Ministry of Education and Research,
or the authors’ host institutions and affiliations.
Here, we document a) how you can use BTTF whois with a whois client, and
b) how to obtain bulk results. Furthermore, we provide an overview over the
returned JSON’s structure.
BTTF whois can be used with a standard whois client. The date format is
YYYYMMDD.
"source": "RIPE"
},
"seen": [
"20180703"
],
"changed": "20180703",
"change_guessed": true,
"orgs": [
{
"org_id": "@family-471",
"org": {
"org_id": "@family-471",
"org_name": "Cloudflare Inc",
"country": "US",
"source": "ARIN,RIPE"
},
"seen": [
"20180703"
],
"changed": "20180703",
"change_guessed": true
}
]
}
]
}
}
}
# READY
{"IP": "[Link]", "QDATE": "20210101", "results": {"DATA_FIRST": [...]
{"IP": "[Link]", "QDATE": "20120101", "results": []}
{"IP": "[Link]", "QDATE": "20210201", "results": {"DATA_FIRST": [...]
# goodbye
]
}
}
References
1. Jsommers, et al.: pytricia: an IP address lookup module for Python, 30 August
2022. [Link] Accessed 30 Aug 2022
2. ARIN. WHOIS for NET-18-32-0-0-1, 7 October 2019. [Link]
net/NET-18-32-0-0-1. Accessed 01 Sept 2022
3. CAIDA. Routeviews Prefix to AS mappings Dataset for IPv4 and IPv6, 30 August
2022. [Link] Accessed 30
Aug 2022
4. CAIDA. The CAIDA AS Organizations Dataset, all dates, 30 Aug 2022. https://
[Link]/data/as-organizations. Accessed 30 Aug 2022
5. Chatzis, N., Smaragdakis, G., Böttger, J., Krenc, T., Feldmann, A.: On the benefits
of using a large IXP as an Internet vantage point. In: Proceedings of the 2013
Conference on Internet Measurement Conference (2013)
6. Chen, E., Stewart, J.: A framework for inter-domain route aggregation. RFC 2519.
IETF, February 1999. [Link]
7. Daigle, L.: WHOIS protocol specification. RFC 3912. IETF, September 2004.
[Link]
8. Farsight Inc., Farsight - Security Information Exchange (SIE). [Link]
[Link]/solutions/security-information-exchange/
9. Fiebig, T., et al.: Heads in the clouds: measuring the implications of universities
migrating to public clouds. arXiv preprint arXiv:2104.09462 (2021)
10. Fiebig, T., et al.: Heads in the clouds? Measuring universities’ migration to public
clouds: implications for privacy & academic freedom. In: Proceedings on Privacy
Enhancing Technologies Symposium, vol. 2 (2023)
11. Giotsas, V., Livadariu, I., Gigis, P.: A first look at the misuse and abuse of the
IPv4 transfer market. In: Sperotto, A., Dainotti, A., Stiller, B. (eds.) PAM 2020.
LNCS, vol. 12048, pp. 88–103. Springer, Cham (2020). [Link]
978-3-030-44081-7_6
12. Giotsas, V., Luckie, M., Huffaker, B., Claffy, K.: Inferring complex AS relation-
ships. In: Proceedings of the 2014 Internet Measurement Conference (2014)
13. Haas, J., Mitchell, J.: Reservation of last autonomous system (AS) numbers. RFC
7300. IETF, July 2014. [Link]
14. Hawkinson, J., Bates, T.: Guidelines for creation, selection, and registration of an
autonomous system (AS). RFC 1930. IETF, March 1996. [Link]
[Link]
15. Hohlfeld, O.: Poster: operating a DNS-based active internet observatory. In: Pro-
ceedings of the 2018 ACM SIGCOMM Conference (SIGCOMM) (2018)
16. Housley, R., Curran, J., Huston, G., Conrad, D.: The internet numbers registry
system. RFC 7020. IETF, August 2013. [Link]
17. Huston, G.: Autonomous system (AS) number reservation for documentation use.
RFC 5398. IETF, December 2008. [Link]
18. Huston, G., Michaelson, G.: Validation of route origination using the resource
certificate public key infrastructure (PKI) and route origin authorizations (ROAs).
RFC 6483. IETF, February 2012. [Link]
A Historic IP Attribution Service for Network Measurement 225
19. Kumari, W., Bush, R., Schiller, H., Patel, K.: Codification of AS 0 processing. RFC
7607. IETF, August 2015. [Link]
20. Liu, S., Foster, I., Savage, S., Voelker, G.M., Saul, L.K.: Who is .com? Learn-
ing to parse WHOIS records. In: Proceedings of the 2015 Internet Measurement
Conference (2015)
21. Livadariu, I., Elmokashfi, A., Dhamdhere, A.: On IPv4 transfer markets: analyzing
reported transfers and inferring transfers in the wild. In: Computer Communica-
tions, vol. 111 (2017)
22. Livadariu, I., Elmokashfi, A., Dhamdhere, A., Claffy, K.: A first look at IPv4
transfer markets. In: Proceedings of the Ninth ACM Conference on Emerging Net-
working Experiments and Technologies (2013)
23. Luckie, M., Hyun, Y., Huffaker, B.: Traceroute probe method and forward IP
path inference. In: Proceedings of the 8th ACM SIGCOMM conference on Internet
measurement (2008)
24. Mitchell, J.: Autonomous system (AS) reservation for private use. RFC 6996. IETF,
July 2013. [Link]
25. Prehn, L., Lichtblau, F., Feldmann, A.: When wells run dry: the 2020 IPv4 address
market. In: Proceedings of the ACM Conference on Emerging Networking EXper-
iments and Technologies (CoNEXT) (2020)
26. Rekhter, Y., Moskowitz, B., Karrenberg, D., Groot G.J.d., Lear, E.: Address allo-
cation for private internets. RFC 1918. IETF, February 1996. [Link]
rfc/[Link]
27. Richter, P., Allman, M., Bush, R., Paxson, V.: A primer on IPv4 scarcity. ACM
SIGCOMM Comput. Commun. Rev. 45(2) (2015)
28. van Rijswijk-Deij, R., Jonker, M., Sperotto, A., Pras, A.: A high-performance,
scalable infrastructure for large-scale active DNS measurements. IEEE J. Sel. Areas
Commun. 34(6) (2016)
29. RouteViews: RouteViews Project, 30 August 2022. [Link]
Accessed 30 Aug 2022
30. Schmidt, M.A., Executive, I.R.: Letter to: to the members of the
MIT community, 20 April 2017. [Link]
e22e50cd52b7dffcf5a4db2b8ea4cce0. Accessed 01 Sept 2022
31. Sediqi, K.Z., Prehn, L., Gasser, O.: Hyper-specific prefixes: gotta enjoy the little
things in interdomain routing. ACM SIGCOMM Comput. Commun. Rev. 52(2)
(2022)
32. Sermpezis, P., Kotronis, V., Dainotti, A., Dimitropoulos, X.: A survey among net-
work operators on BGP prefix hijacking. ACM SIGCOMM Comput. Commun.
Rev. 48(1) (2018)
33. Team Cymru. IP to ASN mapping service. [Link]
services/ip-asn-mapping/
226 F. Streibelt et al.
34. Vohra, Q., Chen, E.: BGP support for four-octet autonomous system (AS) number
space. RFC 6793. IETF, December 2012. [Link]
35. Zhou, L., Kong, N., Shen, S., Sheng, S., Servin, A.: Inventory and analysis of
WHOIS registration objects. RFC 7485. IETF, March 2015. [Link]
rfc/[Link]
Open Access This chapter is licensed under the terms of the Creative Commons
Attribution 4.0 International License ([Link]
which permits use, sharing, adaptation, distribution and reproduction in any medium
or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons license and indicate if changes were
made.
The images or other third party material in this chapter are included in the
chapter’s Creative Commons license, unless indicated otherwise in a credit line to the
material. If material is not included in the chapter’s Creative Commons license and
your intended use is not permitted by statutory regulation or exceeds the permitted
use, you will need to obtain permission directly from the copyright holder.
Towards Diagnosing Accurately
the Performance Bottleneck
of Software-Based Network Function
Implementation
Ru Jia1,2,3(B) , Heng Pan1,4 , Haiyang Jiang1 , Serge Fdida3 , and Gaogang Xie5
1
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
{jiaru,panheng,jianghaiyang}@[Link]
2
University of Chinese Academy of Sciences, Beijing, China
3
Sorbonne University, Paris, France
4
Purple Mountain Laboratories, Nanjing, China
5
Computer Network Information Center, Chinese Academy of Sciences, Beijing,
China
xie@[Link]
1 Introduction
As the size of the network rapidly grows, traditional underlying network that
based on custom hardware, face significant development costs combined the low
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 227–253, 2023.
[Link]
228 R. Jia et al.
flexibility and scalability. In order to resolve the issue, network providers move
hardware middleboxes to software-based network functions (NFs) running on the
commodity servers. The softwareization of NFs improves the operation efficiency
through simpler deployment and upgrade cycles. However, software-based NFs
can lead to a significant performance issue, that is difficult to diagnose.
Compared with hardware platforms, processing packets in software means
complex running environments and more intense resource contention. Moreover,
new large-scale network scenarios, e.g., data centers, cloud platforms and the
5G mobile network [1], introduce complicated functional requirements and make
the code size of software-based NFs largely increased. For example, the popular
software packet processing framework, Cisco VPP [2], contains more than 2000
source files, 450,000 code lines. When performance issue emerge, the complicated
running environments and code structures make it difficult for developers to
quickly locate the issue. As the result, performance diagnosis, the process of
finding and explaining the performance issue in NF program, is difficult and
time-consuming in production environment.
In order to explore the performance behavior and perform NF performance
diagnosis, developers generally use the general-purpose performance diagnosis
tools based on the CPU hardware feature, Performance Monitoring Counter
(PMC). Diagnosis tools based on PMC sampling (Linux Perf [3], Intel VTune
[4]) have been widely used in NF performance diagnosis, since they are easy-to-
use and have low overhead. There are also a lot of research focusing on perfor-
mance diagnosis on High Performance Computing (HPC) and general-purpose
computing programs ( [5–12]). However, NF programs are different from general-
purpose programs. It brings new challenges to performance diagnosis. Therefore
the NF performance diagnosis is re-considered in the work and we present three
requirements critical to the issue:
1) Fine granularity. Due to the queues and batch operations in NFs, the
performance issue are transitive cross packets. To identify the root cause,
NF performance diagnosis solution should be fine-grained enough to perform
packet-level performance tracing.
2) Flexibility. The modular architecture and high-performance requirements
complicate the code structure and execution of NF. NF performance diagnosis
solution should be flexible enough to handle inline functions and libraries
loaded at run-time.
3) Perturbation-free. The performance perturbation caused by measurement
is unavoidable. To reflect the original behavior as accurately as possible, per-
formance diagnosis solution should be perturbation-free as much as possible.
Facing to these requirements for NF performance diagnosis, are these general-
purpose performance diagnosis methods still suitable for NFs? Unfortunately,
there is no evaluation of existing PMC-based performance diagnosis methods in
NF scenarios.
We investigate two major types of PMC-based performance diagnosis meth-
ods, sampling-based and instrumentation-based methods, summarize their capa-
bilities in NF performance diagnosis (Sect. 5). Through theoretical mechanism
Towards Diagnosing Accurately the Performance Bottleneck 229
2 Background
The list of features that need to be supported by network functions has grown
rapidly. New network services appear with new functional requirements, new
network protocols emerge and evolve. Due to the slow development cycle (typ-
ically years), the closed, static and inflexible hardware, can no longer support
the complex, rapidly evolving network functions [33].
Software-based NF was proposed and rapidly gained popularity. In the past
decade, software-based NF has developed from the earliest simple software switch
to complex firewall, IPsec gateway, OpenFlow [20] switch, Network Intrusion
Detection System (NIDS), etc. Compared to hardware, the NF softwareization
introduces complicated performance issue. Resource contention in the complex
run-time environment will make NF performance unstable. The inefficient code
segments also degrade performance. Even though many different diagnosis tools
have been designed, NF performance diagnosis and optimization is still heav-
ily dependent on the experience of programmers. The performance diagnosis
is the bottleneck of the NF development cycle, and full of challenges [13]. In
this section, we first briefly describe the packet processing procedure of network
functions, explain the complexity of NF performance issues in Sect. 2.1. Then
we describe conventional performance diagnosis solutions, which are commonly
adopted in NF performance diagnosis in Sect. 2.2.
Network Functions
DPDK DPDK
Ports OS Ports
Input traffic
Server Output traffic
1) The Traffic Capture module captures packets on the wire. High performance
NF commonly adopts user-space driver frameworks, e.g., DPDK, in the mod-
ule to bypass the kernel’s protocol stack and move the packet processing
entirely into user-space. The packets’ pointers are delivered through queue
data structures and the packets are processed in batch. The procedure con-
tains massive memory operations and address calculations. Any inappropriate
data structure design, e.g., unaligned memory address, is likely to cause seri-
ous performance issue.
2) The Parser module processes the network protocols, e.g., dealing with IP
fragments and TCP reassembly. Due to complex logic processing and plenty
of execution branches, parser module usually face high prediction error rate
and low instruction cache hit rates. For example, if a branch prediction is
wrong, the instructions and the correlative computational results have to be
“flushed” and replaced with correct ones.
3) Then, according to the functional requirements, NF performs rule matching
on packets. Rule matching is critical to most NFs (Table 1). The rule match-
ing is computationally intensive, and usually not a simple one-dimensional
numeric match. All of the rule sets, the traffic patterns, the design and imple-
mentation of algorithms, affect the CPU usage and data cache utilization of
rule matching, which finally will affect the performance.
4) Depending on the result of rule matching, the packet may be rewritten, and
forwarded to specified output port. The memory occupied by the packet will
be modified or copied, and at last released. Frequent memory access put
pressure on the cache. Data locality will seriously affect the utilization of
cache, which will obvious affect the performance.
0 t 2t 3t 3.5t 4t 5t 6.5t
0 1 2 3 4 5
Queue 3
(len = 2) 2 2 4
Classifier 3 3 5 5
(batch=2) 0 1 1 1 2 2 4 4
0 1 2 3 4 5
while the Classifier processing p1 , and lead to its large processing delay (2.5t).
Since the Classifier is blocked by the abnormal processing of p1 , subsequent three
packets p2 , p3 , p4 have to wait in the queue. Combined with the queueing delay,
packet p2 even experience a larger delay (3t) than p1 . In addition, due to the
batching operation of packets p4 and p5 , even if p5 doesn’t need to queue, it
experiences a delay of 1.5t, not t. In this example, only packet p1 triggers the
abnormality, but it affects the delay of the following 4 packets. So we suppose
that p1 is the culprit, and p2 , . . . , p5 are victims.
Performance diagnosis measurements need to be able to distinguish among
the processing of different packets, and find out which packets are culprits, and
which packets are just victims. It means that performance diagnosis should sup-
port packet-level performance tracing. Furthermore, packet processing usually
has many stages, the abnormal event can occur at any stage of the packet pro-
cessing. Finding the culprit packet is not enough to pinpoint abnormal events.
NF performance diagnosis should be able to identify which stage of certain
packet processing the abnormal event occurred at, such as, “the abnormal event
occurred at the classifier stage in the processing of packet p1 ”. It means that
NF performance diagnosis should support function-level performance monitor-
ing. Existing researches [13,14], which are based on queue monitoring and flow
measurement, only approximately identify the culprit packets/flows, and cannot
go deep into each stage of packet processing.
Based on this challenge, we propose the first requirement for NF performance
diagnosis, Fine granularity: NF performance diagnosis should provide packet-
level and function-level performance tracing.
Executable file
[Link] [Link] [Link] [Link] [Link]
Headers Headers
Conf. file [Link]
linking Info. of dyn.
libraries Init codes Lib0 disable
Headers Lib1 enable
Startup Headers
Info. of dyn. Lib2 disable
libraries [Link]
Headers Init codes
Running
Startup [Link]
read_conf(....);
[Link] Init codes load_modules(...); [Link]
Load
Process address Process address
Executable file Part of init. codes Process address space
space space
the NF performance diagnosis. Due to the resource contention in the real run-
ning environment, there is a big gap between the performance behavior of rule
matching in offline algorithm experiments and in online NF processing.
For example, in TupleMerge [22], which is a research on packet classification,
VPP’s TSS algorithm averages 2.93µs of packet classification time with a 256k
rule set in offline. While in the actual forwarding environment with a 4K rule
set, the time of each lookup reach 10µs-12µs. Although TulpeMerge optimizes
specifically for the difference between online and offline environments, compared
with it average classification time of 0.64µs with 256k rule set at offline, its
overhead is already between 0.55µs and 0.7µs with 4k rule set at online.
Since offline performance diagnosis cannot efficiently find performance issues
in real-world system rule matching, online performance diagnosis tools is neces-
sary. However, the interference with the original program by online measurement
is unavoidable. Performance data cannot be captured without extra measure-
ment operations. Although the PMC hardware can count the hardware perfor-
mance metrics of certain process with ignorable overhead, complete performance
diagnosis also requires the extra operations such as reading values from hardware
PMCs to user space, constructing performance data records. All of these extra
operations compete for CPU time, cache and memory, with the NF process,
disrupt the performance behavior of NF execution. For example, let’s consider
the usage of CPU caches. SEPS’17 [30] shows that, the number of mis-predicted
branches and instruction cache misses will significantly increase due to the mea-
surement operations. In NF diagnosis, we usually need to identify the culprit
packet, measurement should remain the differences of the processing of different
packets. The perturbation here is more related to the interference with these
differences. Due to the extra operations and resource contention, the differences
of the processing of different packets will be disrupted.
Based on this challenge, we propose the third requirement for NF perfor-
mance diagnosis, Perturbation-free: Performance perturbation to the packet-
level performance behavior, caused by online measurement should be as small as
possible.
4 Overview
Before stepping into our evaluation, we first introduce our experiment environ-
ment in Sect. 4.1, and evaluation methods in Sect. 4.2.
Because our aim is to analyze and evaluate the general performance diagnosis
methods in NF performance diagnosis scenario, we build a real NF forwarding
environment, as shown in Fig. 4.
We use open source DPDK-based packet processing framework, Cisco VPP,
to construct our NFs. All the NFs are running at a high performance commodity
server, with an Intel Xeon Platinum 8160 CPU @2.10GHz, and 128GB DDR3
Towards Diagnosing Accurately the Performance Bottleneck 237
DPDK DPDK
Phy port Phy port
memory. Each core is equipped with a 32KB L1 data cache and a 1024KB
L2 cache. A 33MB L3 cache is shared among all cores. The CentOS 7.9.2009
operation system with linux kernel 3.10 is installed on the server. The software
packet generator is developed by ourselves based on DPDK, to support high per-
formance packet sending and receiving with hardware timestamps. The packet
generator is running on another server which has the same configuration with
the one running NFs. To avoid the fluctuations in propagation delay, two servers
are directly connected via the high performance Mellanox MT27800 NIC. The
software we used are listed in Table. 4.
Table 4. The information of software used in the paper
We choose three NFs, Firewall (FW), NAT, IPSec gateway (IPsec) as our
target NFs. All of them are built based on Cisco VPP [2], which is one of the
most popular packet processing frameworks. And to make it easy to follow,
we choose the firewall NF as an example throughout the paper. Firewall is a
very typical and common network function, which involves all packet processing
stages described in Sect. 2.1. In Sect. 6, we also give the detailed experimental
results of NAT and IPsec.
NF performance diagnosis. However, can they meet the demand for NF per-
formance diagnosis? Unfortunately, at present, there is a lack of verification of
existing tools in NF scenarios.
In this paper, we validate and evaluate existing PMC-based performance
diagnosis tools from two perspectives in real NF forwarding scenario.
We evaluate two sampling-based methods (Perf [3] and HPCToolkit [5]) and
two instrumentation-based methods (Score-P [8] and TAU [10]) in our NF for-
warding environment. And to make it easy to follow, we choose Perf and TAU as
the representatives of two categories to discuss their details. More explanations
will be given in Sect. 5.2.
The rule sets and traffic patterns are all generated by ClassBench [36], which
has been widely used to evaluate the performance of NF and packet processing
algorithms. ClassBench produces rule sets based on seed files generated from
real rule sets, and sequences of packet headers to exercise them. The software
packet generator read the packet headers, and constructs 64-byte packets with
random payloads.
As mentioned before, we select two typical and popular tools, Perf and
TAU. The underlying mechanism of the same type of diagnosis tools is similar.
The major difference of sampling-based tools is the visual analysis and collec-
tion of metrics (see Table 5). From the perspective of PMC metrics, Perf is a
representative sampling-based tool, which is open source and widely used. In
the part of instrumentation-based tools, the differences appear in the instru-
mentation technology and the construction of probes. However, as we listed in
Table 5, most of them don’t have full support to NF performance diagnosis. For
example, HPCToolkit and Intel VTune only support measuring the functions in
external shared libraries. And Score-P cannot be applied to VPP framework,
since VPP’s complicated building environment. TAU can be directly applied to
VPP framework without any modification and able to measure most of functions,
so we choose TAU to represent instrumentation-based tools.
Sampling-based Perf generates time-aggregated performance data as shown
in Fig. 6. Each rectangle represents a function while the shade of color indi-
cates the number of performance events (CPU cycles) triggered by the func-
tion throughout the entire measurement task cycle. The up-to-down posi-
tions of the rectangles (functions) represent the function call sequence. Indeed,
Perf is able to identify hot spot functions. However, the left-to-right posi-
tions do not reflect the function execution sequence. For example, in the
Towards Diagnosing Accurately the Performance Bottleneck 241
flame graph, dpdk input node f n avx2 → acl in l2 ip4 node f n avx2 →
ethernet input node f n, does not mean their execution sequence. Limited by
the PMC sampling mechanism, Perf cannot achieve per-packet measurement
and analysis. The sampling of PMC is based on the frequency of occurrences
of hardware performance events. Specifically, only when the program triggers N
specified events, a sampling will be performed. Even with the call stack infor-
mation, the results can only reflect the total number of the event triggered by
certain function in a period of time. That said, we cannot distinguish the data of
a specific function execution, let alone the data of a specific packet processing.
clib..
clib.. singl..
multi_acl_match_get_applied_ace_index ac..
hash_multi_acl_match_5tuple acl..
acl_plugin_match_5tuple_inline fill..
acl_fa_inner_node_fn acl_.. m.. et..
mlx5_rx_burst acl_fa_outer_node_fn r.. eth_..
dpd.. rte_eth_rx_burst acl_fa_node_fn d.. t.. eth_.. l..
dpdk_device_input acl_in_l2_ip4_node_fn_avx2 dpdk_devi.. ethe.. l2..
dpdk_input_node_fn_avx2 dispatch_node
dispatch_node dispatch_pending_node
Fig. 6. The flame graph of FW based on part of data captured by Perf (cpu-cycles)
Fig. 7. The performance tracing of FW based on part of data captured by TAU (sys-
time)
also fetch the time-aggregated global hot spots and call relationships via sim-
ple calculation. But instrumentation-based tools usually have two drawbacks as
follow.
– Limited measurable ranges. Even though TAU can leverage the PMC, it can-
not measure the library functions that are dynamically loaded. For example,
ACL and DPDK-related modules are reported as the hot spots in Perf while
TAU cannot measure them.
– High measurement overhead. Compared with sampling-based performance
diagnosis methods, instrumentation-based methods introduce much greater
overall measurement overhead. So far, it is widely believed that it will severely
disrupt the performance behavior of NF, make measurements untrustworthy.
However, does the high overall overhead of the measurement lead to large
performance perturbation?
6 Performance Perturbation
Intuitively, the performance diagnosis tools will introduce extra measurement
overhead. But this leaves a question that whether higher measurement overhead
leads to larger performance perturbation. This section replies to the question.
The key reason is that those abnormal packets will lead to significant cost both
in the two scenarios. That said, the abnormal events can be reserved in the
measurement result. Logically, the measurement result can be viewed as a slow-
down of the original performance. As shown in Fig. 9, though TAU introduces
high measurement overhead increasing the average latency from 5.94µs up to
89.90µs, the similarity is reserved.
We believe that high performance distribution similarity means low perfor-
mance perturbation. This is because it is possible to infer to the original perfor-
mance based on the measurement result.
1) Latency is one of the most important performance metrics for NF, and is able
to clearly characterize NF packet-level performance.
2) For NF diagnosis, we usually focus on the related performance data. For
example, to identify the costliest function, we need to distinguish who has
a relatively high resource consumption. To identify the reason of fluctua-
tions, we need to identify which packet took longer process, and which part
of code was unstable. Therefore, the diagnosis methods should remain the
fine-grained differences (relative state) of the processing. Even if the overall
overhead is high, as long as the relative state can be maintained, we can still
locate the culprit. The global performance metrics, such as throughput, can-
not describe fine-grained performance differences. Packet-level performance
metric, latency and its distribution, can clearly describe the packet-level dif-
ferences of the processing, as we described in Sect. 6.1.
3) Our goal is to evaluate performance perturbation accurately. If the measure-
ment of the metrics used to evaluate the perturbation introduces new pertur-
bation, the results will become unreliable. The measurement of latency can be
executed in the side of packet generator, isolated from the execution of NF.
In contrast, the measurement of PMC hardware performance metrics must be
executed in the same environment with NF. Fine-grained PMC measurement
will introduce new perturbation.
For an input packet set P = {p0 , p1 , . . . , pn }, latM (pi ) is denoted as the pro-
cessing latency of the packet pi when the measurement operations are activated.
Likewise, lat(pi ) refers to the process latency of packet pi without any measure-
ment operation. We normalize the latency data, and use E[latM ] and E[lat] to
respectively refer to their average values. With this basis, for a packet pi , we
define its “latency distance” as follow.
Thus, we further define the coefficient of interference— the sum of each packet
“latency distance”. Specifically, for the input packet set P , the corresponding
CoIP is calculated as follow.
n
CoIP = Di (2)
i=0
It is clear that, for one NF running under the same configurations, the less
the CoI is, the smaller performance perturbation is. Logically speaking, it is
impossible to achieve the ideal zero-performance-perturbation measurement due
to a few complex influence factors, such as cache contention.
Fig. 11. The coefficient of interference of different performance diagnosis tools on dif-
ferent NFs
We apply these tools to 4 NFs constructed by VPP while the results are
shown in Fig. 11. FW 1k is a firewall with 1k rules while FW 3k is the firewall
with 3k rules. Both of the rule sets are generated by ClassBench [36]. NAT is
a simple SNAT that translates the source IP addresses based on the longest
prefix match rules. IPSec provides the authentication of IP packets based on the
Authentication Header protocol (AH) while packets are classified by 100 5-tuple
rules, and then perform authentication with the SHA1-96 cryptographic hash
algorithm.
Each performance diagnosis tool can be configured into two modes: the
measurement with high overhead ( hi) and the measurement with low over-
head ( lo). The overall overhead in the Fig. 11 is calculated as follows. The
Latencywith measurement represents the average latency of NFs with the measure-
ment. And Latency represents the average latency without the measurement.
Towards Diagnosing Accurately the Performance Bottleneck 247
Latencywith measurement
Overhead = −1 (6)
Latency
For all NFs, even running the sampling-based tools with low overhead, the
coefficient of interference is much larger than that of TAU. For example, let’s
consider FW 3k in Fig. 11. The overhead of Perf lo is very low (0.020) comparing
with TAU hi (14.240). On the contrary, the CoI of Perf lo (0.713) is 1.412 times
as much as that of TAU hi (0.505). For all cases, the CoI in TAU is only 7.39%
to 74.31% of that in Perf. In summary, even though instrumentation-based per-
formance diagnosis tools introduce larger overall overhead than sampling-based
tools, their performance perturbation is less than sampling-based tools in NF
scenario.
The Reason for Why Sampling-Based Tools Lead to Large Perfor-
mance Perturbation. Recall that the principle of sampling-based performance
diagnosis (see Sect. 5.1) shows that the PMC sampling is based on the frequency
of the triggered performance events. That said, the location where extra mea-
surement operations are performed in the target program cannot be controlled;
the number of samples happened in each packet processing is unpredictable and
uneven.
Fig. 13. The mapping between the performance tracing data and the RX&TX times-
tamps
diagnosis, but it still fails to measure the library functions that are dynamically
loaded when running NFs. In addition, for multi-core multi-thread programs,
only the PMC on the master core can be manipulated correctly. Thanks to the
modern instrumentation technology, it has already supported the instrumenta-
tion for run-time multi-core multi-threading processes. With this basis, we can
effectively alleviate the limitations of TAU. In addition, the probe also has a
large optimization space, such as replacing the syscall-based perf event with
the user-mode instruction supported by the newer Linux kernels.
Based on the above design principles, we plan to design a packet-level per-
formance diagnosis method based on modern instrumentation technology, and
build well-defined lightweight measurement probes in the future.
8 Related Work
9 Conclusions
References
1. Faqir, Z.Y., Michael, B., Sibylle, S., Fabian, S.: NFV and SDN-Key technology
enablers for 5G networks. IEEE J. Sel. Areas Commun. 35(11), 2468–2478 (2017)
2. Cisco: vector packet processing (2022). [Link]
3. Linux Community: perf: Linux profiling with performance counters (2009). https://
[Link]/[Link]/Main Page
4. Intel Corporation: intel VTune performance analyzer (2022). [Link]
com/content/www/us/en/develop/documentation/vtune-help/[Link]
5. Laksono, A.S.B., Michael, F., Mark, K., Gabriel, M., John, M., Nathan, R.T.:
HPCTOOLKIT: tools for performance analysis of optimized parallel programs.
Concurr. Comput. Pract. Exper. 22(6), 685–701 (2009)
6. Pengfei, S., Shuyin, J., Milind, C., Xu, L.: Pinpointing performance inefficiencies
via lightweight variance profiling. In: Proceedings of the International Conference
for High Performance Computing, Networking, Storage and Analysis, SC2019, pp.
1–19. Association for Computing Machinery, Denver, Colorado (2019)
7. Qidong, Z., Xu, L., Milind, C.: DrCCTProf: a fine-grained call path profiler for
ARM-based clusters. In: Proceedings of the International Conference for High Per-
formance Computing, Networking, Storage and Analysis, SC2020, pp. 1–16. IEEE
Press, Atlanta, GA, USA (2020)
8. Andreas, K., et al.: Score-P: a joint performance measurement run-time infrastruc-
ture for Periscope, Scalasca, Tau, and Vampir. In: Brunst, H., Müller, M., Nagel,
W., Resch, M. (eds.) Tools for High Performance Computing 2011. LNCS, pp.
79–91. Springer, Heidelberg (2011). [Link] 7
9. Markus, G., Felix, W., Brian, J.N.W., Erika, Á’., Daniel, B., Bernd, M.: The
Scalasca performance toolset architecture. Concurr. Comput. Pract. Exper. 22(6),
702–719 (2010)
10. Sameer, S.S., Allen, D.M.: The TAU Parallel Performance System. Int. J. High
Perform. Comput. Appl. 20(2), 287–311 (2006)
11. David, B., et al.: Caliper: performance introspection for HPC software stacks. In:
Proceedings of the International Conference for High Performance Computing,
Networking, Storage and Analysis, SC2016, pp. 550–560. IEEE Press, Salt Lake
City, UT, USA (2016)
12. Nicholas, N., Julian, S.: Valgrind: a framework for heavyweight dynamic binary
instrumentation. In: Proceedings of the 28th ACM SIGPLAN Conference on Pro-
gramming Language Design and Implementation, PLDI2007, pp. 89–100. Associ-
ation for Computing Machinery, San Diego, California, USA (2007)
13. Junzhi, G., Yuliang, L., Bilal, A., Aman, S., Minlan, Y.: Microscope: queue-based
performance diagnosis for network functions. In: Proceedings of the 2020 Confer-
ence of the ACM Special Interest Group on Data Communication, SIGCOMM2020,
pp. 390–403. Association for Computing Machinery, Virtual Event, USA (2020)
14. Yiran, L., Liangcheng, Y., Vincent, L., Mingwei, X.: PrintQueue: performance
diagnosis via queue measurement in the data plane. In: Proceedings of the 2022
Conference of the ACM Special Interest Group on Data Communication, SIG-
COMM2022, pp. 516–529. Association for Computing Machinery, Amsterdam,
Netherlands (2022)
15. Luis, P., Rishabh, I., Arseniy, Z., Jonas, F., Katerina, A.: Automated synthesis of
adversarial workloads for network functions. In: Proceedings of the 2018 Conference
of the ACM Special Interest Group on Data Communication, SIGCOMM2018, pp.
372–385. Association for Computing Machinery, Budapest, Hungary (2018)
252 R. Jia et al.
16. Rishabh, I., Luis, P., Arseniy, Z., Solal, P., Katerina, A., George, C.: Performance
contracts for software network functions. In: 16th USENIX Symposium on Net-
worked Systems Design and Implementation, NSDI2019. USENIX Association,
Boston, MA, USA (2019)
17. Xiaoqi, C., et al.: Fine-grained queue measurement in the data plane. In: Proceed-
ings of the 15th International Conference on Emerging Networking Experiments
And Technologies, CoNEXT2019, pp. 15–29. Association for Computing Machin-
ery, Orlando, Florida (2019)
18. Vimalkumar, J., Mohammad, A., Yilong, G., Changhoon, K., David, M.: Millions
of little minions: using packets for low latency network programming and visibility.
In: Proceedings of the 2014 Conference of the ACM Special Interest Group on Data
Communication, SIGCOMM2014, pp. 3–14. Association for Computing Machinery,
Chicago, Illinois, USA (2014)
19. John, S., Oliver, M., Adam, J.A., Eric, K., Jonathan, M.S.: Scaling hardware accel-
erated network monitoring to concurrent and dynamic queries with *flow. In: Pro-
ceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference,
ATC2018, pp. 823–835. USENIX Association, Boston, MA, USA (2018)
20. Nick, M., et al.: OpenFlow: enabling innovation in campus networks. SIGCOMM
Comput. Commun. Rev. 38(2), 69–74 (2008)
21. Srinivasan, V., Suri, S., Varghese, G.: Packet classification using tuple space search.
SIGCOMM Comput. Commun. Rev. 29(4), 135–146 (1999)
22. James, D., et al.: TupleMerge: fast software packet processing for online packet
classification. IEEE/ACM Trans. Networking 27(4), 1417–1431 (2019)
23. Xinyi, Z., Xie, G., Xin, W., Penghao, Z., Li, Y., Kavé, S.: Fast online packet
classification with convolutional neural network. IEEE/ACM Trans. Netw. 29(6),
2765–2778 (2021)
24. Sorrachai, Y., James, D., Alex, X.L., Eric, T.: A sorted partitioning approach to
high-speed and fast-update OpenFlow classification. In: 2016 IEEE 24th Interna-
tional Conference on Network Protocols, ICNP2016, pp. 1–10. IEEE, Singapore
(2016)
25. Kirill, K., Sergey, I.N., Ori, R., William, C., Patrick, E.: Exploiting order inde-
pendence for scalable and expressive packet classification. IEEE/ACM Trans. Net-
working 24(2), 1251–1264 (2015)
26. Vincent, M.W., Sally, A.M.: Can hardware performance counters be trusted? In:
2008 IEEE International Symposium on Workload Characterization, pp. 141–150
(2008)
27. Dmitrijs, Z., Milan, J., Matthias, H.: Accuracy of performance counter measure-
ments. In: 2009 IEEE International Symposium on Performance Analysis of Sys-
tems and Software, ISPASS2009, pp. 23–32. IEEE, Boston, Massachusetts (2009)
28. Todd, M., Amer, D., Matthias, H., Peter, F.S.: Understanding Measurement Per-
turbation in Trace-based Data. In: 2007 IEEE International Parallel and Dis-
tributed Processing Symposium, IPDPS2007, pp.1–6. IEEE, Long Beach, Cali-
fornia (2007)
29. Matthias, W., et al.: Detection and visualization of performance variations to guide
identification of application bottlenecks. In: 2016 45th International Conference on
Parallel Processing Workshops, ICPPW2016, pp. 289–298. IEEE, Philadelphia,
PA, USA (2016)
30. Lehr, J.-P., Iwainsky, C., Bischof, C.: The influence of HPCToolkit and Score-p
on hardware performance counters. In: Proceedings of the 4th ACM SIGPLAN
International Workshop on Software Engineering for Parallel Systems, SEPS2017.
Association for Computing Machinery, Vancouver, BC, Canada (2017)
Towards Diagnosing Accurately the Performance Bottleneck 253
31. Srikanth, K., Ratul, M., Patrick, V., Sharad, A., Jitendra, P., Paramvir, B.:
Detailed diagnosis in enterprise networks. In: Proceedings of the ACM SIGCOMM
2009 Conference on Data Communication, SIGCOMM2009, pp. 243–254. Associ-
ation for Computing Machinery, Barcelona, Spain (2009)
32. Ben, P., et al.: The design and implementation of open vSwitch. In: 12th USENIX
Symposium on Networked Systems Design and Implementation, NSDI2015, pp.
117–130. USENIX Association, Oakland, CA (2015)
33. Eddie, K., Robert, M., Benjie, C., John, J., Marinus, F.K.: The click modular
router. ACM Trans. Comput. Syst. 18(3), 263–297 (2000)
34. Buck, B., Hollingsworth, J.K.: An API for runtime code patching. Int. J. High
Perform. Comput. Appl. 14(4), 317–329 (2000)
35. Derek, B., Qin, Z., Saman, A.: Transparent dynamic instrumentation. In: Proceed-
ings of the 8th ACM SIGPLAN/SIGOPS conference on Virtual Execution Envi-
ronments, VEE2012, pp. 133–144. Association for Computing Machinery, London,
England, UK (2012)
36. David, E.T., Jonathan, S.T.: ClassBench: a packet classification benchmark.
IEEE/ACM Trans. Networking 15(3), 499–511 (2007)
37. Sangjin, H., Keon, J., Aurojit, P., Shoumik, P., Dongsu, H., Sylvia, R.: SoftNIC:
a software NIC to augment hardware. Technical Report No. UCB/EECS-2015-
155 (2015). [Link]
html
Network Performance
Evaluation of the ProgHW/SW
Architectural Design Space of Bandwidth
Estimation
1 Introduction
Bandwidth estimation (BWE) is an essential functionality used in various net-
work fields ranging from cloud applications to congestion control [3,4,27,33,46].
For example, BWE can improve the performance of Hadoop by optimizing the
bandwidth utilization among a group of virtual machines (VMs) [27]. However,
inaccurate BWE can cause packet loss and degraded throughput [24,46]. There-
fore, how to improve BWE accuracy is an ever-lasting research topic.
While researchers and engineers have conducted many studies and eval-
uations of BWE, most of them focus on either the algorithmic designs and
parameters of BWE [34,41,42,44,48] or the impact of network conditions on
BWE [10,22,42], but the architectural optimizations have not been addressed
adequately. Nevertheless, as networks become faster (e.g., 10 Gbps, 100 Gbps)
and more complex nowadays, the architectural aspects of BWE become increas-
ingly important [31]. Consider time precision as an example. BWE relies heavily
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 257–283, 2023.
[Link]
258 T. Fang et al.
The main cause for the inadequate architectural evaluation is that the evalu-
ation of the ProgHW/SW space is more laborious and challenging than evalua-
tions carried out on pure SW. There are two reasons. First, the evaluation period
of ProgHW/SW often takes much longer time than that of pure SW. Compared
with SW-level compilation, ProgHW-level compilation or ProgHW offloading has
many extra steps, such as netlist synthesis, place and route, and timing, power,
and area constraints. These extra steps are time-consuming. Second, ProgH-
W/SW space provides more combinations than pure SW, which increases the
workload of the architectural comparison. Specifically, for one endpoint, a dif-
ferent BWE algorithm may have different components, and each component can
choose to stay at either ProgHW or SW to make up a different architecture. For
two endpoints, a sender and a receiver can use different architectures to execute
BWE. These diverse combinations increase the number of evaluation cases.
We deal with the ProgHW/SW evaluation challenges by leveraging an insight:
most factors that impact BWE performance can be attributed to a single factor:
inter-packet delay (IPD), and architectures of different BWE algorithms have
Evaluation of the ProgHW/SW Architectural Design Space 259
many common components related to IPD. Therefore, we classify and study dif-
ferent ProgHW/SW architectures based on how they impact IPD. IPD describes
the difference among packet latencies, which differentiates our classification from
other works that focus on the absolute values of packet latencies [21,47]. In
addition, we modularize the BWE components that process and transmit IPD
information and reuse those modules in different architectures, thereby saving
time from ProgHW-level compilation and reducing evaluation difficulty.
Our contributions are summarized as follows:
– In terms of the evaluation object, we build an IPD-based classification to sys-
tematically evaluate different ProgHW/SW architectures of BWE. Further-
more, we study topics that have not been addressed well, such as heteroge-
neous combinations and the comparison between SW-level and ProgHW-level
optimization techniques. In addition, we make several new findings as follows:
• If the limited supply of cloud ProgHW only allows one end (sender or
receiver) to be deployed with a new BWE ProgHW/SW configuration,
then deployment on the receiver alone can achieve a similar effect to the
two-end configuration and can only consume half the ProgHW resources
at the same time.
• Although ProgHW/SW designs and pure SW designs have comparable
performance in low-speed networks, the former show much better per-
formance in high-speed networks. Specifically, in a 100 Gbps network,
ProgHW/SW designs can improve average IPD accuracy by 45% (max:
64%) and average BWE accuracy by 20% (max: 35%).
• Although offloading modules from SW to ProgHW can have better timing
accuracy, offloading more does not necessarily achieve better performance
in BWE. We find that the offloading of modules that directly update IPD
can maximize BWE accuracy.
– In terms of the evaluation methodology, we propose an IPD-modular method
that improves the evaluation efficiency by multiple times.
– We implement BWE modules in ProgHW and make them meet the require-
ments of BWE evaluations and be portable across different architectures.
The paper is organized as follows: Sect. 2 introduces our motivation and the
background of BWE. Section 3 presents our IPD-based classification of archi-
tectures. Section 4 presents the implementation details of our modularization.
Section 5 evaluates ProgHW/SW architectural space for BWE. Section 6 sum-
marizes related works and Sect. 7 concludes the paper.
2.1 Motivation
Format
Symbol(i)l i denotes the i-th packet and l denotes any of four locations: sSW (sender’s
SW), rSW (receiver’s SW), sHW (sender’s NIC port), and rHW (receiver’s
NIC port)
ΔSymbol(i)l Symbol(i + 1)l − Symbol(i)l
Symbol
t(i)l Measured timepoint of i-th packet at location l
t(i)R
l Real timepoint of i-th packet at location l
tdr(i)l Clock drift at l: t(i)l − t(i)R
l
d(i)SW Measured delay of i-th packet from sender to receiver by SW:
t(i)rSW − t(i)sSW
d(i)HW Measured delay of i-th packet from sender to receiver by HW:
t(i)rHW − t(i)sHW
d(i)R Real delay: t(i)R R
rHW − t(i)sHW
IP D(i)l IP D(i)l = Δt(i)l = t(i + 1)l − t(i)l
ΘIP D(i)SW IP D(i)rSW − IP D(i)sSW
ΘIP D(i)HW IP D(i)rHW − IP D(i)sHW
n(i)R
s Real delay from SW to ProgHW on sender
n(i)R
r Real delay from ProgHW to SW on receiver
by cross traffic, then the rest capacity for our usage is called available bandwidth.
While the first two metrics only consider network speed, the achievable through-
put also considers an endpoint’s processing speed and protocols. In a nutshell,
achievable throughput indicates the maximum throughput that a system can
achieve under a given protocol, network speed, and processing speed.
From the perspective of working principles, BWE algorithms can be classified
into the packet-pair type and the packet-train type as shown in Fig. 2. The details
of each algorithm in Fig. 2 can be found in [25]. The two types differ in timing
features but share a basic idea: a sender sends out a set of packets with a pre-
defined timing feature. If the sending rate exceeds the bandwidth, a receiver will
detect a change in the timing feature when those packets arrive. BWE algorithms
iteratively adjust sending rates to find the turning point where the change occurs,
and the turning point is the estimated bandwidth value.
We first define IPD as follows. Examples of IPD such as IP D(i)sSW and
IP D(i)rSW are shown in Fig. 2.
Definition 1. Inter-packet Delay (IPD) is the timing delay between any two
consecutive packets.
For the packet-pair type, a pair of packets are sent out back-to-back or in a
pre-defined IPD value. Then, a receiver captures the receiving IPD and compares
it with the pre-defined IPD to infer bw-capa or avai-bw. The formal expression
is as follows:
Definition 2. A packet-pair BWE algorithm is a function of the difference
between the receiving and sending IPDs. Assume there are n pairs of packets
(i.e., 2×n packets) in transmission, the estimated value BW is as follows:
262 T. Fang et al.
We identify and modularize common BWE components that process and trans-
mit IPD information, and our classification of different ProgHW/SW archi-
tectures is based on different allocations of those modules. Specifically, there
are three modules for the sender: packet generator, IPD modulator, and IPD
transceiver, and three modules for the receiver: IPD gauge, IPD transceiver, and
IPD processor. In a BWE process, a sender uses an IPD modulator to set pre-
defined timing features, and a receiver uses an IPD gauge and an IPD processor
to measure and analyze the change in those timing features. We will illustrate
those modules by using the traditional SW architecture of type 1.
Type 1 (No IPD Optimization). This type does not involve any specialized
optimization to improve IPD accuracy. One feature of type 1 is that most BWE
modules locate in user space. The traditional SW architecture is a representa-
tive of this type, whose architecture is shown in Fig. 3. There are several steps
in a complete procedure of BWE. On a sender, the packet generator generates a
sequence of packets. Then, the sender’s IPD modulator specifies the IPD infor-
mation of packets through a system timer. Next, these packets pass through
the IPD transceiver whose major component is the TCP/IP stack, and they
reach the MAC (Ethernet) TX port. After being transmitted through the net-
work path, these packets reach a receiver. On the receiver, the IPD transceiver
uploads IPD information, and then the IPD gauge module measures the IPD of
packets. Lastly, the measured IPD is used by the IPD processor to infer network
bandwidth.
264 T. Fang et al.
Type 2 (IPD Noise Mitigation). The feature of this type is that both IPD
modulator and IPD gauge are still in user space as type 1, but architectures are
improved to mitigate timing noise. There are several choices to do the mitigation:
TCP/IP stack offloading, kernel bypass, or BWE functions offloading as shown
in Fig. 4. Both TCP/IP stack offloading and kernel bypass aim to reduce the
number of data copies in packet transmission so that timing is more stable,
and performance is better. The difference between these two is that TCP/IP
offloading moves TCP/IP stack down to ProgHW while kernel bypass moves it
up to user space. TCP/IP offloading has several related works [7,39].
BWE functions offloading, to the best of our knowledge, has rarely been
studied, so we implement our custom version of this architecture by offloading
the packet generator and IPD processor modules down to ProgHW. The imple-
mentation details are presented in Sect. 4. BWE functions offloading is based on
TCP/IP offloading, which means that TCP/IP stack is also offloaded to ProgHW
in the BWE functions offloading architecture.
Evaluation of the ProgHW/SW Architectural Design Space 265
Type 3 (IPD HW Modulation). The feature of this type is that both the
IPD modulator and IPD gauge are placed close to NIC port. The purpose of
such placement is to restore IPD information tampered by the timing noise of
the ProgHW-SW transmission path. Specifically, both the IPD modulator and
IPD gauge modules adopt a HW timer rather than a SW timer to improve timing
stability. On the sender, ProgHW modulates the IPD of packets according to
the IPD specification of SW. On the receiver, ProgHW records receiving IPD
and sends the IPD information to SW.
This type uses the combination of stream control signals and the timer of
ProgHW to achieve accurate IPD modulation and gauge [16,20]. Specifically, on
the sender, a HW timer is used to measure the delay of packet transmission.
Then, the delay is compared with a specified IPD. If the delay is smaller, stream
control signals block the following packets until the specified IPD is reached. On
the receiver, a look-up table is dedicated to storing IPD information of packets.
The receiving IPD information is then uploaded to the IPD processor (Fig. 5).
state that the specified or measured relative timing values are equal to the real
values.
Definition 4. Packet-pair timing accuracy requirement: Assume there are n
pairs of packets (2×n packets) in transmission, the requirement is shown below.
Note that clock jitter is different from clock drift. Clock jitter means temporal
timing variation while clock drift means spatial timing variation. According to
the past study [11], the clock jitter of ProgHW is around several picoseconds,
which accounts for less than 0.1% of IPD measurement.
Theorem 1. ∀i ∈ [1, n], k = 2i − 1, if IP D(k)sSW = IP D(k)sHW and
IP D(k)rSW =IP D(k)rHW , then the Packet-pair timing accuracy requirement
(Definition 4) can be satisfied.
268 T. Fang et al.
The One-way delays d(i + 1)SW and d(i)SW can be expressed in the form of
timepoints:
4 Implementation of Modules
According to Sect. 3, different architectures share many modules in common to
process and transmit IPD information, and this section introduces our imple-
mentation details of those modules. We have two requirements for those modules:
(1) they are portable across different architectures, and (2) they are suitable for
BWE evaluations. The IPD transceiver module [7,8] and packet generator [16,40]
of existing works satisfy those requirements, but the IPD modulator, IPD gauge,
and IPD processor do not, so we focus on the last three. We use our IPD pro-
cessor to make up the architecture of BWE functions offloading (type 2), and
we use our IPD modulator and IPD gauge to make up the architecture of IPD
HW modulation (type 3).
FIFO size of these modules. The default size is 8 packets × 64 bytes, which is not
large enough to hold hundreds of packets in some packet-train BWE algorithms,
such as pathload [23]. Thus, we resize the FIFO to 1000 packets × 1500 bytes
which are larger than the maximum value of most BWE algorithms. Second, to
avoid overflow, we set a proper bit width for both sender’s and receiver’s IPD
arrays. These arrays are used to store sending and receiving IPDs. According to
our study, both the packet-train and the packet-pair types spend less than 30 min
doing estimation and the default timing precision of two FPGA boards is 8 ns [2,
50], which means that the width should be at least 38 bits to store IPDs (30 mins
×60 × 109 /8 ns < 238 ). In addition, for the IPD modulator, we use the read-valid
signal of AXI-Stream to make sure that packet transmission follows the specified
IPD. For the IPD gauge, we use the transmission-last signal of AXI-Stream to
record the arrival time of a new packet. We use our implementation of IPD
modulator and IPD gauge to build the IPD HW modulation architecture (type
3). Besides, we use BRAM to implement packet FIFOs for packet transmission
among FPGA modules.
5 Evaluation
We use Alveo U280 FPGA [2] and NetFPGA-SUME [50] to examine different
ProgHW/SW architectures. Alveo U280 FPGA is used for 100 Gbps experiments
in Open Cloud Testbed (OCT) [30]) while NetFPGA-SUME is used for less than
or equal to 10 Gbps experiments in our local testbed built with mininet version
2.2.2. Specifically, we deploy two Alveo U280 FPGA boards on two VMs of
OCT and each is equipped with 32 Virtual CPU cores and 64GB RAM. We
deploy NetFPGA-SUME on a Dell Precision 3630 machine with Intel Xeon-E5
16 Cores and 64GB RAM. The network topology of NetFPGA-SUME is shown
in Fig. 8 where two nodes on the top generate cross traffic and two nodes on the
bottom run BWE algorithms and optionally run other concurrent applications.
The testbed for Alveo U280 is similar to Fig. 8 with two differences. First, the
bottleneck link is a Dell Z9100 100G switch. Second, because a user can create
no more than two nodes in OCT, there are no cross-traffic nodes (i.e., no Node2
and Node3 in Fig. 8). The Operating system is Ubuntu 2020.4 LTS with the
Linux kernel 5.4. To compile ProgHW source code, we use Xilinx Vivado Design
Suite v2020.1. We choose two representative BWE algorithms: bprobe [13] of the
packet-pair type and pathload [23] of the packet-train type for our experiments.
272 T. Fang et al.
are shown in Fig. 9. We use the formula below to define the IPD measurement
error (IP Derr ) where #IPD is the number of measured IPD samples. For each
experiment set, we collect 40 samples. Furthermore, we check the cost efficiency
of each architecture by metrics of offloading workload and ProgHW resources
consumption. ProgHW resources are described by three key metrics: the number
of Look-up Tables (LUTs), Flip-flops (FFs), and BRAMs. The results are shown
in Table 3.
1 D
#IP
IP Derr = · (IP D[i] − IP Dactual )2
#IP D i=1
For both sender and receiver, according to Fig. 9, IPD noise mitigation (type
2) and IPD HW modulation (type 3) have better IPD accuracy than the pure
SW architecture, especially in short IPD (e.g., 120ns). The advantage of type 2
comes from the kernel-bypass effect, which requires less data copy from SW to
ProgHW. However, this effect cannot completely remove the IPD noise of systems.
Type 3 shows better IPD accuracy than type 2, and the former can keep IPD error
within 1%. The main reason is that accessing the ProgHW timer is more stable
than accessing the SW timer. From the experiment, we find that IPD restoration
is more effective than noise mitigation, and this result is consistent with our analy-
sis in Sect. 3.2. In addition, we also find that type 3 does not completely remove IPD
measurement errors. This is because the type 3 design needs 1 clock cycle to read
the measured IPD to a register. In terms of cost effectiveness, we find that more
offloading does not necessarily lead to better performance in ProgHW/SW designs.
As shown in Table 3, although Combov uses 70% fewer resources than BWE func-
tions offloading architecture, it can achieve 20% more accurate IPD than BWE
functions offloading. Thus, we suggest engineers prioritize offloading modules that
can directly update or recover IPD.
274 T. Fang et al.
feature of ProgHW. We also find that relocating TCP/IP alone (e.g., TCP/IP
offloading or DPDK) is not enough to achieve the best BWE accuracy. The rea-
son is that relocating TCP/IP cannot completely remove timing noise in user
space. However, the main advantage of DPDK is that it generally requires less
time for development and deployment since it is a SW-based solution. Further-
more, we find that the packet-pair type gets more performance improvement
than the packet-train type in type 2 and 3. This is probably because the packet-
pair type only uses a single IPD sample rather than multiple samples to estimate
bandwidth, so the packet-pair type is more sensitive to timing noise compared
with the packet-train type.
BW Eacc = 1 − BW Eerr
1 E
#BW
BW Eerr = · (BW Ei − BW Eactual )2
#BW E i=1
276 T. Fang et al.
One possible explanation for this phenomenon is that the duty of the sender
and the receiver is different. The sender aims to saturate avai-bw by continuously
increasing the rate to transmit packets. The receiver uses IPD information to
calculate bandwidth. The SW-based sender uses interrupt coalescing [35] and the
Combov sender uses packet buffering. Both those techniques can achieve fast rate
to saturate avai-bw, so the sender replacement does not have a significant dif-
ference. However, interrupt coalescing on the receiver can damage each packet’s
timing information, which reduces BWE accuracy. If limited ProgHW resources
only allow one endpoint to use ProgHW, then deployment on the receiver may
achieve better performance than on the sender.
Group3 - Summary: The receiver side has a larger impact on BWE per-
formance than the sender side in terms of ProgHW/SW configurations. This
finding is useful to save costs. Specifically, we can assign ProgHW/SW configu-
rations to the receiver alone to achieve comparable performance to the two-end
configurations.
ulate both CPU-intensive and I/O-intensive applications. We set 50% CPU load
and 10% memory load for a CPU-intensive application and 10% CPU load and 50%
memory load for an I/O-intensive application. The results are shown in Fig. 10c.
We find that type 3 can greatly resist the influence of either CPU-intensive
or memory-intensive applications compared with the other two types. One pos-
sible reason is that its IPD restoration mechanism can correct timing errors
caused by concurrent applications in SW. Type 2 can also resist the influence
to some extent. Type 2 offloads components down to ProgHW, so it becomes
less dependent on CPU for network-related operations. In addition, we find that
memory-intensive applications are more impactful than CPU-intensive applica-
tions on SW. Specifically, the former degrades SW performance by 26% while
the latter degrades it by 16%.
than 40% estimation error is produced. One possible reason is that the irregular
insertion of cross packets into probing packets can interfere with BWE.
Group5 - Summary: Type 2 and 3 show better BWE accuracy than type 1
under the influence of cross traffic.
6 Related Work
6.1 Related Evaluations of Bandwidth Estimation
7 Conclusion
In the paper, we provide an IPD-based modular method to systematically clas-
sify and evaluate ProgHW/SW space of BWE. This evaluation method shows
higher efficiency than traditional evaluation methods. Furthermore, we make
some new findings from the architectural evaluation. According to our experi-
ment results, the IPD HW modulation architecture shows the best improvement
in BWE performance. Specifically, it can increase IPD accuracy by 45% and
BWE accuracy by 20–30% in a 100 Gbps network. We also find that the receiver
side affects BWE more than the sender side. In the future, we plan to extend
ProgHW/SW space study to more types of BWE algorithms.
Acknowledgement. The work presented in this paper was supported in part by NSF
CNS-1616087 and CNS-2135539.
References
1. AXI Reference Guide (2020). [Link]
documentation/ug761 axi reference [Link]
2. Alveo U280 Product Brief (2021). [Link]
kits/alveo/[Link]
3. Apache Hadoop (2021). [Link]
4. AWS High Performance Computing (2021). [Link]
5. Intel. Intel DPDK: Data Plane Development Kit (2021). [Link]
6. Lookbusy - a synthetic load generator (2021). [Link]
Evaluation of the ProgHW/SW Architectural Design Space 281
27. LaCurts, K., Deng, S., Goyal, A., Balakrishnan, H.: Choreo: network-aware task
placement for cloud applications. In: Proceedings of the Conference on Internet
Measurement Conference, pp. 191–204 (2013)
28. Larsen, S., Sarangam, P., Huggahalli, R., Kulkarni, S.: Architectural breakdown of
end-to-end latency in a TCP/IP network. Int. J. Parallel Program. 37(6), 556–571
(2009)
29. Lee, K.S., Wang, H., Weatherspoon, H.: Sonic: precise realtime software access and
control of wired networks. In: USENIX Symposium on Networked Systems Design
and Implementation (NSDI), pp. 213–225 (2013)
30. Leeser, M., Handagala, S., Zink, M.: FPGAs in the Cloud. Authorea (2021).
[Link]
31. Li, Y., et al.: HPCC: high precision congestion control. In: Proceedings of the
ACM Special Interest Group on Data Communication, pp. 44–58. Association for
Computing Machinery (2019)
32. Liao, G., Znu, X., Bnuyan, L.: A new server I/O architecture for high speed net-
works. In: The 17th International Symposium on High Performance Computer
Architecture, pp. 255–265. IEEE (2011)
33. Lin, W., Liang, C., Wang, J.Z., Buyya, R.: Bandwidth-aware divisible task schedul-
ing for cloud computing. Softw. Pract. Exp. 44(2), 163–174 (2014)
34. Liu, X., Ravindran, K., Loguinov, D.: Evaluating the potential of bandwidth esti-
mators. In: The 4th New York Metro Area Networking Workshop (NYMAN) (2004)
35. Moreno, V., del Rio, P.M.S., Ramos, J., Garnica, J.J., Garcia-Dorado, J.L.: Batch
to the future: analyzing timestamp accuracy of high-performance packet I/O
engines. IEEE Commun. Lett. 16(11), 1888–1891 (2012)
36. Putnam, A., et al.: A reconfigurable fabric for accelerating large-scale datacenter
services. In: 2014 ACM/IEEE 41st International Symposium on Computer Archi-
tecture (ISCA), pp. 13–24 (2014)
37. Ramos, J., del Rı́o, P.S., Aracil, J., de Vergara, J.L.: On the effect of concur-
rent applications in bandwidth measurement speedometers. Comput. Netw. 55(6),
1435–1453 (2011)
38. Ribeiro, V.J., Riedi, R.H., Baraniuk, R.G., Navratil, J., Cottrell, L.: Pathchirp:
efficient available bandwidth estimation for network paths. In: Passive and Active
Measurement Workshop (2003)
39. Ruiz, M., Sidler, D., Sutter, G., Alonso, G., López-Buedo, S.: Limago: an FPGA-
based open-source 100 GbE TCP/IP stack. In: International Conference on Field
Programmable Logic and Applications (FPL), pp. 286–292. IEEE (2019)
40. Salmon, G., Ghobadi, M., Ganjali, Y., Labrecque, M., Steffan, J.G.: NetFPGA-
based precise traffic generation. In: Proceedings of NetFPGA Developers Work-
shop, vol. 9. Citeseer (2009)
41. Shriram, A., Kaur, J.: Empirical evaluation of techniques for measuring available
bandwidth. In: International Conference on Computer Communications (INFO-
COM), pp. 2162–2170. IEEE (2007)
42. Shriram, A., et al.: Comparison of public end-to-end bandwidth estimation tools on
high-speed links. In: Dovrolis, C. (ed.) PAM 2005. LNCS, vol. 3431, pp. 306–320.
Springer, Heidelberg (2005). [Link] 24
43. Skhiri, R., Fresse, V., Jamont, J.P., Suffran, B., Malek, J.: From FPGA to support
cloud to cloud of FPGA: state of the art. Int. J. Reconfig. Comput. 2019 (2019)
44. Sommers, J., Barford, P., Willinger, W.: Laboratory-based calibration of available
bandwidth estimation tools. Microprocess. Microsyst. 31(4), 222–235 (2007)
Evaluation of the ProgHW/SW Architectural Design Space 283
45. Strauss, J., Katabi, D., Kaashoek, F.: A measurement study of available band-
width estimation tools. In: Proceedings of the 3rd ACM SIGCOMM conference on
Internet measurement, pp. 39–44 (2003)
46. Wang, H., Lee, K.S., Li, E., Lim, C.L., Tang, A., Weatherspoon, H.: Timing is
everything: accurate, minimum overhead, available bandwidth estimation in high-
speed wired networks. In: Proceedings of the 2014 Conference on Internet Mea-
surement Conference, pp. 407–420 (2014)
47. Yasukata, K., Honda, M., Santry, D., Eggert, L.: Stackmap: low-latency networking
with the OS stack and dedicated NICs. In: USENIX Annual Technical Conference
(USENIX ATC 2016), pp. 43–56 (2016)
48. Yin, Q., Kaur, J., Smith, F.D.: Can bandwidth estimation tackle noise at ultra-
high speeds? In: IEEE 22nd International Conference on Network Protocols, pp.
107–118 (2014)
49. Zhou, H., Wang, Y., Wang, X., Huai, X.: Difficulties in estimating available band-
width. In: International Conference on Communications, vol. 2, pp. 704–709. IEEE
(2006)
50. Zilberman, N., Audzevich, Y., Covington, G.A., Moore, A.W.: NetFPGA SUME:
toward 100 Gbps as research commodity. IEEE Micro 34(5), 32–41 (2014)
An In-Depth Measurement Analysis
of 5G mmWave PHY Latency and Its
Impact on End-to-End Delay
1 Introduction
The past few years have seen a rapid commercial deployment of 5G networks.
With enhanced mobile broadband services (eMBB), 5G promises to offer much
higher bandwidth than previous generations of cellular networks to consumers.
Existing measurement studies [10,20,23,29,33] have found that 5G radio tech-
nologies can in general achieve higher throughput performance than 4G LTE.
For example, with line of sight (LoS), mmWave 5G radio can deliver up to sev-
eral Gbps of downlink (DL) bandwidth [20,29,33] and up to hundreds of Mbps
uplink (UL) bandwidth [23], albeit their performance can fluctuate wildly.
Motivations for this Study. From the perspective of new applications which
require mission critical communications, what is perhaps more exciting is the
promise of 5G to offer millisecond (ms) or even sub-millisecond (PHY-layer)
latency support to applications [Sect. 7.5 in [3]]1 e.g., through the so-called
Ultra Reliable Low Latency Communication (URLLC) services [Sect. 7.9 in [3]]
[4,17,27]. These applications include but are not limited to, Autonomous Vehi-
cles (AVs) and drones supported with edge-assisted cooperative driving/flying
intelligence, Augmented/Virtual reality (AR/VR), and “metaverse”, all which
require extreme low latency and very high reliability to make crucial decisions.
such as the placement of the application server and packet payload affect the
latency of 5G PHY-layer and therefore the E2E delay experienced by applica-
tions? We answer these questions through a close look analysis of 5G mmWave
PHY-layer key performance indicators (KPIs) with the aim of quantifying the
impact of various factors and configurations. Our approach is laid out as follows:
First, we aim to quantitatively understand the PHY-layer latency and study
it under the “best-case” scenario (Sect. 4). Second, we quantify the impact of
several factors that impact the PHY-layer latency (Sects. 5 and 6). Lastly, we
explore the latency benefits and drawbacks of deploying services on edge nodes
supported by mmWave 5G (Sect. 7). Based on our knowledge, our paper is the
first to answer the question, “Is sub-millisecond PHY-layer latency achievable
with today’s commercial 5G”? And what impact does several factors like 5G
smartphone radio ON-OFF cycle and server placement have on the PHY-layer
and E2E delays. Next, we summarize our key findings and contributions.
F1. Today’s Best Achievable PHY-layer Delay (Sect. 4). Our analysis
shows that the best achievable mmWave 5G PHY-layer latency is 0.85 ms
which occurs about 2.27% of the time. Sub-millisecond (≤ 1ms) PHY-layer
latency is guaranteed only 4.42% of the time, with PHY-layer latency reach-
ing up to 3.08 ms about 22.36% of the time (Sect. 4.1). This delay is limited
by network side UL scheduling with control overhead contributing to the
largest share (about 81%) compared to data overhead, as a result of schedul-
ing requests and backoffs on the busy shared radio channel (Sect. 4.3).
F2. Impact of Channel Conditions (Sect. 5). A UE periodically (based on
the configurations) reports the DL channel condition to the base station
by calculating the value of the channel quality indicator (CQI), which is a
number from 1 to 15, where 15 indicates the best channel condition. When
the CQI value drops, transmitted data might be corrupted, requiring re-
transmission (ReTx). Our experiments show that: 1) The PHY-layer latency
when exactly one ReTx occurs is 1.33 ms, making sub-millisecond (≤ 1ms)
PHY-layer latency not achievable. 2) As the number of ReTxs increases,
the overhead of the PHY-layer data increases 3.5 times the overhead of the
control (Sect. 5.1). 3) On average, there is a 2ms additional overhead delay
on the PHY-layer when the CQI drops noticeably (Sect. 5.2).
F3. Impact of Mobility and Handovers (HOs) (Sect. 6). As mmWave is
directional, highly susceptible to many impairment factors, and has shorter
coverage ranges, mobility not only affects the channel condition experienced
by a UE, but also causes HOs in some situations. All these further impact the
latency on the PHY-layer. We find that: 1) When a UE is walking with good
channel conditions (i.e., high CQI value) and no HOs occur, the additional
PHY-layer overhead due to mobility is 0.51 ms (Sect. 6.1). 2) When there is
a HO, the minimum additional PHY-layer overhead is 2 ms (Sect. 6.2).
F4. Impact of UE Sleep Cycle (Sect. 7.2). As a way to reduce power con-
sumption on 5G smartphones, 5G supports discontinuous reception (DRX).
The operations of DRX modes depend on the UE’s state. We focus only on
the connected state (CDRX), namely, the UE has established a connection
An In-Depth Measurement Analysis of 5G mmWave PHY Latency 287
with the base station. In such a state, the UE radio antennas go through ON
and OFF cycles (i.e., awake and asleep states). Two scenarios can occur;
1) The DL transmission occurs while the UE is awake, no additional delay is
incurred (best case). 2) The network has data, but the UE is asleep (worst
case). Our results show that there is an additional overhead of 6.4 ms (on
average) to the PHY-layer latency in the worst case.
F5. Impact of Packet Payload Size (Sect. 7.3). We use PING packets to
mimic different application payload sizes. We find that the packet payload
size has little to no impact on the PHY-layer delay. Our results show that the
same time is taken to transmit a ping packet with 100 bytes and 1200 bytes
payload. This is because when the payload size of the PING packet increases,
the network adopts more hybrid ARQ (HARQ) process IDs [1] that work in
parallel to send and receive data between the UE and the base station.
C1. We present an in-depth and thorough analysis which allows for the quanti-
tative revelation of the status quo of today’s mmWave 5G PHY-layer delay,
identifying carrier specific configurations and poor design choices which hin-
ders 5G’s promise of sub-millisecond PHY-layer delay.
C2. We study several factors that impact the latency on the PHY-layer and quan-
tify them, showing that 5G network configurations and server placement deci-
sions can significantly impact the PHY-layer delay and thus E2E latency.
C3. We make all our data as well as other artifacts used in our study publicly
available to enable research continuation within the community: https://
[Link]/FarRoss/5gPHYLatency
Ethical [Link] study was carried out by paid and volunteer stu-
dents. We purchased several dedicated smartphones for experiments only and
several unlimited plans from AT&T and Verizon mmWave 5G carriers. No per-
sonal identifiable information (PII) was collected or used, nor were any human
subjects involved. This study is consistent with the Wireless Network Customer
[Link] work does not raise ethical issues.
and Verizon (VZW)) using Non-Standalone mode (NSA) [5]. NSA adopts a dual
connection mode in which 4G acts as an anchor for the control plane functional-
ity and to ensure continuous data connectivity. On the other hand, Standalone
mode (SA) relies on 5G for all control and data plane activities. Since mmWave
deployments are not continuous and have coverage holes, using mmWave with
SA 5G can lead to loss of connectivity during mobility. Additionally, any future
SA mmWave 5G deployments will most likely use the same 5G RAN technolo-
gies. Thus, we believe that our finding will also be valid for future mmWave SA
5G deployments. Mid-bands (3.3-3.8 GHz) and low-bands (700 MHz, n28) have
not been deployed yet, thus, beyond the scope of this study. Refer to recent
work [22] for a study of the mid-band 5G in Europe. Most of our controlled
experiments are focused specifically on Area 1.
5G UE and Measurement Tools. We use four phones, two S20s (Exynos 990
Qualcomm SM8250 Snapdragon 865 5G) and two S21 Ultras (Exynos 2100 Qual-
comm SM8350 Snapdragon 888 5G) [8]. We believe that these phones represent
the state-of-the-art 5G smartphones at the time we conducted the measurement
study with powerful communication modems, Mali-G77 MP11 and Mali-G78
MP14, respectively. Moreover, smartphone chip-sets do not affect the network
performance at the TCP and application layers [44].
To access the 5G New Radio (NR) stack and PHY-layer KPIs from chip-
set’s diagnostic interfaces (Diag), we use a professional tool called XCAL [6].
XCAL runs on a laptop connected to smartphones via USB or USB-C (Fig. 1).
It monitors, decodes, and deciphers signaling messages and the 5G RAN protocol
stack interactions between the UE and gNB following the 3GPP Rel-15 standard.
For our controlled experiments, we choose traceroute and ICMP-based PING
packets of 32 bytes because of two reasons; 1) It is readily available in Android
smartphones and does not require rooting devices. 2) To avoid any limitations
due to lack of radio resources using bigger packet sizes. However, we also study
the impact of larger packet sizes on PHY-layer and E2E delay (See Sect. 7).
will always be in RRC Connected state when receiving the PING echo reply, as the
length of RRC Connected is 320 ms [33] which is far greater than the worst RTT
(100 ms) observed in our experiments. Before each experiment, we close/stop all
background apps, disable background-app refresh, and turn off the WiFi interface.
To avoid delay overhead during transitions from RRC IDLE or RRC Inactivity to
RRC Connected state, we first play a random YouTube video for 30 s, then immedi-
ately close the YouTube app, wait 2 s, and then start the experiment. This ensures
that the UE is in the RRC Connected state before sending the echo request. To min-
imize the UE-side factors that may affect our measurements, we placed the smart-
phones on a flat surface during stationary experiments and kept them attached to
a car phone holder for driving experiments.
In this section, we introduce the 5G NR, 5G RAN, and zero in on the 5G PHY-
layer, and outline its key operations. The goal is two-fold: 1) introduce the key
PHY-layer interactions used in 5G NR defined by the 3GPP standards that are
most relevant to our study to justify our results and insights; and perhaps more
importantly, 2) dissect the various components of 5G PHY processing, and iden-
tify the major factors which may influence 5G PHY latency, and consequently
the E2E latency experienced by applications running on a UE or a remote server.
3
The primary physical channel for the DL transmissions (base station to UE) is
PDSCH (physical downlink shared channel), and for the UL transmissions (UE to
base station) is PUSCH (physical uplink shared channel).
An In-Depth Measurement Analysis of 5G mmWave PHY Latency 291
bandwidth and latency requirements of applications. The wider SCS not only
allows for higher channel bandwidth, but also enables lower latency through
a shorter slot time, i.e., from 1 ms in 15 kHz down to 0.125 ms in 120 kHz
(mmWave). A slot is defined as the basic (time) unit in which radio transmissions
are commonly scheduled [Sect. 4.3.1 in [12]] (See Sect. 4). Our study focuses
on 5G mmWave, as it can (potentially) provide both high bandwidth and low
latency.
During each slot, one data chunk4 is transmitted over the radio interface
to/from the UE. The scheduling configurations are exchanged via the down-
link control information (DCI)/the uplink control information (UCI) carried
in the Physical DL Control Channel (PDCCH)/Physical UL Control Channel
(PUCCH) respectively, as part of the PHY-layer control signaling (See Fig. 2).
5G mmWave uses time division duplex (TDD) which means both the DL and UL
share the same carrier frequency (physical transport channel) [16]. However, the
transmissions of DL and UL are scheduled at different times, e.g., using different
slots on the same frequency. We expand on these points below.
Slots and Scheduling. The 3GPP standards allow flexible scheduling of which
slots are dedicated for DL vs. UL transmissions [Sect. 5 in [16]]. However, we find
that current commercial 5G deployments still use a “fixed” pattern. For exam-
ple, as illustrated in Fig. 2, VZW mmWave 5G uses a 5-slots pattern, DDDSU
for DL/UL transmission scheduling: The first three slots (“DDD”) are reserved
for DL transmission only, the last slot, (“U”) is reserved for UL transmission
only, while the fourth slot, (“S”) is flexible – it can be used either for DL or UL
transmission, or both. For DL Transmission (data sent from gNB): the schedul-
ing information carried in the DCI specifies which symbols within “D” (and
“S”) slots are used to carry data; it also indicates which symbols in the “U”
(and “S”) slots may be used to carry UL transmissions, including UCI. DCI
is typically carried in the first 1-3 symbols in a “D” or “S” slot, while UCI is
carried in the last symbol in a “U” or “S” slot. While the UE is active in a
“Connected” state, it monitors the physical channels to see if there is DL data
and/or control traffic for it. For UL Transmission (data sent from UE): the UE
first sends a scheduling request in either the “U” or “S” slot which only informs
the network that the UE has data to transmit. The UE later sends the Buffer
State Report (BSR) [Sect. 5.4.5 in [13]], which informs the network the UL data
volume. With the BSR information, the network then explicitly grants the UE
resources. Lastly, the UE prepares and transmits the data using the scheduled
future UL slots. As a result, we can deduce that this configuration enables asym-
metric traffic between UL and DL demands. Thus, UL transmissions likely incur
longer latency than DL, which is also confirmed by our results in Sect. 4.
Channel Conditions (CQI), Modulation and Coding Schemes (MCS).
A UE periodically reports to the gNB the DL channel condition using the channel
quality indicator (CQI), a number from 1 to 15, where 15 indicates the best
4
Assuming no spatial multiplexing, which is the case of VZW 5G mmWave. However,
with spatial multiplexing, at most 2 Transport Blocks can be transmitted per slot.
292 R. A. K. Fezeu et al.
channel condition [Sect. 5.1.6 in [14]]. The gNB uses this CQI value to determine
which modulation (e.g., QPSK, 32QAM, or 64QAM) and coding rate (e.g., the
number of redundant bits) to use to encode the data. This is collectively referred
to as the Modulation and Coding Schemes (MCS) [Sect. 5.1.3 in [14]]. The MCS
value informs a UE on how to decode a DL transmission or how to encode a UL
transmission. The main take-away is the following: higher CQI generally leads to
higher MCS – if there is sufficient data buffered to warrant it; and higher MCS
means more information bits (i.e., more data from the upper layer) is carried
per slot. As the MAC layer multiplexes data from multiple “logical” channels
(e.g., RRC messages, multiple concurrent user sessions), an IP packet from an
application server to a UE (or vice versa) can be segmented into multiple data
chunks, therefore requiring multiple slots for the packet to be delivered to the
user (or server), incurring longer latency even under “ideal” channel conditions.
conditions i.e., CQI values ≥ 12 which indicates high MCS [Sect. 5.2.2 in [14]]
(See Sect. 5) and no ReTxs occur. We summarize all the latency definitions in
Table 1.
PHY-layer latency, TP hy is defined as the time taken to send a PING echo request
in the UL, (TU L ) and receive the corresponding echo reply in the DL, (TDL ) on
the physical layer. i.e., TP hy = TU L + TDL . To compute TP hy , we carefully
trace every PING packet on the UE side down the 5G RAN stack. Based on
the data collected on the different radio channels, we use domain knowledge to:
1) isolate the PING packet from other noisy data such as beam management-
related control plane messages, 2) correlate the different transport channel PING
related messages, and 3) synchronize (and group) the different channel events in
UL and DL. Furthermore, we compute i) the time taken to send the PING data
on the physical transport data channel, TPDatahy and ii) the time taken to send
related control messages on the physical transport control channels, TPCtrl
hy .
Table 1. Summary of the Definitions for the Different Latency Terms Used
D1 + wired delay + D1
TE2E RT T Round Trip Time TE2E RT T = T5G Core+Inet +
from the applications T5G RAN
TP hy RT T Round Trip Time See Fig. 20
from the PHY-layer
294 R. A. K. Fezeu et al.
0.85 3.08
Fig. 3. Best Case TP hy showing min and Fig. 4. Breakdown of TP hy into Control
max achievable delays. and Data delays.
Results. We make the following observations. (1) In the best case, today’s
TP hy delay scale, can be as low as 0.85 ms and as high as 3.08 ms (See Fig. 3).
(2) Interestingly, only 4.43% of all our dataset samples have delays ≤ 1 ms. In
other words, sub-millisecond latency occurs about ≤ 5% of the time. Most delays
fall between 1 ms and 2.5 ms (i.e., 87.69%), and 7.83% have delays between
2.5 ms and 3.08 ms. The maximum best case TP hy latency is largely unsurprising:
previous studies have calculated this delay to be between 2.19±0.36ms [44].
Nevertheless, our results provide insight into today’s expected delay scale, which
can inspire new design opportunities. For example, to ensure that 5G can support
latency-critical applications, sub-milliseconds PHY-layer transmission is a must.
In particular, Rel 15 38.913 [3] standardized the 5G first hop (i.e., PHY-layer)
delay for URLLC to 1 ms. (3) A breakdown of the best case TP hy delay into
the control (TPCtrl Data
hy ) and data (TP hy ) overhead shows that the control overhead
is on average 3.78 times more than the data overhead (See Fig. 4). Thus, it is
clear that, today’s mmWave 5G PHY-layer latency is far from enabling latency-
critical applications. The question now remains, what are the design opportunities
or improvements which can favor the majority of the delay to fall below 1 ms?
To answer this question, we use Fig. 5 to dissect TP hy into DL and UL delays.
DL Latency Results. We find that the best (i.e., min) DL delay TDL is 0.09 ms,
which occurs 1.95% of the time (See Fig. 6a). This implies that D1 , D2 , and D3
can occur within one slot (≤ 0.125 ms), the S slot in DDDSU. However, we can
see that TDL has multiple peaks such as 0.17, 0.22, and 0.45 ms. This is due
to scheduling the 3 predefined tasks D1 , D2 , and D3 across slots and varying
number of OFDM symbols within each slot (refer to Fig. 2). For example, when
TDL = 0.45 ms, D1 , D2 , and D3 span 3.6 slots (i.e., 0.45 ms ÷ 0.125 ms). We
also find that, more than 50% of the time, the network configures the UE to
wait at least 6 slots (0.75 ms) before it can send the ACK control message in
D3 (See Fig. 6c). This time includes the processing delay on the UE side.
irrespective of the packet payload size (See Sect. 7.3) and is due to the fact that;
1) Today’s mmWave 5G implements same slot scheduling, i.e., D1 and D3 are
in the same slot (as shown in Fig. 2) and, 2) the DL control ( D1 and D3 ) and
DL data ( D2 ) messages occupy two-to-eight and one-to-nine OFDM symbols
respectively.
as TUCtrl
L , which can involve multiple unsuccessful scheduling request attempts
due to back-offs on the busy shared channel. Afterward, the UE prepares and
sends the data in step U3 . We refer to this time as TUData
L . The total time TU L
= TUCtrl
L + T Data
UL is the UL delay in the PHY-layer.
UL Ctrl and Data Latency. We further break TU L down and characterize the
cost on each network communication group, i.e., the control (TUCtrl L ) and data
(TUData
L ) overheads. Figure 7b shows that, considering a T UL time of 1.5 ms as an
example, the control overhead TUCtrl
L accounts for approximately 81% i.e., 1.7 ms.
Simply put, the control overhead ( U1 + U2 ) is responsible for the lion share of
the UL delay, unlike the case for DL. This shows that the UL control overhead
TUCtrl
L ( U1 + U2 ) takes much longer than data transmission TU L
Data
( U3 ) in the
UL. This is because of two reasons: 1) We find that the UE takes more time
waiting to be granted access to the busy shared channel ( U1 ) than the actual
grant time ( U2 ) as shown in Fig. 7c. 2) A single UL transport block gets split
into multiple code-blocks [Sect. [Link] in [14]] in the UE MAC layer, which are
then transmitted on the PHY-layer, and reassembled in the gNB MAC layer. In
the “best” case, all the code-blocks are transmitted in one UL transmission cycle
(Tx Cycle), as warranted by the allocation of network resources as specified in
U2 . We define a Tx Cycle as one round of U1 , U2 , and U3 . However, when the
UL-centric apps with heavy UL traffic demands like AR, the cyclic fixed slot
configuration means that, the network is not aware of the UE-side heavy traffic
demands. Therefore, we claim that, offloading some UL functions to the UE will
help cap the lion share control plane overhead and further reduce latency. For
example, introducing a mechanism by which a UE can signal heavy UL traffic
to the network and request a UL specific slot configuration or implementing a
true cross-layer signaling mechanism to anticipate and signal specific application
PHY-layer latency requirements could be ways to achieve this. This might also
help address variations (or instabilities) in latency, although these instabilities
are largely due to channel conditions (see below).
HO areas. Third, despite these measures to ensure no HO, we still observe and
discard experiments with any HO occurrences. As a way to quantify the impact
of the CQI on latency, we divide the CQI values into CQIlow = (6, 9], CQImedium
= (9, 12], and CQIhigh = (12, 15], and refer to it as such hereafter. Note that
even when the UE is in CQIhigh , the CQI value can still change slightly between
12+ and 15, and ReTxs may occur. Thus, during our experiments, we fix the
CQI range, keep all other factors constant, and investigate the impact of slight
CQI changes on the PHY-layer latency.
Fig. 11. 1ms TP hy is defeated with one Fig. 12. Impact of Retransmissions on
ReTx. TDL , TU L , and TP hy .
300 R. A. K. Fezeu et al.
We study the impact of CQIlow , CQImedium and CQIhigh with a fixed number of
ReTxs. Fig. 14 shows that when there is no ReTxs, there is at least an additional
2 ms overhead on TP hy with poor channel conditions (i.e., CQI changes from
CQIhigh to CQIlow ). A similar conclusion is observed for TU L (See Fig. 15). This
overhead is due to a lower MCS when the CQI drops to CQIlow . This will cause
a decrease in the code rate i.e., less useful bits are transmitted per slot, resulting
in more time to transmit an entire transport block. The impact of CQI on TDL
is rather insignificant.
Summary and Implications: The HARQ process is primarily used to speed
up ReTxs. The sender stores all transmitted data in its buffer and discards them
only after receiving an ACK from the receiver. The receiver also stores all erro-
neous packets and uses them to improve decoding [Sect. 5.4.2 in [13]]. This may
cause unavoidable latency overhead, particularly when the channel conditions
change very suddenly from CQIhigh to CQIlow . This is because UL data trans-
mission that is encoded with an MCS value suitable for the current reported CQI
value may not be suitable at a later time when there is a ReTx and the CQI
value drops. This can result in more ReTxs and higher latency, which explains
10
10 Num. ReTx=0 Num ReTx=1 Num. ReTx=0 Num ReTx=1
TPhy (in ms) 8
8
(a) CQI Range (b) CQI Range (a) CQI Range (b) CQI Range
Fig. 14. Impact of CQI and Num. ReTx Fig. 15. Impact of CQI and Num. ReTx
on TP hy . We consider Num. ReTx=0 and on TU L . We consider Num. ReTx=0 and
Num. ReTx=1. Num. ReTx=1.
why exactly one ReTx with PHY-layer latency 1.33 ms occurs about 2.27% of
time. Given this, we conclude that improving the HARQ process to account for
CQI to MCS mismatch, especially when channel conditions drop, can provide a
remedy and perhaps eliminate the additional overhead due to more ReTxs. In
the practical sense, this calls for an extensive re-design of mmWave PHY-layer
operations.
6 Impact of Mobility
In this section, we address two key questions: First, what is the additional PHY-
layer overhead due to UE-side activity (i.e., mobility) in mmWave 5G? and
second, how does mobility influence the PHY-layer latency in UL and DL?
Methodology: Similar to our experimental setup in Sect. 5, we minimize HOs
and conduct clear LoS walking experiments and do not walk beyond identified
potential HO patches. We study the best case i.e., the UE is in CQIhigh with
slight CQI fluctuations and no ReTxs.
the UE is in CQIhigh while walking and stationary and the Fig. on the right
shows TP hy while walking and stationary. We see that, even in CQIhigh , the
CQI values fluctuates frequently when the UE is walking. This is because, as
shown in Fig. 17, the network adopts a lower MCS values during mobility as a
way to minimize the number of ReTxs and meet the target BLER rate of <10%
[Table 8.1.1-1 in [12]]. However, adopting lower MCS increases the best case (i.e.,
min) TP hy from 0.85 ms to 1.36 ms between stationary and walking, respectively,
shown in Fig. 18. A difference of 0.51 ms, about 5 slots.
Fig. 20. Dissecting the E2E RTT into T5G RAN and T5G Core+Inet .
5G RAN upper layers in the UE and ii) the 5G Core + Internet delay, i.e.,
T5G Core+Inet defined as the time from when the UE sends the PING echo request
in U3 to when it receives the PING echo reply from the edge server on the PHY-
layer in D1 . Therefore, TE2E RT T = T5G RAN + T5G Core+Inet (See Fig. 20). To
divide the E2E RTT, we compute TP hy RT T , the physical layer RTT including
T5G Core+Inet as shown in Fig. 20. Then, T5G Core+Inet = TP hy RT T − (TU L +
TDL ). From T5G Core+Inet , we calculate T5G RAN = TE2E RT T - T5G Core+Inet .
Results. As shown in Fig. 21 and Table 2, the 5G RAN delay takes on average
7.32 ms regardless of the server location. However, as the distance between the
UE and the server increases, T5G Core+Inet increases dramatically to be 10 ms,
30 ms, and 35 ms (on average) across the WL, LZ, and RG servers, respec-
tively. This signifies the importance of edge server placement on RTT. Next, we
demonstrate the benefit of deploying applications on the WL, and setbacks of
deploying applications on the LZ and RG servers w.r.t. a UE location.
CDRX Sleep Timers. The CDRX cycle is controlled by the CDRX ON and
the CDRX OFF timers – The CDRX ON timer determines how long the UE will
stay ON and the CDRX OFF timer dictates the duration the UE will stay OFF.
The CDRX ON/OFF duration cycles may be extended further on the basis of the
CDRX Inactivity timer. The CDRX Inactivity timer determines how long the UE
MUST stay ON upon reception of access to the busy shared channel (i.e., U2 ),
which will further extend the duration of the UE ON [1]. We observe that, both
VZW and AT&T configure the CDRX ON and CDRX Inactivity duration as
8 ms and 30 ms, respectively.
An In-Depth Measurement Analysis of 5G mmWave PHY Latency 305
Delay Due to CDRX. Since the CDRX Inactivity timer starts when the
UE acquires access to the busy shared channel, we therefore compute
TCDRX Overhead = TP hyRT T CDRX - 30 (CDRX Inactivity duration), where
TP hyRT T CDRX is the time between when a UE acquires access to the busy
shared channel ( U2 ) and receives the DCI which indicates an echo PING reply
on the PHY-layer ( D1 ), i.e., the time from U2 —>edge server—> D1 in Fig. 5.
We find that, in the WL case, the UE will never go to sleep before receiving the
echo PING reply from the server. This is because, in the WL, TP hyRT T CDRX
<< 30ms (CDRX Inactivity). However, in the LZ and RG cases, the UE goes into
sleep mode (CDRX OFF) about 60% and 97% of the time respectively before
receiving the PING echo reply (See Fig. 22a). We show a detailed illustration
of this behavior for each server in Fig. 23 by showing the arrival time for three
sample PING echo replies w.r.t. the UE status CDRX ON/OFF. We further com-
pute TCDRX Overhead , the additional time taken before the network sends the
PING echo reply to the UE when the UE is asleep (CDRX OFF) because the
CDRX Inactivity timer has expired. We find that, TCDRX Overhead = 6.4 ms (on
average) (See Fig. 22b).
Fig. 22. Impact of CDRX and server placement on PHY-layer. a) [1] In the WL case,
the UE will NEVER go to sleep. [2] In the LZ case, the UE goes to sleep 60% of the
time, while [3] in the RG case, the UE will go to sleep 97% of the time. b) Additional
6.4 ms delay (on average) overhead due to CDRX.
Fig. 23. Detailed illustration of how the CDRX and server placement impact the E2E
RTT. Sever placement causes an additional delay due to CDRX, TCDRX Overhead in
the LZ and RG edger server.
an insignificant increase in the E2E RTT (See Fig. 25). This is because both the
UE and the gNB will take more time to reassemble the data chunks from all
processes before forwarding it to the RAN upper layer for processing.
Summary and Implications. Although the role of CDRX in the manage-
ment of UE power is paramount [24], our experiments show that there is a
trade-off with the E2E latency in the LZ and RG edge nodes. Without devalu-
ing the CDRX benefits, our experiments reveal that, the additional overhead due
to CDRX (i.e., TP hyRT T CDRX = 6.4 ms) is primarily due to the network side
CDRX sleep timer configurations. We claim that adopting dynamic context-aware
CDRX timer configuration may significantly reduce or perhaps even eliminate the
latency effect due to CDRX especially in far edge nodes. For example, increasing
the CDRX Inactivity timer from 30 to 35 or 40 ms can potentially reduce the per-
ceived latency of the E2E application by 6.4 ms on average. Additionally, it will
be beneficial to customers with limited monetary resources as deploying applica-
tions on the closest edges, such as the WL node, is very expensive [45]. However,
achieving this context-aware CDRX timer configurations requires a truly 5G NR
An In-Depth Measurement Analysis of 5G mmWave PHY Latency 307
cross-layer design which perhaps calls for a protocol redesign. This approach is
particularly difficult and have not yet been studied in the literature.
Fig. 24. Impact of Payload size on TP hy . Fig. 25. Payload size has little to no
impact on TPData
hy and E2E RTT.
8 Related Work
We discuss the related work in two categories: Commercial 5G Network
Measurements. Researchers have conducted several studies on commercial 5G
networks since their debut in 2019. Among them, Narayanan et al. examines
for the first time the performance of mmWave 5G on smartphones [29]. The
same team also investigates 5G performance prediction [30], application QoE,
and device power consumption [33]. Xu et al. study the coverage, performance,
and energy consumption of sub-6Ghz 5G in China [44]. Rochman et al. compare
5G deployment in Chicago and Miami [37]. Rischke et al. measure 5G campus
networks [36]. Pan and Claudio et al. examine the 5G performance on high-speed
trains and in public bus transit systems respectively [22,34]. Compared to all
the above studies, our work focuses on the latency of 5G networks in the context
of 5G last-mile latency support for edge computing [7]– an important but under
explored topic.
5G Physical Layer. There are a plethora of works on the PHY-layer founda-
tions of 5G, including mmWave [40,42], signal propagation [41,43], beam form-
ing [18,38], and massive MIMO [39,46], to name a few. Compared to the above
works that solely tackle the E2E latency, [19,23,24,28,29,33,44] also quantify the
PHY-layer UL and DL latency separately, but not both from different points of
view. Almost in line with our work, Xu et al. quantify the latency of 5G mmWave
PHY-layer in China to be 2.19±0.36 ms [44]. However, they do not state or show
whether <1ms PHY-layer latency is achievable with today’s mmWave 5G NR
deployments. Additionally, factors that can further increase PHY-layer latency
were not explored. Thus, to our knowledge, our paper is the first to systemat-
ically study and quantify the impact of several factors on PHY-layer latency,
and the impact of server placement and CDRX on E2E latency. We are also the
first to answer the question “Is sub-millisecond PHY-layer latency feasible with
today’s commercial 5G”. Additionally, our paper provides insights to network
operators to capitalize on which other related works lack on.
308 R. A. K. Fezeu et al.
10 Conclusion
Using a commercial 5G tool to extract detailed physical channel events and mes-
sages, this study presents a first-of-a-kind comprehensive in-depth measurement
study of mmWave 5G latency performance on the PHY-layer. Our findings show
that the current 5G RAN-induced latency is limited by both UL scheduling and
An In-Depth Measurement Analysis of 5G mmWave PHY Latency 309
References
1. 5G NR: Connected Mode DRX. [Link]
[Link]. Accessed Nov 2022
2. Amazon web services (aws). [Link]
3. 5G; study on scenarios and requirements for next generation access technologies
(3gpp tr 38.913 version 14.3.0 release 14) (2017). [Link]
etsi tr/138900 138999/138913/15.00.00 60/tr [Link]
4. [Link] (2019)
5. 5G SA vs 5G NSA: What are the differences? [Link]
5g-nsa-what-are-the-differences/ (2022). Accessed Nov 2022
6. Accuver XCAL. [Link]
ckattempt=2 (2022). Accessed Nov 2022
7. AWS Wavelength. [Link] (2022). Accessed Nov
2022
8. Samsung galaxy S21 5G featuring a Qualcomm snapdragon 888 5G mobile plat-
form. [Link]
s21-5g (2022). Accessed Nov 2022
9. Speedtest by Ookla. [Link] (2022). Accessed Nov 2022
10. T-Mobile hits 3 Gbps 5G speeds without mmWave in world record produc-
tion test. [Link] (2022).
Accessed Nov 2022
11. 3GPP: 5G; NR; Multiplexing and channel coding (3GPP TS 38.212 version 15.2.0
Release 15) (2018). [Link] ts/138200 138299/138212/
15.02.00 60/ts [Link]. Accessed Nov 2022
12. 3GPP: 5G; NR; Requirements for support of radio resource management (3GPP
TS 38.133 version 15.3.0 Release 15) (2018). [Link] ts/
138100 138199/138133/15.03.00 60/ts [Link]. Accessed Nov 2022
13. 3GPP: 5G NR: Medium Access Control (MAC) protocol specification (3GPP TS
38.321 version 15.5.0 Release 15) (2019–05). [Link] ts/
138300 138399/138321/15.05.00 60/ts [Link]. Accessed Nov 2022
310 R. A. K. Fezeu et al.
14. 3GPP: 5G; NR; Physical layer procedures for data (3GPP TS 38.214 version
16.2.0 Release 16). [Link] ts/138200 138299/138214/
16.02.00 60/ts [Link] (2020). Accessed Nov 2022
15. 3GPP: 5G; NR; Radio Resource Control (RRC); Protocol specification (3GPP
TS 38.331 version 16.2.0 Release 16) (2020). [Link] ts/
138300 138399/138331/16.02.00 60/ts [Link]. Accessed Nov 2022
16. 3GPP: 5G; NR; Radio Link Control (RLC) protocol specification (3GPP TS 38.322
version 16.2.0 Release 16) (2021). [Link] ts/138300
138399/138322/16.02.00 60/ts [Link]. Accessed Nov 2022
17. Admin, G.: News & events (2017). [Link]
news/sa1-5g
18. Ahmed, I., et al.: A survey on hybrid beamforming techniques in 5G: Architecture
and system model perspectives. IEEE Commun. Surv. Tutorials 20(4), 3060–3097
(2018)
19. Corneo, L., Eder, M., Mohan, N., Zavodovski, A., BayhanZ, S.: Surrounded by the
clouds. In: The Web Conference (2021)
20. Dinh, P., Ghoshal, M., Koutsonikolas, D., Widmer, J.: Demystifying resource
allocation policies in operational 5G mmwave networks. In: 2022 IEEE 23rd
International Symposium on a World of Wireless, Mobile and Multimedia Net-
works (WoWMoM), pp. 1–10 (2022). [Link]
2022.00016
21. Fang, Z., Wang, G., Xie, X., Zhang, F., Zhang, D.: Urban map inference by perva-
sive vehicular sensing systems with complementary mobility. Proceed. ACM Inter.
Mobile Wearable Ubiquit. Technol. 5(1), 1–24 (2021)
22. Fiandrino, C., Juárez Martı́nez-Villanueva, D., Widmer, J.: Uncovering 5G per-
formance on public transit systems with an app-based measurement study. In:
Proceedings of the 25th International ACM Conference on Modeling Analysis and
Simulation of Wireless and Mobile Systems, pp. 65–73 (2022)
23. Ghoshal, M., et al.: An in-depth study of uplink performance of 5g mmWave net-
works, pp. 29–35. 5G-MeMU 2022, Association for Computing Machinery, New
York, NY, USA (2022). [Link]
24. Hassan, A., et al.: Vivisecting mobility management in 5G cellular networks. In:
Proceedings of the ACM SIGCOMM 2022 Conference, pp. 86–100. SIGCOMM
2022, Association for Computing Machinery, New York, NY, USA (2022). https://
[Link]/10.1145/3544216.3544217
25. Hassan, A., et al.: Vivisecting mobility management in 5G cellular networks. In:
Proceedings of the ACM SIGCOMM 2022 Conference. pp. 86–100. SIGCOMM
2022, Association for Computing Machinery, New York, NY, USA (2022). https://
[Link]/10.1145/3544216.3544217
26. Li, Y., et al.: Experience: a five-year retrospective of mobileInsight. In: Proceedings
of the 27th Annual International Conference on Mobile Computing and Network-
ing, pp. 28–41 (2021)
27. McLaughlin, R.: 5G low latency requirements (2021). [Link]
com/5g-low-latency-requirements/
28. Mohan, N., Corneo, L., Zavodovski, A., Bayhan, S., Wong, W., Kangasharju, J.:
Pruning edge research with latency shears. In: Proceedings of the 19th ACM Work-
shop on Hot Topics in Networks, pp. 182–189 (2020)
29. Narayanan, A., et al.: A first look at commercial 5G performance on smartphones.
In: Proceedings of The Web Conference 2020, pp. 894–905 (2020)
An In-Depth Measurement Analysis of 5G mmWave PHY Latency 311
30. Narayanan, A., et al.: Lumos5G: mapping and predicting commercial mmWave 5G
throughput. In: Proceedings of the ACM Internet Measurement Conference, pp.
176–193. IMC 2020, Association for Computing Machinery, New York, NY, USA
(2020). [Link]
31. Narayanan, A., Ramadan, E., Quant, J., Ji, P., Qian, F., Zhang, Z.L.: 5G tracker:
a crowdsourced platform to enable research using commercial 5G services. In: Pro-
ceedings of the SIGCOMM2020 Poster and Demo Sessions, pp. 65–67 (2020)
32. Narayanan, A., et al.: A comparative measurement study of commercial 5G
mmWave deployments. In: IEEE INFOCOM 2022 - IEEE Conference on Computer
Communications, pp. 800–809 (2022). [Link]
2022.9796693
33. Narayanan, A., et al.: A variegated look at 5g in the wild: performance, power, and
qoe implications. In: Proceedings of the 2021 ACM SIGCOMM 2021 Conference,
pp. 610–625. SIGCOMM 2021, Association for Computing Machinery, New York,
NY, USA (2021). [Link]
34. Pan, Y., Li, R., Xu, C.: The first 5G-LTE comparative study in extreme mobility.
Proceed. ACM Measure. Anal. Comput. Systems 6(1), 1–22 (2022)
35. Ramadan, E., Narayanan, A., Dayalan, U.K., Fezeu, R.A., Qian, F., Zhang, Z.L.:
Case for 5G-aware video streaming applications. In: Proceedings of the 1st Work-
shop on 5G Measurements, Modeling, and Use Cases, pp. 27–34 (2021)
36. Rischke, J., Sossalla, P., Itting, S., Fitzek, F.H., Reisslein, M.: 5G campus networks:
a first measurement study. IEEE Access 9, 121786–121803 (2021)
37. Rochman, M.I., et al.: A comparison study of cellular deployments in Chicago and
Miami using apps on smartphones. In: Proceedings of the 15th ACM Workshop
on Wireless Network Testbeds, Experimental evaluation & CHaracterization, pp.
61–68 (2022)
38. Roh, W., et al.: Millimeter-wave beamforming as an enabling technology for 5G
cellular communications: theoretical feasibility and prototype results. IEEE Com-
mun. Mag. 52(2), 106–113 (2014)
39. Shepard, C., Blum, J., Guerra, R.E., Doost-Mohammady, R., Zhong, L.: Design
and implementation of scalable massive-Mimo networks. In: Proceedings of the 1st
International Workshop on Open Software Defined Wireless Networks, pp. 7–13
(2020)
40. Singh, V., Mondal, S., Gadre, A., Srivastava, M., Paramesh, J., Kumar, S.:
Millimeter-wave full duplex radios. In: Proceedings of the 26th Annual Interna-
tional Conference on Mobile Computing and Networking, pp. 1–14 (2020)
41. Solomitckii, D., Orsino, A., Andreev, S., Koucheryavy, Y., Valkama, M.: Charac-
terization of mmWave channel properties at 28 and 60 GHZ in factory automation
deployments. In: 2018 IEEE Wireless Communications and Networking Conference
(WCNC), pp. 1–6. IEEE (2018)
42. Sur, S., Pefkianakis, I., Zhang, X., Kim, K.H.: Towards scalable and ubiquitous
millimeter-wave wireless networks. In: Proceedings of the 24th Annual Interna-
tional Conference on Mobile Computing and Networking, pp. 257–271 (2018)
43. Sur, S., Venkateswaran, V., Zhang, X., Ramanathan, P.: 60 GHZ indoor network-
ing through flexible beams: a link-level profiling. In: Proceedings of the 2015 ACM
SIGMETRICS International Conference on Measurement and Modeling of Com-
puter Systems, pp. 71–84 (2015)
312 R. A. K. Fezeu et al.
44. Xu, D., et al.: Understanding operational 5G: a first measurement study on its
coverage, performance and energy consumption. In: Proceedings of the Annual
Conference of the ACM Special Interest Group on Data Communication on the
Applications, Technologies, Architectures, and Protocols for Computer Communi-
cation, pp. 479–494 (2020)
45. Xu, M., et al.: From cloud to edge: a first look at public edge platforms,
pp. 37–53. IMC 2021, Association for Computing Machinery, New York, NY,
USA (2021). [Link] [Link]
[Link]/10.1145/3487552.3487815
46. Zhao, R., Woodford, T., Wei, T., Qian, K., Zhang, X.: M-cube: a millimeter-wave
massive mimo software radio. In: Proceedings of the 26th Annual International
Conference on Mobile Computing and Networking, pp. 1–14 (2020)
A Characterization of Route Variability
in LEO Satellite Networks
1 Introduction
Low Earth Orbit (LEO) Satellite networks are emerging as an essential part of
the future of global telecommunication, with pilot networks already in deploy-
ment and more planned [8,10,11,15]. Satellites operating in a low Earth orbit
provide low-latency communication, making them superior to terrestrial net-
works in some scenarios [38,39]. At the same time, operating in low earth orbit
inherently makes the satellite network very dynamic, with satellites orbiting the
Earth every 100 min to maintain their orbits [3]. As a result, a satellite is only
visible for a maximum of a few minutes to any single ground station. The highly
dynamic nature of LEO satellite constellations introduces significant variability
in the underlying network characteristics including the topology of the network,
leading to frequent changes in the path characteristics.
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 313–342, 2023.
[Link]
314 V. Bhosale et al.
Drastic changes are more likely if the travel direction between the communicating
ground stations is not along any of the orbital planes. The impact of such changes is
a function of the total path length (i.e., the longer the path, the smaller the impact
of a change in its building blocks). Thus, we conclude that this structure is deter-
mined by the relative position of the communicating pair of ground stations (i.e.,
the distance and angle of travel between them).
Finally, we show that RTT variability does not necessarily decrease by
increasing the number of deployed satellites (i.e., by increasing path diver-
sity) (Sect. 6). In particular, we measure RTT variability in a constellation as
we add orbital shells. Simply adding orbital shells doesn’t reduce RTT variabil-
ity, with variability depending on the exact configuration of each shell (i.e., the
number of satellites and orbits and the inclination of orbits). To evaluate the
impact of constellation configuration, we compare two Starlink constellation con-
figurations submitted to the FCC. We find that the more recent configuration
introduces more RTT variability than the abandoned configuration.
2 Background
We start with a brief discussion of relevant background on LEO satellite networks
needed to follow the rest of our study.
Fig. 2. An illustration of a LEO satellite network, showing the different types of links.
providing WiFi connectivity in airlines and cruise ships [13,58,83] The projected
success of the current reincarnation of LEO satellite networks is driven by the
reduced cost of building and launching such satellites [4,22,25,59,62].
In a LEO satellite network, satellites are placed in a number of shells, each
consisting of a number of circular orbits or orbital planes at a constant alti-
tude. Orbital planes are characterized by their altitude (the height above sea
level), and their inclination angle (the angle at which they intersect the equa-
tor). An inclination angle of 90◦ refers to a polar orbit. However, most of the
current constellations have smaller inclination angles to provide greater coverage
to densely populated areas [70]. Figure 1 shows an example with three different
orbital planes, each at a different inclination. Similar orbital planes are equally
spaced to form an orbital shell. Table 1 highlights these parameters for three of
the largest proposed constellations.
Links (GSLs) that operate in the Ku/Ka bands. A ground station only commu-
nicates with satellites that are visible above a certain elevation angle above the
horizon, limiting the time traveled by the wave in the earth’s atmosphere to ensure
the quality of the link. Figure 2 shows an illustration of a LEO satellite network.
The LEO satellite network topology is highly dynamic in nature owing to the
rapid motion of the satellites. A LEO satellite travels at about 27,000 km/hr to
maintain its orbit. Thus, a satellite is visible for a maximum of 10–12 min from
any point on earth. However, the satellite has to conform to the elevation angle
bounds required for communication, limiting the accessibility time to a maximum
of 4.5 min. The exact amount of time that a satellite remains visible from a
ground station depends on the altitude of the satellite. At a higher altitude, a
satellite travels at slightly slower speeds, increasing the duration of its visibility.
Currently, satellites do not rely on ISLs and use ground station gate-
ways [5,6]. There have been some environmental concerns regarding laser-based
ISLs, leading to the initial batch of Starlink satellites being launched with-
out ISLs [40]. However, there have been recent launches of satellites with ISLs
onboard, with more launches of similar satellites planned [32]. The current plan
for LEO satellite networks is to carry traffic through ISLs to the ground station
closest to the destination server [38,47]. Relying on ISLs yields high data rates
and low latency, provides better resistance to weather conditions and faces no
regulations and a reduced risk of jamming [41,66]. Therefore, in this paper, we
focus on ISL-based networks, where traffic is routed through the satellite net-
work from the source terminal closest to the sender to the destination terminal
closest to the receiver.
The 2018 FCC filings by Starlink indicate the presence of 4 silicon-carbide
communication components on every satellite [73], with recent work identify-
ing them to be used as ISLs [20,38,46]. ISLs can be dynamically configured
to connect satellites, allowing for the formation of many different topologies.
The setup of an ISL can take between a few tens of seconds [69] to about a
minute [84], during which the link cannot be used. The setup time for Starlink
ISLs might be lower because the inter-satellite distances are smaller in the Star-
link constellation. However, link setup won’t be instantaneous, greatly reducing
the utility of these links if reconfiguration was frequent. Hence, we assume static
ISL configurations that require no ISL reconfiguration operations. In particu-
lar, we assume that satellites are connected following the so-called +Grid con-
figuration where each satellite connects to two satellites in its own orbit, and
with one satellite in each of the adjacent orbits. This configuration has been
selected by the earlier work as the most likely configuration to be used in prac-
tice [20,26,37,38,53,57,67,68,80,81].
The +Grid topology has many configurations, depending on how inter-orbit
links are formed. The configuration of inter-orbit ISLs depends on the phase shift
of orbital planes. The phase shift is a value between zero and one, determining
the relative motion of satellites in adjacent orbits. At zero, all satellites with
the same index in all orbital planes cross the equator at the same time. At one,
satellite n in orbital plane p crosses the equator at the same time as satellite n+1
in orbital plane p + 1. We use a phase shift of 0.5 as it very closely corresponds
318 V. Bhosale et al.
(a) The CDF for the ratio between the path (b) The CDF for the ratio between the highest
lengths achieved using our +Grid variation and lowest latencies observed for the two ISL
and the nearest-neighbor +Grid variations.
Fig. 3. Characterizing the benefits of using the variation of +Grid over the nearest-
neighbor +Grid ISL configuration
to the phase offset parameter used by Starlink [74] and can potentially provide
good coverage of the Earth by uniformly distributing satellites in orbit [19]1 .
Now consider how inter-orbit ISLs should be formed in the presence of a phase
shift between orbital planes. If a satellite is connected to its nearest neighbors in
adjacent orbits in the presence of phase shifts, the resulting topology will be an
inclined grid, providing poor east-west paths [38]. Alternatively, prior work [19,
38,46] has intuitively argued for a slight variation of the +Grid configuration
where a satellite connects with a nearest neighbor in one adjacent orbit and with
a phase shifted neighbor in the other orbit. This slight variation on the +Grid
topology results in shorter east-west paths. Path 2 in Fig. 16 is an example of
paths created by that modified configuration. We verified this intuition with the
following simulation study. We explain our simulation setup in Sect. 3.
We ran a simulation of the two +Grid variations for 100 min using the top 100
cities worldwide as source-destination pairs (total 4950 pairs). We measure the
RTTs for all 4950 pairs every second to observe the variability inherent to the two
choices. We first look at the ratio of the latencies observed for the two different
variations at every second for all these source-destination pairs in Fig. 3a. While
the nearest-neighbor +Grid configuration has shorter paths in more than 40% of
the scenarios, the maximum latency gain is just 43%. On the other hand, when
the nearest-neighbor +Grid configuration has longer paths, its paths can be
more than 5 times longer. To better understand the performance of the nearest-
neighbor +Grid configuration, we look at the ratio of the maximum RTT and the
minimum RTT between every pair of cities during the course of our simulation
(Fig. 3b). We observe that the nearest-neighbor +Grid configuration can have
this ratio greater than 7 compared to about 2.7 for the variation to +Grid
we use. Based on this simulation, we conclude that the nearest-neighbor +Grid
configuration increases the magnitude and variance of the lengths of paths. Thus,
1
Since the phase offset does not impact the stability of the GSL and ISL connections
which we show to be the major reason for route variability in satellite networks, our
results in this paper hold true for any chosen phase offset value.
A Characterization of Route Variability in LEO Satellite Networks 319
for the rest of this paper, we use the +Grid variation employed in earlier work,
leading to the configuration that minimizes path length variability.
3 Study Setup
Our study relies on simulations performed using the Hypatia framework [46] as
the starting point, augmenting it with additional emulators as needed for the
purposes of our study. We leverage Hypatia to generate the Two Line Element
(TLE) information for satellites, a standard representation for satellite orbits
containing the satellite identifier and orbit parameters [1]. Using that informa-
tion, we are able to determine the ISL and GSL connectivity to conform with
all the necessary physical requirements. Hypatia’s routing algorithm selects the
shortest path between a pair of ground stations every 100ms. We use an interval
of 1s to accelerate our simulations.2 We use the Cesium [7] (a javascript library
2
It is telling that we are still able to show the impact of frequent routing even with
a reduced frequency of route updates.
320 V. Bhosale et al.
– the lifetime of a path: the duration a path remains valid (i.e., usable), allowing
two specific communication ground stations to reach each other through the
satellite hops of that path, and
– the usage time of that path: the duration a path is chosen by a routing algo-
rithm to route traffic between the two communicating ground stations.
paths. We also analyze the relationship between the usage time and the lifetime
of paths. In particular, we compute the ratio between the usage time and the
lifetime for all studied paths (Fig. 5). The figure shows that for the three con-
stellations, at least 50% of the paths are used for less than half of their lifetime.
(a) The CDF for the ratio between the (b) The CDF for the ratio between usage
shortest path length and the longest path time and lifetime of the longest shortest
length for a pair of ground stations, high- path for a pair of ground stations, high-
lighting that for 70% of them abandoning lighting that 70% of them are abandoned
a path yields a maximum of 25% reduction for more than than half of their lifetime.
in latency.
Fig. 6. Characterizing the lifetime and benefits of abandoning longest shortest paths,
showing that for the majority of cases they are abandoned for no significant gain.
A Characterization of Route Variability in LEO Satellite Networks 323
RTT reflects the length of the shortest path possible between two ground sta-
tions. The maximum RTT reflects the length of the longest shortest path.
The longest shortest path is the longest path selected by a routing algorithm
to connect a pair of ground stations. Consider that as the topology changes,
the composition (i.e., hops) and length of the shortest path between any two
ground stations changes. The routing algorithm always selects the shortest pos-
sible path. Amongst all these paths, we focus on the longest one, calling it as
the longest shortest path. The ratio between the minimum RTT and the length
of the longest shortest path reflects the highest performance gain that a routing
algorithm can make when abandoning a path (Fig. 6a). The figure shows that the
maximum performance gain is less than 25% for 70% of the source-destination
pairs. Note that this result is fairly conservative since we focus on the best pos-
sible performance high churn can produce. In many cases, abandoning a path
would lead to smaller gains.
We contrast the result in Fig. 6a with the ratio between the lifetime and usage
time of longest shortest paths. Figure 6b shows that 70% of longest shortest paths
are used for less than half of their lifetime (while 70% of them are only 25% longer
than the best possible RTT as shown in Fig. 6a). This implies that even if longest
shortest path were to be abandoned for the maximum possible gain, that gain
in most cases will be modest. Achieving the lowest possible latency matters for
some applications (e.g., High-Frequency Trading [65]). However, it won’t impact
the performance of most applications, especially given that the latency of ISL-
based LEO satellite networks can be around 30% better than the latency of the
terrestrial Internet [38,39].
To better contextualize the result,
we consider a concrete example. In
particular, we consider the Jakarta-
Bogotá route. Figure 7 shows a time
series of RTT values. The thick grey
line shows the actual RTTs that will
be observed using the shortest path
routing policy, whereas the dotted
lines represent the RTT of individ-
ual paths for the time they are valid.
While the first switch takes place at Fig. 7. The RTT of paths between Jakarta
17 s due to the end of the first path, and Bogotá, showing eight path changes in
the second switch takes place 5 s later 200 s. Dotted lines represent the RTT of dif-
due to a difference of 0.005 ms in ferent paths (in different colors). The solid
the latencies of the second and third line represents the achieved RTT.
paths, only to switch back to the second path 8 s later due to the end of the third
path. Such frequent switching leads to four changes in the first 42 s, with the
maximum latency gain of about 3.5 milliseconds (i.e., about 2.5% of the total
RTT).
Takeaway: Path length variability causes a high churn in routes, yet the vari-
ability is very small in most cases that it doesn’t warrant the high churn.
324 V. Bhosale et al.
Fig. 8. An illustration of the path utilization experiment between 2000 ground stations
in New York and 2000 ground stations in London. The illustration shows the available
paths between the two cities.
in this example that all connections are using the same paths. The purple path
is abandoned by all connections for the yellow path for about half a millisecond
lower RTT. All the connections abandon the yellow path for the black path for
a half millisecond gain, only for all of them to reuse the yellow path after 28 s.
These gains are considerably low compared to the RTT of the path (around 1–
2%). In addition, the benefits of these gains may get diluted due to the inability
of the transport layer to keep up with the changes.
Next, we consider the impact of path
churn on the behavior of the transport
layer. We consider the link utilization,
the 95th percentile delay, and the power
defined as the ratio of utilization to the
95th percentile delay exhibited by differ-
ent congestion control protocols. We look
at the route between Pune, India and Fig. 10. The RTT of the route Pune
Lahore, Pakistan. A path change in a and Lahore
LEO satellite network may end up chang-
ing the observed RTT and bandwidth (e.g., switching to a path with a different
number of flows competing for its bandwidth). Thus, we evaluate the impact of
RTT variability and bandwidth variability, separately and combined. Figure 10
shows the delays for the 60 s time interval we use for this experiment. Instead of
assigning a different bandwidth value for every path, we select two bandwidth
levels that we alternate between with every path change to show the impact of
changing bandwidth on TCP algorithms. In particular, the bandwidth changes
between 204 Mbps and 48 Mbps. We choose these values as they closely corre-
spond to the range of bandwidth specifications for Starlink [75]. The results are
shown in Table 2.
Table 2. The utilization (%), 95th percentile one-sided delay (in ms from sender to
receiver), and power (defined as the ratio of utilization to the 95th percentile delay)
for three different scenarios of path variability for five different congestion control
algorithms. Bold reflects the best result in its row and italic reflects the worst result
in the row. No single algorithm optimizes both delay and utilization.
Takeaway: High churn in routes, caused by shortest path algorithms, can cause
poor path utilization and poor performance by congestion control algorithms.
Fig. 11. Heat maps showing the ratio between max RTT and min RTT in paths
between Null Island(0◦ latitude, 0◦ longitude), Darfur, and Kyiv and 2700 nodes uni-
formly distributed around the globe. The redder the point, the higher the ratio, indi-
cating higher variability.
Figure 11 shows that there is a clear structure for ground station placements
that would yield high variability. For example, low latitude source ground sta-
tions (e.g., Null Island and Darfur) observe high variability when communicating
with ground stations placed in a ring-like structure with diagonal ribbons extend-
ing from it. On the other hand, the structure only includes the ribbons for high
latitude stations (e.g., Kyiv). This structure only impacts destinations that are
within 1500–3000 km from the source (geodesic distance). We found these results
to hold regardless of the longitude of the source station and similar structures
repeat for source stations at the same latitude.
We find that a particular route between two ground stations shows high
variability when the makeup of the paths changes drastically as satellites move.
328 V. Bhosale et al.
To better understand this behavior, we study the building blocks for paths and
their characteristics.
Properties of ISLs. We record the lengths of all ISLs in a 100 min time interval
for the first shell of Starlink. We observe that the lengths of ISLs are highly
predictable. Figure 12a shows the CDF of the median length of ISLs, broken
down based on their type.
The results show that there are two types of inter-orbit ISLs: one with a
median length of about 760 km and the other at 1384 km km. The two types of
inter-orbit ISLs occur alternately such that each satellite has one inter-orbit ISL
of both the types. This is an outcome of the phased orbit structure we discussed
in Sect. 2. On the other hand, intra-orbit ISLs are uniform with all their lengths
at about 1970 km km, which is about 150% more than the first cluster of the
inter-orbit ISLs and 50% more than the second. There is little variability in the
lengths of ISLs. Figure 12b shows the CDF of the ratio between the minimum
length and maximum length of an ISL during the 100 min period for the two
types of ISLs. All intra-orbit ISLs and 50% of inter-orbit ISLs exhibit minimal
length change (0.2% and 6%, respectively). The length of the other 50% of
inter-orbit ISLs can change by up to 21% due to the varying distances between
different orbital planes across latitudes [30,45]. The orbits are closer to each
other at higher latitudes compared to the lower ones, and hence the inter-orbit
ISL lengths vary accordingly.
A Characterization of Route Variability in LEO Satellite Networks 329
Takeaway: The length of a single ISL is stable and there are three different
types of ISLs each having significantly different lengths.
Takeaway: Due to the stability of the ISLs, the variability in the lifetime of
GSLs has a greater impact on the variability in the lifetime of paths.
Fig. 15. An example showing that using an intra-orbit ISL when traversing across
orbital planes can increase the path length
However, the impact of that change is also a function of the total length of the
path. For example, replacing a single link in a path made of 20 links will result in
a much lower change in total path length compared to replacing a link in a path
made of two links. To illustrate this point, consider paths from Kyiv to Cairo.
Figure 15 shows two paths between the two cities. Each is the shortest available
path during different time intervals. Figure 15a shows a path made entirely of
three inter-orbit links. Figure 15b shows a path with a smaller hop count but
with a single intra-orbit link, leading to a longer path due to the larger length
of intra-orbit ISLs.
Choosing GSLs. The choice of a GSL is not entirely a routing decision. Recall
that GSLs are wireless links with SNR determining the quality of the link. The
SNR depends on the length of the GSL and weather conditions among other
things. Thus, the choice of GSLs can be made independent of the routing deci-
sion. We explore the impact of that choice on RTT variability. In particular, we
assess the impact of the choice of GSLs, considering the worst case by examining
the first hops that yield the longest paths. In particular, we compute the ratio
between the length of the shortest possible path through the worst case first hop
and the length of the actual shortest path. This metric measures the worst pos-
sible performance based on the choice of the first hop. We measure that metric
Fig. 16. An example showing the value of Intra-orbit links and the downside of not
using them when applicable. Note that Path 2 is valid path but not a shortest path
and is used just for illustration (i.e., never picked by a routing algorithm).
Fig. 17. The impact on the length of a path by using the worst first hop compared to
the best one. The line shows the average of the ratio and the shaded part shows the
standard deviation.
332 V. Bhosale et al.
for four source ground stations on the 85◦ longitude, uniformly spaced between
a latitude of 5◦ and 55◦ . Each source ground station communicates with 2700
points distributed uniformly within the coverage area of Starlink’s first shell.
Figure 17 summarizes the results as a function of the distance between commu-
nicating ground stations. The results show that a poorly selected first hop may
double the RTT. However, with the increase in the distance between the source
and the destination, this impact reduces.
Takeaway: The choice of the first hop is integral to determining and optimizing
overall path length.
Fig. 19. An example of extreme RTT variation showing three paths between Darfur
and Isangi, reflecting time steps 80,81, and 86 in Fig. 18
Fig. 20. An example of modest RTT variation showing two paths between Darfur and
Muynak, reflecting time steps 42 and 43 in Fig. 18
reach the next orbit and then an intra-orbit link to finally reach the destination.
Whereas at time step 80, the first hop can reach the destination with just one
intra-orbit ISL making it shorter.
Since the Darfur-Muynak route is along the orbital planes for the first shell,
the intra-orbit ISLs are used predominantly. We highlight the switch happening
at timestep 43 in Fig. 20. In this case, even swapping out an intra-orbit ISL for an
inter-orbit one doesn’t lead to much change since the links are always traveling
towards the destination.
334 V. Bhosale et al.
Fig. 22. The CDF of the maximum RTT variation observed by paths for different shell
configurations
7 Discussion
Topology Variants. As discussed earlier, our simulations use a specific vari-
ant of the +Grid topology due to its good coverage properties, making it the
most likely topology to be used in practice. However, there are several other inter-
satellite network topologies with desirable properties. For example, extending
inter-orbit ISLs beyond adjacent orbits was shown to reduce latency and improve
network throughput [20]. This approach requires more ISL reconfigurations as
satellites move. In our study, the properties of ISLs were relatively stable. Thus,
we posit that any such added variability in ISLs can increase path churn by adding
3
Not all inclinations we used might be possible due to interference or orbital con-
straints. Our goal is simply to highlight the impact of that parameter.
336 V. Bhosale et al.
another source of variability. Future work should explore the tradeoffs offered by
such topologies, taking into account variability, latency, and throughput.
Even for the +Grid topology, we focused on a single variant. The configura-
tion of other variants will depend on the phase shift in orbits and how inter-orbit
links are formed. For example, a variant can create more uniform inter-orbit ISLs
by connecting satellites to their nearest neighbors in adjacent orbits, harming
latency but reducing path variability. We leave it for future work to explore the
impact of such variants on route variability.
Ground Relays. Our study does not take into account satellite networks that
rely on ground relays. This is driven by the fact that currently deployed Starlink
satellites use ground stations only to connect to the Internet directly and not
to connect to each other. A network with ground relays, where ground stations
and user terminals act as ground relays as described in [39], will likely exhibit a
higher degree of variability than the network we studied owing to the increased
number of GSLs which are a big contributor to route churn.
8 Related Work
Optimizing Delay. The domain of LEO satellite networks has seen an
increased amount of interest from the networking community over the past
few years. The community has been especially excited about the potential of
these networks to outperform terrestrial networks. This has led to topology
design proposals that aim for inter-satellite network topology providing low
latency [20,38]. While these focused on ISL-based networks, others explored
achieving low latency in the absence of ISLs using ground relays [39]. The goal
for all of these studies is to optimize the network for low latency to outperform
terrestrial fiber networks. We offer a different perspective, showing that optimiz-
ing exclusively for delay can be harmful to network utilization and path-adaptive
algorithms. We also show that a slight sacrifice in delay can improve route sta-
bility. We hope that our insights will help design algorithms that can better
navigate these tradeoffs. In addition, there have been multiple efforts [18,33,49]
analyzing the challenges with integrating the LEO satellite network with the
current internet backbone. They look at how satellites can be used to assist with
inter-domain routing. While our work focuses only on intra-domain routing for
the satellite network, it will help the design of such systems by providing better
routing through the satellites.
A Characterization of Route Variability in LEO Satellite Networks 337
Routing. A lot of work has been done in the past looking at different goals
for routing such as reducing propagation delay [30,38,39,46,60], improved load
balancing [76], and energy efficiency [14,45] (Sect. 2). However, most of these
proposals were made for an older generation of satellites and applications. The
current generation consists of a considerably larger number of satellites and also
incorporates many advancements in the satellite communications domain [12].
We hope our work motivates a resurgence in research on routing algorithms in
satellite networks.
9 Conclusion
In this paper, we study the variability in paths rampant in LEO satellite net-
works. We concretely present the amount of route churn and RTT variability,
also highlighting the impact of such variability on path utilization and conges-
tion control. We delve deeper into the reasons why this variability exists by
presenting the building blocks of paths and infer that this variability exhibits a
spatial structure. Our hope is that this work will provide the key insights for the
design of specialized routing and perhaps, congestion control algorithms for LEO
satellite networks, taking into account that when delay is an option significant
gains can be made in overall network performance.
References
1. SGP4 Propagator. [Link] orbitPro
p [Link]
2. SGP4 Propagator (2016). [Link]
orbitProp [Link]
3. Popular Orbits 101 (2017). [Link]
orbits-101/. Accessed 30 Nov 2017
4. Launch Costs to Low Earth Orbit, 1980–2100 (2018). [Link]
net/data-trends/[Link]. ACcessed 12 Dec 2022
5. FCC Selected Application for Space Exploration Holdings, LLC. SES-
LIC2019090601171 (2020)
6. FCC Selected Application for Space Exploration Holdings, LLC. SES-
LIC2019021100151 (2020)
338 V. Bhosale et al.
28. Dong, M., et al.: PCC vivace: online-learning congestion control. In: Proceedings
of USENIX NSDI 2018 (2018)
29. Duplyakin, D., et al.: The design and operation of CloudLab. In: Proceedings of
the USENIX Annual Technical Conference (ATC), pp. 1–14 (2019). [Link]
fl[Link]/paper/duplyakin-atc19
30. Ekici, E., Akyildiz, I.F., Bender, M.D.: A distributed routing algorithm for data-
gram traffic in leo satellite networks. IEEE/ACM Trans. Netw. 9(2), 137–147
(2001)
31. Woollacott, E.: Starlink terminals smuggled into iran - but how effective can they
be? (2022). [Link]
terminals-smuggled-into-iranbut-how-effective-can-they-be/?sh=1c2952561027.
Accessed 28 Oct 2022
32. Foust, J.: SpaceX adds laser crosslinks to polar Starlink satellites (2021).
[Link]
Accessed 26 Jan 2021
33. Giuliari, G., Klenze, T., Legner, M., Basin, D., Perrig, A., Singla, A.: Internet
backbones in space. In: ACM SIGCOMM Computer Communication Review, vol.
50, pp. 25–37. ACM New York (2020)
34. Giuliari, G., Ciussani, T., Perrig, A., Singla, A.: ICARUS: attacking low earth orbit
satellite networks. In: Proceedings of USENIX ATC 2021, pp. 317–331 (2021)
35. Goyal, P., Agarwal, A., Netravali, R., Alizadeh, M., Balakrishnan, H.: ABC: a sim-
ple explicit congestion controller for wireless networks. In: Proceedings of USENIX
NSDI 2020 (2020)
36. Goyal, P., Shah, P., Zhao, K., Nikolaidis, G., Alizadeh, M., Anderson, T.E.: Back-
pressure flow control. In: Proceedings of USENIX NSDI 2022, pp. 779–805 (2022)
37. Handley, M.: Starlink revisions (2018). [Link]
v=QEIUdMiColU&ab channel=MarkHandley
38. Handley, M.: Delay is not an option: low latency routing in space. In: Proceedings
of HotNets 2018, pp. 85–91 (2018)
39. Handley, M.: Using ground relays for low-latency wide-area routing in megacon-
stellations. In: Proceedings of HotNets 2019, pp. 125–132 (2019)
40. Harris, M.: SpaceX Claims to Have Redesigned Its Starlink Satellites
to Eliminate Casualty Risks (2019). [Link]
to-have-redesigned-its-starlink-satellites-to-eliminate-casualty-risks. Accessed 21
Mar 2019
41. Hauri, Y., Bhattacherjee, D., Grossmann, M., Singla, A.: “internet from space”
without inter-satellite links. In: Proceedings of HotNets 2020, pp. 205–211 (2020)
42. Henderson, T.R., Katz, R.H.: On distributed, geographic-based packet routing for
leo satellite networks. In: Proceedings of GLOBECOM 2000, vol. 2, pp. 1119–1123.
IEEE (2000)
43. Hoots, F.R., Roehrich, R.L.: Models for propagation of norad element sets (1980)
44. Hu, M., Xiao, M., Xu, W., Deng, T., Dong, Y., Peng, K.: Traffic engineering for
software defined leo constellations. IEEE Trans. Netw. Serv. Manag. 19, 5090–5103
(2022)
45. Hussein, M., Jakllari, G., Paillassa, B.: On routing for extending satellite service
life in leo satellite networks. In: Proceedings of IEEE GLOBECOM 2014, pp. 2832–
2837. IEEE (2014)
46. Kassing, S., Bhattacherjee, D., Águas, A.B., Saethre, J.E., Singla, A.: Exploring
the “Internet from space” with Hypatia. In: Proceedings of ACM IMC 2020, pp.
214–229 (2020)
340 V. Bhosale et al.
47. Cowing, K.: Euroconsult report addresses challenges and potential of optical
communications for nascent space applications market (2023). [Link]
com/space-commerce/euroconsult-report-addresses-challenges-and-potential-of-
optical-communications-for-nascent-space-applications-market/. Accessed 31 Jan
2023
48. Ancin, K.: 5G + LEO: Verizon and Project Kuiper team up to develop con-
nectivity solutions (2021). [Link]
project-kuiper-team. Accessed 12 May 2022
49. Klenze, T., Giuliari, G., Pappas, C., Perrig, A., Basin, D.: Networking in heaven
as on earth. In: Proceedings of HotNets 2018, pp. 22–28 (2018)
50. Kuiper USASAT-NGSO-8A ITU filing: USA2019-12905 (2018). [Link]
int/ITU-R/space/asreceived/Publication/DisplayPublication/8716
51. Kuiper USASAT-NGSO-8B ITU filing: USA2019-13020 (2018). [Link]
int/ITU-R/space/asreceived/Publication/DisplayPublication/8774
52. Kuiper USASAT-NGSO-8C ITU filing: USA2019-12909 (2018). [Link]
int/ITU-R/space/asreceived/Publication/DisplayPublication/8718
53. LeoSat: Technical Overview. [Link]
[Link]
54. Li, Y., et al.: “internet in space” for terrestrial users via cyber-physical convergence.
In: Proceedings of HotNets 2021, pp. 163–170 (2021)
55. Li, Y., et al.: HPCC: high precision congestion control. In: Proceedings of ACM
SIGCOMM 2019 (2019)
56. Lin, X., Rommer, S., Euler, S., Yavuz, E.A., Karlsson, R.S.: 5G from space: an
overview of 3GPP non-terrestrial networks. IEEE Commun. Stand. Maga. 5(4),
147–153 (2021)
57. Ma, J., Qi, X., Liu, L.: An effective topology design based on LEO/GEO satel-
lite networks. In: Yu, Q. (ed.) SINC 2017. CCIS, vol. 803, pp. 24–33. Springer,
Singapore (2018). [Link] 3
58. Micah Maidenberg, A.S.: Delta Air Lines Tested SpaceX’s Starlink Internet
for Planes Delta CEO Says (2022). [Link]
tested-spacexs-starlink-internet-for-planes-delta-ceo-says-11650316287, Accessed
28 Oct 2022
59. Baylor, M.: With Block 5, SpaceX to Increase Launch Cadence and
Lower Prices (2019). [Link]
increase-launch-cadence-lower-prices/. Accessed 12 May 2022
60. Mohorcic, M., Werner, M., Svigelj, A., Kandus, G.: Adaptive routing for packet-
oriented intersatellite link networks: performance in various traffic scenarios. IEEE
Trans. Wirel. Commun. 1(4), 808–818 (2002)
61. Netravali, R., et al.: Mahimahi: accurate {Record-and-Replay} for {HTTP}. In:
Proceedings of USENIX ATC 2015, pp. 417–429 (2015)
62. Niederstrasser, C.: Small launch vehicles-a 2018 state of the industry survey. In:
Proceedings of AIAA/USU Conference on Small Satellites (2018)
63. Papapetrou, E., Karapantazis, S., Pavlidou, F.N.: Distributed on-demand routing
for leo satellite systems. Comput. Netw. 51(15), 4356–4376 (2007)
64. Perkins, C.E., Royer, E.M.: Ad-hoc on-demand distance vector routing. In: Pro-
ceedings of IEEE Workshop on Mobile Computing Systems and Applications
(WMCSA 1999), pp. 90–100. IEEE (1999)
65. Plac, C.: LEO speed: when milliseconds are worth $Millions... An NSR Insight
(2020). [Link]
worth-millions-an-nsr-insight/. Accessed 31 Oct 2022
A Characterization of Route Variability in LEO Satellite Networks 341
66. Schlicht, A., Marz, S., Stetter, M., Hugentobler, U., Schäfer, W.: Galileo pod using
optical inter-satellite links: a simulation study. Adv. Space Res. 66(7), 1558–1570
(2020)
67. Siddiqi, A., Mellein, J., de Weck, O.: Optimal reconfigurations for
increasing capacity of communication satellite constellations. In: 46th
AIAA/ASME/ASCE/AHS/ASC Structures, Structural Dynamics and Mate-
rials Conference, p. 2065 (2005)
68. Sidibeh, K.: Adaption of the IEEE 802.11 protocol for inter-satellite links in LEO
satellite networks. Ph.D. thesis, University of Surrey (United Kingdom) (2008)
69. Smutny, B., et al.: 5.6 Gbps optical intersatellite communication link. In: Free-space
Laser Communication Technologies XXI, vol. 7199, pp. 38–45. SPIE (2009)
70. SpaceX: FCC Selected Application for Space Exploration Holdings, LLC (2016).
[Link]
71. SpaceX: Spacex Non-Geostationary Satellite System (2019). [Link]
IBFS/SAT-MOD-20190830-00087/1877671
72. SpaceX FCC Filing: SpaceX V-band Non-Geostationary Satellite System (2017).
[Link] key=1190019
73. SpaceX Update: Spacex Non-Geostationary Satellite System (2017). https://
[Link]/myibfs/[Link]?attachment key=1569860
74. SpaceX Update: Spacex Non-Geostationary Satellite System (2020). [Link]
report/IBFS/SAT-MOD-20200417-00037
75. Starlink (2022). [Link]
Accessed 31 Oct 2022
76. Taleb, T., Mashimo, D., Jamalipour, A., Kato, N., Nemoto, Y.: Explicit load bal-
ancing technique for ngeo satellite ip networks with on-board processing capabili-
ties. IEEE/ACM Trans. Netw. 17(1), 281–293 (2008)
77. Telesat: Telesat’s responses - Federal Communications Commission (2018).
[Link] key=1205775
78. Telesat: Application for Modification of Market Access Authorization (2020).
[Link]
79. Wadhwa,V., Salkever, A:: How Elon Musk’s Starlink Got Battle-Tested
in Ukraine (2022). [Link]
musk-satellite-internet-broadband-drones/. Accessed 12 May 2022
80. Vladimirova, T., Sidibeh, K.: Inter-satellite links in leo constellations of small satel-
lites (2007)
81. de Weck, O.L., de Neufville, R., Chaize, M.: Staged deployment of communications
satellite constellations in low earth orbit. J. Aeros. Comput. Inf. Commun. 1(3),
119–136 (2004)
82. Yan, F.Y., et al.: Pantheon: the training ground for Internet congestion-control
research. In: Proceedings of USENIX ATC 2018, pp. 731–743 (2018)
83. Young, C.: SpaceX’s Starlink internet will soon be available aboard cruise ships and
airplanes (2022). [Link]
internet-available-cruise-ships-airplanes. Accessed 28 Oct 2022
84. Zech, H., et al.: Lct for edrs: Leo to geo optical communications at 1, 8 gbps
between alphasat and sentinel 1a. In: Unmanned/Unattended Sensors and Sensor
Networks XI; and Advanced Free-Space Optical Communication Techniques and
Applications, vol. 9647, pp. 85–92. SPIE (2015)
342 V. Bhosale et al.
85. Zech, H., Biller, P., Heine, F., Motzigemba, M.: Optical intersatellite links for nav-
igation constellations. In: Sodnik, Z., Karafolas, N., Cugny, B. (eds.) International
Conference on Space Optics (ICSO 2018), vol. 11180, pp. 370–379 (2019)
86. Zhang, S., Li, X., Yeung, K.L.: Segment routing for traffic engineering and effective
recovery in low-earth orbit satellite constellations. Digital Commun. Netw. (2022)
Topology
Improving the Inference of Sibling
Autonomous Systems
1 Introduction
Autonomous systems (ASes) are the basic constituent elements of the Internet
routing system, managing routing decisions and resources (i.e., IP address pre-
fixes and routers) under a single administrative unit. An AS is uniquely identified
by an Autonomous System Number (ASN), which is assigned by a Regional Inter-
net Registry (RIR) as an identifier in the Border Gateway Protocol (BGP). Each
AS is typically owned by an individual organization, and one organization may
own and operate multiple ASes. ASes owned by the same organization are often
referred to as sibling ASes. AS-to-organization mappings act as a bridge connect-
ing AS-level and organization-level information. An accurate mapping between
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 345–372, 2023.
[Link]
346 Z. Chen et al.
siblings and non-sibling related events compared to CA2O. Our improved AS-to-
organization mappings provide useful context for examining hijacking events and
forensic investigations. Our output dataset is publicly available to the research
community1 .
In addition, RIRs are not responsible for integrating NIRs’ Whois databases into
their RIR Whois database. The structures of Whois databases vary significantly
across RIRs, as they are influenced by the local RIR registration policies. More
details of Whois databases for each RIR are summarized in Appendix A.
Fig. 1. Detection of pools of ASes and organizations that are potentially related accord-
ing to CA2O and PDB.
the involved ASes (e.g., @family-28933). The results demonstrate that CA2O is
so consistent with the Whois mappings that Whois inaccuracies would reflect
directly on CA2O (as described in Sect. 4).
To avoid relying on Whois as the single data source, we leverage PeeringDB
as an extra dataset. We realize that the disagreements between PDB and CA2O
are quite valuable because they help locate potential errors. For example, AS32787
and AS20940 are two famous Akamai ASes with big customer cones (i.e., high AS
ranks), but CA2O does not regard them as siblings and maps them to different
org-objects (Akamai Technologies, Inc and Akamai International B.V ). However,
PDB disagrees with CA2O, where the two ASes are siblings under the PDB orga-
nization Akamai Technologies. The disagreement is a hint directing us to focus
on mappings of the involved CA2O and PDB organizations. Consequently, we can
divide the problem into individual sub-problems based on disagreements, and then
conquer them by manually figuring out the real mappings.
Based on the above observations, the roadmap of our work becomes clear:
discover all the disagreements between CA2O and PDB, dive into the disagree-
ments to figure out the causes of the inaccurate Whois data (which affect the
CA2O dataset), and manually correct the inaccuracies (Sect. 4). Furthermore,
since repeating the manual effort is not scalable, we design an automatic app-
roach, which is able to automatically generate a dataset containing improved
inferences of sibling relations for each new version of CA2O (Sect. 5).
4 Semi-manual Investigation
In this section, we dig into the disagreements on sibling relations between CA2O
and PDB. In Sect. 4.1, we design a pipeline named Pool Detection to automatically
Improving the Inference of Sibling Autonomous Systems 351
locate and categorize disagreements between CA2O and PDB. To identify sibling
relationships and AS-to-organization mappings for each pool, we carry out a man-
ual labeling process as explained in Sect. 4.2. In Sect. 4.3, we identify two pitfalls
of the Whois data, which are the causes of inaccuracies in CA2O. In Sect. 4.4,
we present the results of our investigation and illustrate how the pitfalls influ-
ence the CA2O dataset. Lastly, we briefly introduce a dataset (named reference
dataset) produced by our investigation in Sect. 4.6. We refer to the whole effort as
a semi-manual investigation because it combines the automatic detection of dis-
agreements and the manual labeling process.
We collected both the CA2O and PDB datasets on 2022-07-01. Our dataset
contains 104,153 ASes that were currently allocated by RIRs (i.e., administra-
tively alive [24]) according to the delegation files archived on that day.
PDB does not disagree with CA2O on 19,578 ASes in 6,577 pools. There are
three possible reasons why a pool may lack disagreement: 1) PDB completely
lacks any information on all of the ASes in the pool (9,884 ASes in 3,626 pools);
2) PDB partially agrees with CA2O when PDB only has information on some
of the ASes in a pool (8,588 ASes in 2,475 pools, with PDB having information
on 2,923 ASes), or 3) PDB fully agrees with CA2O (1,106 ASes in 476 pools).
As the primary objective of this study is to address the disagreements in AS
sibling relationships between CA2O and PDB, the remainder of our work centers
on pools where PDB and CA2O have conflicting views on AS sibling relationships.
This includes 961 pools comprising of 9,534 ASes (32.7% of the total sibling
ASes), which are further categorized into the following three mutually exclusive
classes based on the properties of each pool:
i i
Class 1 (1:N): |PORGs.CA2O | = 1 AND |PORGs.P DB | > 1. In this case, the dis-
agreement is that CA2O identifies all the ASes of the pool (two
or more) as siblings, while PDB associates them with different
organizations.
i i
Class 2 (N:1): |PORGs.CA2O | > 1 AND |PORGs.P DB | = 1. In this case, the
disagreement is that PDB identifies two or more ASes as siblings
while CA2O associates them with different organizations.
i i
Class 3 (N:M): |PORGs.CA2O | > 1 AND |PORGs.P DB | > 1. In this case, the
disagreement is due to CA2O finding sibling relationships that
PDB does not recognize and vice versa.
CA2O PDB
Category #Pools #ASes #Orgs #ASes #Orgs
Class 1 (1:N) 544 5,680 544 1,506 1,312
Class 2 (N:1) 337 1,901 791 1,060 337
Class 3 (N:M) 80 1,953 292 817 336
Overall 961 9,534 1,627 3,383 1,985
So far, the Pool Detection locates 961 groups of disagreements where either
CA2O or PDB may contain inaccurate mappings. To determine the root cause of
inaccuracies and correctly establish AS sibling relationships, we must thoroughly
examine each pool individually to identify accurate mappings and sibling rela-
tionships. To this end, we perform a manual labeling process in an attempt at
obtaining ground truth.
– Our labeling process confirms two organizations are under the same entity or
an AS is owned by an organization. For instance, by investigating keywords
of brand names, we recognize that Netflix Inc and Netflix Streaming Services
Inc. are under the same entity.
– Our labeling process finds evidence that the two organizations are owned by
different entities or an AS does not have any relation with an organization. For
example, Skywolf Technology (a [Link]) and LSHIY Network (a [Link])
are in the same pool, where CA2O maps AS7720 (SKYWOLF-AS-AP) to
354 Z. Chen et al.
Class-1. Our study on Class-1 reveals that the APNIC-LIR issue is the sole
cause of disagreement between CA2O and PDB. In other words, CA2O might
wrongly map customer ASes to APNIC LIR organizations but does not miss
siblings. Among the 544 pools, we recognize 26 pools that contain APNIC LIRs,
where 375 ASes are involved. Our manual labeling process corrects the mappings
of 194 out of 375 ASes by associating the ASes to the actual owners (either
[Link] or organizations from descr ), where we confirm the ownership based
on the evidence found by the four-step process above.
As shown in the Class-1 branch of Fig. 3, we first separate the pools whose
i
PASN s contain more than one APNIC-delegated ASes (denoted as candidate
APNIC-LIR pools) to locate the possible APNIC-LIRs (remind we do not have
an official list of APNIC LIRs), because an APNIC LIR must have at least
two APNIC-delegated ASes: one for itself, one for its customer. For the pools
impacted by the APNIC LIR issue, CA2O incorrectly maps all customer ASes,
while PDB is more accurate. Among the 194 mappings that we corrected, 46 ASes
have information in PDB, where 42 of them are accurate. For example, SingTel
Optus is an APNIC LIR, and CA2O considers 63 ASes to be siblings under it.
Though PDB only has information for 3 out of 63 ASes, the AS-to-organization
mappings are all correct: AS9342 (ABCNET-AS-AP) to Australian Broadcasting
Commission, AS9426 (WESTPAC-AS-AP) to Westpac Bank, AS9438 (NETRO-
AS-AP) to Netro. Another important observation is that the descr field con-
tributes more than PDB when correcting the mappings of customer ASes: 152
out of 194 mappings are corrected based on the descr.
The situation is quite different for pools in which [Link] is not an APNIC
LIR. For the other 64 candidate pools (which we confirm the [Link] are not
APNIC LIRs) as well as the other non-candidate pools, CA2O is very accu-
rate while PDB is not. We identify two problems with the PDB data. First,
PDB sometimes over-divides organizations and sibling ASes. For example, PDB
wrongly separates Zettagrid and Conexim Australia as two organizations and
breaks the sibling relation between AS7604 (ZETTAGRID-AS) and AS37996
(CONEXIM-NET-AS-AP). Indeed, Conexim is a subsidiary of Zettagrid, and
CA2O correctly identifies the two ASes as siblings. Second, we discover that
the PDB information could be outdated. For example, CA2O maps AS21461
and AS44700 as siblings under Haendle & Korte GmbH while PDB disagrees
and maps them to two organizations (Haendle & Korte GmbH and Transfair-
Net). After consulting the Internet operator by email, we learned that Haendle
& Korte bought Transfair-Net, and AS21461 would be disabled in near future.
To conclude, if the [Link] of a class-1 pool is an APNIC LIR, the map-
pings of CA2O are problematic for ASes of customer organizations, while PDB
is more accurate. In addition, the descr field in Whois can be a useful source of
information. Otherwise, for the pools without APNIC LIRs, CA2O and Whois
are significantly correct, while PDB tends to be inaccurate.
Class-2. For the pools in Class-2, the APNIC-LIR issue is unlikely to occur, but
the multi-orgID issue often leads to CA2O missing many siblings. Among the
Improving the Inference of Sibling Autonomous Systems 357
337 pools, we correct 306 pools (727 [Link], 1,770 ASes are involved) that
are impacted by the multi-orgID issue. Our manual labeling process merges the
org-objects under the same entity and considers all involved ASes as siblings.
As shown in the Class-2 branch of Fig. 3, PDB is quite accurate for the major-
ity (∼90%) of pools, while CA2O breaks siblings into different org-objects. There
i
are 174 pools where organizations in PORGs.CA2O contain the same brand name
in their org-names (e.g., Netflix Streaming Services and Netflix Inc). In addi-
tion, 32 pools are related to acquisitions or mergers (e.g., Nutrien and Agrium).
For the remaining pools, the [Link] are groups or subsidiaries with differ-
ent brand names. For example, one of the pools contains two [Link] (VIX
Route Server, ACONET ) and one [Link] (University of Vienna). We directly
learned from Vienna University that both VIX and ACONET are owned and
operated by their Computer Center.
For the remaining 10% pools, we consider CA2O to be correct while PDB is
problematic because we do not find any evidence to prove the organizations in
i
PORGs.CA2O operate under the same entity. We also observe that the involved
organizations usually have different websites or LinkedIn profiles. For example,
PDB maps AS24390 (USP-AS-AP) to AARNet, while CA2O maps it to The
University of the South Pacific. We believe CA2O is correct instead of PDB
because AARNet is an ISP providing services to the education and research
communities in the Australian area.
358 Z. Chen et al.
Although CA2O might miss sibling ASes, the existing mappings within each
[Link] (i.e., ASN-org) are significantly accurate: 706 out of 791 (∼90%)
[Link] have ASes all with the same keyword in names. For the other 85
organizations, we also manually verify the correctness. For example, AS29697
(CNS) is correctly mapped to BeeksFX VPS by CA2O, because we found Beeks
group acquired CNS in 2019 [12].
To conclude, the pools of Class-2 tend to be affected by the multi-orgID
issue. For most pools, PDB is correct while CA2O over-divides organizations.
However, PDB is not always accurate, as it is problematic in approximately
10% of pools, while CA2O remains correct in those instances. Furthermore, the
mappings within each [Link] are highly precise, with almost no cases of an
ASN being incorrectly assigned to an organization when it actually belongs to
another organization.
Class-3. The situation in Class-3 is more entangled: the two pitfalls might exist
simultaneously. Moreover, CA2O and PDB might both be inaccurate in one pool.
Our manual labeling process corrects 223 [Link] (1,422 ASes included) that
miss siblings, and corrects 313 mappings of customer ASes where CA2O maps
them to APNIC LIRs.
For 88% of pools without any APNIC LIR (60 out of 68 pools), CA2O misses
some siblings due to the multi-orgID issue. In Class-3, PDB may sometimes over-
divide organizations, unlike in Class-2 where PDB is always correct on the pools
where CA2O is incorrect. We take the pool of Akamai as an example: there are 4
[Link] (Akamai International B.V.; Akamai Technologies, Inc; Linode, LLC
(APNIC); Linode, LLC (RIPE)) and 4 [Link] (Akamai Technologies; Asavie
Technologies; Instart Logic, Inc; Nominum, Inc), where all of the organizations
actually belong to Akamai because of a series of acquisitions.
In addition to the case where all organizations operate under the same entity,
we have observed instances where a pool without APNIC LIRs may contain
two or more completely distinct organizations. Indeed, the involved organiza-
tions operate some ASes together under partnerships, whereas CA2O and PDB
map these ASes differently. For instance, a pool contains 4 [Link] (Arabian
Internet & Communications; Saudi Telecom Company (STC); London Internet
Exchange Ltd ; LINX USA Inc) and 2 [Link] (LINX and Saudi Telecom Com-
pany (STC)), where all 4 [Link] miss siblings: the first two [Link] are
actually under the same entity because of an acquisition, and the latter two are
subsidiaries. Interestingly, the four organizations fall in the same pool because
of AS31177 (JED-IX), which PDB maps to LINX and Whois maps to STC. In
fact, LINX and STC entered into a partnership to form an Internet Exchange
Point called JEDIX in 2018. Our manual labeling process corrects the two orga-
nizations by finding the separated siblings and maps AS31177 to STC because
of the contact email of this AS.
Improving the Inference of Sibling Autonomous Systems 359
In the case of pools that contain APNIC LIRs, the main difference from
Class-1 is that some organizations obtain their ASNs from multiple APNIC
LIRs, which greatly increases the size of the pools. The biggest pool in Class-3
contains 14 [Link], 97 [Link], and 155 ASNs. As an example, YuetAu
Network owns AS147047 and AS138435, which are applied separately through 2
APNIC LIRs (NEXET LIMITED and Aperture Science Limited).
For the 6 pools that are impacted by both issues, we take the xTom pool as
an example: xTom is a hosting provider which provides services in a wide range
of regions. The pool contains 9 subsidiaries of xTom delegated by 4 RIRs except
for LACNIC (e.g., xTom Hong Kong Limited; xTom GmbH ), where the multi-
orgID issue leads to missing sibling relations. Moreover, two of the subsidiaries in
the APNIC region are LIRs, where CA2O also makes mistakes on the customer
ASes. For example, xTom Limited (APNIC) helps Wolf Network Lab to apply for
AS138038 (WOLFLAB-AS-AP), while CA2O wrongly maps AS138038 to xTom.
During our investigation, we identify 8 APNIC LIRs (73 ASes involved), for
which we do not find disagreements in their pools, because none of the customer
ASes maintain any information in PDB. We manually include them to achieve
a more accurate dataset, where details can be found in Appendix C.
Our approach consists of five stages. Initially, we build a graph in which the
nodes are the ASNs, [Link], and [Link] of a pool. Then, we conduct a
three-step data preparation to extract a set of identification features for each
node. To complete the graph initialization, we design and implement different
strategies for different classes, including pre-populating edges and sometimes
adding new organization nodes. Afterwards, we examine each pair of nodes and
populate edges between them if any matching keywords are found in the two
sets of features. Finally, we run a Breadth-First Search algorithm on each graph
and output connected components as clusters of sibling ASes and corresponding
mapped organizations.
Improving the Inference of Sibling Autonomous Systems 361
Data Collection. Given that it is difficult for the automatic approach to take
advantage of the Google search engine and consulting ground truths from Inter-
net operators (as what we do in manual labeling), we partially compensate for it
by leveraging more informative fields. As shown in Table 4, we collect 6 types of
attributes from Whois and PDB: ID, Name, Alias, Descr, Admin, and Website,
where we supplement the website attribute with data from [Link] as well.
For ASN nodes, we collect fields of AS-name, descr, and admin-c from Whois,
where the admin-c field relates to the administrative contact, and the descr field
contains auxiliary information (remind that the two fields could reveal the actual
owners of customer ASes). In addition, we collect alias from the AKA (i.e., also
known as) field of PDB, where some Internet operators record aliases of ASes.
This field might help identify relations between objects involved in acquisitions
or mergers. At last, we collect website URLs from both PDB and [Link]
(18,885 websites from PDB, 9,422 from [Link]), where ASes operated by
different groups of an organization might use the same website URL.
For organization nodes, we collect orgID and org-name from Whois databases
for [Link]. From PDB data, we collect org-name, alias (from AKA field),
and website (13,357 websites) for [Link].
However, there are a few special cases that we should take into considera-
tion. During the manual investigation, we notice 6 APNIC LIRs and 1 APNIC
NIR register all the customer ASes with their own admin-c. To ensure the
accuracy of our final dataset, we do not gather admin-c for ASes in these 7
pools (Appendix C). Such manual input of prior information requires consistent
updating.
In addition, we notice that some website URLs are not up-to-date, as they
automatically redirect to a different domain. Given that new domains may reveal
more information, we employ Selenium in Python to scrape updated website
URLs. For example, AS199422 records [Link] as its website URL
in PDB, however, this URL is redirected to [Link]
because Rezopole was merged to FranceIX in December 2020 [10]. As a result,
we updated the website information of 1,880 ASes and 241 [Link].
Class-1. Our approach only applies to the pools with multiple APNIC-delegated
ASes in Class-1, which are potentially impacted by the APNIC-LIR issue. Given
that the mappings of CA2O on customer ASes are unreliable, we need to indepen-
dently establish relationships between organizations and ASes. Consequently, we
initialize each graph without any AS-organization links from CA2O. In addition,
Improving the Inference of Sibling Autonomous Systems 363
we separate the descr fields from ASN nodes and initialize them as individual
potential organization nodes (i.e., [Link]) according to what we learned from
some APNIC LIRs in Sect. 4.3.
Class-2. The multi-orgID issue is the only potential problem in pools of Class-2,
where CA2O may miss some relations between organizations. During the inves-
tigation, we discovered that the existing mappings within each [Link] are
quite reliable. Thus, we need to find potential relations between different orga-
nization nodes and merge them into bigger clusters. To this end, we keep edges
between AS-organization according to the CA2O mappings. By doing so, each
[Link] and its ASes from CA2O are connected in the initialized graph.
Class-3. Though both pitfalls might exist in pools of Class-3, only the
[Link] with multiple APNIC-delegated ASes are possibly impacted by the
APNIC-LIR issue. For these [Link], we use the same strategy as Class-1
to discard AS-organization links of CA2O and add [Link] as potential orga-
nizations. For the other organizations, we use the same strategy as Class-2 to
keep the CA2O mappings.
So far, we have initialized a graph with four types of nodes and some edges,
where each node is associated with a keyword set. In this stage, we compare
every pair of nodes (i.e., ASN-ASN, ASN-Org, Org-Org) and populate an edge
if there is any same keyword between the two sets. The criterion we used to
compare keywords is a keyword prefix matching: if one word in a keyword set is
equal to or is the prefix of any word in another keyword set, we consider the two
nodes to be related. We do not use simple matching because it might miss some
relations. For example, the keyword prefix matching can find relations between
Internet Systems Consortium, Inc. and AS5277 (ISC-F-AS), since the keyword of
AS5277 (isc) is the prefix of the acronym of the organization (isci). We emphasize
that the risk of mismatching two randomly unrelated organizations is minimized
since the Pool Detection pipeline narrows down the problem scope to related
organizations according to CA2O and PDB.
364 Z. Chen et al.
Table 5. Reconstruction rate of our automatic approach, where A refers to the APNIC
LIR issue and M refers to the multi-orgID issue.
After comparing every pair of nodes by keyword matching, the final step is to
identify clusters of ASes and organizations on the graph that have been created.
We define connected components (CCs) as clusters, where each CC is a set of
nodes that are linked to each other by paths. To find CCs, we run a Breadth-
First Search algorithm on each graph. For each ASN node, the other ASN nodes
in the same cluster are its siblings, and the organizations (from CA2O or PDB
or Descr) in the same cluster are the inferred organizations.
5.7 Evaluation
6
The official romanization system for Standard Mandarin Chinese in China.
Improving the Inference of Sibling Autonomous Systems 365
In this section, we present a case study of BGP hijacking analysis to illustrate the
relevance of our sibling dataset. We focus on Multiple Origin AS (MOAS) events,
which are potentially linked to one type of BGP hijacking attack. A MOAS event
occurs when in BGP, an IP prefix appears to be originated from more than one
ASes [27]. In this context, sibling relationships between involved ASes provide
key information to understand the event, the likelihood of misconfiguration and
to eventually start a forensic investigation. For instance, the sibling relationship
between involved ASes is an important factor when determining if an event is
malicious or not. If no other suspicious behaviors are detected (e.g., the AS is
infiltrated by attackers), the events between sibling ASNs are highly possible to
be non-malicious. As a case study, we collect all 97,975 MOAS events (containing
30,709 pairs of ASNs) monitored by the Global Routing Intelligence Platform
[5] in 2021 and compare the results of using our dataset or CA2O on identifying
events that happened between sibling ASes.
Using our dataset we discover more sibling-related events, also identify several
non-sibling related events which CA2O identifies as sibling-related events. Both
our dataset and CA2O agree on 2,076 pairs of ASes being siblings. However,
our dataset additionally identifies 17% more pairs of sibling ASNs, with a total
of 360 pairs and 4,219 events. We list some examples in Table 6, where the
sibling relationship discovered by our dataset provides more context to the events.
In addition, using our method, we identify 11 MOAS events that happened
between ASes of APNIC Local Internet Registries (LIRs) and customer ASes,
which CA2O considers as sibling-related events. Our dataset provides a more
precise interpretation for these events: it is possible that the LIR serves as the
upstream and originates the prefix in BGP for its customer. For example, our
dataset identifies an event that happened between AS9658 and AS131212, where
366 Z. Chen et al.
Ruralco Holdings Limited by Whois, where Nutrien acquired Ruralco in 2020. How-
ever, since AS137900 is not registered in PDB, and it does not have any sibling
according to CA2O, our Pool Detection isolates AS137900 from the other two
ASes of Nutrien. If the relations between ASes and organizations could be cor-
rectly extracted from the natural language data, the Pool Detection could become
more precise as well as the automatic clustering approach. Towards improving this
problem, leveraging natural language processing methods might be one possible
solution.
Even though the knowledge in PDB is fully leveraged, some information is
still not covered by the datasets we used, especially for mergers and acquisitions.
There are some commercial databases such as Crunchbase and Dun & Bradstreet
which contain plenty of information such as acquisition history, and subsidiary
list. As for the drawbacks, the databases are neither authoritative nor directly
maintained by the operators, and it is hard to validate the information.
Our definitions of ownership and sibling ASes mainly aim for applications related
to Internet behaviors at an organizational level, which may not fit AS-level stud-
ies perfectly. For example, although CenturyLink acquired Level 3 and then
renamed to Lumen in 2020, business types between the two divisions are quite
different: ASes (previously) operated by Level 3 are mainly for transit purposes
(e.g., AS3356), while ASes (previously) operated by CenturyLink are mainly for
residential Internet services (e.g., AS3561). In this scenario, separating ASes of
368 Z. Chen et al.
these two divisions could benefit the AS-level analysis, such as AS-type classi-
fication. One possible solution is a hierarchical structure of AS-to-organization
mappings which also takes the subsidiaries and divisions into account. A pre-
liminary but not verified structure is to organize a tree-like hierarchy for the
pools impacted by the multi-orgID issue, where we place our reference organiza-
tion(s) at the top level, [Link] at the middle level, and ASes at the bottom
as leaves. Consequently, the AS-level analysis only focuses on the middle and
bottom layers, while information on sibling relations between the subsidiaries is
maintained at the top layers. We leave the evaluation of the necessity and effects
of such hierarchical mappings as future work.
8 Conclusion
9 Ethics
APNIC. The bulk Whois data of APNIC is public, while among 7 NIRs, only
JPNIC and KRNIC publish their bulk Whois data. We learned from the APNIC
helpdesk that if NIRs make further assignments within the NIR-maintained
whois database, they may not be reflected in the APNIC Whois database.
20,127 ASes are delegated in the APNIC region including the ones delegated
by the NIRs. aut-num (i.e., autonomous system number) and organisation are
the AS-object and org-object in APNIC Whois, associated with the org field
(i.e., org-id of the organization) in aut-num. However, 8,781 ASes in APNIC
do not have org field (i.e., no related organization objects), where 99.4% of
such ASes are registered in the countries of 7 NIRs. For these ASes, the descr
(i.e., description) field in AS-objects carries the name of the owner organization
Improving the Inference of Sibling Autonomous Systems 369
without association by org-id. The descr field is mandatory [16], and all AS-
objects have such field including the ones with associated organization-objects.
RIPE NCC. The bulk Whois data of RIPE is public. 37,672 ASes are delegated
in the RIPE region, which is the most among the 5 RIRs. RIPE NCC has a similar
structure as APNIC that there are aut-num and organization objects associated
by org-id. Though no NIR exists in the RIPE region, there is still a small amount
of ASes (108 ASNs) without associated organizations, whose holder organization
is in the descr field. Different from APNIC, the descr field is not mandatory and
only 3,962 ASes in RIPE have this field.
AFRINIC. The bulk Whois data of AFRINIC is public. AFRINIC allocates
the least AS numbers among RIRs, where only 2,168 ASes are delegated in the
AFRINIC region. The Whois structure of AFRINIC is similar to APNIC and
RIPE but more consistent: all aut-num objects have org fields associating with
org-objects and the descr field is also mandatory in AFRINIC.
ARIN. The access to ARIN bulk Whois data needs an application (we get access
for this work). 31,446 ASes are delegated in the ARIN region. ARIN uses its own
format of Whois [8]: ASHandle and OrgName are two main objects, associated
by OrgID. AS-objects does not have the descr field and every ASN-object has
an associated org-object.
LACNIC. The access to LACNIC bulk Whois data needs an application (we do
not get access for this work). 12,740 ASes are delegated in the LACNIC region.
To compare CA2O with LACNIC Whois, we conduct a web scraping on the
LACNIC official webpage for Whois to collect the Whois mappings.
For each set of extracted English keywords, we first filter out the words in
the first list. Then we examine if all the remaining words exist in the second list.
If so, we do not use the second list; otherwise, we use the list to filter out part
of the words.
References
1. [Link]. [Link]
2. The CAIDA AS Organizations Dataset. (Downloaded on July 1 (2022)). https://
[Link]/data/as-organizations
3. CAIDA AS Rank. [Link]
4. Daily snapshots of PeeringDB data. (Downloaded on April 4 (2022)). https://
[Link]/datasets/peeringdb/
5. Global Routing Intelligence Platform (GRIP). [Link]
6. GTT acquired Interoute in 2018. [Link]
releases/gtt-to-acquire-interoute/
7. The Internet registry system. [Link]
governance/internet-technical-community/the-rir-system
Improving the Inference of Sibling Autonomous Systems 371
27. Zhao, X., et al.: An analysis of BGP multiple origin AS (MOAS) conflicts. In:
Proceedings of the 1st ACM SIGCOMM Workshop on Internet Measurement, pp.
31–35 (2001)
28. Ziv, M., Izhikevich, L., Ruth, K., Izhikevich, K., Durumeric, Z.: ASdb: a
system for classifying owners of autonomous systems. In: Proceedings of the
21st ACM Internet Measurement Conference, pp. 703–719. ACM, Virtual
Event (2021). [Link] [Link]
10.1145/3487552.3487853
A Global Measurement of Routing Loops
on the Internet
1 Introduction
Routing loops1 are the phenomenon in which packets never reach their desti-
nation because they loop among a sequence of routers. They are the result of
network misconfigurations, inconsistencies, and errors in routing protocol imple-
mentations. In addition to being a pernicious threat to Internet reliability and
reachability [13,29], routing loops can even enable or exacerbate denial of service
attacks [3,26,41].
1
The literature is split between referring to these as “routing loops” [13, 26] or “for-
warding loops” [41]. We use “routing loops” to differentiate them from loops that
arise from application-level redirects [8].
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 373–399, 2023.
[Link]
374 A. Alaraj et al.
Surprisingly little is known about the true global prevalence of routing loops.
Although there have been several large, longitudinal, or distributed Internet
measurements to detect routing loops [2,7,29,42], they tend to make several
simplifying assumptions. For instance, one of the largest studies of routing loops
of which we are aware [41] tracerouted only two IP addresses (.1 and a ran-
dom one) in each of about 5.5M /24 subnets—the implicit assumption being
that addresses within a /24 will largely experience the same routing behavior.
Unfortunately, we are unaware of any prior work validating such assumptions.
In this paper, we perform a straightforward yet illuminating experiment: we
traceroute the entire IPv4 address space from two vantage points. We discover
over 24 million IP addresses with routing loops: over 21× more than two con-
current, distributed traceroute scans [7,33] put together.
What allows us to scan many more addresses than prior work is that we do
not perform full traceroutes of all destination IP addresses. Our insight is that
we only need to use higher TTL values to discover routing loops; lower TTL
values can largely be avoided. We use Yarrp [2], a large-scale network traceroute
tool, to traceroute all 3.7 billion routable IPv4 addresses for a range of 10 TTLs
per IP. To parameterize our scanning rate, we performed experiments to evaluate
routers’ maximum ICMP response rate, to avoid missing routing loops due to
router response rate limits. By using fewer TTLs and more addresses, we are
able to perform, to our knowledge, the most comprehensive study of routing
loops to date.
We analyze our resulting dataset to better understand the nature of routing
loops on today’s Internet. In particular, we explore routing loops’ root causes,
locations, size, and potential impact.
Our results justify full-Internet scanning to discover routing loops: 35% of the
/24 subnets in which we discovered routing loops have at most 10 IP addresses
experiencing routing loops. In addition, scanning just the .1 address in each
/24 would only discover 26.5% of the 320k unique routing-loop containing /24
subnets that we find when we scan every IP. Moreover, we discover that routing
loops are not evenly distributed: within a /24 subnet, routing loops occur more
often at higher last-octet values than low ones.
Collectively, our results demonstrate that the common strategy of scanning
only one or two IP addresses [2,15,16,21,36,41] per /24 subnet is likely to miss
many routing loops. In fact, we find that the common addresses that prior
approaches sample—such as gateways (typically the .1 address of a /24)—have
the least routing loops.
Contributions. We make the following contributions:
– We perform the largest traceroute study of the Internet to date from two
vantage points, tracerouting over 3.7 billion IPv4 addresses.
– We analyze this dataset to understand the prevalence (Sect. 4) and structure
(Sect. 5) of routing loops in today’s Internet.
– We discover that sampling at the /24 subnet granularity is often insufficient
to capture routing loops within the subnet.
A Global Measurement of Routing Loops on the Internet 375
2 Experiment Design
Our general goal and approach are straightforward: identifying IPv4 addresses
with persistent routing loops by running partial traceroutes to every IPv4
address on the Internet. Although conceptually simple, there are several com-
plications to doing this experiment in practice. In this section, we will describe
our methodology and our experiments we used to design this approach.
To issue traceroutes, we use a modified version of Yarrp [2], a traceroute tool
designed to scan large networks. To our knowledge, we are the first to use Yarrp
to scan every IPv4 address; we will describe our minor modifications in Sect. 3
that enable us to use Yarrp in this way.
Many routers have rate limits on how fast they will respond with ICMP Time
Exceeded messages. For tracerouting the Internet from a single vantage point,
this could result in missing data: if we probe hop-sharing routes faster than their
router’s rate limit, those routers may only respond to a subset of our probes,
which would cause our scan to miss hops and undercount routing loops.
To determine rate limits of routers, we wrote a Golang traceroute utility
to send ICMP-eliciting probes to a sample of routers repeatedly at a specific
rate, and observe the speed with which they respond. We selected a sample of
10,000 random IP addresses and sent probes with a TTL limit of 15 hops. We
sent each IP address approximately 50,000 probes per second (about 22 mbps)
for 1 s, and recorded the response rate we received. Our goal is to momentarily
saturate the ICMP response rate limit of each router, allowing us to approximate
its maximum response rate. To minimize interference between responses from
routers, we scanned only 5 routers in parallel during our test.
Of the 10,000 target IPs we probed, 2803 (28%) had an on-path router that
responded with an ICMP Time Exceeded message. In total, we observed 5361
unique router IPs2 . The CDF of the maximum rate we received responses from
these routers is shown in Fig. 2. We observe a median rate limit of about 2,300pps
(about 1.7 mbps), with small but clear steps around likely popular configuration
such as 1,000, 2,500 or 15,000pps. Over 99% of routers that respond at all will
respond to at least 3pps.
To determine our send rate for scanning the Internet, we start with a goal
of averaging 3pps to any individual router (the 99th percentile response rate) to
maximize the chance we receive a response. We observed from preliminary scans
that there are over 35,000 unique router IPs at each hop we scan between 10–25
2
Many targets had multiple routers that responded, due to load-balancing, ECMP,
or the varying paths different probe packets take to reach their destination.
A Global Measurement of Routing Loops on the Internet 377
hops from us. This means that for the 3.7 billion IPs we scan, we expect to hit
each router on average 108,000 times each. If we want to send a probe to each
one no more than 3 times a second, our scan should take approximately 36,000 s,
which equates to a scan rate of 105kpps. We note that it is likely possible to scan
much faster than this, as many routers will not respond to our probes at all, and
higher TTLs will hit even fewer routers. Nonetheless, we conservatively scan at
100kpps for our full IPv4 scans, which takes approximately 10 h per TTL value.
10000 1
Full /24 0.9
8000 Repeated .1
Unique Router IPs
Single .1 0.8
0.7
6000 0.6
CDF
0.5
4000 0.4
0.3
2000 0.2
0.1
0 0
5 10 15 20 25 30 10 100 1000 10000 100000
Hop (TTL) Max send rate (pps)
3 Scanning Methodology
Informed by the experiments in the previous section, here we detail our scanning
methodology. We will issue TCP probes to every IPv4 address with the ACK flag
set to port 80, with TTL values from 246 to 255 (inclusive).
While Yarrp [2] is designed to traceroute large networks, we had to make several
small modifications to allow it to scan every IPv4 address. Internally, Yarrp
has an “Entire Internet” mode, designed to handle scanning large networks.
Unfortunately, the number of bits available in this mode limits the possible
permutation size for scanning to just 1 probe per /24 subnet; outside this mode,
Yarrp’s total permutation size is limited to a domain (IPs × max TTL) of size
232 , still too small to scan the entire IPv4 Internet. We also discovered and
fixed an integer overflow bug that prevented Yarrp from sending probes with
T T L > 127.
To address these limitations, we provided Yarrp with a shuffled list of all
3.7 billion routable IPv4 addresses in a 50 GB file, and had it send probes to a
single TTL at a time. We repeated this 10 times, for each TTL value in (246–
255) inclusive. We limited Yarrp’s sending rate to 100kpps. As Yarrp validates
and logs all of the ICMP Time Exceeded messages it receives, we can construct
partial traceroutes for all IP addresses for the TTL hops we scanned.
A Global Measurement of Routing Loops on the Internet 379
To verify that the results we find from the Yarrp scans are not artifacts of our
modified version of Yarrp, we used the traceroute utility to run TCP ACK
traceroutes. This utility, however, is not designed for Internet-wide scanning: it
has long timeouts, is not designed for scaling to scanning IPs simultaneously,
and generally limited in speed. However, it does provide a useful check against
our results.
To deal with traceroute’s performance limitations, we limited its use to val-
idating the list of 24 million looping IPs we found from VP 1 using Yarrp. We
similarly limited it to the TTL range 246–255, a sending rate of 500 pps, and
a timeout of 500 ms. These parameters allow us to scan the 24 million IPs in a
reasonable time, but note the aggressive timeouts may result in missed hops.
Using our same routing loop definition in Sect. 2.1, traceroute confirms 18 mil-
lion IPs (out of 24 million) appear to be loops. We ran follow-up traceroutes with
longer timeouts on a random sample of IPs found by Yarrp but not by traceroute,
and found about 75% of them do have loops, suggesting that overall traceroute
confirms over 93% of the loops we find. We believe the remaining difference is
largely due to network churn (the traceroute scans happened a month after our
original Yarrp scans).
Overall, we believe that while there may be loops that are transient (only
present temporarily or for short periods of time), the vast majority of the loops
that we identify are indeed persistent routing loops.
Table 2. Scans Results — We count the number of IP addresses with routing loops
along their path. We consider a routing loop to be persistent from a vantage point if it
exists in two consecutive scans from the respective vantage point. Note that 1st Scan
and 2nd Scan are back-to-back consecutive scans run from each vantage point during
the same period of time.
Dataset 1st scan 2nd scan 1st scan ∩ 2nd scan (Jaccard index)
VP 1 26.26 M 26.36 M 24.78 M (0.8898)
VP 2 25.97 M 25.95 M 24.59 M (0.8995)
VP 1 ∩ VP 2 24.46 M 24.47 M 23.28 M (0.9076)
to find persistent loops, and find over 23 million persistent loops in common
between both vantage points. Each vantage point also finds on the order of 5%
of looping IPs that the other does not: VP 1 found 1.5 M loops that VP 2 did
not, while VP 2 found 1.3 M not seen by VP 1.
We find that these relatively small differences are partly due to noise: for
each of the 1.5 M (1.3 M) IPs discovered only by VP 1 (VP 2), we re-ran a
follow-up Yarrp scan from VP 2 (VP 1). Each vantage point found over 400 K
loops that it originally missed, suggesting that network noise (e.g. packet drops
or ICMP rate limits) prevented us from finding these loops originally.
We analyzed routing loops present in the RIPE and CAIDA Ark datasets using
the same routing loop definition in Sect. 2.1. This entails that we enforce the
following requirement: for a destination in the RIPE dataset to be considered
experiencing a routing loop, it must exhibit a looping behavior in two consecutive
scanning windows, where each window is 5 consecutive days long. We enforce
this requirement to exclude transient loops and maintain consistency with our
routing loop definition in Sect. 2.1. Note that we confirmed that over 98% of
destinations probed in the RIPE dataset are probed in both scanning windows
at least once; however, these probes are not necessarily from the same source
nor necessarily traversing the same path. We do not enforce this requirement
on CAIDA Ark dataset since its random nature of scanning does not permit
this kind of validation during the same period of time. This means that some
382 A. Alaraj et al.
Dataset Traceroutes Unique destinations Unique router IPs Looping destinations (%)
This paper 7,400,351,422 3, 700, 175, 711 512, 785 24, 783, 989 (0.66%)
RIPE atlas [33] 977,156,702 848, 348 709, 675 41, 196 (4.85%)
CAIDA ark [7] 213,333,652 206, 007, 571 2, 645, 943 1, 137, 830 (0.55%)
of the routing loops we find in the CAIDA Ark dataset could potentially be
transient loops. Therefore we do not compare the persistent routing loops from
our Internet-wide scans with it. Table 3 compares our dataset with these three
datasets. Note that the number of unique router IPs in the table excludes IPs in
the RFC 1918. Note that the RIPE Atlas probes do not provide any insight on the
types or codes of the ICMP messages in their traceroutes. Therefore we assume
that all ICMP messages in the RIPE Atlas dataset are ICMP Time Exceeded
messages, unless we manually observe otherwise. We find that perspective plays
an important role in finding routing loops: For the scans dataset from VP 1, we
only observe around 43% of the 41,196 routing loops in the RIPE Atlas dataset.
We manually analyze a small random sample of the persistent routing loops
that are found in the RIPE Atlas dataset but not in our Internet-wide scans
for the same period of time. We performed manual traceroutes from VP 1 and
did not observe routing loops for almost the entire sample, suggesting these are
either loops that our scans did not originally identify and yet they resolved after
our scanning window, or merely not visible from VP 1. For the few ones in
the random sample that looped with manual traceroutes but did not exhibit a
looping behavior in our dataset, we find that their neighboring IPs in their /24s
do in fact exist in our dataset, suggesting that we potentially missed these loops
due to ICMP rate limiting or packet loss.
We find over 24 million distinct IPv4 destinations with routing loops from VP 1,
but many of these loops may be caused by a single misconfigured router or
subnet, and could be considered as part of the same loop. For instance, if every
IP address in a /24 subnet has a routing loop involving the same routers, it is
natural to cluster these IPs together into a single loop, and say this single loop
affects 256 addresses, rather than to consider it as 256 unique loops.
In this section, we investigate clustering the loops we found from VP 1 by
grouping destination IP addresses into subnets, and comparing the routers
involved in the loops.
Grouping Destination IPs into Subnets. While ideally a subnet (e.g. a /24) that
has a single loop would show loops for all (e.g. 256) addresses, we must account
for missed packets in our dataset, due to rate limits, packet drops, or other errors.
A Global Measurement of Routing Loops on the Internet 383
For instance, we may find that in a fully-looping /24, we only observed 230 IPs
that looped, but would nonetheless want to cluster these IPs into a single /24
(and not 230 distinct /32s, or some other fragmentation of the subnet).
To do this, we heuristically group looping IPs into the largest-containing
(CIDR prefix) subnet that has more than a threshold fraction of IP addresses
exhibiting a loop. For instance, if we set a threshold of 50%, we would label a /24
as a loop if that subnet contained more than 128 addresses that exhibit looping
behavior.
107 1
> 50% looping Hops
106 > 75% looping Subnets
Number of subnets
5 0.8
10 clustering+grouping
CDF of clusters
104 0.6
103
102 0.4
101
0 0.2
10
0 0
/1
/1
/1
/1
/1
/2
/2
/2
/2
/2
/2
/2
/2
/2
/2
/3
/3
/3
1 10 100 1000
5
6
7
8
9
0
1
2
3
4
5
6
7
8
9
0
1
2
Prefix Number of {Hops,Subnets}
If we were to traceroute only at the subnet level, we would likely miss these
loops, as discuss in Sect. 5.5.
We analyze the scans results from VP 1 and find that we receive at least one
ICMP Time Exceeded message for 29.7 M and 30 M unique destination IPs
probed for the first and second scans, respectively. For the first scan, we label
3.4M reconstructed traceroutes as non-qualifying routing loops (based on our
routing loop definition in Sect. 2.1), of which 2.5M have a single non-repeating
router IP and 3.9K have 10 unique and non-repeating router IPs. For the second
scan, we label 3.7M destination IPs based on their reconstructed traceroutes as
non-qualifying routing loops, of which 2.6M have a single non-repeating router
IP and 19.6K have 10 unique and non-repeating router IPs. These small numbers
(3.9K and 19.6K) of routing loops with a size of more than 10 router IPs suggest
that scanning the Internet for more than 10 TTL hops has diminishing returns
on finding more routing loops.
We adopt the general classification proposed by Xia et al. [41] and further expand
on it. We classify the routing loops we find from VP 1 into two main categories
based on the involvement of the destination AS in the loop. We argue that
this classification helps attribute the root-cause of these loops, because once
packets reach the destination AS, they are subject to the routing policies and
(mis)configurations of the destination AS. We characterize the involvement of
an AS in a routing loop as follows: Given a traceroute for a destination IP for
hops 246–255 that experiences a routing loop, if there exists at least one router
IP address that responds with an ICMP Time Exceeded message, then the AS
of that IP is involved in the loop.
Fig. 5. Examples of three types of routing loops in our dataset. We discovered the root
cause for some routing loops by directly working with network operators of affected
networks.
Loops Involving the Destination AS. We observe over 79% (19.7M) of the
routing loops we find involve the destination AS, indicating that these loops
occur at the edge of the network and closer to their destination IPs. We reached
out to several network operators for networks in which we discovered loops to
ask about their root cause. Between these discussions, and additional follow-up
experiments, we find several types of misconfigurations at the root of the looping
behavior.
386 A. Alaraj et al.
reachable or do not exhibit any looping behavior for TCP SYN packets. To find
the transport-state dependent loops, we run a full Internet scan using Yarrp with
TCP SYN packets. We then calculated the set difference between the persistent
loops in our original ACK scan and this SYN scan. We find over 6M looping IP
addresses that did not exhibit routing loops with SYN packets. These types of
routing loops can only be discovered by ACK scans (or any outstanding TCP
packet that is not a SYN). Figure 5 (b) shows an illustration of this kind of loops.
Fourth, a misconfigured destination in which packets reach the destination
IP but are not delivered due to some unclear misconfiguration at the destination.
We observe 27K destinations in our dataset that exhibit this misconfiguration.
Figure 6 (f) shows an example of this kind of loops.
Loops Not Involving the Destination AS. Over 21% (5M) of the persistent
routing loops we discover from VP 1 do not involve the destination AS. Using
CAIDA’s dataset of AS classification [6], we classify the relationship between
the AS(es) in the loop and the AS of the destination. We do not investigate the
root-cause of these loops and leave this to future work.
No Apparent Relationship. For 5.81% (291K), we cannot find any relation-
ships between the destination AS and the ASes involved in the loop based on the
traceroutes in our dataset. Figure 6 (g) shows an example of this kind of loops.
Customer Relationship. We find 79.3% (3.97M) have at least one router
IP belonging to an AS to which the AS of the destination is a customer. For
4.86% (244K), the destination AS is in the customer cone of one of the ASes
involved. Figure 6 (d) shows an example of this kind of loops. After analyzing
the relationships in that loop, we conjecture that our odd TTL probes expire at
an AS (AS267699) to which the destination AS is a customer. Whereas our even
TTL probes expire at an AS (AS12956) to which AS267699 (the provider AS to
the destination AS) is a customer.
Provider Relationship. For 4.75% (238K), we find that the destination AS is
a provider to one of the ASes involved in the loop. Figure 6 (h) shows an example
of this kind of loops.
Peering Relationship. For 1.51% (76K) of loops, we find that the destina-
tion AS is in a peering relationship with at least one AS in the loop. And for
0.07% (3842), we find one of the ASes in the loop to be in the customer cone of
the destination AS. For instance, the loop to [Link] (AS20485) involves
the router IP [Link] (AS50439) which is in the customer cone of the
destination AS.
No AS Announced. Finally, for around 3.66% (183K), we find that the router
IPs in the loop are not announced by any AS, therefore we cannot infer any
relationships
We note that it is possible that the destination AS of a given looping IP is, in
fact, involved in its routing loop, but the destination AS’s routers limit their
388 A. Alaraj et al.
Topologically. Using the April 17, 2022 RouteViews dataset [28], we analyze
the Autonomous System Number (ASN) of each of the destination IPs as well
as the routers involved in the persistent routing loops we discovered from VP 1.
Although other router ownership inference approaches [23] provide a better accu-
racy, we used the IP-to-ASN mapping approach for simplicity. We exclude router
IPs that are not announced by any AS (which amount to only 0.74% of router
IPs in our dataset).
Figure 7 shows the distribution of how many ASes are involved in each routing
loop. While over 91% of loops contain routers within a single AS, we find many
loops that include multiple ASes. For instance, we find 3 routing loops with
as many as 6 distinct ASes involved in the loop. Figure 6 (e) shows one such
example, involving routers from Akamai (in Miami, FL, USA), multiple ASes of
Telefonica Brazil (in Sao Paulo, BR), and intermediate routers from GlobeNet.
We find that over 25% of all routing loops have at least one of the involved
routers in a different AS than the destination of our probe. Meanwhile, 0.11%
of routing loops involve the destination IP itself in the routing loop, as shown
in Fig. 6 (f). We performed follow up traceroutes on a small sample of these
addresses (e.g. [Link] and [Link]) and confirmed that the round trip
time for each probing packet increases with higher TTL values for these peculiar
routing loop cases, which suggests a looping behavior.
Geographically. Using the April 2022 MaxMind [25] dataset, we analyze the
geolocation of the destination IPs in the persistent routing loops we found from
VP 1. We find 19% are in the US, followed by 6% in each of India, Brazil and
Japan, followed by 5% in China.
We also find that 5% of the loops have a destination IP address in a different
country than at least one of the routers in the loop. However, we note that this
may simply be due to geolocation error: prior work has shown that geolocating
router IPs can be complex and leads to inaccuracies [11]. We leave correcting
this problem to future work, such as potentially using router hostnames to infer
location [20].
A Global Measurement of Routing Loops on the Internet 389
108
Fig. 7. Routing loops sizes and the number of ASes they span — For each
persistent routing loop, we count the number of unique router IPs, and the num-
ber of unique Autonomous System Numbers by resolving the routers’ IP addresses to
Autonomous System Numbers using recent datasets.
Routing loops can range in size from one repeated router to containing many
routers in a repeating loop. However, we measure the size of the routing loops in
our dataset based on the number of unique router IPs, meaning that a routing
loop could have multiple router IPs in it, yet belonging to a lesser number of
physical routers. We do not perform router dealiasing [35] techniques on our
dataset.
The majority (57%) of routing loops we find contain two unique router IP
addresses, a result corroborated by prior work [41]. We also find larger loops—up
to 9 unique IP addresses—in a small percentage of cases (0.03%). Figure 7 shows
the distribution of loop sizes observed.
We find that routing-loop containing /24 subnets are limited to a small number
of looping IPs in them. For each /24 subnet, we counted the number of target
IP addresses that contained a routing loop in the corresponding traceroute.
We expected that if a /24 subnet contained a routing loop to one IP address,
it would also contain routing loops to many other IP addresses in the same
subnet. Figure 8 shows the distribution of looping IP addresses per routing-
loop containing /24 subnet, ranging from 1–256 loops. Over half of the /24
subnets that contained a routing loop had fewer than 25 looping IP addresses
(out of 256). While it is possible that some subnets only responded to a subset
of our probes, our scanning rate should have had each /24 receive a packet from
us every 2 min on average, well under the ICMP rate limits for most routers
(see Fig. 2).
390 A. Alaraj et al.
0.1
Fraction
0.01
0.001
0.0001
1 16 32 64 96 128 192 224 256
Number of loops per subnet
Fig. 8. Loops per /24 subnet — For each /24 subnet, we count the number of des-
tination IP addresses whose path showed a routing loop. 35% of subnets had at most
10 IP addresses that had a loop along their path, and only about 19% had over 200 IP
addresses. This demonstrates that subnet sampling is likely to miss routing loops.
110000
105000
100000
Count
95000
90000
85000
0 16 32 64 96 128 192 224 255
Last octet
Fig. 9. Last octet distribution — We count the number of looping IP addresses for
each last octet (e.g. 1.2.3.x). The .1 address contains the least number of routing loops
(85,254), while the .255 contains the most (106,950), likely due to being a common
broadcast address. (Note the y-axis does not start at 0)
How Many IPs per Subnet are Needed to Find Looping Subnets? A
natural follow-up question is if there is any sample (short of scanning all 256
addresses) of a /24 that can be probed to find most loops?
Our dataset identifies over 320k unique /24 subnets that contain routing
loops. Sending probes to only the .255 last octet in all /24 networks would have
identified just over 33% of these subnets-containing loops. Adding additional
last octets would help find more, but as Fig. 10 shows, it would take over 45 IPs
probed per /24 to discover over 90% of the routing-loop containing /24 subnets
that we find when scanning all IPs. We believe this result justifies scanning all
addresses, and not sampling a handful of IPs per subnet.
Prior work has already shown that routing loops can have an impact on network
attacks. In 2021, researchers discovered that middleboxes could be weaponized to
launch reflected amplification attacks, and that routing loops could be abused to
make the attack more damaging [3]. An attacker launches the attack by spoofing
their source IP address to that of their victim and sending a packet sequence that
contains a request for some forbidden resource. When the middlebox responds
to the seemingly forbidden request (such as by sending a block page), it will
send this response to the victim, effectively amplifying the attacker’s traffic.
The authors found that if a vulnerable middlebox is within a routing loop, the
middlebox can be re-triggered each time the packets circle the routing loop,
significantly improving the amplification factor for the attacker (up to infinite
amplification for infinite routing loops). In this regard alone, we believe that
trying to identify routing loops is an important goal. Note that [3] examined
the top 1 million amplifying hosts (by number of packets sent) from their SYN;
PSH+ACK scan to identify which ones have routing loops along their path. We,
however, approach this differently.
392 A. Alaraj et al.
1
0.9
Fig. 10. Finding Loops by sampling — We measured how many unique /24 subnets
that contain loops can be found as a function of the number of last-octets needed to be
scanned. For instance, by only scanning the .255, .127, .0, and .191 last octets (the
top 4 in our dataset), one could identify over 51% of the /24 subnets that contained
routing loops (compared to scanning all last octets). However, it would take scanning
over 45/256 IPs per /24 to discover over 90% of unique /24s containing loops.
After we scanned the entire IPv4 Internet from VP 1 using ACK packets and
found over 24 M destinations experiencing persistent routing loops, we now ask
the question of how many of these routing loops can be weaponized in a TCP
reflected amplification attack. We discussed in Sect. 5.1 that some routing loops
are transport-state dependant and that scanning the IPv4 Internet with ACK
packets reveals additional 6 M routing loops that cannot be discovered with a
SYN scan.
We performed two forbidden scans 5 on the 24 M looping destinations using
two different packet sequences, namely a single PSH+ACK, and a SYN followed by
PSH+ACK (SYN; PSH+ACK).
We first performed the PSH+ACK single packet scan. We found 273,201 looping
destinations to be true amplifiers with an average amplification rate of 386.73,
a median of 2.46 and a maximum of 1,414,267. We note that routing loops with
an amplification rate of ≥ 100 are concentrated at 6601 destinations from 119
different ASes in 13 different countries.
Second, we performed the SYN; PSH+ACK scan. We found 1,089,457 looping
destinations to be true amplifiers with an average amplification rate of 4.44, a
median of 1.02 and a maximum of 1204.37.
This shows that finding and exploiting the transport-state dependant loops
renders TCP reflected amplification attacks more effective.
In addition to this, we explore an avenue of the potential impact of routing
loops that, to our knowledge, has not been studied before: are there any (named)
services behind the destinations that experience routing loops? To answer this
question, we reached out to a few AS operators for some of the looping IPs in
our dataset to draw their attention to these loops and obtain ground-truth infor-
mation about the root-causes of them. We received a response from a university
5
[Link]
A Global Measurement of Routing Loops on the Internet 393
network engineering team detailing the cause of the loops in their network. We
elaborated on the root-cause of these loops in Sect. 5.1. Based on the nature
of these NAT-related loops, the looping IPs in this case do not have services
that otherwise could have been reachable, since the IPs are used to translate
internal addresses to external, public ones. We do not know how prevalent this
kind of loops is, as they do not have distinguishing characteristics. For our cor-
respondence with other AS operators, we unfortunately have not received any
response.
In an attempt to find whether any domain name resolves to a looping IP in
our dataset, we used ZDNS6 to resolve the top 5 M domain names in the tranco
list [31]. We found 1256 domain names that resolve to 719 looping IPs in the
dataset from our ACK scan from VP 1. Then using ZMap, we ran a SYN scan
on these 719 IPs and found that 400 of them respond with SYN;ACK indicating
that they do not exhibit any looping behavior when scanned with SYN packets.
We also manually browsed a sample of the domains behind these 400 IPs and
were able to reach and view their web content as if no loops existed along their
paths. After further examination, we found that these 400 IPs experience the
transport-state dependant kind of loops that we discussed in Sect. 5.1. Since the
web content for the domains behind the 400 transport-state dependant looping
IPs is reachable and browse-able, the impact of their loops can be dismissed;
however, this kind of loop consumes their network resources unnecessarily and
can be used against these domains to disrupt their service.
6 Related Work
Our work extends prior work primarily by probing many more destination IP
addresses, uncovering many more routing loops than previously found, uncover-
ing new types of routing loops that to our knowledge have not previously been
known, and discovering that IP addresses within the same /24 can experience
different routing loop behavior. We compare to prior work in terms of how we
find routing loops, and what our findings reveal.
Identifying Routing Loops. Our overall approach to detecting routing
loops—looking for repeated entries in a traceroute—is well-established in prior
literature. To name a few: Paxson [29] ran periodic traceroutes between 27 sites
and looked for IP addresses repeated at least three times to infer the presence
of a routing loop. Paxson further differentiated routing loops as being persistent
(they never reached the destination during the traceroute) and temporary (they
did resolve during the traceroute, and ultimately reached the destination). Xia
et al. [41] also look for repeats in traceroutes to identify persistent routing loops,
and Lone et al. [19] similarly use them to infer the absence of ingress filtering.
Other studies have investigated how routing loops can be predicted by
changes in BGP [38,45], how the presence of loops can indicate route leaks [18],
and the dynamics of transient loops [30]. However, all of these works rely on
6
[Link]
394 A. Alaraj et al.
small-scale traceroutes of samples of the Internet, and use these to infer trends
on the larger Internet.
In addition to these active techniques, routing loops have also been detected
passively. Hengartner et al. [13] used packet traces from a tier-1 ISP to detect
routing loops. By comparison, this method does not have as broad a view of
routing loops as ours (they find only 4318 routing loops in total), but is able to
detect transient routing loops—they find that most routing loops last less than
10 s.
Performing Massive Scans. Where we primarily differ from the above is the
sheer scale at which we actively probe for routing loops. Prior work has used
dozens [29], hundreds of thousands [42], and as many as 11M [41] destination IP
addresses to discover routing loops. More recently, FlashRoute [15] presented a
tool capable of tracerouting one IP per /24 subnet in a matter of minutes. Using
this technique, they also discover over 16K prefixes that contain routing loops.
Rüth et al. found 439K prefixes containing routing loops in 2019 using follow-up
traceroutes to ZMap scans, performing 27M traceroutes [34].
In contrast to prior work, our study scans the entire public7 IPv4 address
space: over 3.7 billion IP addresses in total, and we find over 24 million IPs that
contain a routing loop, comprising over 320K /24 subnets.
Prior work also assumed that sampling one or two IP addresses within each
routable /24 would provide representative samples [2,16,21,36,41]. Our work
challenges this assumption by showing that routing loops are not uniform within
/24 boundaries—rather, to obtain a global view of routing loops, far more com-
prehensive scans are necessary.
Prior Findings About Routing Loops. Paxson [29] observed that persistent
routing loops existed as early as 1996. To estimate the scale of persistent routing
loops, Xia et al. [41] issued traceroutes to two IP addresses (.1 and a random
one) in each of about 5.5M /24s. From this, they identified 207,891 /24s with at
least one temporary routing loop; of those, they repeatedly probed the two IP
addresses and found that 135,973 of the /24s had loops for each traceroute. Xia et
al. assumed that if the two candidate addresses experienced routing loops, then
all addresses in the /24 would, as well, resulting in their estimate of 135,973 *
28 ≈ 35M total IP addresses with routing loops. Our work draws this assumption
into question by empirically demonstrating that different IP addresses in the
same /24 can exhibit different looping behavior.
Representative Scanning. Heidemann et al. [10] scanned the entire IPv4
Internet and proposed a method for generating responsive, complete and sta-
ble hitlist, a list of representative, alive IPv4 addresses that should suffice to
represent their respective /24s during an Internet scan. We differ in this matter–
we scan the Internet not to discover live hosts, but rather routing loops. We also
show that the /24 granularity is not sufficient to discover all routing loops on
the Internet.
7
We adopt ZMap’s blacklist, which excludes reserved addresses, private addresses, the
loopback prefix, and the addresses reported to us to opt-out from being scanned.
A Global Measurement of Routing Loops on the Internet 395
Preventing Loops. There is also work on mitigating routing loops and the
harm caused by them, using reconfigurable networks [37], static analysis of the
data plane [22], and real-time detection of transient loops in network traffic [14,
17]. These works primarily focus on fixing or preventing loops in the first place,
rather than measuring them comprehensively across the Internet as we do.
ICMP Rate-Limiting Studies. To parameterize our tool, we evaluated the
maximum rate at which we could probe routers without hitting a rate limit
(Sect. 2). We are not the first to do such a study. Ravaioli et al. [32] sent TTL-
limited ICMP echo requests from 180 PlanetLab hosts at a rate of 1–4,000pps,
with TTLs ranging from one to five. They reported that 60% of the routers exhib-
ited rate limits. By comparison, we explored larger send rates (up to 50,000pps)
and larger TTL values (1–15), but from only a single vantage point. Guo and
Heidemann [12] performed a different experiment, measuring the rate-limiting
of pings (not ICMP Time Exceeded responses), and observed only six out of
approximately 40,000 subnets rate-limiting them on the forward path. In con-
trast, our study focuses on ICMP Time Exceeded messages.
Broad Implications of Loops. Persistent routing loops can be used to perform
DoS attacks against the involved networks directly, by exploiting the simple fact
that a single packet will get relayed multiple times [40]. But more recent work has
shown how routing loops can be used to enable or exacerbate attacks indirectly
as well. For instance, Bock et al. [3] detail how middleboxes can be used for DoS
amplification attacks, and how routing loops around vulnerable middleboxes can
worsen these attacks. Nosyk et al. [27] detail how routing loops can exacerbate
DNS-based DoS attacks, finding 115 routing loops that enable high-amplification
attacks. Attackers can also create or induce their own routing loops in order to
amplify DoS attacks, by misconfiguring content delivery networks [8], leveraging
IPv6 tunnels [26], or ARP spoofing on a wireless network [4]. Finally, Marder et
al. [24] detail how loops complicate inferring outbound addresses in traceroutes,
which is useful for inferring router ownership or organizations.
BGP AS Path Looping Behavior (BAPL). Loops can also be observed at
the BGP level, by looking for loops in the AS paths of BGP update messages [43].
These so-called BGP AS path looping (BAPL) behavior can result in multi-AS
routing loops [44].
7 Ethics
We designed our experiments to have a minimal impact on other hosts. Our
IPv4 Internet-wide scans were designed such that each intermediate router would
receive a packet from us at a rate of 3 packets per second, which should have a
negligible impact on end hosts. To avoid overwhelming destination networks, our
experiments probed addresses randomly, which spreads out traffic to any given
destination network across the length of the scan. Our scanning rate should have
had each /24 receive a packet from us every 2 min on average.
396 A. Alaraj et al.
We follow the best practices for high speed scanning laid out by [9]. Our
scanning machines we used for these experiments hosted a simple webpage on
port 80 to explain the nature of our scans and provides a contact email address
to request exclusion from future scans.
For our highest throughput experiments, we took additional precautions. We
planned our experiment to saturate the ICMP rate limit of specific routers, we
limited the experiment to a relatively small number of IP addresses (10,000).
For each IP address, the experiment was limited to 1 s long and only 22 mbps.
8 Conclusion
Routing loops have long been known to exist, but their prevalence and nature
have long been shrouded behind common but untested assumptions, like that
all destination addresses within a /24 are likely to experience the same routing
loops. In this paper, we perform a straightforward but illuminating experiment:
we look for routing loops to all IPv4 addresses by tracerouting to a limited range
of TTLs to examine the true global prevalence of routing loops. We discover over
24 million IPs with routing loops—over 21× more than three concurrent datasets
combined—comprising over 500K routers. And we discuss their structure and
root causes. Also, we uncover new types of routing loops that can be abused for
TCP reflected amplification attacks.
Our resulting datasets confirm some prior results, but also expose unidentified
biases in prior measurement efforts that may inform future studies. In particular,
we find that scanning only the .1 address per /24 misses 73.5% of the routing
loops we were able to find. Indeed, for 35% of the /24s in our dataset, fewer than
11 of their 256 destination addresses result in loops.
Ultimately, our results motivate full-Internet traceroutes. To assist in future
efforts, we made our code and data publicly available at [Link]
RoutingLoops.
References
1. Augustin, B., et al.: Avoiding traceroute anomalies with Paris traceroute. In: ACM
Internet Measurement Conference (IMC) (2006)
2. Beverly, R.: Yarrp’ing the Internet: randomized high-speed active topology discov-
ery. In: ACM Internet Measurement Conference (IMC) (2016)
3. Bock, K., Alaraj, A., Fax, Y., Hurley, K., Wustrow, E., Levin, D.: Weaponizing
middleboxes for TCP reflected amplification. In: USENIX Security Symposium
(2021)
4. Brown, J.D., Willink, T.J.: A new look at an old attack: ARP spoofing to create
routing loops in Ad Hoc networks. In: Ad Hoc Networks, pp. 47–59 (2018)
A Global Measurement of Routing Loops on the Internet 397
23. Marder, A., Luckie, M., Dhamdhere, A., Huffaker, B., Claffy, K., Smith, J.M.:
Pushing the boundaries with bdrmapIT: mapping router ownership at internet
scale. In: ACM Internet Measurement Conference (IMC) (2018)
24. Marder, A., Luckie, M., Huffaker, B., Claffy, K.: VRFinder: finding outbound
addresses in traceroute. In: Proceedings of the ACM on Measurement and Analysis
of Computing Systems, vol. 4(2) (2020)
25. MaxMind: GeoLite2, October 2021. [Link]
geolite2
26. Nakibly, G., Arov, M.: Routing loop attacks using IPv6 tunnels. In: USENIX Work-
shop on Offensive Technologies (WOOT) (2009)
27. Nosyk, Y., Korczyński, M., Duda, A.: Routing loops as mega amplifiers for DNS-
based DDoS attacks. In: Hohlfeld, O., Moura, G., Pelsser, C. (eds.) PAM 2022.
LNCS, vol. 13210, pp. 629–644. Springer, Cham (2022). [Link]
978-3-030-98785-5 28
28. University of Oregon: Route Views Archive Project, October 2021. [Link]
[Link]/bgpdata
29. Paxson, V.: End-to-end routing behavior in the Internet. In: ACM SIGCOMM
(1996)
30. Pei, D., Zhao, X., Massey, D., Zhang, L.: A study of BGP path vector route looping
behavior. In: IEEE International Conference on Distributed Computing Systems
(ICDCS) (2004)
31. Pochat, V.L., Goethem, T.V., Tajalizadehkhoob, S., Korczyński, M., Joosen, W.:
Tranco: a research-oriented top sites ranking hardened against manipulation. In:
Network and Distributed System Security Symposium (NDSS) (2019)
32. Ravaioli, R., Urvoy-Keller, G., Barakat, C.: Characterizing ICMP rate limitation
on routers. In: IEEE International Conference on Communications (ICC) (2015)
33. RIPE NCC Staff: RIPE Atlas: A global internet measurement network. Internet
Protocol J. 18(3), 2–26 (2015)
34. Rüth, J., Zimmermann, T., Hohlfeld, O.: Hidden treasures – recycling large-scale
Internet measurements to study the Internet’s control plane. In: Choffnes, D.,
Barcellos, M. (eds.) PAM 2019. LNCS, vol. 11419, pp. 51–67. Springer, Cham
(2019). [Link] 4
35. Sherry, J., Katz-Bassett, E., Pimenova, M., Madhyastha, H.V., Anderson, T.,
Krishnamurthy, A.: Resolving IP aliases with prespecified timestamps. In: ACM
Internet Measurement Conference (IMC) (2010)
36. Sherwood, R., Bender, A., Spring, N.: DisCarte: a disjunctive Internet cartogra-
pher. In: ACM SIGCOMM (2008)
37. Shukla, A., Foerster, K.T.: Shortcutting fast failover routes in the data plane.
In: Symposium on Architectures for Networking and Communications Systems
(ANCS) (2021)
38. Sridharan, A., Moon, S.B., Diot, C.: On the correlation between route dynamics
and routing loops. In: ACM Internet Measurement Conference (IMC) (2003)
39. Wang, F., Mao, Z.M., Wang, J., Gao, L., Bush, R.: A measurement study on
the impact of routing events on end-to-end Internet path performance. In: ACM
SIGCOMM (2006)
40. Xia, J., Gao, L., Fei, T.: Flooding attacks by exploiting persistent forwarding loops.
In: ACM Internet Measurement Conference (IMC) (2005)
41. Xia, J., Gao, L., Fei, T.: A measurement study of persistent forwarding loops on
the Internet. Comput. Netw. 51(17), 4780–4796 (2007)
A Global Measurement of Routing Loops on the Internet 399
42. Zhang, M., Zhang, C., Pai, V., Peterson, L., Wang, R.: PlanetSeer: Internet path
failure monitoring and characterization in wide-area services. In: Symposium on
Operating Systems Design and Implementation (OSDI) (2004)
43. Zhang, S., Liu, Y., Pei, D.: A measurement study on BGP AS path looping (BAPL)
behavior. In: International Conference on Computer Communication and Networks
(ICCCN) (2014)
44. Zhang, S., Liu, Y., Pei, D., Liu, B.: Measuring BGP AS path looping (BAPL) and
private AS number leaking (PANL). Tsinghua Sci. Technol. 23(1), 22–34 (2018)
45. Zhang, Y., Mao, Z.M., Wang, J.: A framework for measuring and predicting the
impact of routing changes. In: IEEE Conference on Computer Communications
(INFOCOM) (2007)
as2org+: Enriching AS-to-Organization
Mappings with PeeringDB
Augusto Arturi1 , Esteban Carisimo2(B) , and Fabián E. Bustamante2
1
Universidad de Buenos Aires, Buenos Aires, Argentina
aarturi@[Link]
2
Northwestern University, Evanston, Illinois, USA
{[Link],fabianb}@[Link]
1 Introduction
2
Legacy resources [4]—allocations preceding the creation of RIRs—are also subject
to different regulations [46].
404 A. Arturi et al.
While we argue that the growing popularity and use of PeeringDB can offer
a complementary perspective to traditional WHOIS-based approaches, its use
is not without challenges. For instance, the database is voluntarily and does
not provide complete or uniform coverage across regions which could potentially
introduce biases in AS-to-Organization mappings. We also find that despite PDB
providing an Organization Identifier (OrgID), operators sometimes rely on other
fields to communicate siblings, and that in some cases those siblings do not even
have presence on PeeringDB (e.g., Tigo-AS262206 reports AS26617 as a sibling
as2org+: Enriching AS-to-Organization Mappings with PeeringDB 405
in text fields but this network is not registered in PDB). We also find that this
information is often loosely structured as it is intended to be read by human
operators.
In the following paragraphs we discuss some additional challenges with using
PDB to identify ASes belonging to an organization, and potential approaches to
take advantage of its rich information.
(a) (b)
Fig. 1. PeeringDB adoption as fraction of active ASes (left) and per region (right).
Geographic Bias: PDB adoption rate may vary across countries depending
on peering incentives (e.g., local presence of HGs), consolidated peering ecosys-
tems (e.g., presence of large IXPs), common communication practices among
local operators, etc.. A previous study conducted in 2013 found that RIPE is
3
We refer as active ASes to Autonomous System Numbers visible in BGP routing
tables.
406 A. Arturi et al.
We now analyze the PDB data schema (version 2) to identify elements that
could potentially inform siblings. We investigate whether being an operation-
oriented database could bring a different perspective to the sibling inference
problem compared to WHOIS information, which refers as an organization to
legal entities in a specific RIR. We identify two main ways organizations use to
communicate the set of ASes under their management: (i) use of native features
of PDB data schema (org data structure) (ii) custom use of plain text fields
(e.g., aka, notes).
Among the several data entities available in PDB, we focus on those that
are more relevant for this work: organization (org) and network (net). The
data entity org describes organizations with fields such as name, also known as
(aka), website, address, country, etc.. However, the most important attribute
of these entities is the network field which is a list of network identifiers referring
to net entries administered by the organization The data entity net describes
ASes with fields such as name, also known as (aka), network type, several
network attributes (e.g., number of IPv4 prefixes), peering, and more impor-
tantly the organization field referring to the organization this network belongs
to. By combining both data entities using the list of bidirectional network/or-
ganization identifiers, we can directly generate AS-to-Organization mappings.
7 " notes " : " nLayer / AS4436 has been acquired by GTT
Communications / AS 3257 and is no longer directly
peering . Please refer all peering related inquiries to
peering [ at ] gtt [ dot ] net ." ,
8 " org_id " : 8897 ,
9 " policy_url " : " http : // www . gtt . net / peering /" ,
10 " aka " : " Formerly known as nLayer Communications " ,
11 }}}
Listing 1.1. Example of the net entry for AS4436 in the PDB snapshot of Oct. 2020.
We further investigate whether the content reported in fields of org and net
could provide some information of other ASes operated by the same organization.
Listing 1.1 shows an example of some fields in the net entry of AS4436 (nLayer)
to explain how these fields could provide hints about siblings. In this specific case,
nLayer was acquired by GTT in 2012 [55] and this information is available in the
notes, where both nLayer (AS4436) and GTT (AS3257) ASNs are included. Ten
year later, both networks are under different organizations in WHOIS records.
The use of aka in this example is informative but insufficient to obtain a cluster
with ASNs of both networks. In Appendix A we include an example of a net
entry in which operators used the field aka to report ASes under the same
management.
4 Methodology
The conservative approach only uses org id present in the net data entity
and it does not apply any heuristic to infer siblings. On the other hand, the
aggressive approach applies heuristics to extract self-reported siblings ASNs
embedded in either the aka field or the notes field (or both). In this approach,
we create candidate groups of ASes under the same administration as an output
and later apply filters to improve confidence. The conservative approach is a
zero-risk approach since PDB applies mechanisms to authenticate the ownership
of a network resource (see §3.1), preventing two non-sibling ASNs from being
identified by the same org id . The aggressive approach could potentially include
numbers that are not ASNs under the same managements, though, those false
positives are mitigated by the design of our framework. We give users full control
of the combination of these approaches where they can choose any combination
of features. Next, we describe the implementation of our heuristics.
Before starting our process, we sanitize the data from the selected inputs and
normalize the text (e.g., case).
Our PDB-based inference methodology uses three fields of PDB’s net entity,
the org id field in the conservative approach, and notes and aka fields in the
aggressive approach. In this stage, the conservative approach uses the org id
to group together all ASNs were registered by the same organization while the
aggressive approach combines regular expressions (regexes) to extract groups
of ASNs embedded in these fields. Next, we describe the rules applied to extract
self-reported siblings embedded in these fields.
org id. This feature extraction mechanism leverages the native org id field
in the PDB data schema to group together all ASes registered by the same
organization.
aka. For this field, the framework applies a single regular expression that extracts
numbers with 4 to 8 digits to generate the list of candidate siblings. We suspect
that length constraints of this field (limited to 255 characters [48]) discourage
operators from rich semantic statements and hence, sibling ASNs (sometimes
along with AS names) are directly reported. In Appendix B we show a few
examples of how operators report their networks in the aka field as well as the
output of this regex. We acknowledge that this extraction method can result in
wrong inferences. This rule is not capable of inferring candidate siblings ranging
between AS1 and AS999. However, this impact is limited to missing at most 1%
of the siblings since at the time of this submission more than 100,000 [44] have
already been allocated. To be more specific, this rule lacks the semantic context
of the numbers extracted, potentially leading to infer as candidate sibling strings
such as dates and phone numbers. We apply custom filters (§4.2) to mitigate the
presence of spurious numbers.
notes. We develop 37 regexes to extract candidate lists of sibling ASes embed-
ded in different semantic contexts in the notes field. This is a data rich field
(it is an unlimited plain text field [48]) that allows operators to include details
as2org+: Enriching AS-to-Organization Mappings with PeeringDB 409
that do not fit well in any other field, including detailed descriptions or specific
requirements and procedures to peer with the network. The flexibility of the field
and diversity of data reported (siblings, peering policy, capacities, NOC hours,
etc..) sets challenges to identify a candidate list of siblings. Moreover, there is no
convention to report these features, and the text structure can vary significantly
as these messages are meant to be read by human operators.
We categorize the 37 regexes into two groups: simple rules (21) and complex
rules (16). Simple rules aim to extract ASN from simple patterns that are used
to refer to ASes using prefixes such as AS, ASN, ASNS, ASS and ASES, as
it is shown in Table 1. Complex rules aim to extract ASNs from notes using
more complex semantic expressions. We search for common phrases used (with a
maximum of three words) to report ASes under the same management, including
also manages, we administered, merging, as it is shown in Table 2. Due to a lack
of a common structure, we consider candidate siblings to all numbers after this
template phrase. This decision comes at the risk of including numbers unrelated
to ASNs, such as addresses, RFC numbers, ISO standards and others. We also
acknowledge that complex rules are only capable of extracting siblings of records
written in English. In our implementation users can select using simple, complex
or both rules for sibling inferences. In Sect. 6.3 we evaluate the contribution of
each of these rules.
4.2 Filters
Figure 3 shows the most prevalent numeric expression across all notes of
the snapshot of October 1, 2020. The most prevalent numeric expressions were
extracted from notes describing protocol versions (4 and 6), maximum prefixes
accepted/announced (50 or 100), and popular subnet masks (21, 22, 24 and 30
for IPv4 and 48, 64 and 80 for IPv6). The spurious-number filter includes the
most prevalent number expression appearing in at least 15 notes where a knee
is observed in Fig. 3.
This filter also drops numbers that range between 1970 to 2020 since these
numbers tend to refer to dates such as merging dates, last update, etc.. (e.g.,
number of prefixes, phone numbers, addresses, years, etc..).
as2org+: Enriching AS-to-Organization Mappings with PeeringDB 411
Table 3. Example of the p2c filter to filter out notes containing ASNs not related to
the to network entry.
We release our code4 to allow users to make changes in these rules such as
adding and removing them if they consider it necessary.
Customer-to-Provider Filter. We use AS relationships to remove ASNs that
are not part of the same organization. The aggressive approach could potentially
group together ASNs that do not belong to the same organization but both being
present in the same note. In development stage of the project, we found networks
that use their notes to describe their upstream connectivity rather than listing
other networks of the same organization. We then develop a stage to filter out
clusters based on customer-to-provider (c2p) relationships. In our implementa-
tion, users can specify the maximum c2p relationships allowed between ASes
in the same cluster or skip this stage. In cases where this filter is applied, our
PDB-based inference methodology returns a file containing the list of discarded
clusters. Users manually verify these cases (§4.3) and decide to either include or
exclude them from the final inference.
Table 3 show this rule in action in an example in which only one c2p rela-
tionship is allowed for two different notes. In this case, the inferred cluster for
Maxihost (AS396356) is dropped because this network has more c2p relation-
ships that the maximum allowed in this example. Indeed, as the example shows,
Maxihost (AS396356) is describing its upstream connectivity. On the other hand,
the cluster inferred from Telecom Argentina (AS7303) meets the criteria used
for this example (only one c2p allowed) and it is then preserved.
Sibling relationships generate anomalies in the inference of AS relation-
ships [16,20,40,51]. These anomalies challenge to distinguish customer-provider
relationships is between two independent companies or two companies belong-
ing to the same conglomerate. Given that text fields can indistinctly siblings or
4
as2org+ can be found at: [Link]
412 A. Arturi et al.
To conclude the data extraction process, the framework includes a last stage
for human inspection to manually remove errors that were not filtered out in
the previous automatic stages. This stage also allows users to apply their own
judgment to filter out clusters generated by correctly extracting data, though
from entries with mistakes (e.g., typos).
The lack of authentication of the information given in text fields could be
another source of erroneous inferences that requires human inspection. For exam-
ple, we found that for a short period of time (from 2019 to 2020) a Bangladeshi
provider called Brother Online (AS135131) was using its aka field to report
“AS32934” (Meta’s principal peering network) making our PDB-based inferenc-
ing method to group both networks together (see Appendix C). We are unaware
whether this was an unintended or malicious event, though, this event highlights
the sensitivity of as2org+ to imprecise or unauthenticated data provided in text
fields as well as the need of human inspection to rule out these cases.
After extracting and cleaning the embedded data, the framework groups together
partially overlapping clusters scattered across multiple records, fields and data
sources. There are some cases in which sibling information is scattered in the
same field (e.g., notes) across multiple records rather than being centralized.
This is illustrated in Listing 1.2 where both networks report the same parent
network but none of them reference each other. Another popular case is to find
sibling information scattered across multiple fields (e.g., notes and org id). This
stage concludes combining clusters in our PDB-based approach with clusters in
the AS2Org dataset [8] to create a dataset that we call as2org+.
1 # StarHub AS10091
2 { ’asn ’ : 10091 ,
3 ’ notes ’ : ’ Please refer to as4657 PDb for Contact & Peering
Info . Thanks . ’ , }
4 # StarHub AS38861
5 { ’asn ’ : 38861 ,
as2org+: Enriching AS-to-Organization Mappings with PeeringDB 413
Table 4. Effectiveness of the extraction methods given by true positive (tp), False
Positive (f p), false negative (f n), True Negative (tn), accuracy (A), precision (P) and
Recall (R) values.
’20 45 188 161 338 988 18115 ’19 3 (145) 65 (145) 3 (44) 10 (44) 48 (796) 7 (796)
’21 44 208 186 400 1171 20704 ’20 2 (161) 75 (161) 3 (45) 12 (45) 57 (988) 8 (988)
’22 48 234 214 472 1384 23191 ’21 3 (186) 88 (186) 4 (44) 12 (44) 65 (1171) 8 (1171)
’22 2 (214) 106 (214) 3 (48) 11 (48) 76 (1384) 7 (1384)
contributes 4.79% and 16.42% on average compared to org id during this period.
We expect a more prevalent use of org id to report networks under the same
management since this is a native (and compulsory) field The results also suggest
that aka and notes are used to communicate relationships that are not captured
by the org id.
We further investigate partial overlaps between non-atomic clusters inferred
using different features. We specifically look for cases where a cluster inferred
by a field (e.g., notes) is fully contained in a cluster inferred by another field
(e.g., org id). By meeting this condition, the former field would provide no con-
tribution since that information is available in the latter field. Table 7 shows the
number of clusters inferred by each feature that are fully contained in clusters
inferred by another feature. We observe that clusters inferred using notes and
aka fields are rarely contained in each other. This is notably different when we
compute the overlap between notes and aka with org id where up to half of
those clusters are contained in the org id. We suspect that in these overlaps
attempt to make sibling information available in text format at a glimpse. In
any case, the low fractions in these overlaps suggests that each feature provides
a unique contribution that is not visible by any other way.
We investigated the lack of partial overlap between text fields and the org id
and found that this mostly occurs after mergers and acquisitions. We suspect
that this common practice allows operators to quickly communicate mergers and
acquisitions rather than migrating networks to a different PDB organization. We
also believe that the visibility of text fields may be more effective to inform these
changes to other operators.
the prevalence of aka and notes containing numeric expressions. We then use
this information to compute the average number of ASNs extracted per record
containing numeric expressions
Table 8 shows the number of aka and notes fields with non-empty records
(ē), those containing numeric expressions (n) and the number of ASNs extracted
for a 5-year period. Overall, notes are rarely used—only a fraction from 0.15
to 0.13 contains data—and aka (0.47 to 0.54) is more commonly used, however,
both rarely contain numeric expressions (fractions oscillate around 0.08 and 0.03
respectively). Interestingly, the ratio between fields containing numeric expres-
sions and the total number of ASNs extracted is between 0.3 and 0.5, showing
that on average fields with numeric expressions provide 0.3 to 0.5 ASNs per
field. We also observe that the fraction of non-empty records, those containing
numeric expressions, are stable over time while the number of ASNs embedded
in notes augmented in the same period. This growth suggests that notes are
being more frequently used to report other ASes under the same management.
Table 9 shows a Venn diagram with the overlap between the clusters inferred
with simple rules, complex rules and both using a snapshot collected on Septem-
ber 1, 2022. We observe that simple rules capture 90.4% of the clusters (464/513)
while the remaining clusters are observed when complex rules are applied solo
or in combination of simple rules. This is a remarkable observation since simple
rules have patterns that are less prone to capture spurious numbers (we recall
Table 1) and they are highly successful in extracting embedded siblings. This
finding also shows that despite there being no standard format to report sib-
lings, operators mostly use similar unsophisticated patterns. A final observation
is that simple and complex rules infer some identical clusters that are not visible
when both rules are combined. This behavior is due to the fact that the combi-
nation of rules can create a more rich clusters in the entire dataset and some of
these new enriched clusters eventual merge and create discrepancies.
Considering that notes and aka are free text fields, we investigate the use of
these fields to report siblings that are not registered in PDB.
Table 10 shows the number of ASes registered in PDB, the number of ASNs in
clusters inferred from notes and aka fields and the number of those inferred ASN
that have not been registered in PDB. For the 2018–2022 period, we observe that
the prevalence of unregistered ASNs is more significant in aka than in notes,
ranging between 0.23 and 0.15 and 0.09 and 0.06 respectively. We also note
that both trends have been declining over time, though for aka roughly 15% of
the siblings reported are not present in PDB records. We suspect that opera-
tors sometimes report unregistered ASNs in a single record to reduce manage-
ment overhead associated with registration and maintenance of multiple records.
Despite being convenient, reporting ASNs that are not present in PDB lacks
authentication and it is unclear whether these ASNs are in fact all under the
same management.
418 A. Arturi et al.
We now shift our attention to the c2p filter to investigate the impact of dif-
ferent c2p threshold values. In this analysis, we evaluate the trade-off between
discarding false positive inferences (i.e., clustering ASNs from different organiza-
tions) and discarding correctly inferred clusters (i.e., an inferred c2p relationship
between ASes of the same organization).
Table 11. Impact of the c2p filter on the sibling inferences as the number of filtered
clusters with different threshold values. We use three posible outcomes, (i) positive
(ASNs were not under the same management), (ii) negative (ASNs were under the
same management) and (iii) neutral (ASNs were under the same management but the
same information is available through the org id. Numbers in parenthesis correspond
to the fraction of clusters in that category of a c2p threshold value.
We recall that the c2p filter discards clusters (before the data consolida-
tion stage) when the number of c2p relationships across members exceeds the
threshold value (§4.2). For the evaluation, we apply five different threshold val-
ues (0-5) to the snapshot of September 1, 2022. We conduct human inspection to
assess whether the cluster was successfully removed based on the text provided
in the notes. Table 11 shows the results for this human inspection where filtered
clusters are categorized into three types: (i) positive (ASNs were not under the
same management), (ii) negative (ASNs were under the same management) and
(iii) neutral (ASNs were under the same management but the same informa-
tion is available through the org id). The results show that filtered out clusters
are mostly legit siblings and a small fraction of them contain networks report-
ing their upstream connectivity. The overlap between notes and org id (§6.1)
partially mitigates the impact of removing valid clusters.
We manually examined the filtered clusters that contain upstream providers
under different managements. We found that these networks use their notes
to list their connectivity with several large transit networks (e.g., Level3-3356,
Telecom Italia-6762, GTT-3257) that belong to different corporations (see an
example in Appendix D). This example argues in favor of implementing the c2p
filter as a mechanism to prevent our approach from clustering together high-
profile networks that belong to different organizations.
To summarize this analysis, the c2p filter successfully removes false posi-
tive inferences but with the cost of also discarding clusters containing siblings.
The consequence of this filter is that it introduces a human examination phase
as2org+: Enriching AS-to-Organization Mappings with PeeringDB 419
Non-atomic clusters
# Clusters (AS2Org) # clusters # ASes
Field Year All Unmodif. as2org+ AS2Org as2org+ AS2Org Migrant ASes
notes 2018 71288 70806 5529 5729 20994 21373 1518
2019 75223 74979 5925 5932 22498 22348 759
2020 79126 78870 6407 6424 24529 24385 815
2021 86565 86255 6833 6856 26052 25872 1149
2022 90508 90144 7272 7324 27771 27580 1444
aka 2018 71288 70921 5528 5729 20917 21373 1348
2019 75223 75122 5935 5932 22413 22348 363
2020 79126 79022 6420 6424 24446 24385 367
2021 86565 86454 6849 6856 25936 25872 401
2022 90508 90402 7311 7324 27635 27580 712
org 2018 71288 70382 5526 5729 21261 21373 3168
2019 75223 74358 5946 5932 22906 22348 2659
2020 79126 78154 6438 6424 25002 24385 2991
2021 86565 85474 6865 6856 26561 25872 3613
2022 90508 89251 7338 7324 28387 27580 4150
then narrow the analysis and specifically examine the contribution of as2org+ in
modifying the number of non-atomic clusters. We finally contrast both datasets
from the AS-level perspective and investigate the prevalence of migrant ASes,
ASes that moved into a new cluster after adding the PDB-based inference.
Table 12 shows the contribution of the PDB-based inference approach to
as2org+ when different features are used in the 5-year dataset described in §6.
We observe minor modifications to the number of clusters (including non-atomic
clusters), independent of the snapshot and feature used. It is worth noting that
we expect to see minor changes since the Internet is mostly composed of small
single-AS organizations. The number of clusters in AS2Org before and after
combining it with PDB-based inferences shows minor changes too due, in part,
to the impact of the consolidation stages that groups together clusters when they
partially overlap. Nonetheless, the number of migrant ASes reaches 4150 (≈4%
of the ASes in AS2Org database) using the org id field in the 2022 snapshot.
Despite these changes appearing negligible, it is important to examine what ASes
and organizations are being modified by this contribution. In the next section,
we explore some aspects of the network to put in perspective the impact of these
changes.
(a) Changes in the number of siblings (b) Changes in the number of siblings in
inferred as a function of CAIDA’s AS- organization operating Hygergiants.
RANK.
5
The list is composed of Apple-AS714, Amazon-AS16509, Facebook-AS32934,
Google-AS15169, Akamai-AS20940, Yahoo!-AS10310, Hurricane Electric-AS6939,
OVH-AS16276, LimeLight-AS22822, Microsoft-AS8075, Twitter-AS13414, Twitch-
AS46489, Cloudflare-AS13335 and Edgecast-AS15133.
as2org+: Enriching AS-to-Organization Mappings with PeeringDB 423
8 Related Work
Despite the popularity of both WHOIS and PeeringDB datasets, to the best
of our knowledge, there is no prior work that has combined both datasets to
address the AS-to-Organization mapping problem.
Our work builds on the seminar work by Cai et al. [57] which created an
automated methodology using WHOIS records to generate AS-to-organizations
mappings, and Hyun et al. [29] which discusses of common practices in the use of
multiple ASes for a single organization and introduces the idea of using WHOIS
records to identify ASes under the same administration.
The WHOIS data received notorious attention given that this database offers
information that is not embedded in network protocols interactions. To enable
characterizations of the .com WHOIS data, Liu et al. [36] proposed parse and
structure WHOIS query responses using a conditional random field model. For
a different purpose, Livadariu et al. [37] examined WHOIS records to contrast
the results of IP geolocation services finding partial overlaps in geolocation and
delegated country fields.
A number of research efforts relied on PeeringDB as a source of topologi-
cal data. Lodhi et al. [38] investigated the accuracy and representativeness of
PDB records finding strong correlations between address space, traffic volume
and geographic footprint in these records and other sources of network data.
Bottger et al. [5] used several network features publicly reported in PeeringDB
to identify the most prominent CDNs (Hypergiants). Other research efforts relied
on PeeringDB’s AS-to-facilities lists to detect ASes footprint and facilities out-
ages [23,24]. In a recent work, Carisimo et al. [9] leveraged PeeringDB data to
identify ASNs belonging to the same organization in the context of state-owned
Internet Operators.
Acknowledgements. This work was partly funded by the research grant CNS-
2107392.
Listing 1.3. Example of the net entry for AS7303 in the PDB snapshot of October 1,
2020.
Table 13 shows examples in which operators use the field aka to report siblings
and the results obtained after applying the extraction rules.
Table 13. Examples of regex bieng applied to extract siblings from aka field.
Listing 1.4. net entry of AS135131 in the PDB snapshot of October 1, 2020.
Listing 1.5. net entry of 30081 in the PDB snapshot of October 1, 2020.
References
1. Albert, R., Jeong, H., Barabási, A.L.: Error and attack tolerance of complex net-
works. Nature 406(6794), 378–382 (2000)
2. ARIN: Rdap: Whois for the modern world (2016). [Link]
2016/05/26/rdap-whois-for-the-modern-world/
3. ARIN: Organization identifiers (org ids) (2022). [Link]
guide/account/records/org/
4. ARIN: Organizations holding legacy resources (2022). [Link]
resources/guide/legacy/
5. Böttger, T., Cuadrado, F., Uhlig, S.: Looking for hypergiants in PeeringDB. ACM
SIGCOMM Comput. Commun. Rev. 48(3), 13–19 (2018)
6. CABASE: Instructivo peeringdb (2022). [Link]
wp-content/uploads/2014/09/[Link]
426 A. Arturi et al.
27. Holz, R., et al.: Tracking the deployment of tls 1.3 on the web: A story of exper-
imentation and centralization. ACM SIGCOMM Comput. Commun. Rev. 50(3),
3–15 (2020). [Link]
28. Huston, G.: The death of transit and the future internet. In: ITU Workshop on
Network, vol. 2030 (2018)
29. Hyun, Y., Broido, A., Claffy, k.: Traceroute and BGP AS path incongruities. Tech.
rep., Cooperative Association for Internet Data Analysis (CAIDA) (2003–03)
30. ICANN: Registration data access protocol (RDAP) (2022). [Link]
org/rdap
31. Jin, Y., Scot, C., Dhamdhere, A., Giotsas, V., Krishnamurthy, A., Shenker, S.:
Stable and practical AS relationship inference with ProbLink. In: Proceedings of
USENIX NSDI (2019)
32. Kashaf, A., Sekar, V., Agarwal, Y.: Analyzing third party service dependencies in
modern web services: Have we learned from the mirai-dyn incident? In: Proceedings
of IMC, IMC 2020 pp. 634–647. Association for Computing Machinery, New York
(2020). [Link]
33. Labovitz, C., Iekel-Johnson, S., McPherson, D., Oberheide, J., Jahanian, F.: Inter-
net inter-domain traffic. ACM SIGCOMM Comput. Commun. Rev. 40(4), 75–86
(2010)
34. Laskowski, P., Chuang, J.: Network monitors and contracting systems: competition
and innovation. Proc. ACM SIGCOMM 36(4), 183–194 (2006)
35. Liu, E., Akiwate, G., Jonker, M., Mirian, A., Savage, S., Voelker, G.M.: Who’s
got your mail? characterizing mail service provider usage. In: Proceedings of IMC,
IMC 2021, pp. 122–136. Association for Computing Machinery, New York (2021).
[Link]
36. Liu, S., Ian, Foster, Savage, S., Voelker, G., Saul, L.: Who is. com? learning to
parse WHOIS records. In: Proceedings of IMC, pp. 369–380 (2015)
37. Livadariu, I., et al.: On the accuracy of country-level IP geolocation. In: Applied
Networking Research Workshop, pp. 67–73 (2020)
38. Lodhi, A., Larson, N., Dhamdhere, A., Dovrolis, C.: kc Claffy: Using peeringDB
to understand the peering ecosystem. ACM SIGCOMM Comput. Commun. Rev.
44(2), 20–27 (2014)
39. Lodhi, A.H.: The economics of Internet peering interconnections. Ph.D. thesis,
Georgia Institute of Technology (2014)
40. Luckie, M., Huffaker, B., Dhamdhere, A., Giotsas, V., Claffy, K.: As relationships,
customer cones, and validation. In: Proceedings of IMC (2013)
41. Microsoft: Prerequisites to set up peering with Microsoft (2022). [Link]
[Link]/en-us/azure/internet-peering/prerequisites
42. Moura, G.C.M., Castro, S., Hardaker, W., Wullink, M., Hesselman, C.: Clouding
up the internet: How centralized is dns traffic becoming? In: Proceedings of IMC,
IMC 2020, pp. 42–49. Association for Computing Machinery, New York (2020).
[Link]
43. NCC R: Youtube hijacking: A ripe ncc ris case study (2008). [Link]
net/publications/news/industry-developments/youtube-hijacking-a-ripe-ncc-ris-
case-study
44. Nemmi, E.N., Sassi, F., La Morgia, M., Testart, C., Mei, A., Dainotti, A.: The
parallel lives of autonomous systems: ASN allocations vs. BGP. In: Proceedings of
IMC, pp. 593–611 (2021)
428 A. Arturi et al.
45. [Link]: Collaboration between [Link] and PeeringDB is helping improve Internet
traffic exchange in Brazil (2020). [Link]
between-nic-br-and-peeringdb-is-helping-improve-internet-traffic-exchange-in-
brazil/
46. Nobile, L., Morris, T.: Status and solutions for WHOIS data accuracy (2022).
[Link] Nobile Whois Data [Link]
47. PeeringDB: Approving network (net) objects (2022). [Link]
committee/admin/approval-guidelines/#approving-network-net-objects
48. PeeringDB: Peeringdb api documentation (2022). [Link]
apidocs/#operation/create%20net
49. Quoitin, B., Pelsser, C., Bonaventure, O., Uhlig, S.: A performance eval-
uation of bgp-based traffic engineering. Int. J. Netw. Manag. 15(3), 177–
191 (2005). [Link] [Link]
doi/abs/10.1002/nem.559
50. Spring, N., Mahajan, R., Anderson, T.: The causes of path inflation. In: Proceed-
ings of ACM SIGCOMM, SIGCOMM 2003, pp. 113–124. Association for Comput-
ing Machinery, New York (2003). [Link]
51. Subramanian, L., Agarwal, S., Rexford, J., Katz, R.H.: Characterizing the Internet
hierarchy from multiple vantage points. In: Proceedings of IEEE INFOCOM (2002)
52. Testart, C., Richter, P., King, A., Dainotti, A., Clark, D.: Profiling bgp serial
hijackers: Capturing persistent misbehavior in the global routing table. In: Pro-
ceedings of IMC, IMC 2019, pp. 420–434. Association for Computing Machinery,
New York (2019). [Link]
53. Testart, C., Richter, P., King, A., Dainotti, A., Clark, D.: To filter or not to filter:
measuring the benefits of registering in the RPKI today. In: Sperotto, A., Dainotti,
A., Stiller, B. (eds.) PAM 2020. LNCS, vol. 12048, pp. 71–87. Springer, Cham
(2020). [Link] 5
54. USA T.M: T-mobile completes merger with sprint to create the new t-mobile
(2020). [Link]
55. Wire, B.: Gtt acquires nlayer communications, inc. (2012). [Link]
[Link]/news/home/20120501005550/en/GTT-Acquires-nLayer-
Communications-Inc
56. Wu, J., Zhang, Y., Mao, Z.M., Shin, K.G.: Internet routing resilience to failures:
analysis and implications. In: Proceedings of CoNEXT, pp. 1–12 (2007)
57. Xue, X., Heidemann, J., Krishnamurthy, B., Willinger, W.: Towards an AS-to-
organization map. In: Proceedings of IMC (2010)
RPKI Time-of-Flight: Tracking Delays
in the Management, Control, and Data
Planes
Romain Fontugne1(B) , Amreesh Phokeer2 , Cristel Pelsser3 , Kevin Vermeulen4 ,
and Randy Bush1,5
1
IIJ Research Lab, Tokyo, Japan
romain@[Link], randy@[Link]
2
Internet Society, Reston, USA
phokeer@[Link]
3
UCLouvain, Ottignies-Louvain-la-Neuve, Belgium
[Link]@[Link]
4
LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France
[Link]@[Link]
5
Arrcus, Inc, San Jose, USA
1 Introduction
The Border Gateway Protocol (BGP [1]) is the ubiquitous inter-domain routing
protocol of the Internet. Unfortunately, like the rest of the early Internet, it was
designed with no thought to security. One of the main efforts to secure BGP is
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 429–457, 2023.
[Link]
430 R. Fontugne et al.
AS connected mainly to ASes performing ROV (§ 4), and another one with
ASes surrounded by some, but not all, ASes performing ROV to generalize our
findings (§ 5.1).
A Landscape of the Impact of ROV Adoption on the Internet: With
these experiments, we found that: (1) There was a significant time dispar-
ity across RIRs between the operator’s input and the ROA publication delay
(Table 2 and 3). (2) This observation allowed us to discover some startling
anomalies (since corrected after our notification) at ARIN and LACNIC that
delayed their ROA publication time by up to five hours (§ 4.1). (3) There is
an important disparity in ISPs’ reaction time between ROA creation and ROA
deletion, ranging from minutes to an hour. ISPs take significantly more time to
act on ROA deletion than ROA creation (§ 4.2); (4) We also reported anoma-
lous behavior to a Tier1 network which was quickly corrected. (5) There are
vast differences between the RIRs’ administrative practices seriously complicat-
ing the experiment setup, and highlighting how difficult it can be for operators
to streamline their RPKI management procedures at the different RIRs (§ 6).
Extending the Findings with a Longitudinal Study: We further broaden
our study with an analysis of historical RPKI and BGP data (§ 5.2) showing
that the bugs reported to RIRs have been present for years and that long delays
of ROA creation have been quite stable over the past four years.
Inter-RIR Differences in ROA Payloads: Our analysis of RPKI data also
reveals ROA structural differences between the five RIRs, highlighting RIRs’ dif-
ferent management of RPKI data and explaining some of their disparities (§ 5.3).
2 Background
RPKI prefix allocation follows the IANA allocation hierarchy, and each RIR
maintains a separate trust anchor (TA) for the resources for which they are
responsible. Certificates are issued to their members, which are then used to
sign ROAs. Each RIR operates a public repository in which all RPKI objects
(certs, ROAs, CRLs, manifest files [3,17]) are stored.
Figure 1 depicts the steps performed when a resource holder queries an RIR to
update RPKI information for its prefixes. Then the changes are fetched by oper-
ators performing Route Origin Validation (ROV-enabled ASes, green in Fig. 1)
that use this new information to update their routers. Each step described below
is common to all RIRs and ROV-enabled ASes, but each may perform these steps
at different time intervals and frequency.
ROV-enabled ASes check route validity based on the information contained in
ROAs. To get ROA information, routers need to connect to Relying Party (RP)
software which is in charge of fetching ROAs, cryptographically validating their
content, and feeding routers with Validated ROA Payloads (VRPs). Based on the
VRPs, routers can then classify BGP announcements either as Valid, NotFound,
or Invalid. ROV-enabled ASes typically drop the “Invalid” announcements.
432 R. Fontugne et al.
Fig. 1. Data-flow from creation of a ROA by the prefix holder to the corresponding
BGP updates recorded at the route collectors (RIS / RouteViews). The red labels on
the left show the points at which time measurements were taken. (Color figure online)
Each step in the provisioning process introduces delay. The aim of this study
is to track and quantify these delays across RIRs and some ISPs. For this we
collect timestamps at the following points:
1. User Query : The most common way for resource holders to create ROAs
is to query the RIR that provided the IP prefixes. The queries are either via
the RIR’s web portal or the RIR’s REST API if available.
2. ROA Signing : RIRs collect user queries, verify that they are legitimate,
and pass them to certification authority software which computes ROAs and
corresponding metadata information (i.e., manifest and CRL files) and creates
new signed files.
3. ROA Publication: Then RIRs place new ROAs and metadata files into pub-
lic repositories, called Publication Points (PPs), so that Relying Party (RP)
software can fetch them when desired. This seems a simple step, but RIRs
must ensure that RPKI objects are consistent at all times, hence metadata
files and their corresponding ROAs must be atomically published.
4. Relying Party (RP) Validation: RPs are deployed by ROV-enabled ASes
and their role is to periodically fetch and validate all the objects from the
global RPKI repositories. After validation, they produce a list of Validated
RPKI Time-of-Flight 433
Table 1. Summary for the two experiments presented in Sect. 4 and 5.1.
ROA Payloads (VRPs) which routers use to verify incoming BGP announce-
ments. Larger ISPs often deploy multiple RPs to avoid single points of failure,
and to tune the timing with which each RP visits the Publication Points.
5. BGP Update: ROV-enabled routers accept and advertise the Valid and
NotFound announcements only and drop the Invalid ones based on the VRPs
from the RPs. These changes propagate globally in BGP and the effects may
be seen in BGP collection systems (e.g., RIS or RouteViews [15,18]).
6. Traceroute: These routing changes are reflected in the data plane and can
be observed with measurement platforms such as RIPE Atlas [16].
3 RPKI Beacons
To measure the propagation time of RPKI data from RIRs’ Certification Author-
ities to BGP speaking routers we automated RPKI ROA beaconing at each of
the five RIRs. Each beacon is a prefix for which we switch its RPKI status
daily by creating and deleting ROAs. We announce these experimental prefixes
in BGP from a few locations on the global Internet and measure the beacons’
effects in the management, control, and data planes.
We perform two experiments from diverse ASes to measure the propagation time
of RPKI data (Table 1): For the first experiment (§4), we obtained from each
RIR a pair of IPv4 /24 prefixes and a pair of IPv6 /48 prefixes1 . One prefix from
each pair of prefixes is used as a control, while the other is the test prefix. The
control prefixes are expected to be always reachable, with an always valid RPKI
status. If they are not reachable then we know that the experiment is not valid
for that period. For the test prefixes, the BGP announcements do not change,
but we periodically add and remove a ROA to alternatively validate and invali-
date the origin AS of the test prefixes’ BGP data. We track changes reflected at
the management (RPKI), control (BGP), and data plane (traceroute). The pre-
fixes are announced from AS3970 , which is directly connected to AS3130 (not
implementing ROV), which in turn is connected to two ROV-enabled upstream
providers, NTT (AS2914) and Sprint (AS1239), and peering with a ROV-enabled
route server at a large IXP, and directly with a few non-ROV IXP peers. The
results for this experiment are described in Sect. 4.
1
The list of all prefixes is given in appendix, Table 6.
434 R. Fontugne et al.
For the second experiment (§5.1), we used three /24 prefixes (RIPE-A, RIPE-
B, and RIPE-C) from the RIPE NCC and announced them from three diverse
networks, including an IXP and a national ISP with 149 peer ASes. The three
prefixes are used as test prefixes, meaning that we daily alternate the ROA status
for all of them. The results of this experiment are described in Sect. 5.1.
In order to measure only the impact of ROV and avoid delays caused by other
filtering mechanisms, we configured all filtering with the upstream providers
(e.g., through the creation of Internet Routing Registry (IRR) route objects).
We verified that our providers’ filters accepted our prefixes and then left these
mechanisms untouched.
To toggle the RPKI status of the test prefixes, each was invalidated by Pre-
registering a ROA with the origin AS set to the invalid AS 666. For the first
experiment our AS was primarily connected to upstream networks and IXP
route servers that implement ROV, thus at the initial step, our test prefixes
are dropped by ROV-mechanisms and globally unreachable, as opposed to the
second experiment.
The ROA toggling consists of daily repeating the following steps for each test
prefix:
1. ROA creation. At a random time between 00:00 and 06:00 UTC, we request a
new ROA covering the <prefix, AS> to authorize the route to the test prefix.
2. Convergence phase 1. From 06:00 to 12:00 UTC we give sufficient time for
networks to obtain the new ROA, process it, and update their routing.
3. ROA deletion. At a random time between 12:00 and 18:00 UTC, we delete
the ROA created at the first step, hence letting our test prefix fall back to
Invalid.
4. Convergence phase 2. From 18:00 to 00:00 UTC we again wait for all networks
to converge to the new state.
In order to keep the RPKI beacons running over a long time, we automated
all queries to RIRs. ARIN, RIPE, and recently LACNIC, provide APIs to ease
such interactions with their services. We made all queries to these three RIRs
via their APIs. AFRINIC and APNIC have no APIs for RPKI management; we
could only create and delete ROAs via their web portals. To automate AFRINIC
and APNIC processes we implemented Selenium [19] scripts that log in to these
portals and submit web forms for ROA creation and deletion.
In order to measure the time for the above RPKI operations to propagate over
the management, control, and data planes we collect temporal information from
ROAs’ payload, BGP data, and run traceroutes.
RPKI Time-of-Flight 435
User Query. Delays are measured relative to the user query time, that is the
time we request the RIRs to change RPKI (steps 1 and 3 in Sect. 3.2). This
is logged by an NTP-synchronized host that automates the queries for ROA
creation and deletion. We log the precise time of the confirmation from the RIR
portal or API that the query was received without error.
RIR. We infer RIRs’ signing and publication delays from the RPKIviews
archive [20]. This archive consists of RPKI data snapshots taken every 20 min.
Each snapshot contains the raw ROA files of all RPKI repositories as well as
the output of a relying party software, rpki-client [21]. From this dataset, we
compute the signing, publication, and RP delay (Fig. 1).
The signing delay is computed using the signing timestamp found in the ROA,
more specifically in the Cryptographic Message Syntax (CMS) [22] wrapper of
the signed object. As opposed to the “NotBefore” timestamp found in the ROA
payload, which is used to determine at what time a ROA becomes “valid”, the
signing timestamp conveys the time at which the Certification Authority created
the ROA. Unfortunately, the reliability of both timestamps are disputable as
our results show that some RIRs set the signing and/or NotBefore timestamps
arbitrarily (Sect. 4.1 and 5.3)!
The publication delay estimates the delay for an RIR to make newly created
ROAs available to RPs. We infer the typical publication delay from RPKIviews
snapshots. Since RPKIviews takes snapshots every 20 min and assuming that
the publication of ROA is uniformly distributed over time, new ROAs appear in
RPKIviews on average 10 min after their actual public availability. For ease of
discussion, when reporting RPKIviews median delay publication time in Sect. 4,
we subtract 10 min from the measured RPKIviews median delay. We analyze
only these corrected median values, not individual delays.
Relying Party (RP). Computing Relying Party delays on the Internet is par-
ticularly challenging. The delay of RPs depends on three factors: the frequency
at which they poll for new data from publication points, the downloading time,
and the ROA processing time (i.e., mostly reading and decrypting files). Net-
work operators may increase their RPs’ polling frequency to fetch new data more
quickly, but to reduce the burden on publication points, the recommendations
are to poll for new data no more frequently than 10 min using RRDP (or as low
as 1 min if there is caching infrastructure and the If-Modified-Since header value
is set) and not more than every 30 min if using rsync [23]. Furthermore, past
studies showed that 2 and 10 min are the most common RP polling frequencies
[24] which correspond to respectively RIPE v3 validator and Routinator default
values (rpki-client has no default value). As RIPE v3 validator has since been
deprecated, we assume that 10 min is now a common value used by operators
and attempt to estimate RP delay for RPs polling new data every 10 min.
Similarly to the publication delay, we leverage RPKIviews data to infer the
typical delay experienced by an RP polling data every 10 min. Because the 20-
min frequency of RPKIviews translates into a 10-min median polling delay and
a 10-min polling frequency gives a 5 min median polling delay, when reporting
436 R. Fontugne et al.
results in Sect. 4 we correct the RP delay by subtracting 5 min from the median
delay observed with RPKIviews’ RP.
BGP. The ROA toggle described above affects the global reachability of our
announced prefixes. They become unreachable when corresponding ROAs are
deleted and reachable again when ROAs are re-created. We monitor these shifts
in BGP using the RIPE Routing Information Service (RIS) data [15]. We partic-
ularly look into the BGP update messages sent from routers peering with RIS,
and we record for each peer and each test prefix the time of the first announce-
ment after creating a ROA and the time of the first withdrawal after deleting a
ROA (BGP update in Fig. 1). These represent the first routing changes caused
by each of our RPKI beacon events that we expect to be visible at the collector,
and are accurate within seconds.
Section 4 presents delays for the RIS collectors RRC00 and RRC01. RRC00
has the advantage of being a multi-hop collector, meaning that it receives data
from ASes that are located in very diverse locations. RRC01 collects data only
from ASes peering at the LINX IXP which includes both upstream providers for
the first experiment. Hence RRC01 allows us to investigate BGP signals from
networks that make our prefixes globally reachable. In our preliminary analysis
we have looked at an arbitrary set of RIS collectors (RRC03, RRC06, RRC12),
but given the large amount of data, and that we see little difference across
collectors, we present results only for RRC00 and RRC01. Past research has also
shown a high level of redundancy between different collectors [25] which limit
the benefits of using numerous collectors [26].
Traceroute. To test data plane reachability and delay of the prefixes with toggling
ROAs, we performed traceroutes every 15 min from RIPE Atlas with probes in
6 different ASes. The probes were chosen to be inside the ASes that also share
BGP routes with RIPE RIS at RRC00. We pick these ASes to have close vantage
points for BGP and traceroutes, but there is no guarantee that the Atlas probe
and the BGP collector share the same routes, so there could be some mismatch.
However, we also tried a wider set of RIPE Atlas probes using Atlas geo-diverse
selection of probes and observed similar behaviors, so our analysis focuses only on
the traceroutes obtained with the 6 probes mentioned earlier. The measurements
are public (Table 7) and traceroutes are configured to send three ICMP packets
per hop.
Fig. 2. Time from user query to propagation in BGP, for all RRC00 and RRC01 peers.
ARIN and LACNIC had significantly longer creation delays due to a bug related to
ROAs’ NotBefore timestamps. APNIC delay is typically 10 min longer than AFRINIC
and RIPE. Overall the delays for ROA deletion are higher than for ROA creation.
Fig. 3. Time from user query to BGP propagation (RRC00 and RRC01). Focus on the
two upstream providers of our experimental AS: Sprint (AS1239) and NTT (AS2914).
NTT had more consistent delays than Sprint, and Sprint had sometimes very long
delays to withdraw prefixes with deleted ROAs.
438 R. Fontugne et al.
Our Main Findings are that creation times vary significantly across RIRs,
with medians ranging from a few minutes to over an hour for new ROAs to
reach the publication points. The differences lie in the way ROAs are processed
by RIRs, in batches at specific times of the day, and drastic issues we discovered
at two RIRs (each applied a temporary fix). Second, deletion of ROAs takes
longer to reflect in BGP as routers explore alternate routes that have not yet
been invalidated. The slowest element drives the deletion time. Routers with a
slow pulling cache, or redundant caches, invalidate routes late and are used by
neighbors to reach the invalidated resource. Further, for ROA creation, most of
the delay comes from the Relying Party pulling objects at different intervals.
We investigate ROA creation delay. We explore the disparity across RIRs and
between upstream ASes. Figure 2a shows the per-RIR distribution of the delay
between the query time to create a ROA for our test prefix and the time when
reachability is first reported by each RIS peer in BGP.
AFRINIC and RIPE prefixes are seen most quickly in RIS. The median delay
for an AFRINIC IPv4 prefix is 15 min (16 min for IPv6) and 18 min for RIPE
prefixes for both IPv4 and IPv6. APNIC is consistently slower than AFRINIC
and RIPE. The median delay for an APNIC IPv4 prefix is 26 min (28 min for
IPv6). This 10-min extra delay is due to a 20-min batching process at APNIC
(see Sect. 5.3). ARIN and LACNIC prefixes are susceptible to significant delays.
These are due to the timezone problem described below (Publication delay and
ARIN/LACNIC timezone issues ). We have reported this issue to both RIRs
for which ARIN deployed a workaround on 21 April 2022 and LACNIC on 12
October 2022.
For the other three RIRs delays are less than 1 h in at least 95% of the
cases. We also observe some outlying values: In less than 3% of the cases, for
AFRINIC, APNIC, and RIPE, the BGP delays go over 100 min. These delays
are rarely visible from the two upstream providers, NTT and Sprint (Fig. 3a
and 3c). Both always announce the AFRINIC prefix in less than 100 min. We
cannot find consistent behaviors for the observed long delays, these could be due
to unexpectedly long BGP convergence times [27]. In addition, we noticed that
about half of them are related to very small ASes owned by individuals (network
operators) who are active in testing new deployments (e.g., AS15562, AS35619,
AS5662) so these could be the results of experiments.
We observe a large disparity across RIRs in the time elapsed between ROA
creation and the effect in BGP. The same is true across our upstream ASes.
In the next sections, we track the time along the different steps in Fig. 1 to
understand the elements causing these disparities. We rely on Table 2, where we
show the median delay for each of the steps in Fig. 1.
RPKI Time-of-Flight 439
Table 2. ROA Creation Median Delays. Median delay in minutes from the user
query to the step indicated in each column as observed for the IPv4 prefixes from the
five RIRs (IPv6 results are in parenthesis). As described in § 3, delays shown in this
table are either measured from ROA attributes (*) and BGP data (‡), or inferred from
RPKIviews data (†).
in line with RIPE and AFRINIC at around 5 min median delay (“After fix” line
in Table 2).
For LACNIC the issue is apparently only affecting ROAs created with their
API, not the ones created manually on their portal. LACNIC deployed a similar
fix on October 12, 2022, that sets the NotBefore timestamp at 03:00 UTC on the
day before the user query. Since our LACNIC prefixes were returned on October
25, 2022, we observed the effects of this fix for less than two weeks and found that
LACNIC publication delay fell in line with the other RIRs’ delay (not included
in Table 2 due to the small sample size). As these issues bias results for LACNIC
and ARIN, the remainder of this section focuses on results from the other RIRs.
BGP Updates. We usually observe BGP updates for the newly created ROAs
about 3 min after the estimated RP validation time. This delay includes both
the router’s polling from RPs and BGP propagation time, as we are not able
to measure the RP to router delay alone. As RPs signal routers when to pull,
this delay should be dominated by the data transfer and the router processing
of VRPs. Past work on BGP propagation estimate that a new announcement on
BGP takes usually less than a minute to propagate globally [5,6], hence one can
estimate the RP to routers delay should be no more than 2 min.
To further dissect delays observed at this step, we compute the BGP delay
only for the two upstream providers. The BGP delay distributions of NTT
(Fig. 3a) and Sprint (Fig. 3c) are similar to those observed for other peers
(Fig. 2a), and their median values are all within a 4-min difference. Given that
ROV is still deployed very sparsely [28], these results show that (1) the delay
for ASes that are not along ROV-enabled AS paths is dictated by our upstream
providers, (2) ASes beyond our upstreams that perform ROV slower would inval-
idate new routes. In the latter case, because we are connected to Tier1 networks,
and there are many paths between Tier1 networks and RIS collectors, the effect
of other ROV deployments is rarely observed.
We also compared these distributions with five other networks that are
implementing ROV and announcing our prefixes to RRC00 or RRC01 (AS1299,
RPKI Time-of-Flight 441
AS6939, AS7018, AS9002, AS14907) and found no notable differences in the dis-
tributions, meaning that these networks behave similarly to our upstreams; i.e.,
they are at least as fast as our upstreams to pull new RPKI data. With these data
we cannot distinguish if they can fetch RPKI data faster than our upstreams
as their BGP announcements are bound by the time our upstreams made the
prefixes globally available. We come back to this in Sect. 5.1 with experiments
announcing prefixes from very diverse locations.
A careful inspection of the delays for our two upstreams reveals that NTT
is more consistent, the 10th to 90th percentile range for the IPv4 RIPE prefix
corresponds to 12 and 32 min (Fig. 3a) whereas these same percentiles correspond
to a range twice as large for Sprint, i.e., 7 and 47 min (Fig. 3c). Although the first
quartile delay for Sprint is always better than for NTT (e.g. 12 min vs 15 min for
RIPEv4), the third quartile delay for Sprint is consistently longer by 1 to 10 min
for the AFRINIC, APNIC, and RIPE prefixes. We believe this is the result of a
longer RP polling frequency for Sprint but a shorter RP to router delay, and we
confirmed with network operators that indeed NTT is polling RPKI data more
frequently than Sprint and Sprint is using faster RP software.
Using only the data after ARIN’s fix we confirm the delay for ARIN prefixes
improved significantly (shown in Appendix Fig. 11). The interquartile range cor-
responds to 13 and 33 min for IPv4 (13 and 33 for IPv6) which comes very close
to RIPE and AFRINIC results for the same time period. RIPE’s interquartile is
11 to 32 min for IPv4 (11 to 27 for IPv6) and AFRINIC’s interquartile is 10 to
25 min for IPv4 (9 to 29 for IPv6) between 21 April 21st and May 15th 2022.
Fig. 4. Effects of ROA creation (green dots) and ROA deletion (red dots) on prefix
reachability (cyan dot) and unreachability (black dot) in traceroute. Each line shows
a different Atlas probe/prefix pair. Delay between ROA deletion and unreachability
highly varies depending on the topology. IPv4 only, see Fig. 10 for IPv6. (Color figure
online)
We investigate the ROA revocation timing along the steps in Fig. 1. In addi-
tion to longer deletion than creation due to path exploration, we show that
while APNIC demonstrates longer times for the revocation to be published and
to reach Relying parties, the prefixes disappear from BGP only slightly after the
prefixes with ROAs hosted by other RIRs.
BGP Withdraw. BGP delays are significantly higher for ROA deletion than
for ROA creation (Fig. 2b). The median BGP delay for unreachability goes up
to 51 min for IPv4 and 56 min for IPv6 (Table 3). We rarely observe short BGP
delays (Fig. 2b). At best the BGP delay first quartile corresponds to less than
20 min (AFRINICv4) and at worst less than 39 min (APNICv6).
There are two related causes for these high delays, one is related to BGP and
the other to RPs/routers interactions. At ROA creation a prefix is announced
globally in BGP by one of the prefix’s upstreams as soon as either one of them
fetches the new ROA. But at ROA deletion neighbors must all withdraw the
RPKI Time-of-Flight 443
Table 3. ROA Deletion. Median delay in minutes from user query to the step indicated
in each column as observed for the IPv4 prefixes from the five RIRs (IPv6 results in
parenthesis). These delays are either measured from CRL files (*) and BGP data (‡),
or estimated from RPKIviews data (†).
Fig. 5. ROA deletion. Time from user Fig. 6. Effects of ROA creation/deletion
query to BGP withdraw for Deutsche on the data plane. After each ROA cre-
Telekom (AS3320). IPv4 delays are ation or deletion, we observe BGP path
impacted by Sprint late withdrawing. hunting with AS path changes.
the AS path between the RIPE Atlas probe and the destination changes from
[Source AS, AS174 (Cogent), AS1239 (Sprint), Destination AS] to [Source AS,
AS6762 (Telecom Italia), AS1239 (Sprint), Destination AS] before becoming
unreachable. Since Telecom Italia is not performing ROV [30], our hypothesis
is that Cogent fetches RPKI data faster than Sprint and drops the prefix while
Sprint is still announcing it. BGP path hunting then selects an alternate path via
Telecom Italia, until finally Sprint also drops the route and the prefix becomes
unreachable. For probe#4, the delay before unreachability is similar to probe#1,
but we do not observe an AS path change between ROA deletion and destination
being unreachable, the AS path remaining [Source AS, AS7575 (AARNET),
AS6461 (Zayo), AS1239 (Sprint), Destination AS]. Again, our hypothesis is that
Sprint is slow to drop this route and keeps announcing the route to Zayo, which
does not perform ROV [30], so it announces the prefix until Sprint drops it.
Impact on AS Path. Figure 6 shows the impact of ROA creation and deletion
on the observed paths and illustrates BGP path hunting for one of the Atlas
probe/prefix pairs. The Y axis represents the latency between the Atlas probe
and the destination relative to the minimum RTT observed during the measure-
ment period, and the X axis represents time. The vertical lines show the times
of ROA creation/deletion. Each dot is a traceroute, and every time the AS path
changes, we put a label above the dot with the new AS seen in paths taken by
the traceroute packets.
BGP convergence and path hunting are each illustrated after ROA creation
and deletion. After ROA creation, we observe a first path going through AS1299
(Telia), and then a preferred path (in the sense of BGP) going through AS174
(Cogent) is selected. This suggests that Telia was faster to integrate the new
ROA than Cogent. After ROA deletion, we observe that BGP finds another path
going through AS3257 (GTT), and then the destination becomes unreachable,
as we see the dots stopping a short time after the red lines.
RPKI Time-of-Flight 445
Fig. 7. Time from user query to BGP propagation for prefixes RIPE-A, RIPE-B, and
RIPE-C as observed by RRC00 and RRC01 peers.
More Locations. For this stage we obtained three /24 IPv4 prefixes (RIPE-A,
RIPE-B, and RIPE-C) from RIPE NCC and three topologically diverse oper-
ators generously agreed to announce these prefixes from their networks. These
three networks differ significantly from our experimental AS in the first setup;
the locations are on a different continent and have different upstream providers
including networks that do not implement ROV. Therefore, when running RPKI
beacons for these prefixes their reachability is unaffected along paths that have
no network implementing ROV. Only RIS peers that implement ROV or that
are surrounded by ROV lose reachability to these prefixes.
We measured these prefixes from May 6th to October 5th 2022, and again
observe that BGP delay for ROA deletion is significantly longer than it is for
ROA creation (Fig. 7).
The median BGP delay for ROA creation is shorter than during the first
experiment, the median ranging between 11 and 12 min (Fig. 7a) compared to
the median of 18 min observed previously for our IPv4 RIPE prefix, suggesting
that ROV-enabled networks between these origin ASes and RIS peers are faster
than NTT and Sprint in the previous experiment (see Sect. 5.1).
On data plane reachability, we observe the expected behavior that here some
probes never lose reachability, because they find a route via a provider that does
not enforce ROV, as opposed to the probes in the first experiment (Fig. 4).
Fig. 8. User query to BGP delay for prefixes RIPE-A, RIPE-B, and RIPE-C as
observed by Tier1 networks implementing ROV.
networks fetch RPKI data more frequently than NTT, which we confirmed with
operators of two of these networks. This difference may also explain the 7 to
8 min difference between the median delay for the RIPE prefix in the previous
experiment (Fig. 2a) and the three prefixes used here (Fig. 7a).
The delay at ROA deletion is higher than at ROA creation for all monitored
networks as seen in Fig. 8b. As we expect these networks to drop the prefixes
as soon as they get in sync with RPKI, the slower deletion of ROAs is the
result of RP redundancy, i.e. the ROA deletion is not effective until the last
cache withdraws the ROA (Sect. 4.2). The anomaly in Sprint ROA deletion was
confirmed to be due to the Routinator bug.
few exceptional cases with AFRINIC (< 10%) where the NotBefore time is set
before signing time. This provides confidence that the NotBefore time is usually
a good estimator of the signing time for AFRINIC, RIPE and APNIC but not
for ARIN and LACNIC. We also confirmed from our active measurements that
the NotBefore time for RIPE and AFRINIC is usually within a minute of our
query time and on average 10 min later for APNIC.
Below is the process to calculate BGP delay using historical data:
1. VRP data: We first collect a list of VRPs (Validated ROA Payloads) from
the RIPE RPKI archive [32], which provides historical RPKI data organized
by TA (Trust Anchor). Each repository contains the certificates and ROAs
classified by date and also provides a list of VRPs for each day. We extract
the NotBefore time (t0 ) and route (prefix, origin) for each VRP.
2. RIB files: We select from RIS RRC00 collector a RIB dump on a randomly
selected day in May every year from 2018 to 2022.
3. BGP update messages: we extract the BGP update messages from RIS
update files and look for BGP withdrawals at time t1 that correspond to a
VRP’s prefix and where t1 is between t0 and t0 + 1 h.
4. BGP delay: We calculate the BGP delay as t1 − t0 .
Table 4 provides detail about the volume of longitudinal data processed from
the RRC00 collector and from the RIPE RPKI archive. It shows the total number
of RIB entries, the number of invalid routes and the corresponding number of
withdrawals found in BGP data.
Figure 9a shows an overview of the BGP delay for all data points collected
between 2018 and 2022. There is no major difference in median propagation delay
between IPv4 and IPv6, but there is greater variability in IPv6. We observe that
AFRINIC, APNIC and RIPE had consistently shorter median delays over time
while ARIN and LACNIC had higher delays for IPv4. The reason for higher
delays for ARIN and LACNIC may be caused by the anomaly in the publication
process (see 4.1). However, as we can see from Fig. 9b, the median delay remained
usually around 20 min between 2019 and 2022. The numbers for 2018 are slightly
higher but overall these results suggest that the Certification Authority to BGP
delays at ROA creation have been stable over the past four years.
RPKI Time-of-Flight 449
Table 5. Number of unique ROA objects, routes, and signing timestamps from a
snapshot on December 31st 2021 of ROAs created in 2021.
Finally, this section describes the differences between the ROA payloads gener-
ated by the different RIRs and how these can impact ROA publication delay.
Signing Time Distribution. The first notable difference between the ROA
payloads of different RIRs is the distribution of signing and NotBefore times-
tamps. As mentioned in Sect. 4.1 we found that ARIN is using a hardcoded value
for the signing and NotBefore timestamps. Looking at a snapshot of all ROAs
on December 31st 2021, we found that the 29213 ROA objects that ARIN signed
in 2021 contain only 311 unique signing timestamps (Table 5), which is roughly
equal to the number of days in 2021 minus weekends where we rarely see new
ROAs. We have also confirmed that this behavior is present since ARIN started
its RPKI service in September 2012.
For LACNIC, the results are not as clear. We do observe an abnormally high
number of ROAs with the NotBefore time set to 03:00 UTC but not all. This
is because it affects only the API, which was released in 2021, and thus only
recently used in the LACNIC region.
450 R. Fontugne et al.
6 Discussion
Setting up these experiments and maintaining them over several months was an
eye-opener to the challenges that operators face with RPKI and RIRs.
First, the procedures and requirements to obtain resources, activate RPKI,
and manage ROAs for the five RIRs are all quite different. In addition, the lack of
APIs to manage RPKI resources for APNIC and AFRINIC makes automation
a lot more challenging. We implemented Selenium scripts for AFRINIC and
APNIC beacons, which is not trivial given the security measures employed by
RIR portals (e.g., two-factor authentication and password renewal) and need
adjustments whenever portals are updated.
Second, the need for continued monitoring of the management, control, and
data planes is crucial to ensure proper operation of all components involved
and impacted by RPKI. For example, one of our AFRINIC beacons failed for
multiple days because one of the ROA was left un-revoked by the Certification
Authority, even though our deletion query succeeded and it had disappeared
from the AFRINIC web interface. We only noticed this problem in our data
RPKI Time-of-Flight 451
7 Related Work
8 Conclusion
In this paper, we designed wide-ranging experiments to measure the timing and
effects of the propagation of ROAs on the management, control, and data planes.
This enabled us to track how ROAs are disseminated - starting from the moment
creation is triggered through the RIRs’ API/portals, then signed by the hosted
Certificate Authorities, published at their respective Publication Points, to the
moment they are fetched and validated by RPs, consequentially seeing routers
announcing new routes in BGP, and then affecting delay and reachability on the
data plane. We found ROA management issues for two RIRs and discovered that
RIRs usually publish new RPKI information within 5 min, except APNIC which
is 10 min slower. For ISPs, we observe disparate behaviors in the control and data
planes between when routes are validated or invalidated by a ROA creation or
deletion. At the ISP level, we observed that the reaction time following a ROA
deletion is much longer due to BGP and multi-RP deployment that require
complete ROA withdrawals on all RPs for a route to be withdrawn. Predicting
prefix reachability and the BGP convergence time is getting even harder as it
requires insights about which networks are implementing ROV and how quickly
each reacts to RPKI changes. This study reveals some of the complexity added
by RPKI to basic routing operations.
Ethics. This work does not raise ethical issues. It is focussed on the reachability
of experimental prefixes delegated to us by the RIRs specifically for the time of
the experiment. These prefixes were cleared for advertisements by our providers
and documented in the IRR databases. A webpage describing the experiment was
available throughout the experiment ([Link]
timing). In addition, our work does not involve personal identification data.
A Appendix
A.1 Data Plane Availability in IPv6
Fig. 10. IPv6: Effects of ROA creation (green dots) and ROA deletion (red dots) on
prefix reachability (cyan dot) and unreachability (black dot) in traceroute. Each line
shows a different Atlas probe/prefix pair. Delay between ROA deletion and unreacha-
bility highly varies depending on the topology. (Color figure online)
454 R. Fontugne et al.
Fig. 11. ROA creation after ARIN’s fix. RRC00 and RRC01 peers from April 21st to
May 15th 2022. APNIC and LACNIC are not plotted to improve readability. ARIN’s user
query to BGP delay distributions became similar to the ones of AFRINIC and RIPE.
A.3 Reproducibility
Our experimental data is publicly available in order to make the results of this
work entirely reproducible. Our source code and logs of user query time are
available at [Link]
The list of experimental prefixes obtained from the five RIRs are shown in
Table 6.
References
1. Rekhter, Y., Hares, S., Li, T.: A Border Gateway Protocol 4 (BGP-4). RFC 4271,
January (2006)
2. Lynn, C.: X.509 Extensions for Authorization of IP Addresses, AS Numbers, and
Routers within an AS. Internet-Draft draft-clynn-bgp-x509-auth-00, Internet Engi-
neering Task Force
3. Lepinski, M., Kent. S.: An Infrastructure to Support Secure Internet Routing. RFC
6480, February (2012)
4. Mohapatra, P., Scudder, J., Ward, D., Bush, R., Austein, R.: BGP Prefix Origin
Validation. RFC 6811, January (2013)
5. Mao, Z.M., Bush, R., Griffin, T.G., Roughan, M.: BGP beacons. In: Proceedings
of the 3rd ACM SIGCOMM conference on Internet measurement, pp. 1–14 (2003)
456 R. Fontugne et al.
6. Garcia-Martinez, A., Bagnulo, M.: Measuring bgp route propagation times. IEEE
Commun. Lett. 23(12), 2432–2436 (2019)
7. Al-Musawi, B.: Common pitfalls in RPKI deployment and how to avoid them, Apr
(2021)
8. Hlavacek, T.: DISCO: Sidestepping RPKI’s deployment barriers. In: Network and
Distributed System Security Symposium (NDSS) (2020)
9. Iamartino, D., Pelsser, C., Bush, R.: Measuring BGP route origin registration and
validation. In: Mirkovic, J., Liu, Y. (eds.) PAM 2015. LNCS, vol. 8995, pp. 28–40.
Springer, Cham (2015). [Link]
10. Candela, M.: A One-Year Review of RPKI Operations, RIPE 84, May (2022)
11. Candela, M.: One Does Not Simply “Deploy RPKI”, MANRS blog, July (2022)
12. Sermpezis, P., Kotronis, V., Gigis, P., Dimitropoulos, X., Cicalese, D., King, A.,
Dainotti, A.: ARTEMIS: Neutralizing BGP hijacking within a minute. IEEE/ACM
Trans. Netw. 26(6), 2471–2486 (2018)
13. Kimura, T.: Long Chopsticks in Heaven - When Packets Dropped Using ROA,
May (2019)
14. Gilad, Y., Cohen, A., Herzberg, A., Schapira, M., Shulman, H.: Are we there yet?
on RPKI’s deployment and security. Cryptology ePrint Archive (2016)
15. RIPE NCC. Routing Information Service (RIS), May (2022)
16. RIPE NCC. RIPE Atlas, May (2022)
17. Boeyen, S., Santesson, S., Polk, T., Housley, R., Farrell, S., Cooper, D.: Internet
X.509 Public Key Infrastructure Certificate and Certificate Revocation List (CRL)
Profile. RFC 5280, May (2008)
18. University Oregon. Route Views, September (2022)
19. Selenium. Selenium webdriver, September (2022)
20. Job Snijders. RPKIviews, May (2022)
21. OpenBSD. rpki-client, May (2022)
22. Housley, R.: Cryptographic Message Syntax (CMS). RFC 5652, September (2009)
23. Bush, R., Borkenhagen, J., Bruijnzeels, T., Snijders, J.: Timing Parameters in
the RPKI based Route Origin Validation Supply Chain. Internet-Draft draft-ietf-
sidrops-rpki-rov-timing-06, Internet Engineering Task Force, February 2022. Work
in Progress
24. Kristoff, J.: On Measuring RPKI Relying Parties. In: Proceedings of the ACM
Internet Measurement Conference, IMC ’20, pp. 484–491, New York, NY, USA,
2020. Association for Computing Machinery
25. Alfroy, T., Holterbach, T., Pelsser, C.: MVP: Measuring Internet routing from the
most valuable points. In: Proceedings of the 22nd ACM Internet Measurement
Conference, IMC ’22, pp. 770–771, New York, NY, USA, 2022. Association for
Computing Machinery
26. Fontugne, Romain, Shah, Anant, Aben, Emile: The (thin) bridges of as connec-
tivity: measuring dependency using as hegemony. In: Beverly, Robert, Smarag-
dakis, Georgios, Feldmann, Anja (eds.) PAM 2018. LNCS, vol. 10771, pp. 216–227.
Springer, Cham (2018). [Link]
27. Ongkanchana, P., Fontugne, R., Esaki, H., Snijders, J., Aben, E.: Hunting BGP
zombies in the wild. In: Proceedings of the Applied Networking Research Work-
shop, pp. 1–7 (2021)
28. Cloudflare. Is BGP safe yet? No., May (2022)
29. Routinator.: Changelog (v0.11.2), April (2022)
30. Fontugne, R.: The Routing Game: Hunting Invalid Routes., November (2021)
RPKI Time-of-Flight 457
31. Luckie, M., Huffaker, B., Dhamdhere, A., Giotsas, V., Claffy, KC.: AS relationships,
customer cones, and validation. In: Proceedings of the 2013 Conference on Internet
Measurement Conference, pp. 243–256 (2013)
32. RIPE NCC. RIPE NCC’s RPKI repository archive, May (2022)
33. Lepinski, M., Kong, D., Kent, S.: A Profile for Route Origin Authorizations
(ROAs). RFC 6482, February (2012)
34. Harrison, T.: APNIC Registry API, APNIC blog, March (2022)
35. Reuter, A., Bush, R., Cunha, I., Katz-Bassett, E., Schmidt, T.C., Wählisch, M.:
Towards a rigorous methodology for measuring adoption of rpki route validation
and filtering. ACM SIGCOMM Comput. Commun. Rev. 48(1), 19–27 (2018)
36. Chung, T., et al.: RPKI is coming of age: A longitudinal study of RPKI deployment
and invalid route origins. In: Proceedings of the Internet Measurement Conference,
pp. 406–419 (2019)
37. Gilad, Y., Sagga, O., Goldberg, S.: Maxlength considered harmful to the RPKI.
In: Proceedings of the 13th International Conference on Emerging Networking
EXperiments and Technologies, pp. 101–107 (2017)
38. Hlavacek, T., Jeitner, P., Mirdita, D., Shulman, H., Waidner, M.: Stalloris: RPKI
downgrade attack. In: 31st USENIX Security Symposium (USENIX Security 22),
Boston, MA, August (2022) USENIX Association
39. Bush, R.: Origin validation operation based on the Resource Public Key Infras-
tructure (RPKI). IETF RFC7115 (January 2014)
40. Hlavacek, T.,Herzberg, A., Shulman, H., Waidner, M.: Practical experience:
Methodologies for measuring route origin validation. In: 2018 48th Annual
IEEE/IFIP International Conference on Dependable Systems and Networks
(DSN), pp. 634–641. IEEE (2018)
Security and Privacy
Intercept and Inject: DNS Response
Manipulation in the Wild
Yevheniya Nosyk1(B) , Qasim Lone4 , Yury Zhauniarovich2 , Carlos H. Gañán2,5 ,
Emile Aben4 , Giovane C. M. Moura2,3 , Samaneh Tajalizadehkhoob5 ,
Andrzej Duda1 , and Maciej Korczyński1
1
Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, Grenoble, France
{[Link],[Link],[Link]}@[Link]
2
TU Delft, Delft, The Netherlands
3
SIDN Labs, Arnhem, The Netherlands
4
RIPE NCC, Amsterdam, The Netherlands
5
ICANN, Los Angeles, CA, USA
1 Introduction
The Domain Name System (DNS) [41,42] is one of the core Internet protocols. It
was introduced to translate human-readable domain names (e.g., [Link])
into IP addresses (e.g., 2001:db8::1234:5678), but has gone far beyond this
basic service. It is now a large-scale distributed system comprising millions of
recursive resolvers and authoritative nameservers—the two main components of
the DNS infrastructure. It was designed in a hierarchical manner so that no single
entity stores the data about the entire domain name space. Each authoritative
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 461–478, 2023.
[Link]
462 Y. Nosyk et al.
nameserver is only responsible for a subset of the domain tree and it is the role
of recursive resolvers to follow the chain of delegations and find authoritative
query responses.
Nonetheless, DNS is prone to manipulation. The original specification does
not ensure data integrity or authentication, allowing on-path entities (Internet
Service Providers, national censors, attackers, etc.) to intercept plain-text DNS
traffic and inject responses (whether bogus or not). A latter standard, Domain
Name System Security Extensions (DNSSEC) [55], provides data integrity, but
its usage and deployment remain optional and far from being universal [10,17].
In November 2021, Meta engineers reported that users in Mexico were receiv-
ing bogus A records when querying [Link] and [Link] [15]. A
closer look revealed that those queries were routed to the local China-located
anycast instance of the k-root, despite having several other points of presence
nearby [1]. As the Great Firewall of China (GFW) is known to inject bogus
responses when detecting sensitive domains [28], Mexican users might have expe-
rienced collateral damage from DNS censorship. The k-root operator (RIPE
NCC) later confirmed that a Border Gateway Protocol (BGP) route leak made
the local root server instance globally available. This outage stemmed from a
series of unfortunate events but nevertheless rendered both domains unavailable.
The Internet has already seen similar events in the past [56,59] and researchers
reported on detecting DNS root manipulation [21,32,35,44]. However, the preva-
lence of this phenomenon has not been systematically analyzed across all the root
server letters over a long period of time.
In this paper, our contribution is two-fold. First, we leverage built-in RIPE
Atlas [54] measurements to identify probes affected by the November 2021 route
leak. We show that at least two months prior to be reported by Meta, the Chinese
k-root instance had already been accessible from 32 autonomous systems (ASes)
in 15 countries. Second, we set up a nine-month DNS measurement campaign
and observe the same problem in the wild, although this time mostly over IPv6.
More broadly, the DNS manipulation experienced by Mexican clients motivates
us to study the extent of DNS response injection when contacting root servers.
We reveal that even though less than 1% of queries are affected, they originate
from 7% of probes located in 66 countries.
As of October 2022, there are 1,575 anycast instances accessible either world-
wide (global) or only within a limited range of networks (local) [1]. In the latter
case, local network operators typically limit the propagation of BGP routes by
using NO EXPORT or NOPEER BGP community attributes [37]. DNS queries are
then routed to the nearest anycast locations based on the routing tables. As
BGP is latency agnostic, it may eventually map clients to instances from another
continent, even when closer ones are available [43].
To assist in troubleshooting, root servers support DNS queries that iden-
tify individual anycast instances. This is achieved by one of the CHAOS-class TXT
queries [18] (e.g., [Link], [Link]) or the NSID option [8]. The lat-
ter does not require issuing a separate query because the nameserver identifier
provided in the OPT resource record is stored in the Additional section of a DNS
response packet.
Root traffic manipulation has been previously reported twice. In both cases,
a BGP route leak made China-located local instances globally available. Even
though root servers themselves were legitimately run by their operators, the
GFW or other interceptors were likely present in transit. In the first case
in 2010, clients located in the USA and Chile had their queries for three
domains ([Link], [Link], and [Link]) answered with bogus
IP addresses. Upon further investigation, it was found that original queries were
sent towards the i-root instance in Beijing [59]. Similarly, in 2011 the clients in
Europe and the USA were directed towards the f-root instance in Beijing [56],
although no response injection was reported at that time.
[Link]? @[Link]
K-root Costa Rica A [Link]
CH TXT [Link]? @[Link]
TXT [Link]
Fig. 1. DNS traffic seen on the RIPE Atlas probe from Mexico. A middlebox intercepts
the DNS query and injects the bogus response for [Link]. The CHAOS-class
TXT query for [Link] confirms that the request was routed towards the Chinese
instance of the k-root.
The events described in Sect. 2.2 and Sect. 2.3 demonstrate the cases when
queries directed to certain DNS root servers resulted in response injection. In
this section, we set out to characterize the extent of this phenomenon for all
the root letters. We identify the response injectors and factors influencing such
manipulation. We analyze more than 1 billion DNS RIPE Atlas measurements
issued between February and October 2022.
Figure 2 shows the series of queries we issue on each RIPE Atlas probe, directed
to all the root servers on their IPv4 and IPv6 anycast addresses. We explicitly
request them not to perform recursion by setting the Recursion Desired (RD) flag
Intercept and Inject: DNS Response Manipulation in the Wild 465
Fig. 2. DNS queries issued by each RIPE Atlas probe every 12 h. We send queries to
all root servers to resolve A/AAAA records of [Link], [Link], and [Link]
over two transport protocols (TCP/UDP) and both IP versions (IPv4/IPv6).
to false, even though correctly operating root servers would not do it anyways.
Each root server is requested to resolve A and AAAA records of three domain
names ([Link], [Link], and [Link]) over TCP and UDP. The
first two domains are known to trigger censorship middleboxes [48], while the
third one ([Link]) is a control domain. In addition, we request to include the
NSID string in all the responses to learn which anycast instance (if any) answers
our queries. In total, each available probe performs 312 DNS lookups every 12 h.
Table 1. The number of injected responses per domain name and response type.
We now analyze all the DNS responses received on probes when sending queries
to the 13 root servers. Recall that root servers do not directly answer queries for
second-level domains, such as [Link]. Instead, they point to authoritative
DNS servers of top-level domains. Therefore, we refer to each measurement result
as either non-injected (the answer section of the DNS response is empty) or
injected (the answer section contains the response). In the collected dataset, over
9M responses (0.82%) were injected and contained more than 11M individual
resource records of different types. Table 1 presents the response types received
per domain name:
– SOA: contains administrative information about a DNS zone. One probe from
the USA received SOA records for [Link] queries as if they came from
Facebook’s authoritative nameservers. However, the nameserver and main-
tainer names revealed the true originator—a DNS content filter [20]. As no
valid IP address of [Link] was returned, end users would not be able
to access the domain name.
– CNAME: maps one domain name to another. We found 4,536 aliases that
pointed [Link] to [Link]—the service [25] to
exclude explicit content (e.g., pornography, violence) from search results. It
is configured by adding a CNAME record to local DNS configurations. All the
six affected probes (located in Spain, the USA, the Netherlands, and Russia)
received corresponding A or AAAA resource records along with CNAMEs. Apart
from one probe that received a bogus IP address, others would still access
[Link], although some parts of search results would be filtered.
Overall, DNS injection impacted only 0.82% of all the queries issued during
nine months in 2022. Figure 3 further demonstrates that the weekly ratio of
response injection never exceeded 1%, yet proving that it is constantly present
in the wild. Interestingly, response injection does not necessarily prevent access
to requested domains at the DNS level - the majority of all the injected responses
were not bogus. Therefore, if not coupled with other filtering techniques, such
as HTTP(S) interception or destination IP address blocking, DNS manipulation
would stay transparent to end users.
The injected responses demonstrate that DNS queries originated from RIPE
Atlas probes must have encountered middleboxes on the way to root servers.
Such devices were shown to serve different purposes. Transparent forwarders [45]
only relay incoming DNS requests to alternative resolvers, such as public or net-
work’s internal DNS resolvers. Importantly, they do not inject spoofed responses,
but rather let those alternative resolvers respond to end clients directly. More
intrusive DNS interceptors, such as national censors, impersonate intended query
destinations and actively inject bogus responses [38]. DNS interception is often
accomplished at Customer Premises Equipment [51] and can be detected by
issuing CHAOS-class or other DNS queries with the NSID option enabled.
We thus leverage the nameserver identifier option (NSID) to fingerprint ser-
vices that provided responses to RIPE Atlas probes. We extracted and manu-
ally analyzed more than 12K unique NSID strings from over 1 billion measure-
ments. We consulted the web pages of root server operators, online documenta-
tion, and issued additional DNS queries to validate our assumptions. We then
generated regular expressions to match each identified service. Finally, we con-
tacted root server operators and six of them (including Verisign that manages
a-root/j-root with the same pattern) responded confirming the validity of our
mappings.
468 Y. Nosyk et al.
We leveraged 14,335 RIPE Atlas probes from 177 countries and 4,132 ASes. A
great majority of them did not experience DNS manipulation, but a smaller frac-
tion (1,010 or 7.05%) received injected responses—a substantial increase since
2016 when less than 1% of RIPE Atlas probes were reported to be intercepted
when contacting DNS root servers [44]. We compute the fraction of affected
probes per country in Table 3 and plot the results on Fig. 6 in Appendix. Overall,
the manipulation ratio remains low—113 countries do not host a single probe
experiencing DNS injection. On the contrary, some of the other 66 countries
have a significant ratio of manipulated probes - 97.1% in Iran, 83.15% in China,
Intercept and Inject: DNS Response Manipulation in the Wild 469
66.67% in Palestine, 50% in Yemen, and 50% in Saint Barthélemy. As for the
remaining countries (see Table 3 in Appendix), the ratio of manipulated probes
does not usually exceed 30%. The autonomous system distribution is much more
diverse, but more than half of the ASes host only a single probe. Overall, 5.61%
of ASes only host probes experiencing manipulation, 88.92% only host probes
that do not, and the remaining 5.47% have probes of both types.
RIPE Atlas probes are not constantly available and may get occasionally dis-
connected, which makes it non-trivial to run longitudinal measurements. How-
ever, Fig. 3 shows that the proportion of probes experiencing manipulation to all
the participating probes per week remains stable (the corresponding proportion
of measurements exhibits the same behavior). Figure 4 additionally shows that
roughly 20% of probes (mostly located in Iran, China, the USA, and Russia)
experienced response manipulation during all the weeks of the experiment.
We emphasize that DNS interception and injection may happen anywhere in
transit between RIPE Atlas probes and DNS root servers. Therefore, we refer to
countries and networks as those hosting probes that experience injection. We do
not assume that those entities are necessarily responsible for manipulating with
DNS traffic of their clients.
3.7 Limitations
We acknowledge that the presented measurement study has certain limitations.
Using two sensitive domains would not trigger all the existing censorship mid-
dleboxes. Yet, domains from other categories (e.g., gambling or adult content)
would potentially put the owners of RIPE Atlas probes in danger as those could
break local laws. We thus consider this limitation acceptable and suggest the
470 Y. Nosyk et al.
4 Countermeasures
Some techniques presented below can effectively reduce the risks of DNS injec-
tion. However, a deliberate interceptor, especially when located close to the query
source, is capable of monitoring all the client activity and reacting accordingly:
5 Ethical Considerations
Measurement research must be designed extremely carefully so that it minimizes
any risk for involved parties but maximizes the probable benefits, as outlined
in The Menlo Report [9]. This is especially important in censorship studies,
which usually involve actively generating traffic to trigger censors. A rich body
of research [47,48,50,57,58] performed experiments similar to ours, in particular
using RIPE Atlas infrastructure [2,13], and reported that no evidence suggested
that any harm was caused. We further received a formal approval from the
institutional review board (IRB) of our institution. They judged our research as
the one complying with all the ethical requirements.
Our choice of the measurement platform was dictated by several reasons.
RIPE Atlas is an opt-in service where all the participants accept the Terms and
Conditions [52]. In particular, probe hosts agree that i) the permission to install
probes was obtained (§5.1), ii) other users can perform measurements on probes,
in particular for research (§5.4, §4.5), iii) probes may be disabled one month after
a written request is received (§8.4), and iv) measurement results will be made
public either fully or in the aggregated form (§4.2). Our measurements comply
with these terms and in this paper, we do not expose sensitive information about
individual probes, even when allowed (§4.3).
All the RIPE Atlas probes regularly perform a set of built-in measure-
ments [53] for different protocols. More than half of 242 recurring DNS measure-
ments (running every 4 min to 12 h) are destined to root servers. Moreover, each
probe is also requesting A records of popular domain names every 10 min. Con-
sequently, the traffic we generate for this experiment does not stand out from
the normal operation of RIPE Atlas probes. Two out of three domain names
that we query, namely [Link] and [Link], are the first and the
third most popular worldwide respectively [49]. Thus, queries to such domains
are challenging to link to particular end hosts when observed in the wild.
6 Related Work
DNS middleboxes were previously known to interfere with root DNS traffic.
In 2013, Fan et al. [21] reported that 1.75% of 64K vantage points worldwide
would encounter DNS proxies, rogue root servers, and other unusual behav-
ior when sending requests to the f-root. Moreover, some of the queries for
[Link] would be answered directly. Jones et al. [32] further for-
malized the phenomenon as DNS root manipulation. Measurements towards the
b-root from 8K RIPE Atlas probes revealed 10 DNS proxies and one root server
replica in China. Moura et al. [44] identified 74 RIPE Atlas probes (less that 1%
472 Y. Nosyk et al.
of all the 9K probes at that time) that would have root server queries answered
by third parties. More recently, Li et al. [35] found that some of the queries orig-
inated inside China to k-root root server instances located inside the country
would also result in hijacking. In our paper, we present a nine-month longitudi-
nal study that characterizes the extent of DNS response injection across all the
root server letters. We show that the ratio of affected probes has significantly
risen compared to previous studies.
More broadly, various types of DNS manipulation have been extensively
studied in the literature. Censors [7,23,27,46–48,50,57,58], transparent for-
warders [33,45], rogue DNS servers [19], and middleboxes [16,38,51,60,61]—all
interfere with the normal DNS resolution process. Particular attention has been
paid to the GFW of China [3–5,11,22,28,39], known to intercept DNS traffic
and inject bogus responses. Hoang et al. [28] provided the most complete pic-
ture of DNS manipulation by the GFW to date. The authors issued queries for
534M domain names from outside China to controlled servers inside the country.
Their measurements triggered the GFW, suggesting that it indeed operates on
the traffic coming from outside. As witnessed during the November 2021 route
leak, middleboxes were shown to inject globally routable IP addresses in their
bogus responses. These findings are in line with the previous study of the anony-
mous researcher [3], showing that the GFW is acting on the traffic that is barely
traversing Chinese ASes. Overall, the existing research suggests that clients from
Mexico were affected by the operation of the GFW.
7 Conclusions
In this paper, we explored the November 2021 BGP route leak that resulted in
DNS response injection as a side effect. We identified 32 ASes worldwide that
would reach the Guangzhou instance of the k-root and potentially encounter
injecting middleboxes on the way. While this particular problem was quickly
fixed, our longitudinal measurements revealed that DNS injection is omnipresent.
Queries to DNS root servers are constantly getting intercepted and may result
in injected responses, especially when involving sensitive domain names.
We also revealed that the Guangzhou k-root instance became reachable out-
side mainland China several months before it was reported. As it is crucial to
identify such events early enough before the impact on end users becomes appar-
ent, RIPE NCC deployed BGP community attributes identifying each k-root
server instance. Therefore, such leaks are now detectable.
Appendix
Generalized Linear Mixed-Effects Model
0.70 ***
rr [AAAA]
3.50 ***
transport [UDP]
5.99 ***
domain [[Link]]
4.49 ***
domain [[Link]]
1.14 ***
root [b−root]
1.11 ***
root [c−root]
1.13 ***
root [d−root]
1.10 **
root [e−root]
1.09 **
root [f−root]
1.14 ***
root [g−root]
1.14 ***
root [h−root]
0.99
root [i−root]
1.07 *
root [j−root]
0.72 ***
root [k−root]
0.98
root [l−root]
1.09 **
root [m−root]
0.1 0.5 1 5 10 50
Odds Ratios
Fig. 5. Odds ratios of DNS injection survival. Values above 1 (in blue) indicate that
the corresponding variables increase the chances of DNS injection, while ratios below 1
(in red) decrease the chances of DNS injection. The 95% confidence limits are delimited
by horizontal lines. Those that do not cross the zero line correspond to variables that
affect DNS injection more significantly. (Color figure online)
474 Y. Nosyk et al.
Fig. 6. The ratio (in %) of probes that experienced response injection to all the probes
participating in our measurements. We did not receive any results for countries high-
lighted in grey. (Color figure online)
Intercept and Inject: DNS Response Manipulation in the Wild 475
Table 3. The ratio of RIPE Atlas probes experiencing manipulation to all those hosted
in a particular country. This table only includes countries with at least one probe
experiencing manipulation.
References
1. Root Server Technical Operations Association (2022). [Link]
2. Anderson, C., Winter, P., Ensafi, R.: Global censorship detection over the RIPE
Atlas network. In: USENIX FOCI (2014)
3. Anonymous: The Collateral Damage of Internet Censorship by DNS Injection.
SIGCOMM Comput. Commun. Rev. 42(3), June 2012
4. Anonymous: Towards a Comprehensive Picture of the Great Firewall’s DNS Cen-
sorship. In: USENIX FOCI (2014)
5. Anonymous, Niaki, A.A., Hoang, N.P., Gill, P., Houmansadr, A.: Triplet censors:
demystifying great firewall’s DNS censorship behavior. In: USENIX FOCI (2020)
6. APNIC: Encrypted DNS World Map, January 2023. [Link]
edns
7. Filastò, A., Appelbaum, J.: OONI: open observatory of network interference. In:
USENIX FOCI (2012)
8. Austein, R.: DNS Name Server Identifier (NSID) Option. RFC 5001 (2007)
9. Bailey, M., Kenneally, E., Maughan, D., Dittrich, D.: The menlo report. IEEE
Secur. Privacy 10(02), 71–75 (2012)
10. Bayer, J., Nosyk, Y., Hureau, O., Fernandez, S., Paulovics, I., Duda, A., Kor-
czyński, M.: Study on Domain Name System (DNS) abuse : technical report.
Appendix 1. Publications Office of the European Union (2022). [Link]
10.2759/473317
11. Bhaskar, A., Pearce, P.: Many roads lead to Rome: how packet headers influence
DNS censorship measurement. In: USENIX Security (2022)
12. Bock, K., Alaraj, A., Fax, Y., Hurley, K., Wustrow, E., Levin, D.: Weaponizing
middleboxes for TCP reflected amplification. In: USENIX Security (2021)
13. Bortzmeyer, S.: DNS Censorship (DNS Lies) As Seen By RIPE Atlas, Decem-
ber 2015. [Link] bortzmeyer/dns-censorship-dns-
lies-as-seen-by-ripe-atlas/
14. Bortzmeyer, S., Dolmans, R., Hoffman, P.E.: DNS query name minimisation to
improve privacy. RFC 9156 (2021)
15. Bretelle, M.: [dns-operations] K-root in CN leaking outside of CN, November 2021.
[Link]
16. Chung, T., Choffnes, D., Mislove, A.: Tunneling for transparency: a large-scale
analysis of end-to-end violations in the internet. In: IMC (2016)
17. Chung, T., van Rijswijk-Deij, R., Chandrasekaran, B., Choffnes, D., Levin, D.,
Maggs, B.M., Mislove, A., Wilson, C.: A Longitudinal. USENIX Security, End-to-
End View of the DNSSEC Ecosystem. In (2017)
18. Conrad, D.R., Woolf, S.: Requirements for a Mechanism Identifying a Name Server
Instance. RFC 4892 (2007)
19. Dagon, D., Lee, C., Lee, W., Provos, N.: Corrupted DNS resolution paths: the rise
of a malicious resolution authority. In: NDSS (2008)
20. DNSFilter: DNS Threat Protection (2022). [Link]
21. Fan, X., Heidemann, J., Govindan, R.: Evaluating anycast in the domain name
system. In: IEEE INFOCOM (2013)
22. Farnan, O., Darer, A., Wright, J.: Poisoning the well: exploring the great firewall’s
poisoned DNS responses. In: WPES (2016)
23. Gill, P., Crete-Nishihata, M., Dalek, J., Goldberg, S., Senft, A., Wiseman, G.:
Characterizing Web censorship worldwide: another look at the OpenNet initiative
data. ACM Trans. Web 9(1), 1–29 (2015)
Intercept and Inject: DNS Response Manipulation in the Wild 477
24. Gillmor, D.K., Salazar, J., Hoffman, P.E.: Unilateral Opportunistic Deployment
of Encrypted Recursive-to-Authoritative DNS. Internet-Draft draft-ietf-dprive-
unilateral-probing-02, Internet Engineering Task Force, September 2022. work in
Progress
25. Google: SafeSearch (2022). [Link]
26. Hilton, A., Deccio, C., Davis, J.: Fourteen years in the life: a root server’s perspec-
tive on DNS resolver security. In: USENIX Security (2023)
27. Hoang, N.P., Doreen, S., Polychronakis, M.: Measuring I2P censorship at a global
scale. In: USENIX FOCI (2019)
28. Hoang, N.P., Niaki, A.A., Dalek, J., Knockel, J., Lin, P., Marczak, B., Crete-
Nishihata, M., Gill, P., Polychronakis, M.: How Great is the Great Firewall?
USENIX Security, Measuring China’s DNS Censorship. In (2021)
29. Hoffman, P.E., McManus, P.: DNS Queries over HTTPS (DoH). RFC 8484 (2018)
30. Hu, Z., Zhu, L., Heidemann, J., Mankin, A., Wessels, D., Hoffman, P.E.: Specifi-
cation for DNS over Transport Layer Security (TLS). RFC 7858 (2016)
31. Huitema, C., Dickinson, S., Mankin, A.: DNS over Dedicated QUIC Connections.
RFC 9250 (2022)
32. Jones, B., Feamster, N., Paxson, V., Weaver, N., Allman, M.: Detecting DNS root
manipulation. In: PAM (2016)
33. Kührer, M., Hupperich, T., Rossow, C., Holz, T.: Exit from hell? reducing the
impact of amplification DDoS attacks. In: USENIX Security (2014)
34. Kumari, W.A., Hoffman, P.E.: Running a Root Server Local to a Resolver. RFC
8806 (2020)
35. Li, C., Cheng, Y., Men, H., Zhang, Z., Li, N.: Performance analysis of root anycast
nodes based on active measurement. Electronics 11(8), 1194 (2022)
36. Li, Z., Levin, D., Spring, N., Bhattacharjee, B.: Internet Anycast: Performance,
Problems, & Potential. SIGCOMM (2018)
37. Lindqvist, K.E., Abley, J.: Operation of Anycast Services. RFC 4786 (2006)
38. Liu, B., Lu, C., Duan, H., Liu, Y., Li, Z., Hao, S., Yang, M.: Who is answering
my queries: understanding and characterizing interception of the DNS resolution
path. In: USENIX Security (2018)
39. Lowe, G., Winters, P., Marcus, M.L.: The Great DNS Wall of China. New York
University, Technical report (2007)
40. Lu, C., et al.: An end-to-end, large-scale measurement of DNS-over-encryption:
how far have we come? In: IMC (2019)
41. Mockapetris, P.: Domain names - concepts and facilities. RFC 1034 (1987)
42. Mockapetris, P.: Domain names - implementation and specification. RFC 1035
(1987)
43. Moura, G.C.M., et al.: Old but gold: prospecting TCP to engineer and live monitor
DNS anycast. In: PAM (2022)
44. Moura, G.C.M., et al.: Anycast vs. DDoS: evaluating the November 2015 root DNS
event. In: IMC (2016)
45. Nawrocki, M., Koch, M., Schmidt, T.C., Wählisch, M.: Transparent forwarders: an
unnoticed component of the open DNS infrastructure. In: CoNEXT (2021)
46. Niaki, A.A., et al.: ICLab: a global, longitudinal internet censorship measurement
platform. In: IEEE S&P (2020)
47. Pearce, P., Ensafi, R., Li, F., Feamster, N., Paxson, V.: Towards continual mea-
surement of global network-level censorship. In: IEEE S&P (2018)
48. Pearce, P., et al.: Global measurement of DNS manipulation. In: USENIX Security
(2017)
478 Y. Nosyk et al.
49. Le Pochat, V., Van Goethem, T., Tajalizadehkhoob, S., Korczyński, M., Joosen,
W.: Tranco: a research-oriented top sites ranking hardened against manipulation.
In: NDSS (2019)
50. Raman, R.S., Stoll, A., Dalek, J., Ramesh, R., Scott, W., Ensafi, R.: Measuring
the deployment of network censorship filters at global scale. In: NDSS (2020)
51. Randall, A., et al.: Home is where the hijacking is: understanding DNS interception
by residential routers. In: IMC (2021)
52. RIPE Atlas: Legal (2020). [Link]
53. RIPE Atlas: Built-in Measurements (2022). [Link]
measurements/
54. RIPE Ncc: RIPE Atlas (2022). [Link]
55. Rose, S., Larson, M., Massey, D., Austein, R., Arends, R.: DNS Security Introduc-
tion and Requirements. RFC 4033 (2005)
56. Snabb, J.: [Link] moved to Beijing? [Link]
nanog/2011/Oct/12, October 2011
57. Sundara Raman, R., Shenoy, P., Kohls, K., Ensafi, R.: Censored planet: an internet-
wide, longitudinal censorship observatory. In: CCS (2020)
58. VanderSloot, B., McDonald, A., Scott, W., Halderman, J.A., Ensafi, R.: Quack:
scalable remote measurement of application-layer censorship. In: USENIX Security
(2018)
59. Vergara Ereche, M.: [dns-operations] Odd behaviour on one node in I root-server
(facebook, youtube & twitter), March 2010. [Link]
dns-operations/2010-March/[Link]
60. Weaver, N., Kreibich, C., Nechaev, B., Paxson, V.: Implications of Netalyzrs DNS
Measurements. In: SATIN (2011)
61. Weaver, N., Kreibich, C., Paxson, V.: Redirecting DNS for Ads and Profit. In:
USENIX FOCI (2011)
A First Look at Brand Indicators
for Message Identification (BIMI)
1 Introduction
As promising countermeasure technologies against phishing emails, sender
authentication techniques such as Sender Policy Framework (SPF) [38], Domain-
based Message Authentication, Reporting & Conformance (DMARC) [26],
and DNS-Based Authentication of Named Entities (DANE) [23] have been
standardized and have become widespread. In addition to these technologies,
c The Author(s) 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 479–495, 2023.
[Link]
480 M. Yajima et al.
– This is the first large-scale measurement study of the adoption and operation
of BIMI in the wild.
– Of the one million popular domain names, 3,538 have BIMI records.
– Of the 3,538 domain names with a BIMI configuration, only 11% had a valid
logo image and VMC.
– In domain names that had set up a VMC for BIMI, DMARC was set up in
99.5% of the domain names.
– We found 16 BIMI misconfigurations/violations in BIMI records, 1,224 in
logos, 58 in VMCs, and 14 in the DMARC configuration.
– We found 45 domain names having differences between the images contained
in the VMC and the images provided on the server.
– In this study, we found no cases of attacks exploiting BIMI.
A First Look at Brand Indicators for Message Identification (BIMI) 481
2 Background
In this section, we first review the email security mechanisms. We then describe
the specification of BIMI. For reference, we present the survey results of BIMI
implementations for major mail user agents in Appendix.
Table 5 in appendix summarizes the DNS records that must be set for each
of the security mechanisms described above. “Configure” indicates who needs to
configure the record.
BIMI Record: To enable BIMI for a domain name, the following data must be
added to the TXT record of the domain name of the MX server:
v=BIMI1;l=<logo link>;a=<vmc link>,
where logo link describes the brand logo link and vmc link describes the link
for the VMC. Among these links, only https is allowed as a schema.
Logo Image: The brand logo images used by BIMI must be provided in the
SVG file format defined in RFC 6170 [36]. SVG Tiny P/S, currently proposed
as an Internet Draft [17], sets the following restrictions:
• A title tag must be included (64 characters or less is recommended).
• The following attributes must be set in an svg tag:
xmlns="[Link]
version="1.2",
baseProfile="tiny-ps".
• The inclusion of a desc tag is also recommended.
• The size of the logo is recommended to be less than 32 KB.
VMC: VMC is a digital certificate used to certify the ownership of a logo.
Currently, DigiCert and Entrust are two CAs that can issue a VMC [14].
DMARC: In DMARC, the domain name owner can set a policy regarding what
action should be taken by the email recipient when the source authentication by
SPF or DKIM fails. The three policies are as follows:
• “none” indicates that no specific action will be taken.
• “quarantine” indicates that the email recipient will treat as suspicious email
that fails the DMARC mechanism check. The email recipient must take
action, such as placing the email in the spam folder or conducting further
investigations.
• “reject” indicates that an email that fails the DMARC mechanism check is
rejected.
DMARC allows one domain name and its subdomain names to be indepen-
dently configured. A pct is a field that allows the domain name administrator
to gradually implement the DMARC mechanism. By setting the pct, it is pos-
sible to apply a strong denial policy with a certain probability; otherwise, the
next-strongest denial policy is applied. To use BIMI, domain name administra-
tors must fully implement the DMARC mechanism. When using BIMI, “none”
should not be applied.
Vetting Process. In order to use BIMI, it is necessary to obtain a valid VMC
for the target logo as an email client will test both BIMI record and VMC.
A user wishing to obtain a VMC for their logo submits the trademarked logo
and information verifying the identity of the user to the VMC-issuing CA. The
CA will review the submitted information and also conduct a video conference
with the user. If no problems are found, a VMC associated with the logo will be
issued. The two VMC-issuing CAs clearly describe in their Certification Practice
Statement (CPS) that they meet the official security requirements for issuing
A First Look at Brand Indicators for Message Identification (BIMI) 483
the VMC [6,8,9]. They are subject to an external audit in order to conduct the
business of issuing VMCs. This audit is similar to the external audit that CAs
issuing server certificates in Web PKI undergo.
3 Measurement Method
In this section, we present the list of domain names we target for our analysis
and the data collection methodology.
Level Description
1 Has a valid BIMI record shown in Table 5
2 Has a valid logo available for download
3 Has valid logo and VMC available for download
Fig. 1. Fractions (%) of the domain names with valid BIMI records. 10n represents the
logarithmic rank interval ranging from the 10n−1 + 1 th domain to the 10n th domain.
DKIM: In the DKIM survey, we used default and key1 as the selectors.
DANE: In the DANE study, the domain names listed in the MX records were
targeted. If at least one of the domain names listed in the MX record supports
DANE, the domain name is determined to have adopted DANE.
Table 2. Correlations of the email security mechanisms: BIMI vs. other mechanisms.
The rows indicate other email security mechanisms and the columns indicate the BIMI
setting level. The numerical values in the table indicate the number of domain names.
list. As expected, the higher the ranking of a domain name, the higher the rate
of BIMI adoption; for the top-100 domains, more than 10% of domain names
have configured a valid BIMI record. On the other hand, we can see that a
certain number of domain names with low rankings have also adopted BIMI,
suggesting that the use of BIMI is spreading. For reference, we analyzed the
breakdown of the domain names that have configured BIMI. The results are
shown in Appendix.
We analyzed the correlation between BIMI and other DNS-based email security
mechanisms, i.e., whether they are simultaneously employed. Table 2 presents the
results. “MX-enabled” indicates that the results are restricted to only domain
names for which MX records existed. As described in Sect. 2.1, if an email
recipient retrieves BIMI data for a domain name, the domain name must pass
the DMARC authentication, and the configured policy must be “quarantine”
or “reject.” Therefore, a high percentage of BIMI-enabled domain names have
adopted SPF and DMARC.
We found that the number of domain names configuring BIMI is larger than
those of MTA-STS and TLSRPT. This result suggests that BIMI is attracting
the attention of more domain name administrators despite being a relatively new
security mechanism. If a domain name operates BIMI with Level 3 and DANE,
the domain name has an extremely high security level. We found that only two
domain names meet these criteria. DANE requires DNSSEC [18,28,31,33–35]
settings, which are difficult to configure.
We applied BIMI record lookups on the domain names of phishing emails and
websites, which we describe in Sect. 3.1. We found no BIMI records for 114,915
domain names in the two datasets combined; that is, as of today, we have not
observed any phishing attempts that exploit BIMI records. We expect that this
486 M. Yajima et al.
observation is due to the fact that the trademark registration process contributes
to raising barriers to BIMI record operations. However, there is no assurance that
BIMI-abusing domain names will not appear in the future, and it is therefore
necessary to keep a close watch on this aspect.
We first study the inherent configuration errors we found with respect to the
format of the BIMI records collected. It is meaningful to summarize such infor-
mation and share explicit knowledge of the mistakes that administrators are
prone to make.
Logo Setting: Two of the domain names did not have a field to set the logo. In
one of these two cases, only a link to the certificate existed. In addition, although
11 domain names had a field for setting a logo, the content was empty, where
the empty content in the logo setting field indicates that the domain name in
question explicitly refuses to participate in BIMI.
Use of HTTP: There are five domain names whose logo URLs used http
instead of https. None of the five domain names has a URL for the certificate.
Similarly, one domain name was used http in the URL pointing to the certificate.
The URL pointed to the Let’s Encrypt server and not the certificate.
Typos: Six domain names were incorrectly used I= instead of l= as the field
for setting the logo. The certificate link did not exist for any of the six domain
names.
Unnecessary Parentheses: One domain name existed in which the domain
name was described as l=[<logo link>] when setting the logo. The domain
name in question does not contain a certificate link set.
Invalid String: Two domain names existed, in which invalid character strings
were set in records that should describe the URLs.
These misconfigurations were found in domain names that had set only a
logo or had not set a logo at all.
We analyzed logo images in SVG format retrieved from the URLs listed in the
BIMI records. A total of 3,034 logo images were analyzed. In the following,
we show the cases that violated the mandatory and recommended conditions
described in the Internet Draft [17] of SVG shown in Sect. 2.2. Of the domain
names with VMC configured, only five domain names failed to configure SVG
in the correct format.
Title Tag—mandatory: There were 1,008 (33%) logo images without title tags.
Two images with empty title tags are found.
A First Look at Brand Indicators for Message Identification (BIMI) 487
SVG Tag—mandatory: There were 1,224 (40.3%) logo images that did not
conform to the svg tag format.
Desc Tag—recommended: A total of 2,905 (95.7%) logo images did not contain
a desc tag.
Image Size—recommended: In total, 241 (7.9%) logo images exceeded the rec-
ommended 32 KB.
Aspect Ratio—recommended: Logos displayed on email clients are often circles
or squares. It is therefore recommended that the aspect ratio of the logo be 1:1 [1],
and 496 (16.3%) of the logo images do not have this aspect ratio.
5.3 VMC
We analyzed VMCs obtained from the URLs listed in the BIMI records. The
analysis covered 396 certificates collected from domain names with Level 3 BIMI
settings, as shown in Table 1.
Certificate Issuer: Table 3 shows a breakdown of the issuers of the collected
certificates. Currently, certificates issued by parties other than Entrust and Dig-
icert are invalid for BIMI, among which there are five such cases. These certifi-
cates did not contain logo images, whereas all certificates issued by Entrust and
Digicert contained image data.
Certificate Validity Period: We analyzed the validity period of the collected
certificates. As a result, 13 certificates had expired. One of these is the domain
name [Link], which was used by Entrust. The domain name
redirects [Link] However, BIMI records, logos, and certifi-
cate links are still accessible.
Legitimacy of Images Extracted from the VMC: We verified whether the
391 logo images extracted from the collected VMCs matched the logos collected
from the URLs listed in the BIMI records. We found 45 domain names for which
there was a difference between the two logo images. The differences included
the use of completely different images, the presence of line breaks in the files,
differences in the image size, and differences in the SVG titles.
rows represent the configuration policies for the target domain names, and the
columns represent configuration policies for the subdomain names. In the table,
bold numbers indicate the number of policy violations, 12 of which were present.
6 Discussion
6.1 Current Status of BIMI
6.2 Limitations
Our study has the following three limitations. First, in our study, we sent only a
minimum number of queries (up to three) to avoid overloading the target. This
means that if the target server was offline during our study, the data might not
have been correctly retrieved. Second, our study only investigated the specific
selectors for BIMI and DKIM. Therefore, if the target of our survey is to use indi-
vidual selectors for each sending destination, it may be judged as unsupported
in our study. Finally, our study did not clarify the current status of BIMI from
the viewpoint of administrators and email recipients. To investigate the current
issues in setting up BIMI and the effectiveness of BIMI from the viewpoint of
the recipients, it is necessary to conduct an interview study.
Our measurement study discovered several domain names with incorrect BIMI
settings. As an ethical consideration, we decided to notify the administrators of
those domain names to prevent their misuse. In particular, we are in the process
of making a responsible disclosure to the administrators of domain names with
VMC configured but with some misconfiguration. We also plan to notify the
administrators of domain names that have only SVG configured.
7 Related Work
SPF, DKIM, and DMARC: In 2011, Mori et al. conducted an early study on
SPF implementation by investigating the existence of SPF and the errors found
in SPF policies [30]. In 2015, Durumetric et al. measured email servers sup-
porting SPF, DKIM, and DMARC by analyzing SMTP connections on Google’s
email servers [20]. In 2015, Foster et al. investigated the prevalence of SPF and
DMARC from the perspective of email providers [21]. Hu et al. studied the
states of support for SPF, DKIM, and DMARC in 35 email providers in 2018,
and conducted a phishing email measurement with end-users [24]. Deccio et al.
measured the latest status of SPF, DKIM, and DMARC on several email servers
in 2021 [19]. Tatang et al. continuously investigated the status of SPF, DKIM,
and DMARC support for domain names listed in multiple top lists in 2021 for a
period of 1.5 years [40]. Wang et al. conducted measurements of DKIM deploy-
ments using a 5-year Chinese Passive DNS dataset from 2015 to 2020 and server
logs of an Chinese email provider in 2020 [41].
Others: In addition, measurement studies were conducted to elucidate other
individual protocols (see Sect. 2.1). Scheitle et al. were the first to examine the
number of CAAs deployed in 2018 [37]. In 2020, Lee et al. conducted an exten-
sive study to determine how widely DANEs are spread and managed at both
the server and client sides [27]. Tatang et al. conducted the first large-scale
measurement study of MTA-STS adoption in 2021 [39]. Yajima et al. measured
the adoption rates of DNSSEC, DNS Cookies, CAA, SPF, DMARC, MTA-STS,
DANE, and TLSRPT, which are security mechanisms that can be implemented
in 2021 [42].
None of the studies above mentioned any quantitative results for BIMI, which
is just beginning to spread, and our study is the first BIMI measurement app-
roach as of November 2022.
8 Conclusion
In this study, we conducted the first large-scale measurement of BIMI in the
wild. We investigated the prevalence of BIMI in one million domain names and
found that 3,538 already had BIMI records, despite the BIMI mechanism having
only recently begun to be used. We also found that there are intrinsic miscon-
figuration patterns and specification violations in BIMI records, logos, VMCs,
and DMARCs. In addition, no evidence of BIMI abuse was found during our
investigation. For the coming widespread use of BIMI, future work includes
development of a tool that enables domain name administrators to configure
BIMI settings easily and properly, conducting interviews with both domain name
administrators and email users on the incentives of adopting/leveraging BIMI,
and continuously measure the adoption status of BIMI. We hope that the find-
ings we derived through our measurement study of the BIMI will contribute to
its further spread and help thwart the damages caused by phishing attacks.
A First Look at Brand Indicators for Message Identification (BIMI) 491
Table 6. BIMI adoption status of major MUAs. indicates that the valid BIMI logo
was correctly displayed on the corresponding MUA.
In the following, we summarize the current support status of BIMI by the major
Mail User Agents (MUAs) – both webmail services and application-based email
clients. As webmail services, we adopted Gmail, Fastmail, and Yahoo Mail. We
used Google Chrome to study the BIMI adoption status of these webmail ser-
vices. As email client apps, we adopt Apple Mail, Microsoft Outlook, and Thun-
derbird. For Gmail in particular, we checked Gmail apps that work on iOS and
Android.
We picked up the two popular websites operated with the following BIMI-
compatible domain names.
Note that since our goal is not to expose the level of BIMI operation for spe-
cific institutions, and since the BIMI configuration status is likely to be updated
in the future and is not invariant, we decided to refrain from naming the respec-
tive websites. In addition, since the purpose of this study is to evaluate the BIMI
compatibility of MUAs, the type of website does not matter as long as the BIMI
setting on the domain name side is consistent.
We registered email accounts on the two websites, where we used different
email accounts for each MUA. Emails sent from each website were received by
the MUAs used in the experiment to study the adoption of BIMI by MUAs.
Table 6 presents the results of studying whether or not each MUA displays
the BIMI logo for emails sent from website 1 and website 2. The behavior of a cor-
rectly developed BIMI implementation is to display the logo for website 1, which
has perfectly configured BIMI, and not for website 2, which has registered BIMI
records but has incomplete VMC. The study revealed that, for webmail-based
MUAs, Gmail and Yahoo Mail, accessed with Chrome, correctly implemented
BIMI. Fastmail displays the BIMI logo for correctly configured domain names,
but does not validate the VMC. Considering the risk of the above-mentioned fact
being exploited in a phishing attack, we are currently in the process of making a
responsible disclosure to Fastmail. In the email apps, Apple Mail and the all the
versions of the Gmail apps correctly implemented BIMI. As of November 2022,
Outlook and Thunderbird do not support BIMI.
We have categorized domain names that have adopted BIMI. To this end, we
leveraged SimilarWeb [12], which is a commercial database that collects web
traffic statistics and compiles website information collected from million-order
devices deployed around the world. We made use of SimilarWeb to identify cate-
gories of domain names, both for those with BIMI records present, and for those
with VMC set in addition to BIMI records. Table 7 presents the aggregated
Table 7. Top-10 categories of domain names with BIMI configuration. Level 1 (left)
and Level 3 (right).
Level 1 Level 3
Rank Category Count Category Count
1 Computers and Electronics 697 Finance 66
2 Unknown 406 Computers and Electronics 59
3 Finance 373 Lifestyle 29
4 Business and Consumer Services 269 Business and Consumer Services 28
5 Science and Education 197 E-commerce and Shopping 26
6 Health 158 Arts and Entertainment 26
7 Lifestyle 155 Health 25
8 E-commerce and Shopping 140 News and Media 20
9 Travel and Tourism 133 Food and Drink 17
10 Food and Drink 127 Travel and Tourism 15
A First Look at Brand Indicators for Message Identification (BIMI) 493
results for the top-10 categories for domain names with BIMI configuration of
Level 1 and Level 3. Majority of Level-1 websites were dominated by Comput-
ers and Electronics, Finance, and Business uses. Note that “Unknown” indicates
that the category of the website with that domain name was not identified in
SimilarWeb. For the Level-3, the breakdown of the websites was different from
the above, with Finance topping the list. This observation suggests that since
financial websites are often the target of phishing attacks, there is an incentive
for them to eagerly take measures using BIMI.
References
1. Creating BIMI SVG Logo Files (2020). [Link]
logo-files/
2. SVG Conversion Tools Released (2020). [Link]
tools-released/
3. Fastmail now supports BIMI (2021). [Link]
[Link]
4. BIMI Inspector (2022). [Link]
5. BIMI Record Checker - BIMI Record | EasyDMARC (2022). [Link]
com/tools/bimi-lookup
6. DigiCert Certificate Policy/ Certification Practices Statement for Private
PKI Services (2022). [Link]
[Link]
7. Email Sender Identity Verification, Authentication & Security Solutions | Valimail
(2022). [Link]
8. ENTRUST CERTIFICATE SERVICES Certification Practice Statement (2022).
[Link]
st-certifi[Link]?la=en&hash=EA7E3B4CDEB02433939E7
F7AB2762E60
9. Minimum Security Requirements for Issuance of Verified Mark Certificates Version
1.4 (2022). [Link]
10. OpenPhish (2022). [Link]
11. phash (2022). [Link]
12. SimilarWeb (2022). [Link]
13. Tranco (2022). [Link]
14. VMC Issuer Information (2022). [Link]
15. Barnes, R.: Use Cases and Requirements for DNS-Based Authentication of Named
Entities (DANE). RFC 6394, October 2011. [Link]
16. Blank, S., Goldstein, P., Loder, T., Zink, T., Bradshaw, M., Brotman, A.:
Brand Indicators for Message Identification (BIMI). Internet-Draft draft-brand-
indicators-for-message-identification-01, Internet Engineering Task Force, April
2022. [Link]
identification-01, work in Progress
17. Brotman, A., Adams, J.T.: SVG Tiny Portable/Secure. Internet-Draft draft-
svg-tiny-ps-abrotman-03, Internet Engineering Task Force, April 2022. https://
[Link]/doc/html/draft-svg-tiny-ps-abrotman-03, work in Progress
18. Chung, T., et al.: A longitudinal, end-to-end view of the DNSSEC ecosystem. In:
Proceedings of USENIX Security Symposium (2017)
494 M. Yajima et al.
19. Deccio, C.T., et al.: Measuring email sender validation in the wild. In: Proceed-
ings of the International Conference on emerging Networking EXperiments and
Technologies (CoNEXT) (2021). [Link]
20. Durumeric, Z., et al.: Neither snow nor rain nor MITM...: an empirical analy-
sis of email delivery security. In: Proceedings of the ACM Internet Measurement
Conference (IMC) (2015). [Link]
21. Foster, I.D., Larson, J., Masich, M., Snoeren, A.C., Savage, S., Levchenko, K.:
Security by any other name: On the effectiveness of provider based email secu-
rity. In: Proceedings of the ACM Conference on Computer and Communications
Security (CCS) (2015). [Link]
22. Hoffman, P.E.: SMTP Service Extension for Secure SMTP over Transport Layer
Security. RFC 3207, February 2002. [Link]
23. Hoffman, P.E., Schlyter, J.: The DNS-Based Authentication of Named Entities
(DANE) Transport Layer Security (TLS) Protocol: TLSA. RFC 6698, August 2012.
[Link]
24. Hu, H., Wang, G.: End-to-end measurements of email spoofing attacks. In: Pro-
ceedings of the USENIX Security Symposium (2018). [Link]
conference/usenixsecurity18/presentation/hu
25. Kucherawy, M., Crocker, D., Hansen, T.: DomainKeys Identified Mail (DKIM) Sig-
natures. RFC 6376, September 2011. [Link] https://
[Link]/info/rfc6376
26. Kucherawy, M., Zwicky, E.: Domain-based Message Authentication, Reporting,
and Conformance (DMARC). RFC 7489, March 2015. [Link]
RFC7489
27. Lee, H., Gireesh, A., van Rijswijk-Deij, R., Kwon, T., Chung, T.: A longitudi-
nal and comprehensive study of the DANE ecosystem in email. In: Proceedings
of the USENIX Security Symposium (2020). [Link]
usenixsecurity20/presentation/lee-hyeonmin
28. Lian, W., Rescorla, E., Shacham, H., Savage, S.: Measuring the practical impact of
DNSSEC deployment. In: Proceedings of the USENIX Security Symposium (2013)
29. Margolis, D., Brotman, A., Ramakrishnan, B., Jones, J., Risher, M.: SMTP TLS
Reporting. RFC 8460, September 2018. [Link]
30. Mori, T., Sato, K., Takahashi, Y., Ishibashi, K.: How is e-mail sender authentica-
tion used and misused? In: Proceedings of the Collaboration, Electronic messag-
ing, Anti-Abuse and Spam Conference (CEAS) (2011). [Link]
2030376.2030380
31. Müller, M., Chung, T., Mislove, A., van Rijswijk-Deij, R.: Rolling with confidence:
managing the complexity of dnssec operations. IEEE Trans. Netw. Serv. Manage.
(2019). [Link]
32. Nominum: dnspython (2022). [Link]
33. Rose, S., Larson, M., Massey, D., Austein, R., Arends, R.: DNS Security Intro-
duction and Requirements. RFC 4033, March 2005. [Link]
RFC4033
34. Rose, S., Larson, M., Massey, D., Austein, R., Arends, R.: Resource Records for
the DNS Security Extensions. RFC 4034, March 2005. [Link]
RFC4034
35. Rose, S., et al.: Protocol Modifications for the DNS Security Extensions. RFC 4035,
March 2005. [Link]
36. Santesson, S., Housley, R., Rosenthol, L., Bajaj, S.: Internet X.509 Public Key
Infrastructure - Certificate Image. RFC 6170, May 2011. [Link]
RFC6170. [Link]
A First Look at Brand Indicators for Message Identification (BIMI) 495
37. Scheitle, Q., et al.: A first look at certification authority authorization (CAA).
Comput. Commun. Rev. (2018). [Link]
38. Schlitt, W., Wong, M.W.: Sender Policy Framework (SPF) for Authorizing Use of
Domains in E-Mail, Version 1. RFC 4408, April 2006. [Link]
RFC4408
39. Tatang, D., Flume, R., Holz, T.: Extended abstract: A first large-scale analysis
on usage of MTA-STS. In: Proceedings of the Detection of Intrusions and Mal-
ware, and Vulnerability Assessment (DIMVA) (2021). [Link]
978-3-030-80825-9_18
40. Tatang, D., Zettl, F., Holz, T.: The evolution of dns-based email authentication:
measuring adoption and finding flaws. In: Proceedings of the International Sym-
posium on Research in Attacks, Intrusions and Defenses (RAID) (2021). https://
[Link]/10.1145/3471621.3471842
41. Wang, C., et al.: A large-scale and longitudinal measurement study of DKIM
deployment. In: Proceedings of the USENIX Security Symposium (2022)
42. Yajima, M., Chiba, D., Yoneya, Y., Mori, T.: Measuring adoption of DNS secu-
rity mechanisms with cross-sectional approach. In: Proceedings of the IEEE
Global Communications Conference (GLOBECOM) (2021). [Link]
1109/GLOBECOM46510.2021.9685960
Open Access This chapter is licensed under the terms of the Creative Commons
Attribution 4.0 International License ([Link]
which permits use, sharing, adaptation, distribution and reproduction in any medium
or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons license and indicate if changes were
made.
The images or other third party material in this chapter are included in the
chapter’s Creative Commons license, unless indicated otherwise in a credit line to the
material. If material is not included in the chapter’s Creative Commons license and
your intended use is not permitted by statutory regulation or exceeds the permitted
use, you will need to obtain permission directly from the copyright holder.
A Second Look at DNS QNAME
Minimization
Jonathan Magnusson1(B) , Moritz Müller2 , Anna Brunstrom1 ,
and Tobias Pulls1
1
Karlstad University, Karlstad, Sweden
{[Link],[Link],[Link]}@[Link]
2
SIDN Labs, Arnhem, The Netherlands
[Link]@[Link]
1 Introduction
of research for the past decades, these topics were less prevalent when DNS was
implemented over 35 years ago. The early Internet was relatively small, where
everyone knew each other, and the focus was on getting data from point A to
point B. Similar to other areas of the Internet, this has resulted in multiple pro-
posals to improve DNS privacy and security as an afterthought. By encrypting
DNS traffic with TLS, HTTPS, or QUIC [9–11], it is possible to achieve trans-
port confidentiality. By signing sets of RRs and building chains of trust using
DNSSEC [1], it is possible to prove integrity of the data. Due to the hierarchical
structure of DNS, the top two levels of servers—root and Top-Level Domain
(TLD)—are observing a large portion of non-cached requests on the Internet.
From a privacy perspective, it is important that the information sent here is the
minimum needed for each task. This is known as the fundamental privacy prin-
ciple of data minimization [6]. The DNS resolvers may strip unnecessary labels
for each query in the lookup process. This privacy feature is referred to as query
name minimization (qmin) and is at present standardized in RFC 9156 [4].
The aim of this study is to measure the adoption of qmin. To do so, we build
upon the experiments of De Vries et al. [23], who took a first look at qmin adop-
tion in 2018. We considerably extend the experiments, consider an additional
source of passive measurements, and include an additional open-source resolver
for the controlled experiments. Our contributions are as follows:
1. Extended active measurements from 2018 up until October 2022 show that the
adoption of qmin has increased from 2.5k resolvers used by RIPE Atlas probes
in 2018 to 14k. We also show that the adoption of qmin has increased using
active measurements on open resolvers, from 18k open resolvers categorized
as qmin-enabled in 2018 to 80k in 2022.
2. Extended passive measurements at root and TLD name servers show an
increase of qmin adoption from 0.6% in 2018 to 2.5% at one root server
and from 35.5% in 2019 to 57.3% at the .nl TLD. We also observe that a
significant amount of noise is removed when filtering out invalid labels at
root.
3. Up-to-date performance and error-rate measurements of four open source
resolvers (Bind, Knot, PowerDNS, and Unbound) show that while error-rates
have been significantly reduced for all resolvers, the number of packets have
gone both up (Bind and Unbound) and down (Knot).
4. Grounded in our findings, we discuss the trade-off between privacy and per-
formance of minimizing queries, and propose that a promising solution may
be to set the depth limitation of qmin using a public suffix list.
The rest of the paper is structured as follows. Section 2 provides background
with regards to DNS and qmin as well as related work on qmin. We present
active measurements surveying qmin adoption from the client-side perspective
in Sect. 3. Section 4 presents passive measurements surveying qmin adoption from
the name server perspective. Controlled experiments measuring the performance
and error rates of resolvers with qmin implemented are shown in Sect. 5. Section 6
discusses our findings, focusing on the observed resolver behavior and the depth
limit of minimizing queries. Finally, Sect. 7 concludes the paper.
498 J. Magnusson et al.
of the TLD name servers for .domain it appends one additional label to
the query ([Link]) and sends it to the TLD name server. The
resolver will then receive the address to the authoritative name server for
[Link] and finally append the third and last piece of the domain name
([Link]), send the last request and, hopefully, get the requested
RR in return.
The default implementation of qmin in RFC 7816 has two main challenges.
Firstly, the standardized RR type for queries when using qmin is the NS RR,
which could cause some name servers to not respond or result in an error if
no such RR is found for a minimized query. This behavior is not according to
the standard, but an error on the name server side. Secondly, a domain name
with many labels creates additional traffic, which could be abused for Denial-
of-Service (DoS) attacks. If the DNS zone [Link] contains a RR for
[Link], a qmin-enabled resolver would send multiple queries
to [Link] name servers before asking for the final RR. This has led
to alternative implementations of qmin in the wild, which includes requesting A
RRs instead of NS and iteratively adding multiple labels after the second-level
label instead of one at a time.
In November 2021, RFC 7816 was obsoleted by RFC 9156 [4], which com-
bined the results and recommendations from De Vries et al. with experiences
from implementing qmin in the wild. Updated implementation details were pre-
sented to reduce error rates and improve performance while keeping a reasonable
level of privacy. The NS RR was discarded in favor of A and AAAA RRs when send-
ing minimized queries. Fallbacks for specific error codes were specified and two
tunable parameters for incrementally adding labels were introduced. The RR
types used for queries in standard DNS, RFC 7816 and RFC 9156, respectively,
are shown in Table 1.
Table 1. DNS queries and responses of Standard DNS, RFC 7816 and RFC 9156.
There are two modes called “relaxed” and “strict” when enabling qmin on
a resolver [2,17]. In “relaxed mode”, the resolver will fall back to querying for
the full query name to potentially broken name servers. In contrast, the “strict
mode” will not, and it therefore results in more non-resolved domains.
500 J. Magnusson et al.
2
[Link]
A Second Look at DNS QNAME Minimization 501
3 Active Measurements
The goal of the active measurements is to query resolvers on the Internet in
order to observe the adoption of qmin based on their responses. The active
measurements consist of two parts: resolver adoption over time (Sect. 3.1) and
adoption by open resolvers (Sect. 3.2). The former classifies the resolvers used
by RIPE Atlas probes [21] and the latter classifies resolvers from a list generated
by scanning the IPv4 address space for servers listening on UDP port 53 [20].
The purpose of the first active measurement is to see the adoption trend of
qmin over time and to observe characteristics of the resolvers adopting qmin.
The purpose of the second active measurement is to classify open resolvers and
then use these results in the passive measurements (Sect. 4.1) to enhance the
classification accuracy. We also observe the adoption of qmin on open resolvers
since the previous qmin adoption study by De Vries et al..
3
[Link]
502 J. Magnusson et al.
Results. Figure 3 shows the current trend of qmin adoption from the client-side
(RIPE Atlas probes) perspective. Green represents the number of qmin-enabled
resolvers and orange the number of not qmin-enabled resolvers. The gray in turn
shows the resolvers which are not answering to the qmin measurements, but still
responds to other queries done as part of DNSThought. The RIPE Atlas probes
are churned daily in batches and a bug in the locking system of the measurement
caused newly added probes to not query for qmin. This caused a steady increase
of gray resolvers from early 2020 to early 2022. With the help of NLnet Labs we
contacted RIPE NCC and the bug was fixed on the 6th of April 2022.
In order to see the relative adoption of qmin-enabled resolvers we created
Fig. 4, based on the assumption that the out-churned probes are not correlated
with the qmin adoption of their resolvers. We see that 64% of the resolvers used
by RIPE Atlas probes in 2022 have enabled qmin compared to 10% around the
time of the report of De Vries et al. at the end of 2018. When going back to
Fig. 3, we still see the slight increase of qmin-enabled resolvers in April 2018
that was pointed out by De Vries et al. in the original study and attributed
to Cloudflare enabling qmin on their DNS resolvers by default. There has been
a steady increase of qmin-enabled resolvers since then, until a spike in January
2020 after which the adoption was seemingly slowing down. But looking at Fig. 4,
A Second Look at DNS QNAME Minimization 503
53. For the active measurements, requests for a TXT RR were sent to a zone under
NLnet Labs’ control through each of the resolvers on the list to classify them
as either qmin-enabled or not. The flowchart in Fig. 6 shows the process of the
classification. First a query is sent to the resolver, which will either respond or
timeout. If it does not timeout the answer is checked for errors. If the response
is free from errors the next check is for a correct answer. A correct answer is
a TXT record that contains either “HOORAY” or “NO”. Finally the response is
classified as either of those two. The results from the previous study are included
in Sect. 3.2 to allow for comparison.
For this study we used the Rapid7 list from February 2022 since access to
the list was later restricted.4 The list contains 6 million addresses and was used
in April 2022 to send queries from North Virginia, Tokyo, and Frankfurt using
EC2 instances on Amazon Web Services to see if the geographical location had
any effect on the results. We sent 100 queries to each resolver to collect more
data points and get a more comprehensive view of each resolver. However, the
system described previously has the following limitation. The delegation (i.e.,
the NS records from the referral response) from a test might be cached by a
resolver that performs qmin, such that the outcome of a subsequent test to the
same resolver favors qmin. To overcome this limitation, we developed a custom
authoritative server that behaves similarly to the other system, but additionally
allows custom query names, the referrals for which are synthesized, based on
the query. This allows us to send unique query names in close proximity by
using a wildcard label, avoiding the effects of cached delegations. For example, a
query for [Link] (corresponding to the first iteration
of queries from Tokyo) results in a query of [Link] to
ns1 by a qmin resolver. In response, ns1 is able to refer the resolver to ns2 for
[Link]. When the next iteration of queries from Tokyo
is sent, [Link] will not be found in the cache.
4
[Link]
post/2022/02/10/evolving-how-we-share-rapid7-research-data-2/.
A Second Look at DNS QNAME Minimization 505
Fig. 5. Top ten qmin-enabled resolver ASNs (data source: DNSThought [16]).
Results: Over Time and Location. Table 2 shows the results for our mea-
surements (2022) from North Virginia, Tokyo, and Frankfurt as well as the ear-
lier results from 2018 by De Vries et al.. Our results are calculated from all
100 queries to each of the 6 million resolvers. The values therefore represent the
fraction of queries and not the fraction of resolvers. The columns year and geo
indicate which study and which geographical location the queries were sent from.
The #resolv column shows the size of the list from Rapid7 and the #queries
shows the total number of queries sent. The resp column shows the percentage
of queries that did not timeout. The column named noerr shows the percent-
age of queries from responding resolvers that did not contain any errors (e.g.,
SERVFAIL, NXDOMAIN and REFUSED). The correct column shows the percentage
of noerror responses that contained a correct TXT RR reply, which means that
the response was either “HOORAY” or “NO”. The last column, qmin, shows the
percentage of “HOORAYs” out of the total number of correct responses.
506 J. Magnusson et al.
The share of minimized queries in 2018 was 1.6% and in 2022 this num-
ber increased to about 16%, measured from three different geographical loca-
tions. Comparing the geographical locations we observe minimal deviations in
the results. The Rapid7 scan on UDP port 53 in February 2022 resulted in 6
million addresses, which is a decrease of 25% from 2018, and the active mea-
surements of this study using that list shows that the share of non-timeouts has
increased to almost 71% from 64%. Another interesting observation is that the
share of NOERROR replies had gone down from 32% to around 19% and more
than 90% of the errors are REFUSED. The reason could be that some resolvers
are configured to only handle queries from clients within a specified subnet. The
share of correct TXT responses have increased from 72% to 78% and finally the
share of queries classified as minimized have increased from 1.6% to 16%. So
while we only got correct responses from a small fraction of the open resolvers,
we see a ten-fold increase in the use of qmin also in this data set.
during the 100 queries. So in addition to the list of qmin-enabled resolvers and
the list of not qmin-enabled resolvers, we consider a list of resolvers which some-
times answered HOORAY and at other times NO. This list is called conflicting
resolvers. A resolver is classified as qmin-enabled if at least one query resulted in
a HOORAY and none of the queries resulted in a NO. A resolver is classified as
not qmin-enabled if at least one query resulted in a NO and none of the queries
resulted in a HOORAY. Finally, a resolver is classified as conflicting if at least one
query resulted in a HOORAY and at least one query resulted in a NO.
In the original study by De Vries et al. 0.2% of 8 million resolvers were
classified as qmin-enabled (see Table 3). In this study we classified 1.3% of 6
million resolvers as qmin-enabled. In the original study 13.7% of resolvers were
classified as not qmin-enabled, whereas 8.9% of resolvers were classified as not
qmin-enabled in this study. Additionally 2% were classified as the new category
of conflicting in this study.
Table 4. Share of conflicting resolvers, top 10 countries.
Country CN RU US BR ID AU UA IR PL ZA
Share 26.8% 11.4% 5.5% 4.5% 3.6% 3.4% 3.0% 2.4% 2.3% 2.1%
see that the Google Public DNS resolvers were classified as not qmin-enabled
in the active measurements of open resolvers above. We performed additional
queries to verify this behavior of the [Link] and [Link] Google Public DNS
resolvers using three different zones: [Link],
[Link], and [Link].
The first zone is a set of subdomains under the second-level domain of NLnet
Labs that was used for measuring qmin adoption in the study by De Vries et
al.. The second zone is the official name for measuring qmin after the publi-
cation of the original study.5 This is also the zone used by DNSThought at
NLnet Labs. The third zone was set up in early 2022 for this study, using
the label wildcard cache mitigation technique, which neither of the other two
zones implemented. All of these zones are using the same method when mea-
suring qmin from the client-side. When using Google Public DNS resolvers,
only [Link] responded with HOORAY (see Fig. 7),
which is the name used for the RIPE Atlas qmin adoption measurements in
DNSThought.
Fig. 7. Using dig to query a Google Public DNS resolver ([Link]) for qmin.
By contacting NLnet Labs we were told that Google had reached out in
May 2020 in regards to qmin. They wrote that they had implemented qmin
but with a depth limitation that stops after two labels. This would result in
a partial qmin that sends minimized labels to the roots and TLDs but not
to Second-Level Domains (2LDs) such as [Link]. They also said that they
would like to extend it to public suffix plus one label in the future. As a
result, the Google Public DNS does not show up as minimizing queries on
DNSThought at all, which is why they added an exception to the depth lim-
itation for [Link] to reflect that the resolvers do
minimize queries (up to a point). They did not want to “cheat” the system, but
still get credit for the privacy benefit of minimizing queries at the root and TLD
level. This brings to question what an adequate level of minimizing queries is in
regards to performance and privacy, which is further discussed in Sect. 6.
5
[Link]
wouter_de_vries/making-the-dns-more-private-with-qname-minimisation/.
A Second Look at DNS QNAME Minimization 509
4 Passive Measurements
In this section we show how qmin has evolved in the years between the study by
De Vries et al. and October 2022 on a larger scale. As the previous study, we rely
on data collected at the root servers as well as the .nl ccTLD. Furthermore, we
improve the measurement technique, dive deeper into qmin adoption, showing
who drives qmin and who lags behind, and find that qmin-enabled resolvers can
leak information occasionally.
4.1 Method
Fig. 8. Minimized queries to the .nl ccTLD. Whiskers at 5th and 95th percentiles.
Fig. 9. Minimized queries to the .se ccTLD. Whiskers at 5th and 95th percentiles.
4.2 Results
Since 2019, qmin adoption has improved significantly, at least from the perspec-
tive of TLDs. Figure 10 shows the share of queries sent to the K-Root servers
and to three of the four .nl authoritative name servers. The blue color indicates
the share of queries regardless of whether the domain name exists or not. The
yellow color marks measurements that include only minimized queries to existing
domain names.
A Second Look at DNS QNAME Minimization 511
Fig. 10. Minimized queries to the .nl ccTLD and K-Root over time.
The increase in minimized queries at .nl is clearly visible and has now
reached 64% (compared to 43% in 2019). Adoption raises regardless of whether
or not we filter out queries for non-existing domain names. Occasional drops in
minimized queries are caused by random events, for example crawlers or mis-
configurations. The picture at the root is, however, less clear.
We dive deeper into this phenomenon relying on data collected at the .nl
name servers on October 4 2022.8 We map each IP address to its corresponding
country and autonomous system (AS) using the Maxmind database. On this
date, we observe 2.9B queries from 1.4M unique IP addresses from 236 countries
and 38,469 ASes. We count an IP address as a qmin-enabled resolver if the share
of existing queries is above the threshold defined above. This reflects the top
75% resolvers that have enabled qmin in our ground truth data set (see Fig. 8).
By Country. Only countries from which queries have been received from at least
100 unique IP addresses are taken into account. This leaves us with IPs from
159 countries. Interestingly, the country with the highest share of minimized
queries is Yemen (see Table 5). This high share of minimized queries is mainly
driven by large providers. Here, 10% of resolvers have qmin enabled, but those
are responsible for a large share of queries from this country. We explore the
influence of single networks in the next section. Overall, however, these countries
only account for a small fraction of total traffic at .nl.
Table 5. Countries with most minimized .nl queries.
The adoption rate is lower when we look at the countries from which .nl
receives the most queries. Table 6 summarizes the results. For these countries,
the share of minimized queries vary between 1% for China and 59.5% for Great
Britain. Also the share of qmin-enabled resolvers is the highest in Great Britain,
followed by the Netherlands.
Table 6. Deployment of qmin from the top 5 origins of .nl queries.
8
We also carried out our analysis two months earlier, with similar results.
A Second Look at DNS QNAME Minimization 513
Qmin in the Context of Other Internet Standards. Qmin has became best com-
mon practice. Overall, resolvers that have enabled qmin show more support
for “modern” Internet standards and DNS best practices. We compare resolvers
that send queries to the .nl name servers via IPv6, that indicate support for
DNSSEC10 , and that indicate a EDNS(0) buffer size of 1,232 bytes11 with
resolvers that do not follow these best practices.
Table 7 shows that resolvers that have enabled IPv6, show support for
DNSSEC, and set the recommended buffer sizes also support qmin in most cases.
These resolvers reflect 3.7% of all resolvers observed in our dataset. In contrast,
only a minority of resolvers that do not follow any of these best practices have
enabled qmin (4.8% of all resolvers in our dataset). The largest group of resolvers
do indicate support for DNSSEC, but rely on legacy IPv4, and signal a different
buffer size. Of those, 16.4% send minimized queries (63% of all resolvers in our
dataset).
Table 7. Qmin support per resolver by support of modern standards and best common
practices.
Results: Qmin Imperfections. Already De Vries et al. have shown that qmin-
enabled resolvers send queries with three or more labels to the authoritative
servers of the .nl TLD occasionally. The observations at .nl and .se, as shown
in Fig. 8 and Fig. 9, confirm this finding. When neglecting queries to domain
names and records for which .nl are authoritative (e.g., the domain names of
the .nl name servers [ns1-ns4].[Link]), we find that 77% of qmin-enabled
resolvers send queries with more than two labels occasionally. In 55.1% of these
9
[Link]
10
By setting the DO-flag in the query.
11
As recommended by the 2020 DNS Flag day: [Link]
514 J. Magnusson et al.
cases the queries result in NXDOMAIN responses, signalling that the queried domain
name does not exist. We could not find when exactly resolvers would fall back
to sending the full query name, but this shows that even with qmin enabled,
information about lower labels can leak. We could also observe this behaviour at
resolvers of Google Public DNS and we reached out to their operators for clarifi-
cation. Unfortunately, they could not explain to us what causes these occasional
queries for fully-qualified domain names.
5 Controlled Experiments
The purpose of the controlled experiments is to look at the most recent versions
of popular open source resolvers and look at how they handle minimized queries
in regards to performance. In the controlled experiments by De Vries et al., four
open source resolvers were considered due to their popularity: Bind, Unbound,
Knot Resolver, and PowerDNS. Only the first three resolvers had implemented
qmin in their most recent version at the time, which meant that PowerDNS was
excluded. PowerDNS has since then implemented and enabled relaxed qmin by
default in version 4.3.0 and is therefore included in this study.12 The versions of
each resolver for the controlled experiments in this study were: Unbound 1.14.0,
Bind 9.16.24, Knot Resolver 5.4.4 and PowerDNS 4.6.0. DNSSEC was turned
off and the resolvers were configured to have the same size of caches.
5.1 Method
Just as in the original study, we use the Cisco Umbrella Top 1M list [5]
for domains to query in the performance and error-rate measurements. The
list contains the most popular queries based on passive DNS usage on their
Umbrella global network. This list does not only contain browser-based HTTP
user requests, but takes other protocols and non-end-users into account. Like in
the study by De Vries et al., domain names were aggregated from a timespan of
two weeks (April 4th until April 19th, 2022) to avoid daily and weekly fluctu-
ations and patterns. This resulted in 1.3M domain names with a mean of 3.26
labels, a median of 3 labels, a min of 1 and a max of 104 labels. This list of
domains was sorted in four different orders to even out caching effects.
As mentioned in Sect. 2, there are two modes when enabling qmin on resolvers
(i.e., relaxed and strict). These modes dictate whether the resolver should fall
back to full query names when receiving NXDOMAIN or other unexpected responses
from potentially broken name servers. For the controlled experiments in this
study, both Unbound and Bind had the option to turn qmin off, run it in relaxed
mode, or run it in strict mode. PowerDNS could either turn qmin off or turn it
on in relaxed mode. Knot did not have any option to turn off qmin and only
runs in relaxed mode. We used all the possible configurations in the experiments
and compared the results to the results from the controlled experiments by De
Vries et al..
12
[Link]
A Second Look at DNS QNAME Minimization 515
5.2 Results
When looking closer at the traffic for Unbound and Bind we discovered that
Unbound always resolves all name servers received and Bind is establishing TCP
connections to the root servers. The error-rates for Bind have also decreased
and there is no difference between strict and relaxed mode, which is unexpected.
We have not been able to identify why relaxed mode showed no advantage.
Knot is running relaxed qmin without any option to turn it off. The number of
packets and the error-rate have decreased significantly in version 5.4.4 compared
to version 3.0.0. The large decrease in number of packets is the opposite trend
of Unbound and Bind. For PowerDNS the number of packets were similar when
comparing relaxed qmin and no qmin, which was unexpected since qmin should
produce more packets. Since PowerDNS enables qmin with the relaxed mode,
the similar error-rates seen with and without qmin enabled was expected. This
is true for Unbound and Bind as well. As the error-rates of all resolvers have
516 J. Magnusson et al.
decreased since 2018, regardless of qmin and mode, this could be a change on
the name server side and not necessarily on the resolvers themselves.
The controlled experiments showed that the error-rate decreased for all
resolvers compared to the original study, but the number of packets and the error-
rate varied depending on the specific resolver and mode used. The qmin feature is
tightly correlated with the number of packets, as a fully complete domain name
typically requires fewer queries to resolve than a minimally built query that is
iteratively resolved. Our results suggest that the performance of recursive DNS
resolvers with qmin enabled has improved since the previous study, but further
investigation is needed to fully understand the effects on each resolver. The num-
ber of packets can be an important factor in evaluating the performance of a
resolver, as it can indicate the resources and communication required to process
queries and retrieve responses. While a lower number of packets may indicate
efficiency, other metrics such as latency and response accuracy should also be
considered when assessing the performance of a resolver.
6 Discussion
In this section we summarize the results from our measurements and analyze the
general adoption of qmin. Then we look at the improvements of the measurement
methods as well as discuss the balance between performance and privacy.
Table 9. Results of RIPE Atlas probe resolvers, open resolvers, K-Root and .nl ccTLD
For the passive measurements at the root and the TLD we filtered out the
invalid labels that affected the results of De Vries et al.. Here we also see a pos-
itive trend of minimized queries which matches the relative growth in the RIPE
Atlas active measurements, a sixfold increase. The .nl TLD passive measure-
ments also show an increase of qmin, but already in 2019 the share of incoming
minimized queries was high. The resolvers querying domains at .nl are likely
less representative of all DNS resolvers on the Internet, and instead point to the
early adoption of privacy features in dutch DNS infrastructure. The results from
the active and passive measurements show a clear and consistent increase in
the adoption of qmin-enabled resolvers when comparing to the previous study.
While the actual level of qmin adoption varies between the measurements, as
they capture the behavior of different sets of resolvers from different vantage
points, this is a positive development for Internet privacy.
open resolvers using a separate domain in Sect. 3.2. This was unexpected since
the resolvers have been minimizing queries since 2020 according to DNSThought.
Additional queries using different domains showed that the Google Public DNS
resolvers were consistently responding differently based on domain. This was
because Google wanted to get credit for minimizing queries at the root and
TLD level, which originally did not show on DNSThought statistics.
Even though a resolver is not implementing qmin beyond the 2LD, a lack of
data minimization within an authoritative DNS zone is less serious compared to
fully disclosed query names at the root and TLD level. An organization regis-
tering a 2LD is most likely aware of their subdomains, so no harm would come
from exposing those labels to their own name servers. Some organizations reg-
ister domains under e.g., .[Link] or .[Link] and it is therefore not as simple as
to only minimize until the 2LD. We propose setting the depth limit using the
Public Suffix List (PSL) [15] with one additional label (PSL+1). The PSL is a
list maintained by Mozilla mainly used for restricting cookie setting. It contains
effective TLDs (e.g., .com and .net) including those with more than one label
(e.g., .[Link] and [Link].) A qmin-enabled resolver using the PSL+1 approach
and looking up a RR for [Link] should send uk to the root, [Link]
to the .uk ccTLD and then [Link] to the name server of .[Link]. The
resolver would then stop minimizing and send [Link] to the name
server of [Link] which is most likely the authoritative DNS zone. Since
[Link] is in the PSL, we refer to [Link] as PSL+1.
7 Conclusion
our work has helped improve the probing for qmin adoption. With some commu-
nication with NLnet Labs and RIPE Atlas, a bug was fixed where new probes
used by DNSThought were not querying for qmin. When looking closer at the
Google Public DNS we observed that client-side active measurements using these
resolvers seem to be highly dependent on specific domain names. With the help
of NLnet Labs we found out that Google’s resolvers have a qmin depth limita-
tion, except for the domain used by DNSThought. This exception was done in
order to get credit for minimizing at the root and TLD levels. In the controlled
experiments using four open source resolvers we observed that the error-rates
are decreasing. This is likely due to RFC 9156 which switched RR query types,
specified fallbacks on certain errors, and added labels more dynamically and thus
obsoleted the previous implementation of qmin in RFC 7816. But it could also
be a change on the name server side. The number of packets are going up for
two of the resolvers while it is decreasing or rather low for the other two, and it
is still unclear why. We discussed the adequate level of minimizing query names
in regards to both performance and privacy, where we argue that the privacy
risks of leaking sensitive subdomains decrease after the authoritative DNS zone.
We therefore look at the Public Suffix List as a possible resource for configuring
the minimization depth limit.
Research Artifacts
To enable a third look at qmin in the future we provide whatever scripts used
([Link] beyond what was already available by
De Vries et al. [24]. The RIPE Atlas measurements by NLnet Labs are also
accessible at RIPE [21].
Ethical Considerations
In this work we thought carefully about the ethical aspects of our measurements
and disclosure. We used a list of open resolvers from third-party scans instead of
doing the scan of the IPv4 address space on our own, thus avoiding adding more
unnecessary load on the networks. We also spread out our active measurements
in a round-robin style to not put too much load on single resolvers in a short
span of time.
References
1. Arends, R., Austein, R., Larson, M., Massey, D., Rose, S.: DNS security intro-
duction and requirements. RFC 4033, RFC Editor, March 2005. [Link]
[Link]/rfc/[Link]
2. Bind: Bind documentation: options. [Link]
[Link]#options-statement-definition-and-usage. Accessed June 2022
3. Bortzmeyer, S.: DNS query name minimisation to improve privacy. RFC 7816,
RFC Editorm March 2016
4. Bortzmeyer, S., Dolmans, R., Hoffman, P.: DNS query name minimisation to
improve privacy. RFC 9156, RFC Editor, November 2021
5. Cisco: Cisco umbrella top 1m list. [Link]
static/[Link]. Accessed 12–25 Feb 2022
6. Cooper, A., et al.: Privacy considerations for internet protocols. RFC 6973, RFC
Editor, July 2013
7. [Link]: Measuring qname minimisation support. [Link]
reports/adam/qname-minimisation-en/. Accessed Nov 2021
8. Google: Issue 1090985: Disable Intranet Redirect Detector by default. https://
[Link]/p/chromium/issues/detail?id=1090985. Accessed June 2020
9. Hoffman, P., McManus, P.: DNS queries over HTTPS (DoH). RFC 8484, RFC
Editor, October 2018
10. Hu, Z., Zhu, L., Heidemann, J., Mankin, A., Wessels, D., Hoffman, P.: Specification
for DNS over transport layer security (tls). RFC 7858, RFC Editor, May 2016
11. Huitema, C., Dickinson, S., Mankin, A.: DNS over Dedicated QUIC Connec-
tions. RFC 9250m May 2022. [Link] [Link]
[Link]/info/rfc9250
12. ICANN: M3: DNS root traffic analysis. [Link]
html. Accessed Mar 2022
13. Mockapetris, P.: Domain names - concepts and facilities. STD 13, RFC Editor,
November 1987. [Link]
14. Mockapetris, P.: Domain names - implementation and specification. STD 13, RFC
Editor, November 1987. [Link]
15. Mozilla Foundation: Public suffix list. [Link] Accessed 5 June
2008
16. NLnet Labs: DNSThought. [Link]
Accessed 14 Oct 2018
17. NLnet Labs: Unbound documentation: qmin strict. [Link]
[Link]/en/latest/manpages/[Link]?highlight=relaxed
%20qname#term-qname-minimisation-strict-yes-or-no. Accessed May 2021
18. Postel, J.: Internet protocol. STD 5, RFC Editor, September 1981. [Link]
[Link]/rfc/[Link]
19. Randall, A., et al.: Trufflehunter: cache snooping rare domains at large public DNS
resolvers. In: Proceedings of the ACM Internet Measurement Conference, pp. 50–64
(2020)
20. Rapid7 Labs: UDP scans. [Link] Accessed Jan
2022
21. RIPE Atlas: RIPE Atlas measurement. [Link]
8310250/. Accessed 20 Apr 2017
22. Verisign: Chromium’s Reduction of Root DNS Traffic. [Link]
domain-names/chromiums-reduction-of-root-dns-traffic/. Accessed Jan 2021
A Second Look at DNS QNAME Minimization 521
23. de Vries, W.B., Scheitle, Q., Müller, M., Toorop, W., Dolmans, R., van Rijswijk-
Deij, R.: A first look at QNAME minimization in the domain name system. In:
Choffnes, D., Barcellos, M. (eds.) PAM 2019. LNCS, vol. 11419, pp. 147–160.
Springer, Cham (2019). [Link]
24. de Vries, W.B., Scheitle, Q., Müller, M., Toorop, W., Dolmans, R., van Rijswijk-
Deij, R.: A first look at qname minimization in the DNS, datasets. https://
[Link]/wiki/[Link]/Traces#A_First_Look_at_QNAME_
Minimization_in_the_Domain_Name_System. Accessed Oct 2022
Open Access This chapter is licensed under the terms of the Creative Commons
Attribution 4.0 International License ([Link]
which permits use, sharing, adaptation, distribution and reproduction in any medium
or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons license and indicate if changes were
made.
The images or other third party material in this chapter are included in the
chapter’s Creative Commons license, unless indicated otherwise in a credit line to the
material. If material is not included in the chapter’s Creative Commons license and
your intended use is not permitted by statutory regulation or exceeds the permitted
use, you will need to obtain permission directly from the copyright holder.
DNS
How Ready is DNS for an IPv6-Only
World?
Abstract. DNS is one of the core building blocks of the Internet. In this
paper, we investigate DNS resolution in a strict IPv6-only scenario and
find that a substantial fraction of zones cannot be resolved. We point out,
that the presence of an AAAA resource record for a zone’s nameserver does
not necessarily imply that it is resolvable in an IPv6-only environment
since the full DNS delegation chain must resolve via IPv6 as well. Hence,
in an IPv6-only setting zones may experience an effect similar to what
is commonly referred to as lame delegation.
Our longitudinal study shows that the continuing centralization of the
Internet has a large impact on IPv6 readiness, i.e., a small number of
large DNS providers has, and still can, influence IPv6 readiness for a large
number of zones. A single operator that enabled IPv6 DNS resolution–by
adding IPv6 glue records–was responsible for around 20.3% of all zones
in our dataset not resolving over IPv6 until January 2017. Even today,
10% of DNS operators are responsible for more than 97.5% of all zones
that do not resolve using IPv6.
1 Introduction
With the recent exhaustion of the IPv4 address space, the question of IPv6
adoption is gaining importance. More end-users are getting IPv6 prefixes from
their ISPs, more websites are reachable via IPv6, hosting companies start billing
for IPv4 connectivity or give discounts for IPv6-only hosting and IoT devices
further push IPv6 deployment. Yet, one of the main entry-points for Internet
services—the DNS—is suffering from a lack of pervasive IPv6 readiness. While
protocols such as Happy Eyeballs [41,45] help to hide IPv6 problems, they com-
plicate detection and debugging of IPv6 issues. Indeed, the threat of DNS name
space fragmentation due to insufficient IPv6 support was already predicted in
RFC3901, over 18 years ago [18]. Hence, in this paper, we measure the current
state of IPv6 resolvability in an IPv6-only scenario.
glue records, in Jan. 2017 one single provider fixed the IPv6-only name reso-
lution of more than 45.6 M domains (20.3% of the domains in the dataset).
– Resilience mechanisms often hide misconfigurations. For example, broken
IPv6-delegation is hidden by the combined efforts of DNS resilience and
Happy Eyeballs. Correctly configuring ones own DNS zone is not sufficient
and dependencies are often non-obvious.
– Additionally, we conduct a thorough validation of our methodology. We assess
the coverage of the Farsight SIE data in comparison to available ground-
truth zonefile data, finding it to provide sufficient coverage for our analysis.
Furthermore, we cross-validate our passive measurement results using active
measurements, again finding our results to be robust.
– We implemented a DNS measurement tool instead of using, e.g., ZDNS [29], as
we need IPv6 support which ZDNS does not (yet) support. The dataset from
our active measurements and an implementation of our scanning method-
ology, including a single-domain version operators can use to evaluate IPv6
support for their own domains, are publicly available at:
[Link]
in an IPv6-only scenario. Other issues where a zone does not resolve due to,
e.g., DNSSEC problems or unresponsive nameservers, i.e., the strict definition
of “lame delegation” (see RFC8499 [26]) are out-of-scope. The issues we discuss
can also occur in IPv4 DNS resolution, but are usually quickly discovered given
the currently still large number of sites with IPv4-only connection to the Internet,
that will not be able to resolve the affected zones.
For a zone to be IPv6-resolvable —i.e., resolvable using IPv6-only— the zones
of the authoritative nameservers have to be resolvable via IPv6 and at least one
nameserver must be accessible via IPv6. This has to be the case recursively,
i.e., not only for all parents of the zone itself but also for all parents of the
authoritative nameservers in such a way that at least for one1 of the authoritative
nameservers of a zone a delegation chain from the root zone exists, that is fully
resolvable using IPv6. We identify the following misconfigurations which can
cause broken IPv6-delegation in an IPv6-only setting:
– No AAAA records for NS names: If none of the NS records for a zone in their
parent zone have associated AAAA records, resolution via IPv6 is not possible.
– Missing GLUE: If the name from an NS record for a zone is in-bailiwick,
i.e., the name is within the zone or below [26], a parent zone must contain an
IPv6 GLUE record, i.e., a parent must serve the corresponding AAAA record(s)
as ADDITIONAL data when returning the NS record in the ANSWER section.
– No AAAA record for in-bailiwick NS: If an NS record of a zone points to a
name that is in-bailiwick but the name lacks AAAA record(s) in its zone, IPv6-
only resolution will fail even if the parent provides GLUE, when the recursive
server validates the delegation path. One such example is Unbound [35] with
the setting harden-glue: yes–the default.
– Zone of out-of-bailiwick NSes not resolving: If an NS record of a zone is
out-of-bailiwick, the corresponding zone must be IPv6-resolvable as well. It
is insufficient if the name pointed to by the NS record has an associated AAAA
record.
– Parent zone not IPv6-resolvable: For a zone to be resolvable via IPv6
the parent zones up to the root zone must be IPv6-resolvable. Any non-IPv6-
resolvable zone breaks the delegation chain for all its children.
The above misconfigurations are not mutually exclusive. For example, if the
NS sets between parent and child differ, a common misconfiguration [42], the NS
in the parent may not resolve due to missing GLUE (as they are in-bailiwick)
but also the NS in the child may not resolve due to having no AAAA for their
names, if they are out-of-bailiwick. In this paper we investigate the prevalence
of these misconfigurations to evaluate the IPv6 readiness of the DNS ecosystem.
Fig. 2. Zone coverage of Farsight data and number of zones used for the evaluation.
We used available zone files to determine the share of covered second level domains
by Farsight’s dataset. Please note the dip in the graph from February to August 2019,
where our zone file collection was limited, i.e., we only collected few zones with high
coverage (February - April and July, including .com), or no data at all (May and June).
dataset with ground-truth data, i.e., the names extracted from available zone
files. Specifically, we are comparing to .com, .net, and other gTLD (generic TLD)
zone files starting from mid of 2016. Additionally, from April 2017 onward, we
also obtained CZDS (ICANN Centralized Zone Data Service) zone file data for
all available TLDs. Moreover, we use publicly available zone file data from .se,
.nu, and .ch for the coverage analysis. In total, this allows us to compare Far-
sight’s data to more than 1.1k zones as of August 2022.
Looking at coverage over time, we find a significant overlap between the Far-
sight dataset and the number of actually delegated zones based on zone files, see
Fig. 2. Coverage averages above 95% from 2019 onwards, with especially since
May 2021, our coverage reaches over 99%. Furthermore, we find a reduced aver-
age coverage in the beginning of 2017. A closer investigation revealed that these
relate to the introduction of various vanity gTLDs with an overall small size,
i.e., below 100 delegated zones in the TLD. This implies that missing coverage
for just a few zones would lead to a significant reduction in aggregate coverage.
Nevertheless, our analysis shows that a significant share of zones is covered in the
Farsight dataset. Hence, we the Farsight dataset–especially due to the historic
perspective it provides–is ideal to investigate our research questions.
Despite this high coverage, we still face the drawback of the Farsight dataset
relying on real-world usage. As such, a missing record in the passive dataset does
not necessarily indicate non-existence. Hence, we independently corroborate all
major findings with data from TLD zone files for a specific period to check for
missing glue records in the zone file, see Sect. 5.4.
There are many ways to cluster DNS domains into subgroups. For example, one
may look only at the Top Level Domains as specified by ICANN [28], or use
the Public Suffix List (PSL) provided by the Mozilla Foundation [34] to identify
second level domains. The PSL is used by browser vendors to decide if a domain
How Ready is DNS for an IPv6-Only World? 531
is under private or public control, e.g., to prevent websites from setting a super-
cookie for a domain such as .[Link]. Based on matching monthly samples of the
ICANN TLDs and the PSLs we identify TLDs as well as 2 nd Level Domains,
and Zones Below 2 nd Level, i.e., all zones below 2nd Level Domains.
Another way of grouping DNS domains is to use the Alexa Top-1M list [3].
Using, again, matching monthly samples, we distinguish between the Top 1K,
Top 1K–10K, Top 10K–100K, and Top 100K–1M domains. We note that there
are limitations in the Alexa Top List [39,40], but compared to other toplists such
as Tranco [31], the Alexa list is available throughout the measurement period.
Here, we describe how we identify whether zones can be resolved only via IPv4,
only via IPv6, via IPv4 and IPv6, or not at all from the dataset.
1. Per Zone NS set Identification: We first identify all zone delegations
by extracting all entries with rrtype = NS. Next, for all names used in these
delegations, we find all associated IPs by extracting all A and AAAA records. We
do not consider CNAMEs since they are invalid for NS entries, see RFC2181 [20].
We then iterate over all zones, i.e., names that have NS records, to create
a unique zone list. In this process, we record the NS records for each bailiwick
sending responses for this zone observed in the dataset, and for each NS name
all AAAA and A type responses, again grouped by bailiwick from which they were
seen. This also captures cases where parent and child return different NS sets.
2. Per Zone DNS Resolution: We consider a zone to be resolvable via IPv4
or IPv6 if at least one of the NS listed for the zone can be resolved via IPv4 or
IPv6 respectively. Hence, to check which zones can be resolved using which IP
protocol version we simulate the DNS resolution, starting at the root, i.e., we
assume the Root zone . to be resolvable by IPv4 and IPv6. We then iterate over
the zone set with attached NS and A/AAAA data. For each zone, except the root
zone, we initialize an empty state marking the zone as not resolving.
We then attempt to resolve each zone. For that, we first check if the zone’s
parent has been seen.
If so we check for each NS of the zone we are trying to resolve as listed in the
parent whether its name resolves via IPv4 and/or IPv6. This is the case if:
1. The NS is outside the zone we are trying to resolve, the NS’ zone has been
recorded as resolving in the zone state file (via IPv4 and/or IPv6), and there
are A/AAAA records with that zone’s bailiwick for the NS.
2. The NS is in the zone we are resolving and there is an A/AAAA glue record for
the name with the bailiwick of the zone’s parent (only if an in-bailiwick NS is
listed in the parent).
532 F. Streibelt et al.
To ensure full resolution, we also have to check that the NS listed in the child
resolve. For NS with names under the zone this is the case if the NS listed for
this zone in the parent can be reached via IPv4/IPv6, see above, and they have
A/AAAA records with the bailiwick of the zone itself. For out-of-bailiwick NS, this
is again the case if their own zone resolves and they have A/AAAA records.
A single iteration of this process is not sufficient, as zones often rely on out-
of-bailiwick NS. Hence, we continue iterating through the list of zones until the
number of unresolved zones no longer decreases. For a simplified pseudo-code
description, see Algorithm 1.
the DNS tree. From there, we query all authoritative nameservers recorded in
the parent on each layer of the DNS hierarchy using IPv4 and IPv6 where pos-
sible for the NS of that zonelayer. Furthermore, we try to obtain any possibly
available GLUE (A and AAAA) for in-bailiwick NS. For out-of-bailiwick NS, we
try to resolve the NS, again starting from the root. If there is an inconsistency
between parent and child, i.e., if we discover additional NS when querying the
NS listed in the parent, we also perform all queries for this layer against these,
noting that they were only present in the child.
To limit the amount of queries sent to each server, our implementation follows
the underlying principles of QNAME minimization as described in RFC7816 [5].
By using the NS resource record type to query the parent zones we can directly
infer zonecuts and store GLUE records from the additional section, if present.
Note that RFC8020 [6] is still not implemented by all nameservers, thus we can-
not rely on NXDOMAIN answers to infer that no further zones exist below the
queried zone. Our measurement tool will retry queries using TCP on truncation
and disable EDNS when it receives a FORMERR from the upstream server.
To further limit the number of queries sent, all responses, including error
responses or timeouts, are cached. We limit the number of retries (4) as well as
the rate (20 s wait time) at which they are sent. To further enrich the actively
collected dataset, we query all authoritative nameservers of a zone for the NS,
TXT, SOA and MX records of the given zone as well as the version of the used
server software using the [Link] in the CHAOS class. Queries and replies
are recorded tied to the NS that provided them.
We ran these measurements between October 10th to 14th and 22nd to 24th
2022 against the Alexa Top1M from August 15th 2022 containing 476,242 zones,
collecting responses to a total of 32M queries sent via IPv4 and 24M queries
sent via IPv6. Our active measurement dataset (101GB of json data), and a
tool implementing our measurement toolchain are publicly available at: https://
[Link]/mutax/dns-v6-readyness.
the IP as well as on the rDNS name. We also focused our measurements on the
Alexa Top 1M, i.e., sites for which the impact of additional requests at the scale
of our measurements is not significant, while also limiting repeated requests using
caching. During our active measurements, we did not receive any complaints. In
summary, we conclude that this work does not raise any ethical issues.
4 Results
Here, we first provide an aggregate overview of the Farsight dataset. Subse-
quently, we present the results of our analysis of broken IPv6-delegation based on
passive measurement data. Finally, we validate our passive measurement results
against active measurements run from 10th to 24th of October 2022.
Fig. 3. Per month: # of zones (gray line–right y-axis) and IPv4/IPv6 resolvability in
% (left y-axis).
536 F. Streibelt et al.
still receives oversight by NICs, e.g., regarding the RFC compliant use of at least
two NS in different networks [21], while zones below 2nd level domains can be
freely delegated by their domain owners. Also, for sub-domains, we observe three
distinct spikes in Fig. 3d which correspond to the spikes seen for all domains,
recall Fig. 3a. These spikes occur when a single subtree of the DNS spawns
millions of zones. These are artifacts due to specific configurations and highlight
that lower layer zones may not be representative for the overall state of DNS.
Finally, comparing PSL 2nd level domains, see Fig. 3d, to the Alexa Top-1K
domains, see Fig. 3e, we find that IPv6 adoption is significantly higher among
popular domains, starting from 38.9% in 2015 and rising to 80.6% in 2021.
There are two notable steps in this otherwise gradual increase, namely January
2017 and January 2018. These are due to a major webhoster and a major PaaS
provider enabling IPv6 resolution (2017), and a major search engine provider
common in the Alexa-Top-1K enabling IPv6 resolution (2018).
Comparison with Active Measurements: Evaluating zone resolvability from
our active measurements, see Sect. 3.4, we find that 314,994 zones (66.14%) sup-
port dual stack DNS resolution, while 159,166 zones (33.42%) are only resolvable
via IPv4.
A further 2066 zones (0.43%) could not be resolved during our active mea-
surements, and 16 zones (≤0.01%) were only resolvable via IPv6. In comparison
to that, our passive measurements–see also Fig. 3f–map closely: We find 66.18%
(+0.04% difference) of zones in the Alexa Top 1M resolving via both, IPv4 and
IPv6, and 32.23% (−1.19% difference) of zones only resolving via IPv4. Similarly,
1.16% (+0.73%) of zones do not resolve at all, and 0.42% (+0.42% difference)
of zones only resolve via IPv6 according to our passive data. Hence, overall, we
find our passive approach being closely aligned with the results of our active
measurements for the latest available samples. The, in comparison, higher val-
ues for non-resolving and IPv6 only resolving zones are most likely rooted in the
visibility limitations of the dataset, see Sect. 5.4. Nevertheless, based on the low
deviation between two independent approaches at determining IPv6 resolvability
of zones we have confidence in the results of our passive measurements.
Next, we take a closer look at zones that show some indication of IPv6 deploy-
ment, yet, are not IPv6-resolvable. These are zones where an NS has an AAAA
record or an AAAA GLUE. To find them we consider NS entries within the zone
as well as NSes for the zone in its parent. In Fig. 4 we show how their abso-
lute numbers evolve over time (gray line) as well as the failures reasons (in
percentages).
How Ready is DNS for an IPv6-Only World? 537
Fig. 4. Per month: # of zones not IPv6-resolvable with AAAA or GLUE for NS (gray
line–right y-axis) and causes for IPv6 resolution failure in % (left y-axis).
We find that for all four subsets of zones shown—all zones, ICANN TLDs,
Alexa Top-1K, Alexa Top-10K–100K—the most common failure case is missing
resolution of NS in the parent. This occurs mostly when the NS is out-of-bailiwick
and does have AAAA records, but the NS’s zone itself is not IPv6-resolvable. Fur-
thermore, there is a substantial number of zones per category—especially in the
538 F. Streibelt et al.
Alexa Top-1K—where the NS in the parent lacks AAAA while the NS listed in the
zone has AAAA records, commonly due to missing GLUE. We also observe the
inverse scenario, i.e., GLUE is present but no AAAA record exist for the NS within
the zone itself. Both cases can also occur if NS sets differ between the parent
and its child [42].
We see a major change around January 2017, i.e., a sharp increase in zones
that are IPv6-resolvable, which is also visible in Fig. 3: For all zones as well as for
the Alexa Top 10K–100K, we observe that several million zones not resolving via
IPv6 since the start of the dataset but having NSes with AAAA records, now are
IPv6-resolvable. The reason is that a major provider added missing glue records.
Interestingly, we do not see this in the Alexa Top 1K.
In the Alexa Top 1K, and to a lesser degree in the Alexa Top 10K-100K,
we observe a spike of zones that list AAAA records for their NS but are not
IPv6-resolvable in Oct. 2016. This is the PaaS provider mentioned before, first
rolling out AAAA records for their NS, and then three months later also adding
IPv6 GLUE. Operationally, this approach makes sense, as they can first test the
impact of handling IPv6 DNS queries in general. Moreover, reverting changes
in their own zones is easier than reverting changes in the TLD zones–here the
GLUE entries. Again, the major webhoster is less common among the very pop-
ular domains, which is why its effect can be seen in Figs. 4a and 4d, but not
in Fig. 4b. Also, this operator had AAAA records in place since the beginning
of our dataset, as seen by the plateau in Fig. 4d. These observations have been
cross-confirmed by inspecting copies of zonefiles for the corresponding TLDs and
time-periods.
Fig. 5. Per month: # of zones not IPv6-resolvable (gray line–right y-axis) and distri-
bution of zones over NS sets in % (left y-axis).
needed glue, i.e., they were in-bailiwik NS for their own. Among these, 19,310
NS were dual-stack, while 94,192 only had A records associated with them, and a
further 108 NS only had associated AAAA records. Furthermore, 85,213 (90.47%)
of A-only NS needing glue had correct glue set. For dual-stack configured NS,
14,072 (72.87%) have complete (A and AAAA) glue. A further 3,932 (20.36%) NS
only has A glue records, while 24 (0.12%) NS only have AAAA glue, despite gener-
ally having a dual-stack DNS configuration. Finally, of the 108 NS records only
having AAAA records associated, 70 (64.81%) NS have correctly set AAAA glue.
Moving on to the reachability of these NS, we find that of the total number of
NS that have an A record (169,547) and are reachable is at 164,255, i.e., 96.88%
actually responds to queries. For IPv6, these values are slightly worse, with
30,193 of 32,285 NS (93.52%) responding to queries via IPv6. This highlights a
potential accuracy gap of 3–6% for research work estimating DNS resolvability
from passive data. Notably, this gap is larger for IPv6.
5 Discussion
In this section, we first state our key-findings, and then discuss their implications.
Centralization is one of the big changes in the Internet over the last decade.
This trend ranges from topology flattening [4,7] to the majority of content being
served by hypergiants [8] and—as we show—also applies to the DNS. An increas-
ing number of zones are operated by a decreasing number of organizations. As
such, an outage at one big DNS provider [44]—or missing support for IPv6—can
disrupt name resolution for a very large part of the Internet as we highlight in
Sect. 4. In fact, out-of-bailiwick NS not being resolvable via IPv6 is the most com-
mon misconfiguration in our study, often triggered by missing GLUE in a single
zone. Given that ten operators could enable IPv6 DNS resolution for 24.8% of
not yet IPv6 resolving zones, we claim that large DNS providers have a huge
responsibility for making the Internet IPv6 ready.
5.4 Limitations
Since our dataset relies on DNS cache misses, we are missing domains that are not
requested or not captured by the Farsight monitors in a given month. Moreover,
our use of monthly aggregates may occlude short-term misconfigurations. To
address this, we support major findings on misconfigurations with additional
ground-truth data from authoritative TLD zone files.
Similarly, we use the Alexa List with its known limitations [39,40]. Thus, we
cluster the Alexa list into different rank tiers, which reduces fluctuations in the
higher tiers. Furthermore, we only assess zones’ configuration states, and not
actual resolution, i.e., “lame delegation” for other reasons is out of scope.
Furthermore, we cannot make statements on whether the zones we measure
actually resolve, e.g., if there is an authoritative DNS server listening on a con-
figured IP address returning correct results. Still, we have certainty that zones
we measure as resolvable are at least sufficiently configured for resolution. Sim-
ilarly, we can not assess the impact of observed DNS issues on other protocols,
e.g., HTTPs. To further address this limitation of our passive data source, we
conducted active measurements, which validated the observations from our pas-
sive results and added further insights on the actual reachability of authoritative
DNS servers for zones.
Naturally, our active measurements also have several limitations that have to
be recorded. First, we conducted our measurements from a single vantage point.
Given load balancing in CDNs via DNS [43], this may have lead to a vantage
point specific perspective. Nevertheless, we argue that misconfigurations [14]
are likely to be consistent across an operator, i.e., the returned A or AAAA
records may change, but not the issue of, e.g., missing GLUE. Furthermore,
DNS infrastructure tends to be less dynamic than A and AAAA records.
Second, our measurements were only limited to the Alexa Top 1M and asso-
ciated domains. We consciously made this choice instead of, e.g., running active
measurements on all zones in the Farsight dataset to reduce our impact on the
Internet ecosystem.
In summary, our study provides an important first perspective on IPv6 only
resolvability. We suggest to complement our study with active measurements of
IPv6 only DNS resolution and the impact of broken IPv6-delegation on the IPv6
readiness of the web due to asset dependencies as future work.
542 F. Streibelt et al.
6 Related Work
Our related work broadly clusters into two segments: i) Studies on IPv6 adoption
and readiness, and ii) Studies about DNS and DNS misconfigurations.
6.3 Summary
We expand on earlier contributions regarding IPv6 adoption. We provide a more
recent perspective on the IPv6 DNS ecosystem and take a more complete app-
roach to asses the IPv6 readiness in an IPv6-only scenario. This focus on IPv6
is also our novelty in context to earlier work on DNS measurements and DNS
misconfigurations, which did not focus on how IPv6 affects DNS resolvability.
Additionally, our active measurements for validating our passive measurement
results also highlight that the presence of AAAA records does not necessarily imply
IPv6 resolvability. Instead, to measure IPv6 resolvability, the resolution state of
provided IPv6 resources has to be validated.
How Ready is DNS for an IPv6-Only World? 543
7 Conclusion
In this paper, we present a passive DNS measurement study on root causes
for broken IPv6-delegation in an IPv6 only setting. While over time we see an
increasing number of zones resolvable via IPv4 and IPv6, in August 2022 still
44.9% are not resolvable via IPv6. We identify not resolvable NS records of the
zone or its parent as the most common failure scenario. Our recommendations to
operators include to explicitly monitor IPv6 across the entire delegation chain.
Additionally, we conducted a dedicated validation of our results using active
measurements. This validation broadly confirmed our results from the passive
measurements and further highlighted the importance of not only relying on the
presence of specific records, as nameservers for which IPv6 addresses are listed
in the DNS may not actually be responsive.
We plan to provide an open-source implementation of our measurement
methodology along with the paper. Furthermore, we will provide a reduced
implementation of our measurement toolchain which will enable operators to
explicitly check a given zone or FQDN for IPv6-resolvable. Similarly, we will
provide the results of our active measurements as open data.
For future work we suggest to systematically expand our active measure-
ment campaign to assess resolvability, e.g., for websites including all web assets.
Using active measurements, one can explicitly resolve a hostname and run active
checks on the delegation chain, validating the responses of all authoritative name-
servers and find inconsistencies not only between a zone and its parent but also
within the NS set. We conjecture that–especially given the widespread use of
subdomains for web assets–the reduced IPv6 resolvability we observe may have
a significant impact on the IPv6-readiness of the web, i.e., a website using assets
on domains that do not resolve via IPv6 is not IPv6 ready.
Fig. 6. Total number of zones in the dataset per month (gray line) and resolvability
How Ready is DNS for an IPv6-Only World? 545
Fig. 7. Zones unable to resolve using IPv6, but with AAAA records in GLUE or zone
apex (gray line), by resolution failure.
546 F. Streibelt et al.
References
1. Akiwate, G., et al.: Unresolved issues: prevalence, persistence, and perils of lame
delegations. In: Proceedings of the Internet Measurement Conference (IMC), pp.
281–294. ACM (2020). [Link]
2. Allman, M., Paxson, V.: Issues and etiquette concerning use of shared measurement
data. In: Proceedings of the Internet Measurement Conference (IMC), pp. 135–140.
ACM (2007). [Link]
3. [Link] Inc: Alexa Top Sites. [Link]
4. Arnold, T., et al.: Cloud provider connectivity in the flat Internet. In: Proceed-
ings of the Internet Measurement Conference (IMC), pp. 230–246. ACM (2020).
[Link]
5. Bortzmeyer, S.: DNS query name minimisation to improve privacy. RFC 7816
(Experimental), March 2016. [Link] obso-
leted by RFC 9156
6. Bortzmeyer, S., Huque, S.: NXDOMAIN: there really is nothing underneath.
RFC 8020 (Proposed Standard), November 2016. [Link]
[Link]
7. Böttger, T., et al.: Shaping the internet: 10 years of IXP growth. arXiv (2019).
[Link] [Link]
8. Böttger, T., Cuadrado, F., Tyson, G., Castro, I., Uhlig, S.: A hypergiant’s view of
the internet. ACM Comput. Commun. Rev. (CCR) 47(1) (2017)
9. Calder, M., Fan, X., Hu, Z., Katz-Bassett, E., Heidemann, J., Govindan, R.: Map-
ping the expansion of Google’s serving infrastructure. In: Proceedings of the Inter-
net Measurement Conference (IMC), pp. 313–326. ACM (2013). [Link]
10.1145/2504730.2504754
10. Chhabra, R., Murley, P., Kumar, D., Bailey, M., Wang, G.: Measuring DNS-
over-HTTPS performance around the world. In: Proceedings of the Internet Mea-
surement Conference (IMC), pp. 351–365. ACM (2021). [Link]
3487552.3487849
11. Chung, T., et al.: Understanding the role of registrars in DNSSEC deployment. In:
Proceedings of the Internet Measurement Conference (IMC), pp. 369–383. ACM
(2017). [Link]
12. Colitti, L., Gunderson, S.H., Kline, E., Refice, T.: Evaluating IPv6 adoption in
the internet. In: Krishnamurthy, A., Plattner, B. (eds.) PAM 2010. LNCS, vol.
6032, pp. 141–150. Springer, Heidelberg (2010). [Link]
642-12334-4 15
13. Czyz, J., Allman, M., Zhang, J., Iekel-Johnson, S., Osterweil, E., Bailey, M.: Mea-
suring IPv6 adoption. In: Proceedings of the 2014 ACM SIGCOMM Conference
(SIGCOMM), pp. 87–98. ACM (2014). [Link]
14. Dietrich, C., Krombholz, K., Borgolte, K., Fiebig, T.: Investigating system opera-
tors’ perspective on security misconfigurations. In: Proceedings of the 25th ACM
SIGSAC Conference on Computer and Communications Security (CCS), pp. 1272–
1289. ACM (2018)
15. Doan, T.V., Fries, J., Bajpai, V.: Evaluating public DNS services in the wake of
increasing centralization of DNS. In: IFIP Networking Conference (2021). https://
[Link]/10.23919/IFIPNetworking52078.2021.9472831
16. Doan, T.V., Tsareva, I., Bajpai, V.: Measuring DNS over TLS from the edge:
adoption, reliability, and response times. In: Hohlfeld, O., Lutu, A., Levin, D.
(eds.) PAM 2021. LNCS, vol. 12671, pp. 192–209. Springer, Cham (2021). https://
[Link]/10.1007/978-3-030-72582-2 12
548 F. Streibelt et al.
Open Access This chapter is licensed under the terms of the Creative Commons
Attribution 4.0 International License ([Link]
which permits use, sharing, adaptation, distribution and reproduction in any medium
or format, as long as you give appropriate credit to the original author(s) and the
source, provide a link to the Creative Commons license and indicate if changes were
made.
The images or other third party material in this chapter are included in the
chapter’s Creative Commons license, unless indicated otherwise in a credit line to the
material. If material is not included in the chapter’s Creative Commons license and
your intended use is not permitted by statutory regulation or exceeds the permitted
use, you will need to obtain permission directly from the copyright holder.
TTL Violation of DNS Resolvers
in the Wild
1 Introduction
The Domain Name System (DNS) provides a scalable name resolution service.
It uses extensive caching to improve its resiliency and performance with a time-
to-live (TTL) value that specifies how long a DNS record can be cached before
being discarded [22]; the TTL value is assigned by the DNS authoritative servers.
DNS consumers (e.g., DNS resolvers) can cache the DNS responses during the
TTL so that the future requests can be fulfilled locally without sending extra
DNS queries to the DNS authoritative server.
Due to its resiliency and efficiency, DNS has evolved from simply providing
a mapping between human-readable names and network-level IP addresses, to
providing security features for other protocols (e.g., MTA-STS [19], TLSA [13], and
BIMI [4] for email protocols) or better performance by delegating its control
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 550–563, 2023.
[Link]
TTL Violation of DNS Resolvers in the Wild 551
to another entity (e.g., CDN). For example, an email server can publish its
certificate information as a DNS record (i.e., TLSA) so that a sender can cross-
check the certificate. Thus, the service operators have to manage the DNS records
and their security information in a synchronous way. When they update (i.e.,
rollover) their credential information such as public key, they usually publish the
updated DNS records in advance [13,19] and wait at least for the TTL (or twice
of TTL), expecting that DNS resolvers clear the old cache after then. However,
it is unclear how DNS clients follow such practice; for example, a DNS resolver
may cache DNS responses longer than its TTL to reduce DNS requests towards
authoritative servers. This may bring a negative impact on both security and
performance; for example, DNS resolvers that extends the TTL value may impair
the performance of CDNs, which typically uses a lower TTL value to improve
their resilliency and responsiveness [9].
However, it is challenging to understand how such TTL violations exist in
the wild without access to devices or users in affected networks; for example, it is
not straightforward to understand if a local DNS resolver in an ISP extends TTL
without deploying a vantage point in the ISP. To address this challenge, there
have been several successful prior approaches to measuring DNS TTL violations
by using datasets collected from DNS authoritative servers [15], residential net-
works [7], or using active probes such as RIPE Atlas [20]. While these approaches
have identified a number of resolvers that violates the TTL value in DNS records,
but it is typically difficult for others to replicate and often to scale [16,20], and
mostly focus on public DNS resolvers [16].
In this paper, we explore an alternative approach to detecting DNS TTL
violations of resolvers using a residential proxy service called, BrightData, which
allows us to achieve measurements from over 274,570 end hosts and their 27,131
resolvers across 9,514 ASes in 220 countries. We discovered the TTL violation
is prevalent; for example, we find 745 (8.74%) resolvers that extends TTLs.
Furthermore, we find that another form of TTL violation that can happen to
DNSSEC-signed records; we find that 285 DNSSEC-validating resolvers that
return expired DNSSEC-signed responses when the TTL does not expire yet.
We make our analysis code and data public to the research community at
[Link]
2.2 BrightData
In this work, we use BrightData, a residential proxy service, to character-
ize the behavior of resolvers. BrightData, formerly known as Luminati, is the
paid HTTP/S proxy service that routes traffic via residential nodes (called exit
nodes), who installed Hola Unblocker [14]. In order to route traffic, the client
needs to send a HTTP request to a BrightData server, called the super proxy;
the super proxy then forwards the request to an exit node. The exit node can
perform the HTTP request and return the response back to the client via the
super proxy.
BrightData offers options that can be passed with HTTP request to control
exit nodes. Figure 1 shows the overview of how the BrightData platform works.
Exitnode Persistence: The client can find the hash of the exit node’s IP address
in the HTTP response header, x-luminati-ip. By adding the -ip-XX option to
the HTTP request (where XX is the hash of the IP address), the client can use
the same exit node if available. This option is extremely useful to measure the
TTL violation behavior of the resolvers; we can still measure the same resolvers
TTL Violation of DNS Resolvers in the Wild 553
Fig. 1. Timeline of a request in BrightData: the client sends an HTTP request to the
super proxy ①; the super proxy makes a DNS request for the sanity check and forward
the request to an exit node ②–③; the exit node uses its DNS resolver and fetch the
HTTP response ④–⑥ and forward it back to the super proxy ⑦, which return it to the
client ⑧. Brightdata controls the Super Proxy and exit nodes (shown with blue boxes).
(Color figure online)
(used by the same exit node) by finding the same exit node after a TTL with a
longer period of time (e.g., 60 min) expires.
DNS Request Location: By default, DNS resolution is done and cached at the
super proxy’s end; however, the client can specify the dns-remote option to the
HTTP request to make DNS resolution done by the exit node (using the exit
node’s DNS server). In out experiment, we do so as we want the resolution to
be done at the client’s end.
With these options, we use BrightData to let exit nodes send HTTP requests
to our domains; the exit nodes will also send DNS requests to our DNS author-
itative server through their resolvers, which gives an opportunity to understand
their behavior. In the following section, we introduce our experiment methodol-
ogy and its challenges.
3.1 Methodology
At first glance, measuring and identifying resolvers that extend the TTL seems
straightforward: we pick one exit node and request it fetch the domain that
resolves IP1 . After its TTL expires, we update its A record to IP2 and let the
same exit node fetch the same domain to see if they connect to IP1 . However,
in practice it more difficult, because the client may use multiple DNS resolvers
that have many upstream resolvers, thus it may receive multiple DNS responses;
this behavior is common mainly to improve the performance; modern public
DNS servers usually have multiple caches with complex caching hierarchy [1,28].
However, we are not allowed to see which DNS response the exit node actually
554 P. Bhowmick et al.
Fig. 2. Our methodology that extends the BrightData platform in Fig. 1. We control
the DNS authoritative and web server (shown with green boxes); for the same qname,
our DNS authoritative server now returns a different A record to each different resolver
so that we can infer which DNS response the exit node used by monitoring the incoming
IP address of HTTP request. (Color figure online)
used, making it hard for us to identify who has extended the TTL. To address
these issues, we return a different A record to each different resolver so that we can
identify which DNS resolver’s response the exit node has used by monitoring the
incoming IP address of the HTTP request for a certain domain. More specifically,
we proceed our experiments as follows as illustrated in Fig. 2.
(a) As the first phase (P1 ), we first let an exit node fetch a unique subdo-
main, [Link] We extract the x-luminati-ip value from
the HTTP response header so that we can choose the same exit node after
the TTL expires.
(b) At our authoritative nameserver, for each resolver that looks up the same
qname, we pick an IP address that has never been used for the qname and
dynamically generate an A to serve the request. Then, we create an entry
that maps a tuple of qname and the resolver’s IP address to the served IP
address and insert it to the mapping table. If we observe more DNS resolvers
for the same qname than N , we discard the exit node from further analysis.
(c) From the webserver, we examine the destination of the IP address of the
HTTP request to find the matched DNS resolver’s IP address in the mapping
table. This allows us to find the DNS resolver that the exit node used.
(d) Then, we immediately retract all DNS entries from the DNS authoritative
name server to ignore all subsequent DNS requests, and we wait for T T L
to let the cached DNS responses expire.
(e) Once T T L expires, we set our authoritative name server to serve A that
points to the IP address (IPnew ) that has never been assigned to any DNS
resolver. We then use -ip-XX option in the HTTP request to choose the
same node and let it fetch [Link] again. We call this step
the second phase.
Ours
Honor. Ext.
Direct scan Honor. 197 0
Ext. 0 16
Proxy rack Honor. 381 1
Fig. 3. CDF of the fraction of the Ext. 0 62
exit nodes that use the the stale
response for each resolver.
.
3.2 Results
During our measurement period, we are able to send 2,068,686 unique HTTP
queries served by 274,570 unique exit nodes and their 27,131 resolvers1 in 9,514
ASes across 220 countries.
We also run our experiment with five different TTLs (1, 5, 15, 30, and 60 min)
to investigate how TTL values impact on a DNS resolver’s caching behavior.
Identifying potential TTL-extending resolvers is straightforward; when we
observe an exit node that still connects to the webserver with the old IP address,
we can find it by looking up the mapping table and label it as a potential TTL-
extending resolver. However, when we observe an exit node that uses the new
IP address (IPnew ), we cannot simply mark all resolvers in the second phase as
TTL-honoring resolvers because some resolvers only show up in the first phase.
Thus, we mark resolvers as the potential TTL honoring resolvers only when they
appear in both the first and second phase.
Since we use exit nodes as a proxy to understand DNS resolvers behavior, we
cannot blindly use the results to characterize the DNS resolvers; for example, a
1
Since we are only permitted to observe the only egress resolver IPs querying our
authoritative servers, we label each querying IP as a resolver.
556 P. Bhowmick et al.
stub resolvers on the exit node may extend the TTL making their resolvers look
like TTL-extending ones. Thus, we focus on the resolvers where we have at least
5 exit nodes, this sample size allows us to draw strong inferences to characterize
their behaviors; this leaves us 9,031 resolvers, and their 234,605 exit nodes across
the different TTL setups. Then, for each resolver, we calculate the fraction of
the exit nodes that connect to the IPold ; Fig. 3 shows the results and we make
a number of observations. First, we find that there is a clear separation between
TTL-extending resolvers and TTL-honoring resolvers. For example, when we
set our TTL values to 60 min, 4,147 (92.5%) resolvers perfectly honor the TTL
while 14 (0.31%) resolvers extend TTL; the others (7.1%) show mixed behav-
iors, which could be due to the stub or other frontend resolvers that we could
not measure. Second, we find that the number of TTL extending resolvers con-
stantly grows as we decrease the TTL value; for example, the percentage of TTL
extending resolvers increases from 14 (0.31%) to 129 (2.53%), 161 (2.87%), 414
(6.53%), and 745 (8.74%) as we decrease the TTL value from 60, 30, 15, 5 and
1 min. Surprisingly, we also find that the set of TTL extending resolvers that we
measured with T T Lx is always a subset of what we measure with T T Ly if T T Lx
is less than T T Ly . This strongly suggests that some resolvers use a default min-
imum TTL value; this could be due to reduce their resolution load; for example,
popular DNS software such as PowerDNS [26], KnotDNS [17] and Unbound [29]
has an option for this.
3.3 Cross-validation
We now attempt to cross validate our methodology by focusing on the resolvers
that always honor (or extend) the TTL with 1 min, which leaves us 7,160
resolvers. For each resolver, we first attempt to directly send DNS queries to
test whether they respond, and if so, if they extend the TTL by looking up our
domain twice with a time gap of TTL. Since this is likely to allow us to mea-
sure public resolvers, we also leverage another residential proxy service, Prox-
yRack [27] to cover local resolvers as well; the coverage is limited (less than
2+ million residential proxies), but it permits to send an arbitrary UDP traffic
so that we can send a DNS request to its local resolver. For each of the rest
of the resolvers, we attempt to find exit nodes that share the same AS with
the resolver and send DNS requests. Table 1 shows the result; surprisingly, we
find that all 212 resolvers that we could measure show the consistent behaviors
our observation. From the ProxyRack experiment, all 443 resolvers except one
show the consistent behavior; we found one resolver that our methodology and
a ProxyRack experiment disagree, but could not find the cause.
Interestingly, when we consider the number of TTL extending resolvers, we
see more TTL-violating resolvers in ProxyRack experiments than the direct scan-
ning (i.e., 14.0% vs. 7.5%). Since direct scanning only allows us to measure the
public resolvers, it may indicate that the local resolvers are more likely to extend
TTLs; we will explore this in the following section. In summary, we confirm that
our methodology can accurately find the TTL-extension policy of DNS resolvers.
TTL Violation of DNS Resolvers in the Wild 557
Table 2. Top fifteen countries sorted by the fraction of exit nodes that use TTL-
extending resolvers
Table 3. Table showing the top 15 local resolvers that extend TTLs.
responses. The majority of these ISPs and DNS resolvers are in Russia and China;
for example, we measure 13 local resolvers in China Telecom, all of which extend
TTL; this strongly suggests that the TTL extension is imposed by the ISP.
$ dig [Link]
...
;; ANSWER SECTION:
[Link]. 3600 IN CNAME [Link].
[Link] 60 IN A [Link]
TTL Violation of DNS Resolvers in the Wild 559
2
Our methodology can miss domains that delegate its name server to CDNs by replac-
ing their NS records with CDN’s ones. We could potentially identify them by checking
whether both of their web server and DNS server are managed by the same CDN.
However, some companies (e.g., Alibaba and Google) also provide VPS hosting ser-
vice, which will cause false-positive (e.g., the domain owner manages both servers
within the same VPS), thus we only focus on the CNAME expansion information.
560 P. Bhowmick et al.
5 Concluding Discussion
TTL Shortening in the Wild: A DNS resolver may cache the DNS response
shorter than the TTL; unlike TTL extension, however, caching DNS records
shorter than the TTL is not any violation of the DNS standard since RFC
2181 [10]. We can use a similar methodology to detect resolvers that cache DNS
records shorter than the TTL set by the authoritative server; for example, some
resolvers may have a parameter that determines the maximum TTL mainly not
to trust very large TTL values for security purpose [29] [5]. However, resolvers
can also decide to evict the cached DNS response depending on its cache size and
562 P. Bhowmick et al.
eviction policy, which makes it a bit hard to consistently capture DNS resolvers
that always cache shorter than the TTL. By making the second request earlier
than the TTL, we are able to measure 49 (0.99%, out of 4,965) resolvers that
always shorten the TTL and 4653 (93.7%) resolvers that always preserve the
original TTL, but we also find 263 (5.3%) resolvers showing mixed behaviors,
which suggests that their eviction policy might have impacted on, and eventually,
makes it hard for us to further investigate.
References
1. Amit, K., Haya, S., Michael, W.: Counting in the Dark: DNS caches discovery and
enumeration in the internet. IEEE Comput. Soc. DSN (2017)
2. Arends, R., Austein, R., Larson, M., Massey, D., Rose, S.: DNS security introduc-
tion and requirements. RFC 4033, IETF (2005). [Link]
txt
3. Alzoubi, H.A., Rabinovich, M.I., Spatscheck, O.: The anatomy of LDNS clusters:
findings and implications for web content delivery. In: WWW (2013)
4. Blank, S., Goldsten, P., Loder, T., Zinkn, T., Bradshaw, M.: Brand indicators for
message identification (BIMI). In: IETF (2021)
5. BIND max-cache-ttl. [Link] 18 7/[Link]?
highlight=max-cache-ttl
6. Chung, T., et al.: A longitudinal. End-to-end view of the DNSSEC ecosystem, In:
USENIX Security (2017)
7. Callahan, T., Allman, M., Rabinovich, R.: On modern DNS behavior and proper-
ties. CCR 43(4) (2013)
8. CAIDA ASOrganizations Dataset. [Link]
9. DNS based load-balancing. [Link]
what-is-dns-load-balancing/
10. Elz, R., Bush, R.: Clarifications to the DNS specification. RFC 2181, IETF (1997)
11. Edge and Browser Cache TTL. [Link]
edge-browser-cache-ttl/
12. Flavel, A., Mani, P., Maltz, D.A.: Re-evaluating the responsiveness of DNS-based
network control. In: LANMAN (2014)
13. Hoffman, P., Schlyter, J.: The DNS-based authentication of named entities (DANE)
transport layer security (TLS) protocol: TLSA. RFC 6698, IETF (2012)
14. Hola VPN. [Link]
15. Jeffrey, P., Aditya, A., Anees, S., Balachander, K., Srinivasan, S.: On the respon-
siveness of DNS-based network control. In: IMC (2004)
16. Kyle, S., Tom, C., Michael, R., Mark, A.: On measuring the client-side DNS infras-
tructure. In: IMC (2013)
17. cache-min-ttl in KnotDNS. [Link]
[Link]
TTL Violation of DNS Resolvers in the Wild 563
18. Mario, A., Alessandro, F., Diego, P., Narseo, V.-R., Matteo, V.: Dissecting DNS
stakeholders in mobile networks. In: CoNEXT (2017)
19. Margolis, D., Risher, M., Ramakrishnan, B., Brotman, A., Jones, J.: SMTP MTA
strict transport security (MTA-STS). RFC 8461, IETF (2018)
20. Moura, G.: DNS TTL violations in the wild - measured with RIPE atlas.
[Link] moura/dns-ttl-violations-in-the-wild-
measured-with-ripe-atlas
21. Moura, G., Heidemann, J., Schmidt, R.D.O., Hardaker, W.: Cache me if you can:
effects of DNS time-to-live. In: IMC (2019)
22. Mockapetris, P.: Domain Names - Concepts and Facilities. RFC 1034, IETF (1987)
23. Nygren, E., Sitaraman, R.K., Sun, J.: The Akamai network: a platform for high-
performance internet applications. OSR 44(3) (2010)
24. OpenINTEL. [Link]
25. Pochat, V.L., Goethem, T.V., Tajalizadehkhoob, S., Korczyński, M., Joosen, W.:
TRANCO: a research-oriented top sites ranking hardened against manipulation.
In: NDSS (2019)
26. minimum-ttl-override option in PowerDNS. [Link]
[Link]#minimum-ttl-override
27. ProxyRack. [Link]
28. Randall, A., et al.: Trufflehunter: cache snooping rare domains at large public DNS
resolvers. In: IMC (2020)
29. Cache-min-ttl, Cache-max-ttl option in Unbound. [Link]
ation/unbound/[Link]/
Operational Domain Name Classification:
From Automatic Ground Truth
Generation to Adaptation to Missing
Values
Abstract. With more than 350 million active domain names and at least
200,000 newly registered domains per day, it is technically and econom-
ically challenging for Internet intermediaries involved in domain regis-
tration and hosting to monitor them and accurately assess whether they
are benign, likely registered with malicious intent, or have been compro-
mised. This observation motivates the design and deployment of auto-
mated approaches to support investigators in preventing or effectively mit-
igating security threats. However, building a domain name classification
system suitable for deployment in an operational environment requires
meticulous design: from feature engineering and acquiring the underlying
data to handling missing values resulting from, for example, data collec-
tion errors. The design flaws in some of the existing systems make them
unsuitable for such usage despite their high theoretical accuracy. Even
worse, they may lead to erroneous decisions, for example, by registrars,
such as suspending a benign domain name that has been compromised at
the website level, causing collateral damage to the legitimate registrant
and website visitors.
In this paper, we propose novel approaches to designing domain name
classifiers that overcome the shortcomings of some existing systems. We
validate these approaches with a prototype based on the COMAR (COm-
promised versus MAliciously Registered domains) system focusing on its
careful design, automated and reliable ground truth generation, feature
selection, and the analysis of the extent of missing values. First, our clas-
sifier takes advantage of automatically generated ground truth based on
publicly available domain name registration data. We then generate a
large number of machine-learning models, each dedicated to handling a
set of missing features: if we need to classify a domain name with a given
set of missing values, we use the model without the missing feature set,
thus allowing classification based on all other features. We estimate the
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 564–591, 2023.
[Link]
Operational Domain Name Classification 565
importance of features using scatter plots and analyze the extent of miss-
ing values due to measurement errors.
Finally, we apply the COMAR classifier to unlabeled phishing URLs
and find, among other things, that 73% of corresponding domain names
are maliciously registered. In comparison, only 27% are benign domains
hosting malicious websites. The proposed system has been deployed at two
ccTLD registry operators to support their anti-fraud practices.
1 Introduction
Attackers have traditionally used domain names to spread malware, ensure reli-
able communication between malicious command-and-control (C&C) servers and
botnets using domain generation algorithms (DGAs), or to launch spam or phish-
ing campaigns. A domain name can be registered for a legitimate purpose by
a benign registrant or with malicious intent by an attacker. A benign domain
name can also be compromised at the hosting, domain, or website level, and
involved in malicious activities later in its lifetime.
With more than 350 million active domain names1 and at least 200 thousand
newly registered domains per day,2 it is technically and economically challeng-
ing for top-level domain (TLD) registries and registrars to scrutinize them at
the time of registration and accurately assess whether they are benign or likely
registered with malicious intent. Furthermore, once a domain name is involved
in a malicious activity, and the abusive URL is blacklisted, or reported to the
operator’s helpdesk, an investigator must gather evidence on whether the domain
name is attacker-owned (i.e., registered by a malicious actor) or has been compro-
mised (and possibly how) before deciding on the type of the mitigation action.
While a maliciously registered domain name can be suspended, a benign and
subsequently hacked domain name generally cannot be blocked because it may
cause collateral damage to the harmless domain name owner and regular vis-
itors of legitimate websites available under the benign domain name. Instead,
the webmaster or the hosting provider should remove the malicious content (e.g.,
malware or phishing website) from the server and patch the vulnerable applica-
tion to prevent future intrusions [48].
The problem of DNS abuse and domain names being a vehicle for delivering
malicious content [5,27,28,36,47] motivates the development and implementa-
tion of automated methods to support investigators in assessing domain name
maliciousness as well as appropriate and prompt mitigation of security threats.
To address these challenges, several research studies proposed domain name rep-
utation systems based on machine learning (ML) to distinguish between benign
1
[Link]
2
[Link]
566 J. Bayer et al.
After formulating the classification problem and the outcomes (i.e., labels), one
of the starting points in designing any domain name classifier is the selection
of data sources and features for distinguishing between two groups of domain
names (e.g., benign and malicious).
The primary criterion for selecting data sources is their availability. Datasets
used in DNS reputation systems can be either publicly or non-publicly avail-
able. The privileged or commercial sources such as the passive DNS data used
in the Exposure [6,7], Notos [4], or Predator [20] are only limited to those who
have access to such data. Historical data raise a similar problem (e.g., histori-
cal WHOIS data used in the takedown of Avalanche [31]). Furthermore, repro-
ducibility and performance validation of the systems relying on non-publicly
available data by independent researchers may be difficult or impossible.
On the other hand, systems based on publicly available data sources do not
have the problems raised by non-public data sources and can still achieve high
accuracy. Moreover, they are more likely adopted by the involved operators, not
only DNS intermediaries but also, for example, law enforcement agencies [31].
The Mentor [24] and Domain Classifier [29] systems used public data sources
and demonstrated high accuracy. De Silva et al. [10] combined public and non-
public (passive DNS data from Farsight [14]) data sources to achieve 97.2% accu-
racy. COMAR [34] used both publicly and non-publicly available data to distin-
guish between compromised and maliciously registered domain names and con-
cluded that when removing the non-publicly available passive DNS, it achieves
an accuracy of up to 97%.
In Sect. 3.2, we critically revisit relevant features and select those that do not
use privileged or commercial data sources.
Feature importance refers to techniques that assign a score to input features (e.g.,
domain name popularity, domain name age, etc.) based on the extent to which
they contribute to the prediction of the target variable (e.g., classification of
benign versus maliciously registered domain names). Ranking features according
to their importance shows which features are irrelevant and can be omitted. It
reduces the dimensionality of the model, its complexity, the need to collect data,
and makes it possible to estimate the impact of missing features on the system.
In the DNS reputation systems we reviewed, only Hao et al. [20], Maroofi et
al. [34], and Le Pochat et al. [31] documented feature importance of the proposed
models. Note that even the most important feature, if it is missing from the
dataset (and its value cannot be estimated), cannot contribute to the prediction
of the target variable. Therefore, we analyze the extent of missing values resulting
from measurement errors in Sect. 3.4, discuss feature importance in Sect. 3.6, and
show how missing values of selected features affect the classification of domains
using scatter plots.
Operational Domain Name Classification 569
novel approach to automatically generate ground truth data for such systems.
It consists of measuring the mitigation actions on abusive domain names by
TLD registries, registrars, and hosting providers and can be applied to different
domain classification problems.
3 Methodology
In this section, we discuss in detail the methodology to generate ground truth
data automatically, the prototype implementation of the classifier, and the prac-
tical approach to overcoming the problem of missing values.
the server, such as a malware download or a phishing site, and patch the vulner-
able application to prevent future intrusions. Based on these generally accepted
mitigation practices [11], we design the measurement setup to automatically
distinguish between compromised and malicious domains.
original registrant. If it is not the case, it might be possible that such a domain
was blacklisted, removed from the zone, later became available for registration,
and re-registered.
3
[Link]
4
[Link] providers.
Operational Domain Name Classification 573
### Example - Homepage ### and thus, the page was automatically excluded
from the compromised dataset even though the title still contains the original
string.
1. Bing search engine results. As discussed by the authors of the COMAR sys-
tem, it is a paid service, therefore, we exclude it.
2. Features depending on passive DNS. The access to passive DNS data is priv-
ileged and related features proved to have a negligible impact on the perfor-
mance [34].
3. TLD maliciousness index.5 It is calculated by Spamhaus [44] and is not avail-
able for commercial use.
4. The relationship between the domain name and the hosted content. Original
COMAR extracts keywords from the domain name and generates their syn-
onyms using a commercial API. They then determine if the domain name is
related to its content based on the occurrence of the keywords and their syn-
onyms in the text of the home page. Since the API is not publicly available
at the time of writing, we decided to remove this feature.
5. Quantcast ranking system is not publicly available anymore.
6. TLS certificate price. It is not trivial to distinguish between free and paid
certificates since some certificate authorities (e.g., Comodo CA1) offer both
paid and free certificates.
7. Presence of a TLS certificate. We exclude the presence of the Transport Layer
Security (TLS) certificate from our features since the use of TLS certificates
among malicious and benign but compromised domains used in phishing is
comparable (see Sect. 4.1).
8. Valid TLS certificate. For shared hosts, if a certificate is not valid (e.g., wrong
host error), we cannot conclude if the malicious actor issued a wrong certifi-
cate or if the certificate belongs to another domain on the same (shared)
hosting service. Therefore, this feature is not suitable for operational deploy-
ments.
9. TLD price. The TLD price is not unique among all registrars and resellers,
and changes over time. In addition, special offers from registrars or domain
resellers can drastically reduce the price for a specific TLD. It is also difficult
to collect such data at scale.
Based on this analysis, we remove 13 features and train the model with the
remaining 25 features using Logistic Regression. We analyze the coefficients and
remove features that are not important for the model. We present the final set
5
[Link]
574 J. Bayer et al.
of the 17 remaining features in Table 2. Note that the is in alexa (F14) feature
was only available before the termination of service announced by Alexa in May
2022.6 Since this feature is unavailable for only three of the twenty months of phish-
ing data collected, we keep it and use the method of handling missing values as
explained in Sect. 3.3. However, it can be replaced by the Tranco top sites rank-
ing [30] in future work. For has famous brand name (F2), we used the list of target
brand names provided by PhishTank. We consider this binary feature to be true
if a domain name contains one of these trademarks. We use a similar method to
the one proposed by Kintis et al. [25]. However, as our work does not only focus
on combo squatting, we consider this feature to be true even for some of the five
typosquatting models of Wang et al. [51] (e.g., the value of this feature for domain
[Link] with trademark facebook would be true even if it would not be
marked as combosquatted by the method proposed by Kintis et al. [25]).
as DNS-related features. However, some features suffer from missing values either
due to measurement or parsing errors, or the unavailability of data.
For instance, the difference between domain creation (registration) and
domain blacklisting time, also referred to as domain name age [34] (F6), is a
feature derived from WHOIS/RDAP data and is missing for 22.29% of domain
names. However, for some TLDs (e.g., .de TLD), there is no information about
the domain registration date in WHOIS. For other TLDs, it is not feasible to
collect WHOIS information at scale since either there is no conventional WHOIS
server (e.g., for .gr TLD) or the access is restricted to authorized IP addresses
(e.g., .es TLD). Moreover, extracting WHOIS information relies on manual cre-
ation of parsing rules and templates for individual registrars, and is by nature
limited in scope and susceptible to changes in data representation [32]. The
RDAP protocol [38] overcomes the problem of parsing but it is not universally
deployed.
While the features related to the page content play an essential role in classi-
fication [34], our results show that the values for these features are often missing
(up to 47.5% for Hyperlinks). The reason is that data collection related to
web content requires significant resources (e.g., browser emulation in our case)
and its results highly depend on the page load time and implementation. For
instance, poorly maintained websites may result in timeout or measurement
errors. Domain redirection is another reason for the missing values of content-
related features. URLs that use domain redirection (i.e., HTTP 3XX status code
576 J. Bayer et al.
or JavaScript redirection) will load the content of the destination domain name
(i.e., different from the original domain). In such cases, we consider these features
as missing.
Fig. 1. Boxplot showing the MCC of models grouped by the number of removed fea-
tures at a time. Triangles and horizontal lines represent the mean and median of MCC
for each group of models, respectively.
cost compared to systems based on one model only, but overall, it remains low.
More importantly, the cost of training is mainly related to the time required to
generate ground truth, which is low compared to manual labeling. Therefore, the
here-proposed method significantly reduces the overall cost and can label more
samples for training. Given the automated approach for ground-truth generation,
we could consider regular active learning. However, it would require future work
to evaluate if static models exhibit high performance over time.
Table 4. Distribution of top ten models (combinations of removed feature sets) with
the highest coverage of samples in the unlabeled dataset.
During the following evaluation, we use common metrics to evaluate the per-
formance of models. For details, we refer the reader to Appendix A. Figure 1
shows a boxplot summarizing the distribution of Matthews Correlation Coeffi-
cient (MCC) for the 511 models. The MCC of the full model () 1 is 0.87 with
a 93.67% accuracy. The model without the WHOIS feature set () 2 is the most
significant outlier (MCC: 0.78, FNR: 14.2%) for the models with one removed
feature set at a time. This result confirms that the domain age at the time of
blacklisting is a strong feature and its absence causes a significant decrease in
performance. Similarly, the model without the FS2 and FS10 feature sets () 3
Operational Domain Name Classification 579
is the most significant outlier (MCC: 0.71, FNR: 17.9%) for the group of models
with two removed features at a time. If we remove the WHOIS (FS2) and Wayback
machine (FS10) feature sets, the performance of the model is highly impacted.
As expected, one of the worst models 4 (MCC: 0.09, FNR: 98%) lacks all previ-
ously discussed feature sets (FS2, FS10), but also the remainder of the important
features (FS3, FS5, FS6, FS8, and FS9). We discuss the implications of these
findings for operational classification later in this section.
Note that even if MCC is a suitable method to evaluate the performance of
binary classification with an unbalanced distribution of classes, it is still essential
to consider other metrics, such as the false negative and false positive rates. For
instance, we carefully monitor the false negative rate to avoid compromised
domains incorrectly classified as malicious, which may lead to the blocking of a
benign domain causing collateral damage to the legitimate registrant. Therefore,
model 5 that only uses the domain name age (F6) calculated based on WHOIS
has a high MCC (0.78) but it has to be used with caution as its FNR is high
(26.1%).
As described in Sect. 3.3, our system consists of models trained on combi-
nations of incomplete feature vectors and handles the classification of domain
names with missing values. Table 4 shows ten of such combinations that appeared
the most in our unlabeled dataset. For instance, 10.2% of domains that have miss-
ing values for FS9 (Hyperlinks) can still be classified using a model with good
overall performance (MCC: 0.87, FNR: 10%), similar to the model with no miss-
ing values. We observed that 36.5% of domain names in our unlabeled dataset
(see Sect. 4) do not have any missing values and therefore, they can be classified
using the complete model with all 10 feature sets present. The remaining 63.5%
of domain names have at least one missing value and cannot be classified using
a single-model approach (assuming that other methods to handle missing val-
ues are not implemented). The 501 remaining models cover 12.1% of samples.
These models are necessary for the system to handle missing values resulting
from measurements related to each feature set.
Fig. 3. Scatter plot of probability changes between the full model and the model with-
out DNS-related features (FS7).
visualization of the feature importance using scatter plots based on similar prin-
ciples as in their methods. Out of the 511 models, we select those trained with
only one feature set missing (9 models as we do not consider lexical features as
possibly missing). We choose a sample of 10,000 domain names with no miss-
ing values and we classify them first with the complete model and then with 9
models with one removed feature at a time. We present the results for selected
feature sets in Figs. 3 and 4. Each point in the graph represents one domain
name. The y-axis is the predicted probability of a domain being compromised
when classified with the complete model. The x-axis is the probability predicted
by a model without one of the feature sets. Each point is colored based on the
category (content length or domain age). The output probability of points lay-
ing on the line x = y remained unchanged after a feature set elimination. Points
for which x > y (increase in output probability), became “more compromised”
after feature set removal. Similarly, points for which x < y (decrease in the out-
put probability), became “more malicious”. Note that the red zone at the top
left and bottom right corner highlight the points that could potentially change
labels. For instance, Fig. 3 shows that the output probability of a small fraction of
domains was impacted by removing the DNS-related features (FS7). Therefore,
this feature set does not have a high impact on the classification results.
Fig. 4. Scatter plot of probability changes between the full model and the model with-
out feature set FS6 (web technologies).
On the other hand, Figs. 4, 10, and 11 (in Appendix B) demonstrate that
feature sets FS2 (WHOIS), FS8 (Alexa), and FS9 (Hyperlinks) strongly influ-
ence the output of our system as many data points moved horizontally after
eliminating a feature set (i.e., became more malicious or more compromised).
582 J. Bayer et al.
4 Classification Results
In this section, we apply the prototype classifier to 218,806 unlabeled unique
domain names from APWG [3], OpenPhish [39], and PhishTank [49] URL black-
lists collected between January 2021 and September 2022. We study four selected
characteristics of the domain names of malicious URLs and analyze their distri-
bution across different types of TLDs.
The overall classification results show that 73% of phishing domain names
were registered for malicious purposes, and 27% were classified as registered by
benign users but have been compromised. If the domain names were compro-
mised at the hosting rather than at the DNS level, they should not be blocked
by TLD registries or registrars.
We now explain how the compromised and maliciously registered domain names
distinguished by our system differ in terms of four selected features: popular
terms in domain names, the number of web technologies used, the domain name
age, and the use of TLS certificates.
The features indicating that a cybercriminal (rather than a benign user)
has registered a domain name include specific keywords such as ‘verification’,
‘payment’, ‘support’, or brand names (e.g., [Link]). Figure 6
presents a word frequency analysis of the phishing dataset for both domain names
automatically classified as maliciously registered (orange) and those classified as
compromised (blue).
We can observe that cybercriminals tend to incorporate such words into
domain names to lure victims into entering their credentials. The most frequently
used keywords by malicious actors are ‘online’, ‘secure’, ‘bank’, ‘support, ‘info’,
‘login’, and ‘help’. On the other hand, the domain name of compromised sites
rarely contains such specific keywords.
One of the used features is the number of web technologies (F12): a count
of the JavaScript, Cascading Style Sheets (CSS), or Content Management Sys-
tem (CMS) frameworks and plugins used to build the homepage of a registered
domain name. The higher number of technologies used for developing a website
could reflect the amount of effort and time its designer spent to create a fully-
functional website. While this is true for benign (compromised) domain names,
malicious actors tend to put little effort into deploying multiple technologies
when designing websites on maliciously registered domain names, as it is not
critical to the success of phishing attacks. Figure 7 shows the results for compro-
mised and maliciously registered domain names. As many as 52.2% of compro-
mised domains use more than five different (potentially vulnerable) technologies,
frameworks, and plugins to build the website. In comparison, 66.1% of the mali-
ciously registered domain names have no specific technology on their homepage.
We have noticed that many maliciously registered domains either have no home-
page (showing the default directory index served by the web server), redirect to
another domain (e.g., the landing page of a phishing attack), or display a custom
error message (e.g., forbidden page). Instead, they frequently serve the phishing
page either on a URL path or a subdomain level.
The age of a domain name (F6), defined as the time between the registration
of the domain name and its appearance on the blacklist, is one of the important
features of our classifier. Intuitively, the older the domain name, the more likely
it is to have been registered by a benign user but subsequently compromised.
On the other hand, cybercriminals tend to use a domain name for malicious
activities soon after registration. Figure 8 shows the age of domain names for all
TLDs that provide the registration date as part of their WHOIS data: “0” means
that registration and blacklisting occurred on the same day, “1” – the difference
between the registration date and the blacklisting date is at most one year, and
“>5” means that the difference between the domain registration date and the
blacklisting date is greater than five years. For 93.6% of maliciously registered
domain names, the difference between the domain registration date and the
blacklisting date is less than a year, and for 11.3% of them, the domains were
blacklisted on the same day the domain was registered. For compromised domain
names, about 51.4% of them were registered at least six years before being
blacklisted. A possible explanation for this phenomenon is that websites hosted
on older domain names are more likely to use outdated technologies or content
management systems (e.g., vulnerable versions of CMS such as WordPress),
making them easier to compromise.
In some sporadic cases, malicious actors may “age” registered domains, wait-
ing weeks or sometimes months before abusing them, or compromise domain
names shortly after their registration [34]. However, as our system is fully auto-
mated and performs classification based on multiple features (the domain name
age is just one of them), it is more resistant to manipulation (e.g., domain aging).
While we explain in Sect. 3.2 why we avoid using TLS certificate features, we
analyze their use by owners of compromised versus maliciously registered domain
586 J. Bayer et al.
names. According to a PhishLabs report [40], three quarters of all phishing sites
used HTTPS (HTTP over TLS) in 2020 “to add a layer of legitimacy, better
mimic the target site in question, and reduce being flagged or blocked from some
browsers.” However, the report conflates compromised and maliciously registered
domain names. Therefore, to establish whether cybercriminals increasingly use
TLS certificates, we need to distinguish between compromised and maliciously
registered domain names and analyze the use of TLS only in the latter group.
Otherwise, it is unclear whether the TLS certificate was issued at the request of a
criminal for a maliciously registered domain to enhance the website’s credibility
or at the request of a legitimate domain owner for a benign domain name that
was later compromised and abused by a criminal.
Figure 9 shows the statistics of TLS certificates issued for malicious and
benign (and later compromised) domains involved in phishing attacks. The use
of TLS certificates is less widespread among phishers than for benign (but com-
promised) domain names. 63.9% of phishing attacks using compromised domains
take advantage of TLS certificates issued at the request of benign domain own-
ers while 55.2% of maliciously registered domains use TLS certificates delib-
erately deployed by malicious actors to lure their victims. Surprisingly, 15.5%
of maliciously registered domains used most likely paid TLS certificates. We
further investigate these domains and find that 48.7% of them had a TLS cer-
tificate issued by Sectigo [42]. The majority of these domains were registered
with Namecheap [37] which offers a cheap all-in-one hosting package including a
domain name registration, hosting, and a one-year valid TLS certificate issued
by Sectigo [42].
Acknowledgments. We thank Benoı̂t Ampeau, Marc van der Wal (AFNIC) and the
anonymous reviewers for their valuable feedback, Anti-Phishing Working Group, Open-
Phish, and PhishTank for providing access to their URL blacklists. This work has been
carried out in the framework of the COMAR project funded by SIDN, the .NL Reg-
istry and AFNIC, the .FR Registry. It was partially supported by the Grenoble Alpes
Cybersecurity Institute (under contract ANR-15-IDEX-02), and the French Ministry of
Research (PERSYVAL-Lab project under contract ANR-11-LABX-0025-01, and DiNS
project under contract ANR-19-CE25-0009-01).
588 J. Bayer et al.
Appendix
TP + TN
Accuracy =
TP + TN + FP + FN
FP FN
FPR = FNR =
FP + TN FN + TP
TP × TN − FP × FN
M CC = , (2)
(T P + F P )(T P + F N )(T N + F P )(T N + F N )
where TN, TP, FN, and FP represent the numbers of true negative, true positive,
false negative, and false positive, respectively. We refer to compromised domains
as positives and to maliciously registered ones as negatives. Accuracy is the
proportion of correctly predicted labels among all samples. We also make use of
a Matthews Correlation Coefficient (MCC) as defined in Eq. 2 [35]. This metric
was developed to evaluate the quality of a binary classification and its values
vary between –1 and +1, where +1 means perfect prediction (the best score), 0
is equivalent to random results, and –1 shows that all samples were misclassified
(the worst score). In contrast to accuracy, MCC provides a more realistic metric
for imbalanced datasets such as ours.
Fig. 10. Scatter plot of probability changes between the full model and the model
without features related to WHOIS data (FS2).
Operational Domain Name Classification 589
Fig. 11. Scatter plot of probability changes between the full model and the model
without features related to hyperlinks (FS9).
References
1. Alowaisheq, E., et al.: Cracking the wall of confinement: understanding and ana-
lyzing malicious domain take-downs. In: Proceedings of NDSS (2019)
2. Amazon: Alexa: SEO and Competitive Analysis Software (2022). [Link]
[Link]/
3. Anti-Phishing Working Group: Global phishing survey: Trends and domain
name use in 2016 (2016). [Link] Global Phishing
Report [Link]
4. Antonakakis, M., Perdisci, R., Dagon, D., Lee, W., Feamster, N.: Building a
dynamic reputation system for DNS. In: Proceedings of USENIX Security, p. 18
(2010)
5. Bayer, J., et al.: Study on domain name system (DNS) abuse: technical report.
arXiv preprint arXiv:2212.08879 (2022)
6. Bilge, L., Kirda, E., Kruegel, C., Balduzzi, M.: EXPOSURE: finding malicious
domains using passive DNS analysis. In: Proceedings of 18th NDSS (2011)
7. Bilge, L., Sen, S., Balzarotti, D., Kirda, E., Kruegel, C.: Exposure: a passive DNS
analysis service to detect and report malicious domains. ACM Trans. Inf. Syst.
Secur. 16(4) (2014)
8. Corona, I., et al.: DeltaPhish: detecting phishing webpages in compromised web-
sites. arXiv:1707.00317 (2017)
9. Daigle, L.: Whois protocol specification. Technical report, RFC Editor (2004)
10. De Silva, R., Nabeel, M., Elvitigala, C., Khalil, I., Yu, T., Keppitiyagama, C.:
Compromised or attacker-owned: a large scale classification and study of hosting
domains of malicious URLs. In: Proceedings of USENIX Security, pp. 3721–3738
(2021)
11. DNS Abuse Framework. [Link]
590 J. Bayer et al.
12. Donders, A.R.T., van der Heijden, G.J., Stijnen, T., Moons, K.G.: Review: a gentle
introduction to imputation of missing values. J. Clin. Epidemiol. 59(10), 1087–1091
(2006)
13. Emmanuel, T., Maupong, T., Mpoeleng, D., Semong, T., Mphago, B., Tabona, O.:
A survey on missing data in machine learning. J. Big Data 8 (2021)
14. Farsight Security: Passive DNS Historical Internet Database: Farsight DNSDB
(2022). [Link]
15. Felegyhazi, M., Kreibich, C., Paxson, V.: On the potential of proactive domain
blacklisting. In: Proceedings of 3rd USENIX LEET (2010)
16. Frosch, T., Kührer, M., Holz, T.: Predentifier: detecting botnet C&C domains from
passive DNS data. In: Zeilinger, M., Schoo, P., Hermann, E. (eds.) Advances in IT
Early Warning, pp. 78–90. AISEC (2013)
17. Google: Certificate Transparency. [Link]
18. Google Safe Browsing. [Link]
19. Halvorson, T., Der, M.F., Foster, I., Savage, S., Saul, L.K., Voelker, G.M.: From.
academy [Link]: an analysis of the new TLD land rush. In: Proceedings of IMC,
pp. 381–394 (2015)
20. Hao, S., Kantchelian, A., Miller, B., Paxson, V., Feamster, N.: PREDATOR: proac-
tive recognition and elimination of domain abuse at time-of-registration. In: Pro-
ceedings of ACM SIGSAC, pp. 1568–1579 (2016)
21. Hollenbeck, S.: Extensible Provisioning Protocol (EPP) Domain Name Mapping.
RFC 3731, RFC Editor (2004)
22. ICANN: EPP Status Codes — What Do They Mean, and Why Should I Know?
[Link]
23. Internet Archive: Wayback Machine. [Link]
24. Kheir, N., Tran, F., Caron, P., Deschamps, N.: Mentor: positive DNS reputation
to skim-off benign domains in botnet C&C blacklists. In: Cuppens-Boulahia, N.,
Cuppens, F., Jajodia, S., Abou El Kalam, A., Sans, T. (eds.) SEC 2014. IAICT,
vol. 428, pp. 1–14. Springer, Heidelberg (2014). [Link]
642-55415-5 1
25. Kintis, P., et al.: Hiding in plain sight. In: Proceedings of ACM SIGSAC (2017)
26. Kohavi, R.: A Study of Cross-Validation and Bootstrap for Accuracy Estimation
and Model Selection. In: Proceedings of 14th IJCAI, vol. 2, pp. 1137–1143 (1995)
27. Korczyński, M., Tajalizadehkhoob, S., Noroozian, A., Wullink, M., Hesselman, C.,
van Eeten, M.: Reputation metrics design to improve intermediary incentives for
security of TLDs. In: Proceedings of IEEE Euro SP (2017)
28. Korczyński, M., et al.: Cybercrime after the sunrise: a statistical analysis of DNS
abuse in new gTLDs. In: Proceedings of ACM ASIACCS (2018)
29. Le Page, S., Jourdan, G.-V., Bochmann, G.V., Onut, I.-V., Flood, J.: Domain
classifier: compromised machines versus malicious registrations. In: Bakaev, M.,
Frasincar, F., Ko, I.-Y. (eds.) ICWE 2019. LNCS, vol. 11496, pp. 265–279. Springer,
Cham (2019). [Link] 20
30. Le Pochat, V., Van Goethem, T., Tajalizadehkhoob, S., Korczyński, M., Joosen,
W.: Tranco: a research-oriented top sites ranking hardened against manipulation.
In: Proceedings of NDSS. Internet Society (2019)
31. Le Pochat, V., et al.: A practical approach for taking down avalanche botnets
under real-world constraints. In: Proceedings of 27th NDSS (2020)
32. Liu, S., Foster, I., Savage, S., Voelker, G.M., Saul, L.K.: Who [Link]? learning to
parse WHOIS records. In: Proceedings of IMC, pp. 369–380 (2015)
Operational Domain Name Classification 591
33. Ma, J., Saul, L.K., Savage, S., Voelker, G.M.: Beyond blacklists: learning to detect
malicious web sites from suspicious URLs. In: Proceeding of 15th ACM SIGKDD
ICKDDM, pp. 1245–1254. KDD (2009)
34. Maroofi, S., Korczyński, M., Hesselman, C., Ampeau, B., Duda, A.: COMAR: clas-
sification of compromised versus maliciously registered domains. In: Proceedings
of IEEE EuroS&P, pp. 607–623 (2020)
35. Matthews, B.: Comparison of the predicted and observed secondary structure of T4
Phage Lysozyme. Biochimica et Biophysica Acta (BBA) - Protein Struct. 405(2),
442–451 (1975)
36. Moura, G.C.M., Müller, M., Davids, M., Wullink, M., Hesselman, C.: Domain
names abuse and TLDs: from monetization towards mitigation. In: Proceedings of
IFIP/IEEE, pp. 1077–1082 (2017)
37. Namecheap. [Link]
38. Newton, A., Hollenbeck, S.: Registration data access protocol (RDAP) query for-
mat. Technical report, RFC Editor (2015)
39. OpenPhish. [Link]
40. PhishLabs: Abuse of HTTPS on Nearly Three-Fourths of all Phishing Sites
(2020). [Link]
of-all-phishing-sites/
41. PhisLabs: [Link]
42. Sectigo Limited: Sectigo®Official - SSL Certificate Authority & PKI Solutions.
[Link]
43. SiteAdvisor, M.: [Link]
44. Spamhaus. [Link]
45. Spooren, J., Vissers, T., Janssen, P., Joosen, W., Desmet, L.: Premadoma: an
operational solution for DNS registries to prevent malicious domain registrations.
In: 35th ACSAC, pp. 557–567 (2019)
46. SURBL. [Link]
47. Tajalizadehkhoob, S., Böhme, R., Gañán, C., Korczyński, M., Eeten, M.V.: Rotten
apples or bad harvest? what we are measuring when we are measuring abuse. ACM
Trans. Internet Technol. 18(4) (2018)
48. Tajalizadehkhoob, S., et al.: Herding vulnerable cats: a statistical approach to
disentangle joint responsibility for web security in shared hosting. In: Proceedings
of ACM SIGSAC, pp. 553–567 (2017)
49. Ulevitch, D.: PhishTank Join the fight Against Phishing (2006). [Link]
org/
50. URIBL. [Link]
51. Wang, Y.M., Beck, D., Wang, J., Verbowski, C., Daniels, B.: Strider typo-patrol:
discovery and analysis of systematic typo-squatting. In: Proceedings of USENIX
Association, vol. 2, p. 5 (2006)
52. Zhang, P., et al.: CrawlPhish: large-scale analysis of client-side cloaking techniques
in phishing. In: Proceedings of IEEE S&P, pp. 1109–1124 (2021)
Web
A First Look at Third-Party Service
Dependencies of Web Services in Africa
1 Introduction
The websites we use everyday offload critical services such as name resolution
(DNS), content distribution (CDN), and certificate issuance/revocation (CA) to
third parties for key services e.g., AWS Route 53 for DNS, Akamai for CDN,
DigiCert for CA. As a result, the availability and security of these websites, and
thus of our data and operations, depend on the availability and security of those
third parties. The effects of such dependencies are routinely observed in the
Internet today. For example, a dependency on DNS resulted in the downtime of
multiple websites (more than 100K) for several hours together with their DNS
provider (Dyn) which was attacked by a Mirai Distributed Denial of Service
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 595–622, 2023.
[Link]
596 A. Kashaf et al.
(DDoS) attack [24]. Similarly, users of multiple websites lost access to their
accounts for weeks, because a single CA issued an incorrect revocation of a
certificate in 2016 [22].
To gauge the security risk that such dependencies entail, one needs to under-
stand the prevalence of third-party dependencies across the websites that are
important for users all over the world. While such studies exist [25,26,28,33,47,
55], their target users/websites are particularly skewed towards North America
and Europe. The geographical bias of the datasets used in previous studies of
third-party dependencies creates a critical gap as distinct regions exhibit unique
characteristics, needs, and opportunities that are effectively ignored. Naively
assuming that observations generalize across regions, entails risks as it under-
estimates the practicality of certain attacks and creates false assurance of the
security of critical region-specific websites (e.g., those related to government
or health insurance in those countries). This is also recognized by the Internet
Society’s Measuring Internet Resilience in Africa (MIRA) project [46].
To bridge this gap, in this paper we study third-party dependencies of web-
sites in Africa. Our study is motivated by the increasing number of DDoS attacks
in Africa [21], the increasing popularity of third-party services, the low cyber
readiness of African users and businesses [40]. These are exemplified by various
recent attacks. For example, in July 2022, Afrihost, one of the major hosting
and DNS providers in South Africa, went down for 30 h due to load shedding
which caused a cooling equipment failure in one of Afrihost’s datacenters. More-
over, the relative scarcity of local providers urges website operators to rely often
solely on global service providers such as Amazon, Akamai, and Cloudflare whose
outages also affect users, and websites from Africa.
Beyond raising awareness of the unique security challenges that African users
and operators face, our study contributes to the resilience of the Internet in
Africa. Concretely, we aim to provide stakeholders and operators with more tai-
lored insights, to help them avoid common pitfalls in using third-party depen-
dencies, understand their attack surface, and optimize their defense strategies
towards the most pressing needs.
To investigate third-party dependencies in African websites, we focus on
websites which are Africa-centric: websites that are popular in Africa (Africa-
visited ), or predominantly targeted towards Africans (Africa-dominant), or are
hosted in Africa (Africa-hosted ), or are operated in Africa (Africa-operated ).
We investigate their dependencies using four measurement vantage points in
Africa (Nigeria, Rwanda, South Africa, and Kenya). Specifically, our measure-
ment study focuses on answering the following questions: First, how prevalent are
third-party dependencies in the Africa-visited , Africa-hosted , Africa-operated ,
and Africa-dominant websites? Second, how centralized are third-party depen-
dencies among providers used in Africa-visited , Africa-hosted , Africa-operated ,
and Africa-dominant websites? Finally, how does the dependence on third parties
in Africa compare to the US? Since prior work [28] studies third-party dependen-
cies from a US vantage point, hence, in this work, we use the US as a baseline.
A First Look at Third-Party Service Dependencies 597
2 Preliminaries
When a website uses another entity for a particular service (e.g., DNS), we say that
the website has a third-party dependency on that service provider, making it
a third-party provider as opposed to having a private provider which belongs to
the website itself as defined by Kashaf et al. [28]. We illustrate this in Fig. 1. Here,
[Link] uses an entity other than itself for a particular service (here
DNS and CDN). Therefore, [Link] has a third-party DNS depen-
dency on Cloudflare and Dyn DNS, and it has a third-party CDN dependency on
KeyCDN. [Link] in Fig. 1 uses only a single CDN provider. Hence,
it has a critical dependency on KeyCDN. However, since [Link]
uses two DNS providers, it is redundantly provisioned with respect to DNS and
does not have a critical dependency on Cloudflare or Dyn DNS.
For DNS and CDN, we measure critical dependency by analyzing if a given
website is redundantly provisioned or not. However, in the case of CA depen-
dency, a website is critically dependent on a CA if it does not support Online
Certificate Status Protocol (OCSP) stapling. If OCSP stapling is enabled, the
user accessing a given website does not have to contact the OCSP server to check
the website certificate for revocation. Instead, an OCSP response signed by the
certificate authority comes stapled from the website server itself, thus removing
the dependence on OCSP server [4].
Concentration of a Service Provider. The number of websites dependent
on a service provider gives the concentration of that service provider.
Impact of a Service Provider. This gives the number of websites critically
dependent on a service provider.
Table 1. We consider three sets (categories) of websites for our analysis which differ
in the location of their users (usage), the location in which they are hosted (hosting),
and their audience.
Who uses it? Who operates it? Who hosts it? Who is it for? Website sets
uC – – – WC
visited
– hC – – WC
hosted
operated
– – oC – WC
– – - dC WC
dominant
Fig. 2. The figure shows the relationship between different website sets for all four
countries. The visited set is the super-set of all the other sets according to our method-
ology described in Sect. 3.
3 Dataset
websites, we start from the popular websites in each African country which
constitutes the country-visited set WCvisited . This helps in identifying websites
that can impact the African Internet users, website operators, and the Internet
economy of African countries the most.
We use the Chrome User Experience Report (CrUX) dataset [23] to get
the top 10K popular websites in each selected African country. This dataset is
curated monthly by aggregating browsing data of Chrome and Chromium users
who have opted in for browser history and usage statistic [Link] opt-in
requirement may introduce bias and the list may not truly reflect popular web-
sites in a region, however, prior work has evaluated the Google CrUX dataset and
found it to be quite reliable with respect to popularity [53]. Moreover, Chrome
and Chromium browsers constitute more than 80% traffic in our countries of
interest [56]. CrUX is ranked by the number of completed page loads.
The CrUX dataset is aggregated by web origin (e.g., [Link] For
DNS analysis, we need domain names, and using web origin may result in mul-
tiple entries for the same domain. Hence, we normalize this dataset by grouping
web origins by domain names and choosing the smallest rank value as the rank
for each domain. This same normalization technique has been previously done in
prior work [53] and is shown to be accurate at capturing popular websites. [53]
also shows that CrUX is better at capturing popular websites than other top
lists as defined by visit and visitor metrics. In addition, most top lists only give
popular websites in the world. However, for this analysis, we need regional popu-
lar websites and found that the CrUX dataset is a good source for that. We build
our website sets based on the CrUX dataset for August 2022 and the definition
of website sets can be found in Table 1.
Dataset for Country-Visited Websites: We use the CrUX dataset of the
top 10K websites, for NG, RW, KE, ZA, and US. We normalize this dataset for
each country, by grouping web origins by domain names as mentioned above.
This gives us the country-visited WCvisited dataset for each country.
The country-dominant, country-hosted, and country-operated website sets are
built from this dataset with the relationship shown in Fig. 2 for all countries. We
describe our methodology below:
Dataset for Country-Dominant Websites: As defined in Sect. 2, country-
dominant websites are made for users located in the corresponding African coun-
try. A naive approach to collecting such a list is to filter websites by their country
code top-level domain (ccTLD) [8]. However, this approach would result in many
false positives because some domain registrars give .ccTLD domains to anyone.
For example, [Link] has the Libyan ccTLD, but the website is not made for
or visited by Libyan users2 . Therefore, we combine multiple heuristics to collect
the country-dominant websites. Concretely, a website belongs to the country-
dominant set, if it belongs to country-visited set which is the top 10K visited
websites and satisfies one of three requirements. First, we pick websites with
ccTLDs that belong to that particular country. Observe that this filtering is
2
[Link]
602 A. Kashaf et al.
different from the previous heuristic as we require that the website is popular
in that country and has the ccTLD of the same country. Second, the website
hostname contains Africa or the name of an African country. Again, while this
heuristic would alone cause false positives e.g.,ancient-egypt.org3 intersecting
it with the popular sites in Africa considerably decreases those cases. Finally, we
look at the website content of the landing page, and the website URLs referred
to in the landing page to get the phone number associated with the website.
We only consider a website to belong to a particular country if all the phone
numbers mentioned on it have the country code of that country. This tech-
nique reduces false positives resulting from websites containing multiple phone
numbers, not necessarily belonging to the website. For example, the website via-
[Link] contains phone numbers of multiple countries including South Africa
but its dominant country is actually Brazil4 . To conclude, we define country-
dominant websites Wdom−af r as:
4 Methodology
We are interested in measuring the third-party dependencies of Africa-centric
websites on authoritative Domain Name Servers, Content Delivery Networks,
and Certificate Authorities for revocation information (OCSP servers and cer-
tificate revocation list (CRL) distribution points).
3
[Link]
4
[Link]
A First Look at Third-Party Service Dependencies 603
a Connection Refused error, then it means the website does not support HTTPS.
Next, we initiate an HTTPS connection with it and fetch the SSL certificates.
In the NG-visited websites, 94.0% support HTTPS, 95.7% support HTTPS in
KE-visited , and 94.3% support HTTPS in the RW-visited , 95.2% in ZA-visited
and 94.6% in US-visited . We observed 22 distinct CAs for NG, 26 distinct CAs
for RW, 24 distinct CAs for KE, 23 distinct CAs for ZA, and 23 distinct CAs
for US. We classify the CAs as a third party, again using TLD matching, SAN
list, and SOA records [28].
Certain private CAs issue certificates and provide revocation checking for
their own domains only, e.g., Microsoft, etc. Therefore, we use the same heuris-
tics as mentioned for DNS to classify whether OCSP servers and CDPs are
private or third parties as in [28]. Particularly, we classify a CA as private if the
SLD+TLD of the website matched the SLD+TLD of the OCSP server, or if the
SLD+TLD of the OCSP server exists in the SAN list of the website. Moreover,
we classify the CA as third-party if the SOA of the OCSP server and the website
differ. Next, to see if a website has a critical dependency on OCSP servers, we
check if it has enabled OCSP stapling using OpenSSL [61]. If enabled, the cer-
tificate’s revocation status comes stapled from the webserver when a user visits
the website, requiring no online revocation check from the OCSP server.
CDN Measurements: To find CDNs used by a website, we look at the canon-
ical name (CNAME) redirects for the internal resources of a webpage. If the
website is using a CDN for a particular resource, the CNAME of that resource
will point to the CDN. First, we render the landing page of the website using
Puppeteer [49] and record the URL of all the resources retrieved by a website.
Then, if the SLD + TLD of the resource matches that of the website or it exists
in the website’s SAN list, we classify it as an internal resource [28]. Then, we
query the CNAME record for all internal resources of the webpage and use the
CNAME-to-CDN map from the prior work [28], which we verified and extended
to include African CDNs. Then we classify a CDN as a private or third party by
using the same SLD+TLD matching, SAN Lists, and SOA records as done in
the case of DNS and CA and in [28]. We find that 18.5%, 23.9%, 19.6%, 22.0%,
and 40.4% use CDN for NG-visited , RW-visited ,KE-visited , ZA-visited and US-
visited . We observe 56 CDNs for NG, 59 CDNs for RW, 59 CDNs for KE, 55
CDNs for ZA, and 60 CDNs for the US.
4.1 Limitations
– Our analysis considers only four vantage points in Africa. It is possible that
the dependencies in countries for which we do not have more vantage points
vary greatly. While we accept this limitation, however, getting vantage points
in some of the African countries is extremely hard due to the lack of mature
Internet infrastructure, including VPN server presence.
– We only look at popular websites. While this may overlook certain websites,
studying all possible websites is not feasible. We argue that this is a reasonable
compromise as popular websites will be the ones that impact African users
and businesses the most.
A First Look at Third-Party Service Dependencies 605
– Our heuristics for Africa-dominant websites may have false positives and
negatives. However, the correct way to find Africa-dominant websites would
be to choose the websites which have the largest traffic share from Africa. We
had to use these heuristics because the data for per country traffic share of a
website is not available.
– We inherit the limitations of Kashaf et al. [28] as we use their methodology.
– We describe OCSP revocation checking as a critical dependency from a web-
site’s point of view. However, the online revocation behavior of browsers
differs. For example, many existing browsers circumvent online revocation
checking by using other mechanisms like CRLsets in Chrome [12]. Similarly,
Safari performs online revocation checking in case of revoked certificates.
Many browsers also consider failures in revocation checking as a soft-fail.
Note that some browsers allow users to enable online revocation checking.
Moreover, the system’s TLS stack at the user in some cases always performs
online revocation checking no matter what browser is used [12]. Hence, we
focus on the dependency from the website side while keeping in mind that at
the user end, there may be other accommodations that make the dependency
not critical.
5 Findings
Fig. 3. (a) Critical DNS dependency for top 10K US-visited sites when measured from a
US vantage point is 5% to 7% less than the top 10K Africa-visited websites. This gap in
critical dependency increases to 6% to 10% in the more popular (top 1K) websites. (b)
The percentage of websites that are redundantly provisioned is slightly higher (2%) in
the US-visited websites as compared to the Africa-visited websites. However, when we
look at more popular websites (top 1K), for US-visited , the percentage of redundantly
provisioned websites is 5% to 7% higher than the Africa-visited websites. (c) Critical
CDN dependency for the top 10K US-visited sites is similar to the top 10K Africa-
visited websites. However, for more popular websites, US-visited sites are 4% to 15%
less critically dependent than Africa-visited sites. (d) Critical CA dependency for the
top 10K US-visited sites, when measured from a US vantage point, is 7% to 12% less
than the top 10K Africa-visited websites. This gap in critical dependency increases to
20% to 25% in the more popular (top 1K) websites.
to the popular websites in African countries, making African Internet users more
vulnerable.
Figure 3b illustrates the percentage of redundantly provisioned websites in
DNS. We observe that there is not much difference (2%) between US-visited
websites and Africa-visited websites. However, when we look at more popular
websites (top 1K), the gap increases by 5% to 7% from 2%. At the same time, we
find that the use of private DNS is only 3% to 4% higher in US-visited websites
(not shown) and becomes 2% to 5% when we look at more popular websites
(also not shown). This means that critical dependency in more popular US-
visited websites is reduced because of an increase in redundancy instead of the
A First Look at Third-Party Service Dependencies 607
use of Private DNS. However, for Africa-visited , there is not much significant
increase in redundancy for more popular websites, except South Africa.
In case of CDN dependency, 22%, 18%, 23% and 19% websites use a CDN in
ZA-visited , NG-visited ,RW-visited and KE-visited websites respectively, while
in US-visited , 40% websites use CDN (not shown here). Fig. 3c compares the
critical CDN dependency in US-visited with Africa-visited websites. In the top
10K, critical CDN dependency in US-visited is comparable to the Africa-visited
websites. We find the number of redundantly provisioned websites is also similar
(not shown here). When we look at more popular websites (top 1K), the critical
CDN dependency in US-visited is 4% to 14% less than Africa-visited websites
while the CDN adoption in the top 1K websites is almost double in the US
(44.6%) than African countries (20% to 27%). The use of private CDN remains
negligible in US-visited and Africa-visited websites (not shown here). Moreover,
the percentage of redundantly provisioned websites in the top 1K is 5% to 15%
higher for the US-visited as compared to the Africa-visited websites. The reduced
critical dependency as we move towards more popular websites in US-visited
websites is because of an increase in redundancy.
Figure 3d shows the percentage of websites critically dependent on a CA in
the US-visited and Africa-visited websites. The number of websites that support
HTTPS is similar in US-visited and Africa-visited websites (not shown). Recall
that for CAs, critical dependency is measured in terms of whether a website
supports OCSP stapling or not. We find that US-visited websites are 6% to 12%
less critically dependent on CAs compared to Africa-visited . Moreover, as we
move to more popular websites (top 1K), the gap in critical dependency between
US-visited and Africa-visited websites further increases to 20%–25%. This low
adoption of OCSP stapling may be an indicator of low cyber readiness in Africa.
Furthermore, in the US there have been many efforts to promote OCSP stapling,
particularly by popular CDN providers such as Cloudflare, Amazon Cloudfront,
and Akamai. Since the adoption of CDNs in Africa-visited websites is low, this
could explain the lower adoption of OCSP stapling.
To further investigate the results of Figs. 3a and 3b, Fig. 4a also shows critical
dependency and redundancy of websites in a third-party DNS provider but distin-
guishes them between visited, hosted, dominant, and operated website sets. For
the set of visited websites, the critical DNS dependency is very high 91% to 93%,
and stable across countries. This shows that users in Africa from these countries
are equally vulnerable to the side effects of DNS third-party dependencies. If we
look at the hosted websites, the NG-hosted websites are less critically dependent
compared to other African countries. Concretely, the third-party DNS dependency
is only 84% in NG-hosted websites. This is due to two key reasons. First, many
608 A. Kashaf et al.
Fig. 5. (a) We show the percentage of websites that use CDN in different website sets
for each country. CDN usage is less in the specialized sets such as hosted, dominant,
and operated as compared to the visited set except for ZA. (b) We show the percentage
of critically dependent websites on third-party CDN providers with the percentage of
redundantly provisioned websites stacked on it. The height of the bar stack shows the
percentage of websites using a third-party CDN provider. Critical dependency on CDNs
for Africa-centric websites is less prevalent as compared to critical DNS dependency.
dependency decreases across all website sets for both ZA and NG. This is partly
because of an increase in the number of websites using Private DNS (not shown).
For example, for ZA, third-party dependency decreases by 4% for ZA-visited ,
and ZA-dominant. For ZA-hosted it decreases by 8%, while for ZA-operated it
remains the same. In addition to the increase in private DNS, we also observe
an increase in redundantly provisioned websites. For example, in the case of ZA,
redundantly provisioned websites increase from up to 4% in the top 10K, to
6%-12% in the top 1K. We observe a similar trend in NG, KE, and RW. While
the increase in redundancy for more popular websites is encouraging, it is still
far from ideal. Even for more popular websites, third-party dependencies are
highly prevalent. Across different website sets, we see more encouraging trends.
For example, the hosted websites in the top 1K are far less critically dependent
than the other website sets. However, this trend is only for ZA and NG and does
not appear in KE and RW where it is more similar to the other sets. In NG, this
decrease in critical dependence is primarily because of the use of Private DNS.
For ZA, however, this is because some of the websites using global providers are
using multiple providers, and also because all the websites using TENET South
Africa as DNS, are redundantly provisioned.
Fig. 6. (a) For each website set, we show the change in critical dependency as we
move from more popular (top 1K) websites to less popular (top 10K) ones for ZA and
NG. Critical CDN dependency is lower for more popular websites, as compared to the
less popular ones. (b) The percentage of HTTPS support in websites is very high in
Africa-centric websites, with the exception of the RW-hosted set.
dominant sets have reduced critical dependency compared to the visited set,
while the operated set has increased critical dependency.
Figure 6a shows the change in critical CDN dependency as we move from
more popular (top 1K) websites to less popular websites (top 10K). For example,
for ZA, the critical dependency for more popular websites is 8% to 10% lower
than less popular ones (except the operated set). We observe a similar trend for
RW and KE. This reduction in critical dependency for more popular websites is
because they are more redundantly provisioned. The use of private CDN remains
negligible for the top 1K and top 10K websites (not shown here).
Figure 6b shows the number of websites that support HTTPS. HTTPS adop-
tion is in general very high in Africa-centric websites, which is encouraging.
However, there are a few notable exceptions. For example, HTTPS adoption is
low particularly in the RW-hosted websites. It is also low for NG-hosted and
KE-hosted when compared to the visited websites. For RW, the RW-dominant
website set also has lower HTTPS adoption as compared to other countries.
612 A. Kashaf et al.
6 Provider Concentration
In this section, we first look at the concentration among providers for Africa-
visited websites and use US-visited websites as a baseline. Then we closely look
at Africa, for different website sets.
Fig. 8. The CDF of websites against the number of DNS, CDN, and CA providers for
African countries and the US is shown. (a) Concentration of DNS providers in ZA-
visited and KE-visited is slightly higher than RW-visited , NG-visited and US-visited
websites. (b) The concentration of CDN providers in Africa-visited and US-visited
websites is largely similar, with the concentration in US-visited websites being slightly
higher. (c) The concentration of CA providers in Africa-visited websites is slightly
higher than the US-visited websites.
Fig. 9. Fig. 9a shows the dependency graph of the ZA-visited websites on third-party
DNS providers, Fig. 9b shows the dependency graph of NG-visited websites on third-
party CDNs, and Fig. 9c shows the dependency graph of KE-visited websites on third-
party CAs. The size of a node in the dependency graph is proportional to its in-degree
(signifying a dependency on the provider). We label the concentration C and impact
I of the top 5 providers in terms of the percentage of total websites. (a) Cloudflare
and Amazon serve most of ZA-visited websites and have higher concentration and
impact than other third-party DNS providers. (b) Amazon Cloudfront and Akamai
have a slightly higher concentration and impact as CDN providers for NG-visited . (c)
DigiCert and Let’s Encrypt serve the largest number of KE-visited websites and have
a higher concentration and impact than other CA providers.
Fig. 10. (a) Africa local providers like Afrihost and Xneelo show up in the top 5
DNS providers for ZA-dominant websites. (b) Kenya Education Network provides DNS
service for the largest number of KE-hosted websites. (c) Sectigo, Let’s Encrypt, and
DigiCert provide CA services to almost the same number of NG-dominant websites.
The three providers also have similar concentration and impact.
Figure 9 shows the dependency graph for Africa-visited websites. The size of
a node is proportional to its in-degree which is the number of websites depen-
dent on it. We also label the concentration (C) and impact (I) of each provider
A First Look at Third-Party Service Dependencies 615
7 Discussion
8 Related Work
A huge body of work exists that performs dependency analysis. Some of those
analyze dependencies on the country, or/and ISP. For example, Simeonovski et
al. analyzes dependencies with respect to global scale threats where bad actors
can be a country, an autonomous system, or a service provider like an Email
server, DNS etc.. [55]. Similarly, NSDMiner discovers network service depen-
dencies such as ISPs, from passively observed network traffic [43]. Zembruzki et
al. [62] looks at centralization among hosting providers. Hsiao et al. [25] analyzes
the cyber-autonomy of government websites of the G7 countries. Dell et al. [16]
studies third-party DNS dependency using a passive DNS dataset. WebProphet
measures the internal backend infrastructure of websites for performance [35].
Similarly, Ikran et al. studies dependency chains in third-party web content [26].
618 A. Kashaf et al.
9 Conclusion
10 Availability
Our code is publically available5 . Our work does not raise any ethical concerns.
5
[Link]
A First Look at Third-Party Service Dependencies 619
References
1. Ager, B., Mühlbauer, W., Smaragdakis, G., Uhlig, S.: Web content cartography. In:
Proceedings of the 2011 ACM SIGCOMM Conference on Internet Measurement
Conference, pp. 585–600 (2011)
2. Akanho, Y., Alassane, M., Houngbadji, M., Phokeer, A.: African nameservers
revealed: characterizing DNS authoritative nameservers. In: Zitouni, R., Phokeer,
A., Chavula, J., Elmokashfi, A., Gueye, A., Benamar, N. (eds.) AFRICOMM 2020.
LNICST, vol. 361, pp. 327–344. Springer, Cham (2021). [Link]
978-3-030-70572-5 20
3. Arouna, A., Phokeer, A., Elmokashfi, A.: A first look at the African’s ccTLDs
technical environment. In: Zitouni, R., Phokeer, A., Chavula, J., Elmokashfi, A.,
Gueye, A., Benamar, N. (eds.) AFRICOMM 2020. LNICST, vol. 361, pp. 305–326.
Springer, Cham (2021). [Link] 19
4. Bock, H.: The problem with OCSP stapling and must staple and why certificate
revocation is still broken (2017). [Link]
Problem-with-OCSP-Stapling-and-Must-Staple-and-why-Certificate-Revocation-
[Link]
5. Brinkman, I., Merolla, D.: Space, time, and culture on African/diaspora websites:
a tangled web we weave. J. Afr. Cult. Stud. 32(1), 1–6 (2020)
6. Butkiewicz, M., Madhyastha, H.V., Sekar, V.: Understanding website complex-
ity: measurements, metrics, and implications. In: Proceedings of the 2011 ACM
SIGCOMM Conference on Internet Measurement Conference, pp. 313–328 (2011)
7. Calandro, E., Chavula, J., Phokeer, A.: Internet development in Africa: a content
use, hosting and distribution perspective. In: Mendy, G., Ouya, S., Dioum, I.,
Thiaré, O. (eds.) AFRICOMM 2018. LNICST, vol. 275, pp. 131–141. Springer,
Cham (2019). [Link] 13
8. Country domains: a comprehensive ccTLD list. [Link]
digitalguide/domains/domain-extensions/cctlds-a-list-of-every-country-domain/
9. Chavula, J., Phokeer, A., Calandro, E.: Performance barriers to cloud services in
Africa’s public sector: a latency perspective. In: Mendy, G., Ouya, S., Dioum, I.,
Thiaré, O. (eds.) AFRICOMM 2018. LNICST, vol. 275, pp. 152–163. Springer,
Cham (2019). [Link] 15
10. Chege, K.G.: Measuring internet resilience in Africa, November 2020. [Link]
[Link]/blog/2020/11/measuring-internet-resilience-in-africa/
11. Choffnes, D., Wang, J., et al.: CDNs meet CN an empirical study of CDN deploy-
ments in china. IEEE Access 5, 5292–5305 (2017)
12. Chromium, G.: Crlsets. [Link]
crlsets/
13. Chung, T., et al.: Measuring and applying invalid SSL certificates: the silent major-
ity. In: Proceedings of the 2016 Internet Measurement Conference, pp. 527–541
(2016)
620 A. Kashaf et al.
14. Chung, T., et al.: Is the web ready for OCSP must-staple? In: Proceedings of the
Internet Measurement Conference 2018, pp. 105–118 (2018)
15. Comment, D.S.: Load shedding in South Africa causes cooling system failure
at MTN data center, July 2022. [Link]
load-shedding-in-south-africa-causes-cooling-system-failure-at-mtn-data-center/
16. Dell’Amico, M., Bilge, L., Kayyoor, A., Efstathopoulos, P., Vervier, P.A.: Lean on
me: mining internet service dependencies from large-scale DNS data. In: Proceed-
ings of the 33rd Annual Computer Security Applications Conference, pp. 449–460
(2017)
17. Dig: DNS lookup utility. [Link]
18. Durumeric, Z., Adrian, D., Mirian, A., Bailey, M., Halderman, J.A.: A search
engine backed by internet-wide scanning. In: Proceedings of the 22nd ACM
SIGSAC Conference on Computer and Communications Security, pp. 542–553
(2015)
19. Durumeric, Z., Wustrow, E., Halderman, J.A.: ZMAP: fast internet-wide scanning
and its security applications. In: Presented as Part of the 22nd USENIX Security
Symposium (USENIX Security 13), pp. 605–620 (2013)
20. ExpressVPN: High-speed, secure and anonymous VPN service — expressVPN
(2016). [Link]
21. Global, S.: Latest research shows DDoS attacks up by 300% in Africa since (2019).
[Link]
2019/
22. Globalsign certificate revocation issue, 13 October 2016. [Link]
com/en/status. Accessed 23 May 2020
23. Google: Chrome UX report. [Link]
24. Hilton, S.: Dyn analysis summary of Friday October 21 attack, 26 October
2016. [Link]
Accessed 23 May 2020
25. Hsiao, H.C., et al.: An investigation of cyber autonomy on government websites.
In: The World Wide Web Conference, pp. 2814–2821 (2019)
26. Ikram, M., Masood, R., Tyson, G., Kaafar, M.A., Loizon, N., Ensafi, R.: The
chain of implicit trust: an analysis of the web third-party resources loading. In:
The World Wide Web Conference, pp. 2851–2857 (2019)
27. Comprehensive IP address data, IP geolocation API and database - [Link].
[Link]
28. Kashaf, A., Sekar, V., Agarwal, Y.: Analyzing third party service dependencies in
modern web services: have we learned from the mirai-dyn incident? In: Proceedings
of the ACM Internet Measurement Conference, pp. 634–647 (2020)
29. Kotzias, P., Razaghpanah, A., Amann, J., Paterson, K.G., Vallina-Rodriguez, N.,
Caballero, J.: Coming of age: a longitudinal study of TLS deployment. In: Pro-
ceedings of the Internet Measurement Conference 2018, pp. 415–428 (2018)
30. Krishnamurthy, B., Wills, C.: Privacy diffusion on the web: a longitudinal perspec-
tive. In: Proceedings of the 18th International Conference on World Wide Web,
pp. 541–550 (2009)
31. Krishnamurthy, B., Wills, C., Zhang, Y.: On the use and performance of content
distribution networks. In: Proceedings of the 1st ACM SIGCOMM Workshop on
Internet Measurement, pp. 169–182 (2001)
32. Kumar, D., Ma, Z., Durumeric, Z., Mirian, A., Mason, J., Halderman, J.A., Bailey,
M.: Security challenges in an increasingly tangled web. In: Proceedings of the 26th
International Conference on World Wide Web, pp. 677–684 (2017)
A First Look at Third-Party Service Dependencies 621
33. Kumar, R., Asif, S., Lee, E., Bustamante, F.E.: Third-party service dependen-
cies and centralization around the world (2021). [Link]
2111.12253, [Link]
34. Lerner, A., Simpson, A.K., Kohno, T., Roesner, F.: Internet jones and the raiders
of the lost trackers: an archaeological study of web tracking from 1996 to 2016. In:
25th USENIX Security Symposium (USENIX Security 16) (2016)
35. Li, Z., Zhang, M., Zhu, Z., Chen, Y., Greenberg, A.G., Wang, Y.M.: Webprophet:
automating performance prediction for web services. In: NSDI, vol. 10, pp. 143–158
(2010)
36. Liu, Y., et al.: An end-to-end measurement of certificate revocation in the web’s
PKI. In: Proceedings of the 2015 Internet Measurement Conference, pp. 183–196
(2015)
37. Livadariu, I., et al.: On the accuracy of country-level IP geolocation. In: Proceed-
ings of the Applied Networking Research Workshop, pp. 67–73 (2020)
38. Matic, S., Tyson, G., Stringhini, G.: Pythia: a framework for the automated anal-
ysis of web hosting environments. In: The World Wide Web Conference, pp. 3072–
3078 (2019)
39. Maxmind, L.: Geoip country database
40. Moyo, A.: Africa found wanting on cyber crime preparedness, December 2019.
[Link]
41. Mueller, T., Klotzsche, D., Herrmann, D., Federrath, H.: Dangers and prevalence
of unprotected web fonts. In: 2019 International Conference on Software, Telecom-
munications and Computer Networks (SoftCOM), pp. 1–5. IEEE (2019)
42. Mutiso, R., Hill, K.: Why hasn’t Africa gone digital? Scientific American, August
2020. [Link]
43. Natarajan, A., Ning, P., Liu, Y., Jajodia, S., Hutchinson, S.E.: NSDMiner: auto-
mated discovery of network service dependencies. IEEE (2012)
44. Nikiforakis, N., et al.: You are what you include: large-scale evaluation of remote
Javascript inclusions. In: Proceedings of the 2012 ACM Conference on Computer
and Communications Security, pp. 736–747 (2012)
45. Phokeer, A.: The Gambia’s internet outage through an internet resilience
lens, January 2022. [Link]
outage-through-an-internet-resilience-lens
46. Phokeer, A., Chege, K., Chavula, J., Elmokashfi, A., Gueye, A.: Measuring internet
resilience in Africa (Mira). Internet Soc. (2021)
47. Podins, K., Lavrenovs, A.: Security implications of using third-party resources in
the world wide web. In: 2018 IEEE 6th Workshop on Advances in Information,
Electronic and Electrical Engineering (AIEEE). pp. 1–6. IEEE (2018)
48. PrivateVPN: Privatevpn: The world’s most-trusted private VPN provider. https://
[Link]/
49. Puppeteer: Puppeteer, May 2022
50. Rakshit, S.: geograpy3: Extract countries, regions and cities from a URL or text,
October 2022. [Link]
51. Rao, N.: Bandwidth costs around the world, August 2016. [Link]
com/bandwidth-costs-around-the-world/
52. Roesner, F., Kohno, T., Wetherall, D.: Detecting and defending against third-
party tracking on the web. In: Presented as Part of the 9th USENIX Symposium
on Networked Systems Design and Implementation (NSDI 12, pp. 155–168 (2012)
53. Ruth, K., Kumar, D., Wang, B., Valenta, L., Durumeric, Z.: Toppling top lists:
evaluating the accuracy of popular website lists. In: Proceedings of the 22nd ACM
Internet Measurement Conference, pp. 374–387 (2022)
622 A. Kashaf et al.
Abstract. Web cookies have been the subject of many research studies
over the last few years. However, most existing research does not consider
multiple crucial perspectives that can influence the cookie landscape,
such as the client’s location, the impact of cookie banner interaction, and
from which operating system a website is being visited. In this paper,
we conduct a comprehensive measurement study to analyze the cookie
landscape for Tranco top-10k websites from different geographic loca-
tions and analyze multiple different perspectives. One important factor
which influences cookies is the use of cookie banners. We develop a tool,
BannerClick , to automatically detect, accept, and reject cookie banners
with an accuracy of 99%, 97%, and 87%, respectively. We find banners
to be 56% more prevalent when visiting websites from within the EU
region. Moreover, we analyze the effect of banner interaction on different
types of cookies (i.e., first-party, third-party, and tracking). For instance,
we observe that websites send, on average, 5.5× more third-party cook-
ies after clicking “accept”, underlining that it is critical to interact with
banners when performing Web measurements. Additionally, we analyze
statistical consistency, evaluate the widespread deployment of consent
management platforms, compare landing to inner pages, and assess the
impact of visiting a website on a desktop compared to a mobile phone.
Our study highlights that all of these factors substantially impact the
cookie landscape, and thus a multi-perspective approach should be taken
when performing Web measurement studies.
1 Introduction
Web cookies serve various purposes, like keeping the user logged in or storing a
user’s website settings. However, other than their originally intended use, cook-
ies have been exploited for commercial activities like user tracking and advertise-
ment targeting [1,4,17,18,59]. As a consequence, various data protection laws have
been enacted in the past few years, e.g., the General Data Protection Regulation
(GDPR) [19] in the EU and the California Consumer Privacy Act (CCPA) [8] to
regulate the use of cookies.
Numerous studies shed light on the complex ecosystem of sharing users’ per-
sonal information across various third parties [6,26,43,44,64] and to what extent
GDPR mitigates such abuse [74]. However, most of this research was conducted
from a single or a limited number of vantage points (VPs). Thus, in this work,
we characterize the cookie landscape from diverse geographic locations spanning
six continents—North America, South America, Europe, Africa, Asia, and Aus-
tralia. We complement the existing research by globally analyzing the following
aspects of the cookie landscape:
Interaction with Cookie Banners: Most research involving GDPR does not
consider interaction with cookie banners (e.g., clicking accept/reject buttons)
[1,18,45,74]. Thus, we develop the automated tool BannerClick to automatically
detect, accept and reject cookie banners with an accuracy of 99%, 97%, and 87%,
respectively (see Sect. 3). With BannerClick we automatically detect banners on
about 47% of the Tranco top-10k websites in the EU region whereas in non-EU
regions we find banners on less than 30% of websites (see Sect. 4). Furthermore,
we analyze the difference in the number of cookies before and after interacting
with a cookie banner and find an increase of 5.5× for third-party cookies.
Impact of Geographic Locations: To assess the effectiveness of GDPR, we
compare observed cookies (especially third-party and tracking cookies) between
EU and non-EU vantage points (cf. Section 5). We find that without banner
interaction, 43% of websites send more tracking cookies when accessed from non-
EU regions compared to the EU. Even after accepting a banner, 83% of websites
send more tracking cookies in non-EU countries compared to EU countries. This
percentage increases to 96% when rejecting banners. Our findings indicate a
positive impact of GDPR on reducing the number of TP and tracking cookies.
Consistency of Websites: For cookie analysis, it is essential to observe that
when a website is accessed multiple times, it sets a consistent number of cookies.
If the variation in the number of cookies is high, then one cannot have statisti-
cally significant deductions about cookie characteristics (e.g., number of third-
party cookies). Thus we perform two statistical tests: First, we use the coefficient
of variation to test for intra-location consistency, i.e., how consistent the cookie
landscape is when visiting a website multiple times from the same location. Sec-
ond, we use the Mann-Whitney U test [47] to test for inter-location consistency,
i.e., how consistent is the cookie landscape when visiting a website from differ-
ent locations. Our results show that websites are more consistent within the EU
and that we find the most statistically significant differences between EU and
non-EU countries (cf. Section 6).
Cookie Differences Between Landing and Inner Pages: We also explore
the difference in cookies between the landing and inner pages of a website (see
Sect. 7). As shown by previous work, the structure and content of landing pages
differ substantially from inner pages [3]. Similarly, some websites may not send
cookies on landing pages but may send them on inner pages. Hence, we quantify
Exploring the Cookieverse 625
the difference between cookies on the landing and the inner pages of a website.
For instance, at our United States VP, we find that 32% of websites send more
third-party cookies on the landing compared to inner pages. Similarly, 29.7%
of websites send more third-party cookies on inner pages when accessed from
Germany. Overall, we find that 27.4% and 15.7% of websites exhibit different
third-party and tracking cookie behavior on all VPs. Thus, studies analyzing
only the landing pages may not present the full picture of the cookie landscape.
Cookie Differences When a Website is Accessed from Desktop and
Mobile Browsers: As mobile Web browsing is becoming more popular and
overtaking desktop browsing [23,70], it is important to study its cookie differ-
ences. This is underlined by the fact that websites often have mobile-specific ver-
sions that could lead to a difference in cookies. Thus, we conduct measurements
to quantify the cookie differences between mobile and desktop (cf. Section 8). For
instance, our US East VP sees more third-party cookies on desktop compared to
mobile for 28% of all websites. Contrarily, when accessing websites from Brazil,
28% set more third-party cookies on mobile. Overall, 14.6% and 9% of websites
show different third-party and tracking cookie behavior on all VPs. Therefore,
future research investigating cookie behavior needs to take desktop as well as
mobile websites into account.
Additionally, we analyze the impact of the Brazilian and Californian privacy
laws [8,65] on Web cookies. Since these laws came into effect recently (i.e., in
2020), the analysis of their impact is still in its early days [9,53]. Following
California’s privacy law, other US states are also considering adopting online
privacy laws [78]. Thus it becomes necessary to draw insights from the enactment
of these existing laws on the cookie landscape. In Sect. 9, we show that CCPA
does not have a direct positive impact on Web cookies. Instead, we find that
websites publicly adhering to CCPA tend to send more third-party and tracking
cookies compared to others.
Overall, our measurement study highlights that factors like banner interac-
tion, client location, landing vs. inner pages, and desktop vs. mobile substantially
impact Web cookies. Thus, future research should incorporate these factors when
analyzing the cookie landscape. To encourage reproducibility, we open-source our
code [58] and release our data and analysis scripts [57] at [Link].
2 Background
In this section, we provide background information on different privacy laws and
Web measurement platforms.
The only exception is for “strictly necessary” cookies that are essential for a web-
site’s operation, e.g., storing user credentials. According to the GDPR, websites
must obtain users’ consent concisely and transparently. This results in websites
showing cookie banners, informing users about the cookies being collected by the
websites and third parties. Some banners explicitly ask for users’ consent (e.g.,
with accept or reject buttons), and some assume users’ continued website use
as implied consent. In this research, we study the impact of GDPR on cookie
characteristics across the globe.
California Consumer Privacy Act (CCPA): CCPA is a state statute
enacted by the California state assembly in June 2018. CCPA has similar goals
as GDPR: it intends to protect the privacy of the residents of California. CCPA
enables California residents to know what personal data is being collected (e.g.,
their IP address), whether it is being sold to third parties, and the right to
refuse to share their data. All companies operating in California with at least
an annual revenue of $25 million must comply with the law. Importantly, even
if these companies are not headquartered in California (or even the US), they
still come under the purview of the CCPA.
Brazil’s General Personal Data Protection Law (LGPD): Similar to
the EU, Brazil also introduced a privacy law “Lei Geral de Proteção de Dados
Pessoais” (LGPD) [39,65] that was enforced on September 2020. LGPD is again
similar to GDPR. It also focuses on personal data and users’ rights. Moreover,
it states that website publishers must obtain consent before storing the personal
data of clients (in the form of cookie banners).
To the best of our knowledge, in this work, we take the first step to empirically
quantify the impact of CCPA and LGPD on Web cookies.
Mumbai (India), São Paulo (Brazil), Cape Town (South Africa), and Sydney
(Australia). We select these vantage points to have two VPs inside GDPR coun-
tries (Germany and Sweden), two VPs in the US (of which one is in the CCPA
state California), one in Brazil (that has LGPD), one in Africa, one in Australia,
and one in Asia.
In our measurement study, we use the global Tranco top-10k [42] as target
websites for our analysis. The popularity of these websites is measured con-
sidering the actual Web traffic of users [63]. Other counterparts like the Cisco
Umbrella list [31] and the Majestic Million list [46] are created using indirect
sources like DNS queries and URLs embedded in website ads.
Additionally, for some experiments that require repeated measurements (e.g.,
consistency tests), we use a subset of Tranco top-10k websites; we select three
sets of websites: Tranco top-100, 1001–1100, and 9901–10k. These sets include
websites from the top, middle, and bottom of the Tranco top-10k websites and
hence represent different website tiers. We call this subset the “tiered Tranco
list”. In order to identify a suitable OpenWPM configuration, we perform mul-
tiple small-scale test runs. Table 1 shows an overview of our final large-scale
measurement runs. The longest measurement takes 20 days, in which the Web
can change substantially. In order to keep results comparable, we ensure that
each website is crawled at a similar time from all vantage points. In the case of
failure in one vantage point the website would be excluded from the final result.
Moreover, we run OpenWPM in stateless mode and ensure that the browser does
not block tracking when accessing websites [54].
As already mentioned, we completely automate our measurement campaign
and access the Tranco websites using OpenWPM. We now explain our approach
to detecting and interacting with cookie banners on our target websites.
Due to the EU ePrivacy Directive [20] and GDPR [19], many popular websites
explicitly show cookie banners when accessed from within the EU [48]. These
banners must inform the user about what user data will be collected by the
website (using cookies). Moreover, they must provide a clear choice to users on
whether to accept or reject these cookies.
628 A. Rasaii et al.
To test whether websites respect the users’ consent or not, (1) we detect the
banners, (2) interact with them (e.g., accepting/rejecting the banner policies),
and (3) throughout the whole process collect all cookies. We completely auto-
mate this process by developing our tool called “BannerClick ”. We now explain
how our tool detects and interacts with cookie banners using Selenium browser
automation [67].
To detect banners, we first create a corpus of English words that very likely
exist in banners by manually inspecting 50 random websites from Tranco top-
100 domains. The corpus has eight unique English words i.e., cookies, privacy,
policy, consent, accept, agree, personalized, and legitimate interest. We translate
these words into 11 different languages (German, Swedish, Spanish, Italian, Por-
tuguese, Chinese, Russian, Japanese, French, Turkish, and Persian) and append
the translated words to the corpus, increasing the corpus size to 80 words. We
later show that with these words, we achieve an accuracy of about 99% for
detecting banners.
websites show banners. Using BannerClick , we are able to correctly detect ban-
ners on 513 websites. Therefore, only 5 websites show a banner, but BannerClick
fails to detect them. The reasons include the presence of a shadow DOM [50]
on the website ([Link]) and banners having words not present in our cor-
pus ([Link]). Similarly, only 4 websites do not show any banner, but
BannerClick incorrectly detects a banner. For example, [Link]
has cookie-related words in its DOM, but does not show a banner. Overall, Ban-
nerClick detects banners with more than 99% accuracy and extremely low FPR
(0.008) and FNR (0.009).
2
One can simply detect the <button> tags and search for words inside them. However,
we observe that banner buttons are not always implemented in this manner. Instead,
many websites use other types of tags like <input> or <div> to implement buttons.
630 A. Rasaii et al.
list [52] to identify the domain of (1) the website and (2) the URL in the domain
attribute of the cookies. Then for each of the received cookies, we compare its
domain with the website’s domain. On a successful match, we classify the cookie
as first-party; otherwise, we consider it a third-party.
Next, similar to Götze et al. [28], we use the justdomains blocklist [36] to
identify tracking cookies. This list contains entries from various popular tracking
lists viz. EasyList, EasyPrivacy, AdGuard, and NoCoin filter lists, only if the
complete domain is identified as tracking. If the cookie domain matches one of
the domains in the justdomains list, we classify it as a tracking cookie. To ensure
the correct classification of tracking cookies, we perform a small-scale validation:
We identify the top 100 websites sending the most tracking cookies and then
we manually inspect the tracking cookie domain. We confirm that well-known
tracking domains are indeed sending these cookies (e.g., [Link]).
4
Selenium timeout indicates the duration that Selenium waits for a website to be
loaded by the browser.
5
OpenWPM timeout forces the current website crawl to stop upon expiration. That
is useful, as Selenium freezes during the loading of some websites (e.g., [Link]).
632 A. Rasaii et al.
Fig. 2. Cookie differences between no interaction, accept, and reject from the Germany
VP.
Fig. 3. CMP distribution depending on the Tranco rank from the Germany VP.
noticeable. As for the rejection impact on first-party cookies, we can also see a
slight increase in the number cookies. This might be because of cookies that are
being set to keep the state of rejection for future website access. This is further
corroborated, as we do not see this trend for third-party cookies. Furthermore,
we see that the number of tracking cookies is quite low (near zero) when the
banner is not accepted, which indicates the effectiveness of GDPR to reduce
tracking. Overall, we find that banner interaction has a large influence on the
number of cookies, and it is therefore imperative to use tools like BannerClick
to take banner interactions into account.
While accessing these websites with BannerClick , we also analyze the distri-
bution of Consent Management Providers (CMPs). CMPs are platforms that offer
cookie consent handling as a service, i.e., websites can include a ready-to-use, yet
configurable banner instead of developing their own cookie banner solution. The
IAB Europe Transparency and Consent Framework (TCF) is a GDPR-compliant
consent solution that specifies the overall behaviors of CMPs [33]. As mentioned in
the specification of TCFv2 [32], all CMPs need to implement a __tcfapi() func-
tion which allows third parties to have access to the users’ selected preferences and
act accordingly. In BannerClick we use this function to record the name of the
CMP while crawling a website. We observe that—contrary to the specification—
not all websites with CMP banners actually implement the __tcfapi() function.
This specification violation is not limited to a specific CMP. To obtain a bet-
ter and more comprehensive distribution for CMPs, we additionally incorporate
results from the Never-Consent browser add-on [60] into our data. Never-Consent
leverages custom APIs which some CMPs implement in addition or instead of
__tcfapi(). These custom APIs allow for interaction with CMPs to fetch user-
related data or can even trigger a reject all event.
In Fig. 3 we show the cumulative market share of different CMPs for the
Tranco top-10k websites. As we can see, in total within the top-1k websites
around 13% of websites use CMPs. The CMP deployment remains almost con-
stant with increasing rank, hinting at a consistent CMP deployment between
ranks 2k and 10k. The CMP ecosystem is dominated by four companies
(OneTrust, Quantcast, Sourcepoint, and Google) which are responsible for more
634 A. Rasaii et al.
than half of all CMP banners. Interestingly, we can not find a single website in the
top 46 websites using a CMP and there is a generally much lower CMP deploy-
ment among top-ranked websites (see zoomed-in figure). This can be attributed
to the fact that large Internet companies tend to avoid relying on third parties
for handling privacy-sensitive data.
Throughout our study, we see a slight increase in CMP usage: From 95 web-
sites out of the top-1k in July 2021 to 107 websites in January 2022. Therefore, it
seems that CMPs will continue to play an important role in the cookie ecosystem,
which future research should take into account.
As for other VPs, we see fewer CMPs detected on average. This is due to
some CMPs not implementing their APIs (i.e., __tcfapi() or custom ones),
when they do not show a banner, which happens more for non-EU VPs. There is
also an increase in the share of CMPs in the category “Others”, which underlines
that popular CMPs are less likely to provide APIs if no banner is shown.
Finally, we also compare our CMP results to previous work [29]. Their results
for CMPs following TCFv1 are similar to our results for the new TCFv2 standard.
6
The slightly lower number of rejects in Sweden compared to Germany is due to a
lack of Swedish reject-related words in our corpus.
Exploring the Cookieverse 635
Fig. 5. ECDF plot with the average number of TP (left) and tracking (right) cookies
for websites on which BannerClick is able to click accept only in the EU.
Reject Mode: For the reject mode analysis, we again select websites that again
show banners only in the EU, and for which we are able to click the reject button
(i.e., 23.7% of the total). We find that 87% and 96% of these, set fewer TP and
tracking cookies respectively in the EU after rejecting the banner compared to
the no interaction mode at non-EU VPs. We observe a similar trend for FP
cookies: 72% of these websites set fewer FP cookies in the same scenario.
Overall, our results indicate that GDPR has a positive impact on reducing
the number of TP and tracking cookies, but we do not find any measurable
effect of other privacy laws (i.e., LGPD and CCPA) on TP and tracking cookies.
This observation holds good for banner detection as well; we detect a maximum
number of banners in the EU countries.
From each of the VPs (in eight countries), we measure the intra-location
consistency using the coefficient of variation (CoV) as a metric. The CoV is
calculated by dividing the standard deviation by the mean. The smaller the
CoV, the more consistent the cookie behavior is, when looking at it from each
VP separately. We visit each website of the tiered Tranco list from each location
and then calculate the CoV based on the number of cookies the website sends.
Figure 6 (a) shows the ECDF of CoV for third-party cookies. We can clearly
see two groups of websites in the plot: EU (Germany and Sweden) on the top
and non-EU below that. It seems that when visiting websites from within the
EU, they exhibit a more consistent cookie behavior. However, this difference is
influenced mainly by the number of websites that send exactly zero third-party
cookies which result in a CoV of zero: More websites when visited from within the
EU send exactly zero third-party cookies, compared to when visited from a non-
EU VP. This in turn leads to the ECDF curves of EU countries starting higher
than non-EU countries, exhibiting a shifted, but the similar curve and later even
merging. This is another indicator of the effect of the VP’s geographical location
in combination with GDPR on cookie behavior, as pointed out in Sect. 5. Overall,
we find that 75–80% of websites are consistent with a CoV of less than 0.1 (i.e.,
the standard deviation is at most 10% of the mean). For first-party cookies (not
shown) we see a more similar picture across VPs.
Inter-location Consistency: To find statistically significant differences in the
number of observed cookies depending on the VP location we use the Mann-
Whitney U (MWU) test [47]7 . Again, we crawl websites from the tiered Tranco
list 100 times for each interaction (no interaction, accept, reject) from each VP.
7
The MWU test is a statistical post hoc test, i.e., it allows to find differences in the
cookie distribution between all pairs of VP locations. Our setup fulfills the MWU
assumptions, i.e., all test samples from both groups are independent of each other,
the samples are ordinal. The distributions of both populations are identical under
H0 and not identical under H1 .
638 A. Rasaii et al.
Then we apply the MWU test with Holm p-value correction [30] and choose a p-
value of 0.05 to determine statistical significance. In Fig. 6(b) we show a heatmap
depicting the statistical differences. In the figure, we see two main clusters, i.e.,
EU vs. non-EU and non-EU vs. non-EU. We find that the majority of differences
occur between EU (bold label) and non-EU locations, with more than half of all
website-interaction tuples showing a statistically significant difference. On the
other hand, if both locations are either in the EU or both outside the EU, we
see fewer differences. Moreover, we also confirm that the Tranco rank tier does
not affect the differences. An example of such a website is [Link], which
sends on average 5 TP cookies when visited from Germany or Sweden, 10 TP
cookies from Brazil, and more than 80 TP cookies from other countries.
In conclusion, when visiting a website from a GDPR country compared to a
non-GDPR country, there is a significant difference in third-party cookies being
sent by most websites. For first-party cookies (not shown) we see a similar picture
across VPs, although with fewer differences in total.
remaining ones. Finally, we stop searching for inner pages when either 10 inner
pages are found or a total of 50 links (present on the landing page) have been
tested. We repeat the same process for all tiered Tranco websites.
In total, we obtain 2273 inner pages corresponding to 300 Tranco websites.
We access the set of landing and inner pages from all VPs. Like our other exper-
iments, we visit each webpage (landing and inner) five times in each mode (no
interaction, accept, reject) and record the average number of cookies per web-
page. Figure 7 shows the ECDF of the difference of average TP and tracking
cookies from the ten inner pages compared to the corresponding landing page
(in the no interaction mode). The negative difference on the x-axis (left part of
the figure) corresponds to the fraction of websites where we observe more cookies
on a landing page than on inner pages (shown as Landing > Inner). Zero means
the same number of cookies is found for both categories. Positive values (right
part of the figure) correspond to the fraction of websites where more cookies
are sent on inner pages than the landing page (represented as Inner > Landing).
Figure 7 depicts this difference for three VPs i.e., US East, Brazil, and Germany.
We show only these three VPs because we observe nearly the same trend for US
East and US West; observations in Brazil are quite similar to India, South Africa,
and Australia; the trend in EU countries is almost the same.
At all of our VPs, we find that 12.7% and 8% of websites set more TP
and tracking cookies, respectively, on the landing page than on the inner page
(e.g., [Link], [Link], and [Link]). Looking at VPs separately, the
proportion of such websites is the highest in US East (32% TP and 24% tracking)
and the lowest in Sweden (21% TP) and Germany (12.3% tracking). Moreover,
our analysis reveals that 87% of these websites set at least 10 more TP cookies
on average on the landing page at all locations. One possible explanation for
this trend could be that many websites show more content on the landing page,
include more third-party content, and thus set more TP cookies.
640 A. Rasaii et al.
Similarly, we observe that 14.7% and 7.7% of websites set more TP and track-
ing cookies respectively on inner pages across all VPs (e.g., [Link], [Link]
and [Link]). When investigating each VP separately, the proportion of
such websites is the highest in Germany (29.7% TP) and South Africa (19.3
tracking), and the lowest in US East (22% TP) and Brazil (15.3% tracking). It
is interesting to note that, although GDPR discourages the use of third parties
without consent, a substantial fraction of websites prioritize setting TP cookies
on inner pages. This could also facilitate user profiling [1] as third-party ser-
vices could better characterize users’ viewing habits and choice of content at a
more fine-grained granularity. Overall, our results indicate that studying only
the landing page provides a partial picture of the TP cookies a user might get.
In total, 49.3% and 27.3% of websites set a different number of TP and tracking
cookies respectively on landing and inner pages at all our VPs.
Banners on Inner Pages: We check for banner presence as a potential contribut-
ing factor. Although we find a small number of websites with different banner
behavior (e.g., [Link]/map), we generally see a similar number of
banners on landing and inner pages. Overall, using BannerClick , we detect ban-
ners on 22% (US East), 51% (Germany), and 30% (Brazil) of the landing pages
of the tiered Tranco list. Correspondingly, we detect banners on 25% (US East),
50% (Germany), and 31% (Brazil) of the inner pages.
8
Desktop: “Mozilla/5.0 (X11; Linux x86_64; rv:95.0) Gecko/20100101 Firefox/95.0”;
mobile: “Mozilla/5.0 (Android 12; Mobile; rv:68.0) Gecko/68.0 Firefox/93.0”.
9
Desktop: 1366 × 768; mobile: 340 × 695.
10
In some cases this also changes the URL, e.g., by prepending m. or mobile. to the
domain name.
Exploring the Cookieverse 641
for US East and US West. The data from the VPs in the EU are alike, and the
data from the remaining VPs are similar to each other. Hence, we plot the TP
and tracking cookies per website for US East, Germany, and Brazil representing
their respective classes.
At all VPs, we find that 7.3% and 2.7% of websites set more TP and tracking
cookies, respectively, when visited from a desktop (e.g., [Link], [Link]).
On investigating VPs independently, we find that the proportion of such websites
is the highest in US East (28% set TP and 17% set tracking cookies) and the
lowest in Brazil (17% set TP cookies) and Sweden (9% set tracking cookies).
From our analysis, we note that 7% of websites set at least 10 more TP cookies
when being visited from a desktop from US East. These facts can be attributed
to some websites having more content and hence more embedded third parties
on desktop than on mobile. Many websites, when designed for mobile, decrease
the number of advertisements and limit the content to what is visible without
scrolling. This reduces data usage and improves the user’s viewing experience.
We also observe that 7.3% and 6.3% of websites set more TP and tracking
cookies, respectively when viewed from the mobile environment across all VPs
(e.g., [Link], [Link]). Distinct VP analysis shows that the pro-
portion of such websites is the highest in Brazil (28% set TP cookies, 22% set
tracking cookies) and the lowest in Sweden (15% set TP cookies) and in Ger-
many (10% set tracking cookies). Our analysis shows that 4% of websites set at
least 10 more cookies when visited from mobile from non-EU VPs. As users are
increasingly spending more time on their mobile devices [23], some third parties
seem to be prioritizing placing more cookies when sites are visited from mobile
for better targeting. It becomes imperative that measurements from mobile envi-
ronments also be considered for a real-world analysis of cookies.
Overall, we observe that 14.6% websites set a different number of TP cook-
ies when accessed from desktop and mobile environments at all our VPs.
642 A. Rasaii et al.
Furthermore, our findings show a higher degree of similarity between desktop and
mobile compared to previous work [81], which did not consider banner detection
or interaction at all.
Banners on Websites Browsed from Mobile: We check for banner presence as a
potential contributing factor in this experiment as well. Using BannerClick we
detect a similar number of banners on websites when visited from desktop and
mobile (≈ 21% US East, 46% Germany, and 26% Brazil).
9 Impact of CCPA
The California Consumer Privacy Act (CCPA) came into effect in January 2020.
In the context of CCPA, selling personal information in the form of TP cookies
has been a widely debated topic [5]. Thus, we take the first step to studying how
CCPA-compliant websites deal with third-party cookies. To analyze the cookie
landscape of such websites, we first need to find which websites are overtly
complying with CCPA. For this, we use a straightforward approach. Websites
covered by CCPA must include a conspicuous hyperlink on their homepage with
the text “Do Not Sell My Personal Information” (DNSMPI) [78]. We crawl the
tiered Tranco list and identify websites that contain this hyperlink.11
Out of 300 tiered Tranco websites, we identify that 39 websites contain
DNSMPI links from our US West vantage point, 29 websites from US East,
and 21 from Germany. This indicates that a user’s location impacts whether or
not the DNSMPI link is shown. Interestingly, this applies to different locations
within the US as well, i.e., we see 11 websites that only show the DNSMPI link
to clients from California but not when visiting the website from the US East.
To observe the impact of CCPA on TP cookies, we compare the TP cookies of
websites containing DNSMPI links with websites that do not include said links.
We select our US West (i.e., California) VP for this analysis. First, we classify
the 39 websites with DNSMPI links into three sets belonging to Tranco top-
100, 1001–1100, and 9901–10k, respectively. For instance, we obtain 12 websites
that belong to the first set. Thus, to have a fair comparison, we randomly select
the same number of websites without a DNSMPI link from the Tranco top-
100 websites only. We repeat the same process for the other two sets as well.
In the end, we compare websites in the same Tranco rank tier. In total, we
compare 39 websites with DNSMPI links with the same number of websites
without DNSMPI links. This approach ensures that differences in TP cookies
are not due to differences in Tranco rank.
Similar to previous experiments, we crawl each website five times and record
the number of TP cookies. Figure 9 illustrates the variation in average TP cookies
for DNSMPI and non-DNSMPI websites (without cookie banner interaction).
We can see that websites without DNSMPI (blue line) set a lower number of
TP cookies than the websites with DNSMPI (orange line). For example, 42%
11
We use 8 different phrases for searching DNSMPI hyperlinks (e.g., “do not sell my
info”) as suggested by Van Nortwick et al. [78].
Exploring the Cookieverse 643
10 Discussion
Cookie Banner Automation: Since GDPR [19] and similar privacy legisla-
tion came into effect, cookie banners have become more and more prevalent
on the Web. Moreover, during our measurements, we also see a wide variety
of different banners. This not only makes automated detection and interaction
more challenging for research purposes, but it also hinders browser and exten-
sion developers to effectively interact with banners in an automated fashion.
These often rely on manually curated rules, do not have the option to reject
cookie consent [38], or are no longer maintained [60]. Efforts to offer a gen-
eral easy-to-use mechanism to refuse all tracking cookies such as HTTP’s “Do
Not Track” header [49], have not been adopted by the advertising industry and
were therefore abandoned. The deployment of Consent Management Platforms
(CMPs) could be leveraged as a standardized API for application developers to
automate banner interaction. Unfortunately, we confirm previous findings [29]
that many CMP websites do not properly implement these standardized APIs,
which makes it difficult to make use of them. Moreover, CMPs are almost non-
existent for very popular websites, which again leads to a lack of standardization
644 A. Rasaii et al.
potential for websites most visited by users. Additionally, many cookie banners
make it purposefully difficult for people to reject all cookies [68]. As a prominent
example, Google has been fined 150 million € for not providing users a choice to
reject all cookies and was consequently forced to update their cookie banner [15].
All these factors hinder effective banner automation and it is unlikely that the
situation will improve without a joint push by browser developers, advertising
companies, and lawmakers.
Looking Ahead: In order to improve user privacy, browser vendors have
recently started to block third-party cookies at various degrees. Mozilla intro-
duced “Enhanced Tracking Protection” in 2019 [80] and is now moving towards
completely isolated cookie stores per website [51]. Apple has introduced by-
default TP cookie blocking in 2020 [24,71]. Google has long touted its desire
to get rid of TP cookies and proposed a myriad of different possible replace-
ments [10,25,66,72,77]. Getting rid of TP cookies is likely not the end of user
tracking, as different techniques such as Local Storage, IndexedDB, Web SQL,
or browser fingerprinting [41] can easily replace TP cookie functionalities [12].
Finally, privacy regulations such as GDPR are not specifically limited to cookies,
but require informed consent for any shared user data, irrespective of the used
technology. Cookie banners will therefore likely remain a prominent sight in the
future, even if the underlying technology might change.
Limitations: Even though we cover a wide range of factors in our work, there are
natural limitations to our approach. First, since our banner detection approach
leverages words from 12 languages, we might not be able to detect banners on
websites using other languages. Second, we use OpenWPM which uses the Firefox
browser to access websites. Websites could exhibit different cookie behavior when
being accessed from a different browser, such as Chrome or Safari. Third, we
solely focus on HTTPS when accessing websites. Since many browsers use an
HTTPS-first approach and most websites do support HTTPS [22], we think this
focus is warranted. Websites can also be accessed via QUIC, which is not yet
widely deployed [82], and we thus do not consider it in our study. Fourth, to
classify third-party cookies as tracking cookies, we rely on tracking cookie lists.
In order to limit false positive tracking classifications, we use the conservative
approach by Götze et al. [28]. Therefore, our identified tracking cookies serve as
a lower bound. Fifth, to obtain the mobile version of the websites, we modify the
OpenWPM user agent and screen size (see Sect. 8). Although for most websites,
we see the mobile version, for some websites these simple changes are not enough
to load the mobile version [81].
11 Related Work
To regulate the use of cookies, various data protection laws such as the GDPR
[19] in the EU or CCPA [8] in California have been enacted in the last years.
A large body of previous work attempts to quantify the efficacy of such laws.
Dabrowski et al. [13] reported less persistent cookie usage for EU users in com-
parison to US users with Alexa top-100k websites as targets. On the contrary,
Exploring the Cookieverse 645
Sanchez et al. [61] claimed that the US appears to approach cookie regulations
similar to the EU. We do, however, observe a lower number of TP cookies in the
EU when compared to non-EU VPs (see Sect. 5).
Furthermore, to check whether website publishers adhere to the EU cookie
laws, Trevisan et al. [74] developed the tool “CookieCheck” [75]. They reported
that half of the websites they tested (≈ 35k) from an Italian VP, violate the
law i.e., they install profiling cookies12 before the user’s consent. In contrast,
we observe that in the no-interaction mode, “only” about 30% of websites set
tracking cookies at our EU VPs. This might indicate that website publishers are
adhering more to privacy laws over time.
While studying tracking, Iordanou et al. [34] identified the geographic loca-
tions of the tracking servers. They found that around 90% of the tracking flows
originating in the EU terminate at tracking servers hosted within the EU itself.
Additionally, there are multiple measurement studies that highlight how trackers
use cookies for user profiling [6,17,21,26,43,44,64]. As an example, Englehardt
et al. [18] demonstrated that adversaries could reconstruct up to 73% of a user’s
browsing history using only the collected cookies.
Linden et al. [45] took a different direction; they conducted a longitudinal
study to assess privacy policies adopted by website publishers before and after
GDPR went into effect. They reported that GDPR has a positive impact on
privacy policies. Post-GDPR, not only the visual (and textual) representation of
policies have improved, but the coverage of important topics e.g., data retention,
has also increased. Degeling et al. [14] also made similar observations i.e., after
GDPR, many websites have added and updated their privacy policies and now
show cookie banners to the users. Sørensen et al. [69], rather than analyzing
the privacy policies, found that after the introduction of GDPR, the number of
third parties on EU websites has declined. They noted, however, that it cannot be
concluded with certainty that this decline is solely due to GDPR. Kretschmer et
al. [40] conducted a comprehensive survey of the existing research (> 70 research
papers), describing the legal as well as technical aspects of GDPR. They report
that the enactment of GDPR has resulted in a decline in third-party tracking,
increase in cookie banners, and privacy policies in the EU region.
Santos et al. [62] studied cookie banners to analyze how clearly they explain
privacy policies. They manually analyzed 400 cookie banners on English language
websites that are popular in the EU. They report that 61% of banners used
vague language and violated the specificity purpose. Utz et al. [76] rather than
only focusing on the text of the banners, also studied other factors that could
influence user consent decisions (e.g., positioning of the banners on the website).
The authors partnered with an e-commerce website in Germany and reported
that changing the position of the banner or the text has a significant impact on
the users’ consent decisions. For instance, if the banner is shown in in the lower
left part of the screen, users are more likely to interact with it.
12
These are cookies that are managed by Web trackers to identify users and are clearly
subject to explicit consent according to the GDPR.
646 A. Rasaii et al.
More recently, Chen et al. [9] conducted a user survey of Californian con-
sumers to study, to analyze how well they understand privacy policies of popular
websites. They reported a significant variance in how websites interpret CCPA.
Thus, privacy policy disclosures (mandated by CCPA) seem ambiguous to end-
users. To this end, Connor et al. [53] performed a study to specifically analyze
how websites implement “right to opt-out of the sale of users’ personal informa-
tion”. They observed that websites implement this mandate in ambiguous ways,
which deters the users’ motivation to opt-out.
Finally, other research specifically analyzes cookie banners themselves e.g.,
how clearly they specify privacy policies [62] or the impact of banner location on
user consent [76]. Jha et al.’s [35] work is closest to our research. Similar to our
work, the authors also attempted to interact with the banners in an automated
manner to observe differences in cookies. However, their tool only accepts the
privacy policies (of the banner), whereas our tool BannerClick has the capability
to accept as well as reject a banner’s consent.
12 Conclusion
In this paper, we performed a multi-perspective analysis of Web cookies. We
developed BannerClick to automatically detect, accept, and reject cookie ban-
ners with an accuracy of 99%, 97%, and 87%, respectively. Then we ran mea-
surements from 8 geographic locations on 5 continents and identified substantial
differences between these vantage points. We found 56% more banners on web-
sites when visited from an EU vantage point. Moreover, we quantified the effect of
banner interaction: websites sent 5.5× more third-party cookies on average after
clicking “accept”. Accordingly, we observed a similar trend for tracking cookies
as well. Finally, we also identified differences in cookies depending on the vis-
ited page on a website (inner vs. landing) and the client platform (desktop vs.
mobile).
DOM for accept-related words and, on a successful match, attempts to click the
element containing the word. As a result, it can encounter multiple failures before
actually clicking the desired accept button on the banner. On the contrary, Ban-
nerClick first detects the banner and searches for words contained within the
banner. Third, BannerClick can click on accept related elements in 12 popu-
lar languages whereas, Priv-Accept only searches for English words. There are
other differences, e.g., BannerClick looks for banners within the iframes, but
Priv-Accept ignores iframes.
We compare both tools on the Tranco top-1k websites. With Priv-Accept,
we can click accept on 451 websites, whereas with BannerClick , the number is
430. Websites where Priv-Accept could click accept but not BannerClick are
66, and vice-versa 59 websites. The vast majority of the former set are web-
sites that do not show an explicit accept option. These are not considered to
be explicit accepts by BannerClick , however Priv-Accept considers them. Addi-
tionally, Priv-Accept also clicks on the incorrect accept button for 11 websites.
The latter group contains websites where Priv-Accept is unable to identify the
correct button, BannerClick detects banners in iframes, or the website is in a
non-English language.
References
1. Acar, G., et al.: The web never forgets: Persistent tracking mechanisms in the wild.
In: CCS 2014
2. Acar, G., et al.: FPDetective: dusting the web for fingerprinters. In: CCS 2013
3. Aqeel, W., et al.: on landing and internal web pages: the strange case of Jekyll and
Hyde in web performance measurement. In: IMC 2020
4. Bangera, P., Gorinsky, S.: Ads versus regular contents: dissecting the web hosting
ecosystem. In: IFIP Networking 2017
5. Bateman, R.: CCPA: does Using Third-Party Cookies Count as Selling Personal
Information? [Link]
personal-information/
6. Cahn, A., et al.: An empirical study of web cookies. In: WWW (2016)
7. Chameleon Crawler contributors: Chameleon crawler. [Link]
rds/chameleon
8. Chau, E., Hertzberg, R.: California consumer privacy act. [Link]
[Link]/faces/[Link]?bill_id=201720180AB375
9. Chen, R., et al.: Fighting the fog: evaluating the clarity of privacy disclosures in
the age of CCPA. In: WPES 2021
10. Chromium blog: potential uses for the privacy sandbox. [Link]
org/2019/08/[Link]
11. Common crawl: common crawl. [Link]
12. Cookiebot: google ending third-party cookies in Chrome. [Link]
com/en/google-third-party-cookies/
13. Dabrowski, A., et al.: Measuring cookies and web privacy in a post-GDPR world.
In: PAM 2019
14. Degeling, M., et al.: We value your privacy... now take some cookies: measuring
the GDPR’s impact on web privacy. In: NDSS 2019
Exploring the Cookieverse 649
15. Dillet, R.: Google to update cookie consent banner in Europe following fine.
[Link]
europe-following-fine/
16. Durumeric, Z., et al.: ZMap: fast internet-wide scanning and its security applica-
tions. In: USENIX Security 2013
17. Englehardt, S., Narayanan, A.: Online tracking: a 1-million-site measurement and
analysis. In: CCS 2016
18. Englehardt, S., et al.: Cookies that give you away: The surveillance implications
of web tracking. In: WWW 2015
19. European Commission: the general data protection regulation (GDPR) in EU.
[Link]
20. European Parliament: European ePrivacy directive. [Link]
dir/2009/136/2020-12-21
21. Falahrastegar, M., et al.: The rise of panopticons: examining region-specific third-
party web tracking. In: TMA 2014
22. Felt, A.P., et al.: Measuring HTTPS adoption on the web. In: USENIX Security
2017
23. Gibbs, S.: Mobile web browsing overtakes desktop for the first time. https://
[Link]/technology/2016/nov/02/mobile-web-browsing-desktop-
smartphones-tablets
24. GlobalData thematic research: apple block on third party cookies will change dig-
ital media forever. [Link]
25. Goel, V.: Get to know the new topics API for privacy sandbox. [Link]
products/chrome/get-know-new-topics-api-privacy-sandbox/
26. Gonzalez, R., et al.: The cookie recipe: untangling the use of cookies in the wild.
In: TMA (2017)
27. Google: CLD3 on GitHub. [Link]
28. Götze, M., et al.: Measuring web cookies in governmental websites. In: WebSci
(2022)
29. Hils, M., et al.: Measuring the emergence of consent management on the web. In:
IMC (2020)
30. Holm, S.: A simple sequentially rejective multiple test procedure. Scand. J. Statist.
6, 65–70 (1979)
31. Hubbard, D.: Cisco umbrella 1 million. [Link]
14/cisco-umbrella-1-million/
32. IAB Europe: What is TCF v2.0? [Link]
33. IAB Europe: what is the transparency & consent framework (TCF)? https://
[Link]/transparency-consent-framework/
34. Iordanou, C., et al.: Tracing cross border web tracking. In: IMC (2018)
35. Jha, N., et al.: The internet with privacy policies: measuring the web upon consent.
TWEB 16(3), 1–24 (2021)
36. Justdomains: Domain-only filter lists. [Link]
37. Kenneally, E., Dittrich, D.: The Menlo report: ethical principles guiding informa-
tion and communication technology research. SSRN (2012). [Link]
2139/ssrn.2445102
38. Kladnik, D.: I don’t care about cookies. [Link]
eu/
39. Koch, R.: What is the LGPD? Brazil’s version of the GDPR. [Link]
gdpr-vs-lgpd/
40. Kretschmer, M., et al.: Cookie banners and privacy policies: measuring the impact
of the gdpr on the web. TWEB 15(4)
650 A. Rasaii et al.
41. Laperdrix, P., et al.: Browser fingerprinting: a survey. TWEB 14(2), 1–33 (2020)
42. Le Pochat, V., et al.: Tranco: a research-oriented top sites ranking hardened against
manipulation. In: NDSS (2019)
43. Lerner, A., et al.: Internet jones and the raiders of the lost trackers: an archaeo-
logical study of web tracking from 1996 to 2016. In: USENIX Security (2016)
44. Li, T.C., et al.: Trackadvisor: taking back browsing privacy from third-party track-
ers. In: PAM (2015)
45. Linden, T., et al.: The privacy policy landscape after the GDPR. PoPETS (2020)
46. Majestic: the majestic million. [Link]
47. Mann, H.B., Whitney, D.R.: On a test of whether one of two random variables is
stochastically larger than the other. Annal. Math. Stat. 18(1), 50–60 (1947)
48. Matte, C., et al.: Do cookie banners respect my choice? measuring legal compliance
of banners from IAB Europe’s transparency and consent framework. In: S&P (2020)
49. Mayer, J., et al.: Do not track: a universal third-party web tracking Opt Out.
[Link]
50. Mozilla: MDN: using shadow DOM. [Link]
Web/Web_Components/Using_shadow_DOM
51. Mozilla: new year, new privacy protection for firefox focus on android. https://
[Link]/en/mozilla/new-privacy-protection-for-firefox-focus-on-android/
52. Mozilla: public suffix list. [Link]
53. O’Connor, S., et al.: (Un) clear and (In) conspicuous: the right to opt-out of sale
under CCPA. In: WPES (2021)
54. OpenWPM: OpenWPM not using tracking blocking. [Link]
openwpm/OpenWPM/issues/101
55. OpenWPM: openWPM stateful vs stateless crawls. [Link]
OpenWPM/blob/master/docs/Confi[Link]#stateful-vs-stateless-crawls
56. Partridge, C., Allman, M.: Ethical considerations in network measurement papers.
CACM 59(10), 58–64 (2016)
57. Rasaii, A.: Analysis scripts and raw data for BannerClick web measurements.
[Link]
58. Rasaii, A.: BannerClick on GitHub. [Link]
59. Razaghpanah, A., et al.: Apps, trackers, privacy, and regulators: a global study of
the mobile tracking ecosystem. In: NDSS (2018)
60. Robin, M.K.: Never-Consent on GitHub. [Link]
Consent/
61. Sanchez-Rola, I., et al.: Can i opt out yet? GDPR and the global illusion of cookie
control. In: CCS (2019)
62. Santos, C., et al.: Cookie banners, what’s the purpose? analyzing cookie banner
text through a legal lens. In: WPES 2021 (2021)
63. Scheitle, Q., et al.: A long way to the top: significance, structure, and stability of
internet top lists. In: IMC (2018)
64. Schelter, S., Kunegis, J.: Tracking the trackers: a large-scale analysis of embedded
web trackers. In: ICWSM (2016)
65. Schreiber, A.: Right to privacy and personal data protection in Brazilian law. In:
Data Protection in the Internet (2020)
66. Schuh, J.: Building a more private web. [Link]
chrome/building-a-more-private-web/
67. Selenium: browser automation using selenium. [Link]
68. Soe, T.H., et al.: Circumvention by design-dark patterns in cookie consent for
online news outlets. In: NordiCHI (2020)
Exploring the Cookieverse 651
69. Sørensen, J., Kosta, S.: Before and after GDPR: the changes in third party presence
at public and private European websites. In: WWW (2019)
70. Statista: percentage of mobile device website traffic worldwide from 2015 to 2021.
[Link]
from-mobile-devices/
71. Statt, N.: Apple updates Safari’s anti-tracking tech with full third-party
cookie blocking. [Link]
intelligent-tracking-privacy-full-third-party-cookie-blocking
72. Temkin, D.: Charting a course towards a more privacy-first web. [Link]
google/products/ads-commerce/a-more-privacy-first-web/
73. Trevisan, M.: Priv-Accept on GitHub. [Link]
74. Trevisan, M., et al.: 4 years of EU cookie law: results and lessons learned. PoPETS
2019
75. Trevisan, M., et al.: Cookiecheck tool on github. [Link]
CookieChecker/CookieCheckSourceCode
76. Utz, C., et al.: (un) informed consent: Studying GDPR consent notices in the field.
In: CCS (2019)
77. Vale, M.: Privacy, sustainability and the importance of “and”. [Link]
products/chrome/privacy-sustainability-and-the-importance-of-and/
78. Van Nortwick, M., Wilson, C.: Setting the bar low: are websites complying with
the minimum requirements of the CCPA? In: PoPETS 2022
79. WebTAP at Princeton University: studies using OpenWPM. [Link]
[Link]/software/
80. Wood, M.: Firefox blocks third-party tracking cookies and Cryptomining by
default. [Link]
party-tracking-cookies-and-cryptomining-by-default/
81. Yang, Z., Yue, C.: A comparative measurement study of web tracking on mobile
and desktop environments. In: PoPETS (2020)
82. Zirngibl, J., et al.: It’s over 9000: analyzing early QUIC deployments with the
standardization on the horizon. In: IMC (2021)
Quantifying User Password Exposure
to Third-Party CDNs
1 Introduction
Content Distribution Networks (CDNs) [37,45] play an important role in improv-
ing the performance and security of web services. A CDN caches web pages
at servers near end users to reduce retrieval latency. It also blocks malicious
requests to defend a web server against various attacks [20]. Currently, many
websites employ CDNs provided by third-party companies such as Akamai [1],
Cloudflare [3], and Fastly [4].
However, third-party CDNs introduce a considerable security and privacy risk
when they serve websites that enable HTTPS [15,17]. HTTPS uses a certificate
to certify the domain name of a website. Thus, to make the web pages appear
as if they come from the original site, a website has to share its TLS private key
[15] or TLS session keys [51] with the CDN. In both cases, a third-party CDN
can observe the content of all connections between a website and its users.
In this work, we aim to raise awareness of this security and privacy risk
and quantify its severeness from a user’s perspective. We choose to measure the
extent to which users’ website login passwords are exposed to CDNs due to the
HTTPS key sharing practice. Although prior research has shown that private
c The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 652–668, 2023.
[Link]
Quantifying User Password Exposure to Third-Party CDNs 653
key sharing is prevalent on the Internet [15] and HTTPS termination weakens
connection security of a great portion of the Internet [17], it is not clear whether
websites have taken preliminary countermeasures such as client-side encryption
(see Sect. 2) to protect users’ passwords in the case of a passive attacker.
We conduct a measurement on Alexa top 50K sites [2] to quantify pass-
word exposure to CDNs during the user login procedures. We also measure the
deployment of client-side password encryption on websites to understand web-
sites’ treatment of users’ passwords. Such a large-scale measurement is techni-
cally non-trivial, because we need to automate the login procedures on websites
with diverse structures to inspect login requests. Thus, we design and implement
a framework for automatic login. The framework can detect login elements on a
website and collect login requests when it submits credentials to websites.
Our main contributions and findings can be concluded as the following:
– We propose an open-source framework for automatic login1 , which can be
applied to other research such as the measurement of authentication methods.
– Our measurement presents that 33.0% of websites that employ CDNs and
contain login entrances expose users’ passwords in plaintext to their CDNs.
– We find that two popular CDN providers, Cloudflare and Akamai, can observe
users’ passwords from 44% and 25% of their customers, respectively.
– We find prevalent password exposure in most website categories, including
websites whose user accounts should be carefully protected, such as web-
sites related to finance and health. Retail websites substantially benefit from
CDNs, but most of them (58%) expose passwords to CDNs.
– Our result shows that less than 17% of the websites encrypt users’ passwords
when transferring login requests to CDNs, and the top 1,500 websites are
more likely to adopt client-side password encryption.
Overall, our measurement points out potential security issues caused by pass-
word exposure to CDNs. Even though websites trust CDNs, users may concern
about their privacy when CDNs can monitor their private data including pass-
words. Moreover, CDNs have never been secure enough. Prior work has shown
that an attacker can trick some CDNs to cache and reveal other users’ private
data [19,38,39]. Thus, private data leakage to CDNs may turn into a disaster
when attackers or malicious insiders exploit vulnerabilities of CDNs.
2 Background
In this section, we briefly introduce CDNs and HTTPS, and we analyze the
security issues when a website with HTTPS employs a CDN. We also discuss
two countermeasures adopted by websites in practice to address such issues.
responds to the client with cached content. If the requested content is not cached,
the edge server may fetch the content from the origin server which is hosted by
the website (the CDN’s customer ) and is the initial source of all content. CDNs do
not cache private data, as they are usually dynamic.
Modern CDNs are used not only to speed up page loading but also to pro-
vide an effective shield against attacks such as DDoS and code injections [20].
A CDN enlarges the serving capability of its customers to prevent volumetric
DDoS attacks. It also applies techniques such as IP blocking and rate limiting
to block attacks when DDoS happens. For example, Akamai protected its cus-
tomers from 38,905 separate DDoS attacks from 2014 to 2019 [50]. CDNs also
inspect the content of requests and use Web Application Firewall (WAF) to filter
out malicious requests such as XSS injection [59] and SQL injection [24].
Unfortunately, CDNs have become a source of vulnerabilities in the HTTPS
ecosystem in recent years [15,17]. If a website employs a CDN to represent it
to respond to clients’ HTTPS requests, it has to share its private key with the
CDN. With the private key, the CDN can build HTTPS connections with clients,
and clients cannot differentiate between the CDN and the origin server. When a
client requests for private data, the CDN will forward the request by terminating
the HTTPS connection and building another HTTPS connection with the origin
server. Therefore, the CDN becomes a man in the middle when a user’s private
data are transmitted between the client and the origin server [15].
in Sect. 3. Besides, secure public key delivery is non-trivial when HTTPS con-
nections are already intercepted by a CDN [35]. Delivering another certificate
differing from the HTTPS certificate is useless, because a website has to use
JavaScript to conduct encryption in current browsers, and the JavaScript code
cannot obtain the root certificates of a client to verify a certificate. Without
a certificate, if the public key is delivered by a CDN, a CDN with an active
attacker (defined in Sect. 3) inside can launch the man-in-the-middle attacks by
replacing the public key. If the public key is delivered by the origin server, the
origin server is exposed to the public and under the threat of DDoS. In practice,
websites use an asynchronous JavaScript call [6] to request for a public key from
the origin server and encrypt passwords by JavaScript code.
Despite the defects of these two methods, they preserve users’ privacy to
some extent. Moreover, if the origin server builds its own DDoS defense or a
CDN is assumed to be a passive attacker, these two countermeasures can provide
sufficient protection. However, it is unclear about the deployment of these two
countermeasures on websites. Thus, we investigate the password exposure to
provide a profile of their deployment.
3 Threat Model
We use the threat model proposed by the prior work [35]. We consider the pri-
vate data in a website as the data can only be accessed by a authenticated user.
The users can be authenticated by the traditional password, one-time password
(OTP), OAuth [25], certificates, etc. The credentials for authentication are con-
sidered as private data as well. We focus on the measurement of the traditional
password in this paper.
We considered two types of attackers defined in the prior work [35].
– Passive attacker: A CDN behaves honestly to serve the requests, but an
attacker inside the CDN may eavesdrop on the transmitted messages. For
example, a malicious administer of a CDN cannot change the CDN’s behavior
but may peek at the transmitted traffic and record users’ passwords. Client-
side encryption can protect users’ password under a passive attacker.
– Active attacker: An attacker insider CDN may launch arbitrary attacks
including eavesdropping and tampering. Thus, it is more capable than a pas-
sive attacker. For example, a CDN may modify or corrupt the cached HTML
or JavaScript to disable the client-side encryption so that it can observe users’
passwords in the login requests. This may happen when attackers exploit a
vulnerability of a CDN. As previously mentioned, CDN-bypassing can defend
against an active attacker inside a CDN, but it introduces the vulnerability
of DDoS to the origin server.
4 Method
To detect the password exposure, we should inspect a website’s login request and
the destination. Thus, we need a framework for automatic login in a large-scale
656 R. Xin et al.
5 Password Exposure
We only consider HTTPS-enabled websites because a website without HTTPS
apparently contains major vulnerabilities. In Alexa top 50K sites [2], 42,502 of
them enable HTTPS. We run the framework to automatically log into these
websites. If the framework submits the fake credentials to a website, we consider
it performs a login. The framework performs 17,111 logins in total. In this paper,
we focus on these 17,111 websites and call them “login-detected websites”.
We detect CDNs employed by these websites according to Sect. 4. Our result
shows that 12,451 websites employ CDN service, and we call them “CDN-enabled
websites” in this paper. By inspecting their login procedures, we find that 4,114
Quantifying User Password Exposure to Third-Party CDNs 657
websites send the login requests with users’ passwords in plaintext or Base64
encoding to CDNs. We denote these websites as “password-exposed websites”.
We discovered that 33% of CDN-enabled websites expose users’ passwords to
CDNs, demonstrating a potential privacy issue. In this section, we present the
results in detail.
Since our framework may fail to detect the login forms of some websites, the
dataset of login-detected websites is a sample set of all websites that enable
logins. We first investigate the distribution of these samples over rankings.
Figure 1a shows the distribution of login-detected websites. A linear relation-
ship between the CDF and ranking shows a uniform distribution of the websites.
Therefore, the logins detected by our framework are unbiased in the rankings.
To investigate the relationship between websites’ rankings and their prefer-
ence for password exposure, we divide the rankings into 100 intervals. For an
interval Ij , it contains 500 websites ranking in the range of [1+500∗(j−1), 500∗j].
For each interval, we count the password-exposed websites and the CDN-enabled
websites, and we compute the percentage of password-exposed websites in CDN-
enabled websites.
Figure 1b presents the percentage variation across the intervals. Given the
result of unbiased detection in Fig. 1a, we can examine the distribution of pass-
word exposure on website rankings through Fig. 1b. Even though some fluctua-
tions exist, the percentages are overall above 20%, meaning that the password
exposure is common across all rankings. Besides, we can find that the most pop-
ular websites in the first two intervals have relatively low password exposure
percentage. It is because that the top websites are more likely to deploy defense
mechanisms, which can be justified by our analysis in Sect. 6.
658 R. Xin et al.
Table 1. Distribution across CDN providers (a) and website categories(b). The “Per-
cent” column denotes the percentage of password-exposed websites in CDN-enabled
websites. We mark notable data with red color.
We also consider how password-exposed websites are distributed among the CDN
providers. Table 1a presents the number of password-exposed websites in each
CDN provider. As shown in the table, Cloudflare and Akamai are the two most
popular CDNs in the world, and they observe the most users’ passwords from
their customers’ requests. More than 40% of Cloudflare’s customer websites in
our dataset share users’ passwords to Cloudflare, and Akamai observes passwords
from 25% of its customers. Besides, 66% of websites that use Incapsula expose
passwords to the CDN. Some CDNs only observe a small fraction of sensitive
traffic, such as Highwinds and Edgecast.
Compared to the other CDN providers, a much larger portion of Cloud-
flare and Incapsula customers are affected by password exposure. For Cloud-
flare, the reason may be the difference in request redirection methods. Cloud-
flare uses anycast for request redirection by default [14], while the other CDNs
use DNS redirection [37,45]. As discussed in [35], to enable anycast redirec-
tion, a website needs to use Cloudflare as the DNS provider. Such a practice
will transfer a website’s all DNS records to Cloudflare DNS service, includ-
ing the resolution to the domain of the login request (e.g. DNS A record of
[Link]). Cloudflare will conduct anycast redirection for the trans-
ferred domains by default. Therefore, the login request is very likely to be ter-
minated by Cloudflare. We verify this inference by checking the DNS provider
Quantifying User Password Exposure to Third-Party CDNs 659
6 Countermeasures
In this section, we first present the measurement of the countermeasures against
password exposure used by current websites. We also discuss possible counter-
measures that websites and users can adopt.
We need further research on the solutions. As presented in Sect. 2.2, the pre-
liminary strategies of CDN bypassing and client-side encryption can be eas-
ily deployed but contain vulnerabilities. Proposed techniques such as Keyless
SSL [18,40,51], certificate delegation [34], and mcTLS [42] are ineffective in
preserving user privacy. The SGX-based solutions [26,41] can provide compre-
hensive protection, but it is hard to be deployed on CDNs. InviCloak [35] can
achieve the goal of DDoS defense, privacy protection, and instant deployment
simultaneously, but it disables the Web Application Firewall (WAF) of CDNs.
Therefore, further research on this area is critical to a more secure Internet.
We recommend users adopt two-factor authentication. Two-factor authentication
provides additional protection for an account even when the password is stolen
by a hacker. Adopting OAuth is debatable as it may lead to the single point of
failure although it prevents password exposure as discussed in Sect. 6.2.
Websites should adopt preliminary defense. The results shows that many web-
sites do not apply the minimal defense against password exposure. Despite the
preliminary strategies are vulnerable to some attacks, they provide basic protec-
tion for users’ privacy. Since it is acceptable to assume a passive CDNs in most
cases, the client-side encryption usually provides a sufficient protection.
CDN providers should involve in developing and deploying advanced solutions.
The widespread of Keyless SSL on Cloudflare demonstrates that a CDN provider
plays an important role in the security community [51]. Cooperation from CDN
providers can validate researchers’ ideas and advance further research. CDNs
can also guide their customers to deploy a defense mechanism.
This paper presents the preliminary results of password sharing to third-party
CDNS. We propose the following directions as the future work.
1. Augment the existing CDN discovery method to differentiate the hosting
service and the CDN service of a cloud provider, as mentioned in Sect. 4.
2. Quantify the adopted or available countermeasures besides the client-side
encryption in websites, including CDN bypassing, OAuth, one-time password,
two-factor authentication, etc., as mentioned in Sect. 6.
3. Measure private data leakage in websites to understand the security impact
of TLS private key sharing from users’ perspectives, as mentioned in Sect. 6.
4. Survey the users and website developers to understand their awareness of
private data leakage to thrid-party CDNs. Such a survey helps to figure out
the reason why countermeasures are not widespread.
8 Related Work
security of password input fields among the Alexa top 100K sites, and they
found that 62.8% of the websites with a login page are vulnerable to basic man-
in-the-middle attacks [53]. Bonneau et al. surveyed the proposals for replacing
passwords and pointed out the difficulty of replacing passwords [12]. Peng et al.
explored how passwords are spread after they are divulged by phishing sites [47].
In addition, many prior works investigated the prevalence of the password reuse
problem [28,46,49,57] and its countermeasures [55].
CDN security. Researchers have shown the existence of a wide range of vulnera-
bilities in CDNs. Mirheidari et al.’s measurement shows that private data can be
divulged by CDNs through web cache deception [19,38,39]. Nguyen et al. pre-
sented an attack of poisoning CDN cache with error pages, and five CDN services
were vulnerable to such an attack [44]. Besides CDN cache, researchers also pre-
sented approaches to disclosing the IP addresses of origin servers hidden behind
CDNs, demonstrating insufficient DDoS protection of CDNs [30,54]. Moreover,
attackers may utilize a CDN to launch DoS to an origin server or to the CDN
itself [16,23,52]. In addition, Durumeric et al.’s measurement shows that the
HTTPS interception on CDNs may downgrade the TLS version or cipher suites
and thus reduce connection security [17].
Solutions to TLS key sharing. A line of research focuses on building keyless
CDNs. Cloudflare, Akamai, and Modadugu et al. proposed similar solutions
called “Keyless SSL”, respectively [18,40,51]. Certificate delegation [34] and
mcTLS [42] enable a client to recognize the CDN as a delegation of the web-
site. Wei et al. [58] and Ahmed et al. [11] adopted Trust Executive Environment
(TEE) on CDNs for private key management. However, these strategies only
prevent the TLS private key sharing, while users’ private data are still visible
to CDNs. Phoenix [26] and mbTLS [41] extend TEE solutions to fully protect
users’ private data. However, deploying TEE-based solutions on CDNs may take
a long time as it requires upgrades of hardware and operating systems. Invi-
Cloak [35] protects users’ private data with an additional encryption channel
and low overhead, but its adoption by websites in the future remains unclear.
9 Conclusion
Appendix
We present the detail of our auto-login framework in this section. For each web
page, the framework applies four steps to the HTML elements: filtering, classify-
ing, scoring, and submitting credentials. The framework first filters the elements
based on tag names and locations. Then it uses keyword frequency as the features
to classify filtered elements into three classes: login entrances, account inputs,
and password inputs. In each class, it assigns a score to each element according
to features extracted from the HTML code. Finally it fills and submits creden-
tials if the login form is found, or it clicks on the login entrance to visit the login
page. The elements to interact with are chosen by their scores in each class. The
followings paragraphs introduce each step in detail.
1. Filtering: When the framework arrives at a page, it starts with filtering out
elements that are considered irrelevant to login. Specifically, it selects elements
containing one of the following tag names: “input”, “button”, “label”, “a” and
“iframe”. To reduce element candidates, we assume that a login entrance or a
login form should be shown within the area of one and a half of the viewport
height from the top of a web page. The rationale of this assumption is that a
website should place login elements at positions that are easily accessible to
users.
2. Classifying: To classify an element into the classes mentioned above, the
framework extracts strings from HTML properties and the inner text of the
element. It then splits strings into words by camel case and non-word charac-
ters. It computes the frequencies of some keywords in the string. The keyword
frequencies are regarded as a feature of the element. The framework classifies
the element based on these features and heuristic rules. We manually select
eleven keywords and construct rules for classification after examining Alexa
top 100 sites. One example of the rules is that a login entrance should con-
tain at least one of the keywords related to “login”, “account”, or “email”. To
improve the detection accuracy, we also apply some deprecation keywords
such as “user guide” and “policy”. An element is discarded if it contains any
of the deprecation keywords.
3. Scoring: While a website usually contains only one login entrance, the frame-
work may classify multiple elements into the login class. Thus, our framework
assigns scores to elements. For each element, the framework extracts other
features besides keywords, such as the length of inner text and the visibility
of element. The framework uses the features to assign a score to each element
according to the rules we construct manually. For example, in the class of
login entrance, a visible and interactive element receives a higher score than
ones that are not. The frequency of a keyword in an element is also factored
in the scores. Finally, the framework sorts elements in each class according
to their scores.
Quantifying User Password Exposure to Third-Party CDNs 665
Overall, our framework uses heuristic rules to detect login entrances and
input fields of credentials. We implement the framework by using Selenium Web-
Driver [5] to control Chrome. We test our framework on 100 random-selected
websites of which 52 enable the login. The results show that our framework suc-
cessfully submits credentials to 45 of 53 websites, meaning a recall of 84.9%.
The framework ignores all 47 websites without a login entrance, meaning a false
positive of 0%. The overall detection accuracy is (45+47)/100=92.0%.
Existing automatic login frameworks: Browsers such as Chrome and Firefox can
help users automatically fill in the credentials on some web pages. We do not use
this function because it relies on the existence of the “autocomplete” attribute in
HTML elements, and thus it cannot handle the websites that do not enable this
attribute in HTML. Besides the automation of browsers, Peng et al. implemented
a framework to log into phishing websites automatically [47]. Our framework can
handle issues that are common in legitimate sites but rare in phishing sites, such
as confusion caused by sign-up forms and pop-ups. Jonker et al.. also proposed
a framework for post-login security analysis [31]. Our framework shares many
similarities with theirs but adds the capability to operate in the presence of
HTTP Authentication and reCAPTCHA.
References
1. Akamai (2020). [Link]
2. Alexa Top Sites (2020). [Link]
3. Cloudflare (2020). [Link]
4. Fastly (2020). [Link]
5. SeleniumHQ Browser Automation (2020). [Link]
6. AJAX (2022). [Link]
7. Certifications and Compliance Resources (2022). [Link]
trust-hub/compliance-resources/
8. Global CDN and Optimizer - Introduction (2022). [Link]
bundle/cloud-application-security/page/introducing/[Link]
9. PAM2023-CDNPassword (2022). [Link]
ssword
10. Security Measures (2022). [Link]
11. Ahmed, R., Zaheer, Z., Li, R., Ricci, R.: Harpocrates: giving Out Your Secrets and
Keeping Them Too. In: Proceedings of IEEE/ACM Symposium on Edge Comput-
ing (SEC), pp. 103–114. IEEE (2018)
12. Bonneau, J., Herley, C., Van Oorschot, P.C., Stajano, F.: The quest to replace
passwords: a framework for comparative evaluation of web authentication schemes.
In: Proceedings of S&P, pp. 553–567. IEEE (2012)
666 R. Xin et al.
13. Boyko, V., MacKenzie, P., Patel, S.: Provably secure password-authenticated key
exchange using diffie-hellman. In: Preneel, B. (ed.) EUROCRYPT 2000. LNCS,
vol. 1807, pp. 156–171. Springer, Heidelberg (2000). [Link]
540-45539-6_12
14. Calder, M., Flavel, A., Katz-Bassett, E., Mahajan, R., Padhye, J.: Analyzing the
performance of an anycast CDN. In: Proceedings of International Media Conference
(IMC), pp. 531–537 (2015)
15. Cangialosi, F., et al.: Measurement and analysis of private key sharing in the
HTTPS ecosystem. In: Proceedings of Computer and Communications Security
(CCS), pp. 628–640. ACM (2016)
16. Chen, J., et al.: Forwarding-loop attacks in content delivery networks. In: Proceed-
ings of the Network and Distributed System Security Symposium (NDSS). ISOC
(2016)
17. Durumeric, Z., et al.: The security impact of HTTPS interception. In: Proceedings
of Network and Distributed System Security Symposium (NDSS). ISOC (2017)
18. Gero, C.E., Shapiro, J.N., Burd, D.J.: Terminating SSL Connections without
Locally-Accessible Private Keys (2013). u.S. Patents, No. 9,647,835
19. Gil, O.: Web Cache Deception Attack (2017). [Link]
02/[Link]
20. Gillman, D., Lin, Y., Maggs, B., Sitaraman, R.K.: Protecting websites from attack
with secure delivery networks. Computer 48(4), 26–34 (2015)
21. Green, M.: Let’s Talk About PAKE (2018). [Link]
com/2018/10/19/lets-talk-about-pake/
22. Guo, R., et al.: Abusing CDNs for fun and profit: security issues in CDNs’ origin
validation. In: Proceedings of 2018 IEEE 37th Symposium on Reliable Distributed
Systems (SRDS), pp. 1–10. IEEE (2018)
23. Guo, R., et al.: CDN Judo: breaking the CDN DoS protection with itself. In: Pro-
ceedings of 2020 Network and Distributed System Security Symposium (NDSS).
ISOC (2020)
24. Halfond, W.G., Viegas, J., Orso, A., et al.: A classification of SQL-injection attacks
and countermeasures. In: Proceedings of International Symposium on Secure Soft-
ware Engineering. vol. 1, pp. 13–15. IEEE (2006)
25. Hardt, D.: The OAuth 2.0 authorization framework. Internet Engineering Task
Force (IETF) (2012)
26. Herwig, S., Garman, C., Levin, D.: Achieving Keyless CDNs with Conclaves. In:
Proceedings of Security Symposium, pp. 735–751. USENIX (2020)
27. Huang, C., Wang, A., Li, J., Ross, K.W.: Measuring and evaluating large-scale
CDNs. In: Proceedings of the 8th ACM SIGCOMM conference on Internet mea-
surement (IMC), pp. 15–29. ACM (2008)
28. Ion, I., Reeder, R., Consolvo, S.: “... no one can hack my mind”: Comparing expert
and non-expert security practices. In: Proceedings of Symposium On Usable Pri-
vacy and Security (SOUPS), pp. 327–346. USENIX (2015)
29. Jarecki, S., Krawczyk, H., Xu, J.: OPAQUE: an asymmetric PAKE protocol secure
against pre-computation attacks. In: Nielsen, J.B., Rijmen, V. (eds.) EURO-
CRYPT 2018. LNCS, vol. 10822, pp. 456–486. Springer, Cham (2018). https://
[Link]/10.1007/978-3-319-78372-7_15
30. Jin, L., Hao, S., Wang, H., Cotton, C.: Your remnant tells secret: residual resolution
in DDoS protection services. In: Proceedings of 2018 48th Annual IEEE/IFIP
International Conference on Dependable Systems and Networks (DSN), pp. 362–
373. IEEE (2018)
Quantifying User Password Exposure to Third-Party CDNs 667
31. Jonker, H., Karsch, S., Krumnow, B., Sleegers, M.: Shepherd: a generic approach to
automating website login. In: Proceedings of NDSS Workshop on Measurements,
Attacks, and Defenses for the Web. ISOC (2021)
32. Krishnamurthy, B., Wills, C., Zhang, Y.: On the use and performance of content
distribution networks. In: Proceedings of International Media Conference (IMC),
pp. 169–182. ACM (2001)
33. Levy, A.: CDNs and Privacy Threats: A Measurement Study. Ph.D. thesis, Prince-
ton University (2017)
34. Liang, J., Jiang, J., Duan, H., Li, K., Wan, T., Wu, J.: When HTTPS meets CDN:
a case of authentication in delegated service. In: Proceedings of IEEE Symposium
on Security and Privacy (S&P), pp. 67–82. IEEE (2014)
35. Lin, S., Xin, R., Goel, A., Yang, X.: InviCloak: an end-to-end approach to privacy
and performance in web content distribution. In: Conference on Computer and
Communications Security (CCS). ACM (2022)
36. Lu, B., Zhang, X., Ling, Z., Zhang, Y., Lin, Z.: A measurement study of authen-
tication rate-limiting mechanisms of modern websites, In: Proceedings of the 34th
Annual Computer Security Applications Conference (ACSAC), pp. 89–100 (2018)
37. Maggs, B.M., Sitaraman, R.K.: Algorithmic nuggets in content delivery. ACM
SIGCOMM CCR 45(3), 52–66 (2015)
38. Mirheidari, S.A., Arshad, S., Onarlioglu, K., Crispo, B., Kirda, E., Robertson, W.:
Cached and confused: web cache deception in the wild. In: Proceedings of Security
Symposium, pp. 665–682. USENIX (2020)
39. Mirheidari, S.A., Golinelli, M., Onarlioglu, K., Kirda, E., Crispo, B.: Web cache
deception escalates! In: Proceedings of Security Symposium. Boston, MA, pp. 179–
196. USENIX (2022)
40. Modadugu, N., Goh, E.J.: The Design and Implementation of WASP: A Wide-Area
Secure Proxy. Stanford University, Tech. rep. (2002)
41. Naylor, D., Li, R., Gkantsidis, C., Karagiannis, T., Steenkiste, P.: And then there
were more: secure communication for more than two parties. In: Proceedings of
CoNEXT, pp. 88–100 (2017)
42. Naylor, D., et al.: Multi-context TLS (mcTLS): enabling secure in-network func-
tionality in TLS. In: Proceedings of Proceedings of the 2015 ACM Conference on
Special Interest Group on Data Communication (SIGCOMM), pp. 199–212. ACM
(2015)
43. Newton, A., Hollenbeck, S.: RFC7482: registration data access protocol (RDAP)
query format. Internet Engineering Task Force (IETF) (2015)
44. Nguyen, H.V., Iacono, L.L., Federrath, H.: Your cache has fallen: cache-poisoned
denial-of-service attack. In: Conference on Computer and Communications Security
(CCS), pp. 1915–1936. ACM (2019)
45. Nygren, E., Sitaraman, R.K., Sun, J.: The akamai network: a platform for high-
performance internet applications. SIGOPS OSR 44(3), 2–19 (2010)
46. Pearman, S., et al.: Let’s go in for a closer look: observing passwords in their
natural habitat. In: Proceedings of Conference on Computer and Communications
Security (CCS), pp. 295–310. ACM (2017)
47. Peng, P., Xu, C., Quinn, L., Hu, H., Viswanath, B., Wang, G.: What happens after
you leak your password: understanding credential sharing on phishing sites. In:
Proceedings of the 2019 ACM Asia Conference on Computer and Communications
Security (AsiaCCS), pp. 181–192 (2019)
48. Senol, A., Acar, G., Humbert, M., Borgesius, F.Z.: Leaky Forms: a study of email
and password exfiltration before form submission. In: Proceedings of Security Sym-
posium. Boston, MA, pp. 1813–1830. USENIX (2022)
668 R. Xin et al.
49. Shay, R., et al.: Encountering stronger password requirements: user attitudes
and behaviors. In: Proceedings of Symposium On Usable Privacy and Security
(SOUPS), pp. 1–20. USENIX (2010)
50. Sparling, C.: 5 Years of Fighting DDoS with the Power of Akamai (2019). https://
[Link]/2019/07/5-years-of-fighting-ddos-with-the-power-of-akamai.
html
51. Sullivan, N.: Keyless SSL: The Nitty Gritty Technical Details (2014). [Link]
cloudfl[Link]/keyless-ssl-the-nitty-gritty-technical-details/
52. Triukose, S., Al-Qudah, Z., Rabinovich, M.: Content delivery networks: protection
or threat? In: Backes, M., Ning, P. (eds.) ESORICS 2009. LNCS, vol. 5789, pp.
371–389. Springer, Heidelberg (2009). [Link]
1_23
53. Van Acker, S., Hausknecht, D., Sabelfeld, A.: Measuring login webpage security.
In: Proceedings of the Symposium on Applied Computing (SAC), pp. 1753–1760.
ACM (2017)
54. Vissers, T., Van Goethem, T., Joosen, W., Nikiforakis, N.: Maneuvering around
clouds: bypassing cloud-based security providers. In: Conference on Computer and
Communications Security (CCS), pp. 1530–1541. ACM (2015)
55. Wang, K.C., Reiter, M.K.: How to end password reuse on the web. In: Proceedings
of Network and Distributed System Security Symposium (NDSS). ISOC (2019)
56. Wang, X.S., Choffnes, D., Gage Kelley, P., Greenstein, B., Wetherall, D.: Measur-
ing and predicting web login safety. In: Proceedings of SIGCOMM Workshop on
Measurements up the Stack, pp. 55–60 (2011)
57. Wash, R., Rader, E., Berman, R., Wellmer, Z.: Understanding password choices:
how frequently entered passwords are re-used across websites. In: Proceedings
of Twelfth Symposium on Usable Privacy and Security (SOUPS), pp. 175–188.
USENIX (2016)
58. Wei, C., Li, J., Li, W., Yu, P., Guan, H.: STYX: a trusted and accelerated hierarchi-
cal SSL key management and distribution system for cloud based CDN application.
In: Proceedings of the 2017 Symposium on Cloud Computing (SoCC), pp. 201–213.
ACM (2017)
59. Weinberger, J., Saxena, P., Akhawe, D., Finifter, M., Shin, R., Song, D.: A sys-
tematic analysis of XSS sanitization in web application frameworks. In: Atluri, V.,
Diaz, C. (eds.) ESORICS 2011. LNCS, vol. 6879, pp. 150–171. Springer, Heidelberg
(2011). [Link]
60. Wu, T.D., et al.: The secure remote password protocol. In: Proceedings of Network
and Distributed System Security Symposium (NDSS). vol. 98, pp. 97–111. Citeseer
(1998)
Correction to: An In-Depth Measurement
Analysis of 5G mmWave PHY Latency and Its
Impact on End-to-End Delay
Correction to:
Chapter “An In-Depth Measurement Analysis of 5G mmWave
PHY Latency and Its Impact on End-to-End Delay”
in: A. Brunstrom et al. (Eds.): Passive and Active Measurement,
LNCS 13882, [Link]
In the originally published version of chapter 13, the presentation of Jaideep Chan-
drashekar’s affiliation was misleading. This has been corrected.
A F
Aben, Emile 461 Fang, Tianqi 257
Agarwal, Yuvraj 595 Farhan, Syed Muhammad 71
Alaraj, Abdulrahman 373 Fdida, Serge 227
Ammar, Mostafa 129 Feldmann, Anja 525
Apostolaki, Maria 595 Ferlin-Reiter, Simone 191
Arturi, Augusto 400 Fezeu, Rostand A. K. 284
Ashiq, Md. Ishtiaq 550 Fiebig, Tobias 209, 525
Fontugne, Romain 429
B
G
Bayer, Jan 564
Gañán, Carlos H. 209, 461, 525
Belova, Margarita 595
Gasser, Oliver 18, 525, 623
Benjamin, Ben Chukwuemeka 564
Gavrilovska, Ada 313
Bhardwaj, Ketan 313
Gosain, Devashish 623
Bhosale, Vaibhav 313
Gürses, Seda 209
Bhowmick, Protick 550
Bischof, Zachary S. 345
Bock, Kevin 373 H
Bronzino, Francesco 3 Hassan, Ahmad 284
Brouer, Jesper Dangaard 191 He, Jia 129
Brunstrom, Anna 191, 496 Hesselman, Cristian 564
Bush, Randy 429 Høiland-Jørgensen, Toke 191
Bustamante, Fabián E. 400
J
Jia, Ru 227
C Jiang, Haiyang 227
Carisimo, Esteban 400
Carle, Georg 110
K
Carlsson, Niklas 160
Kashaf, Aqsa 595
Chandrashekar, Jaideep 284
Khan, Etienne 46
Chen, Zhiyi 345
Korczyński, Maciej 461, 564
Chiba, Daiki 479
Korenek, Jan 85
Chung, Taejoong 71, 550
L
D Lee, Myungjin 284
Dainotti, Alberto 345 Levin, Dave 373
Deccio, Casey 550 Lichtblau, Franziska 525
Dou, Jiachen 595 Lin, Shihan 652
Drago, Idilio 3 Lindorfer, Martina 209
Duda, Andrzej 461, 564 Lone, Qasim 461
© The Editor(s) (if applicable) and The Author(s), under exclusive license
to Springer Nature Switzerland AG 2023
A. Brunstrom et al. (Eds.): PAM 2023, LNCS 13882, pp. 669–670, 2023.
[Link]
670 Author Index
M Sperotto, Anna 46
Maghsoudlou, Aniss 18 Srisa-an, Witawas 257
Magnusson, Jonathan 496 Streibelt, Florian 209, 525
Maroofi, Sourena 564 Sundberg, Simon 191
Minneci, Benjamin 284
Mohammadinodooshan, Alireza 160 T
Mori, Tatsuya 479 Tajalizadehkhoob, Samaneh 461
Moura, Giovane C. M. 461 Testart, Cecilia 345
Müller, Moritz 496 Trevisan, Martino 3
N V
Narayanan, Arvind 284 van der Ham, Jeroen 46
Nosyk, Yevheniya 461 van Rijswijk-Deij, Roland 46
Vermeulen, Kevin 429
P Vermeulen, Lukas 18
Pan, Heng 227
Patel, Jay 257 W
Pelsser, Cristel 429 Wabeke, Thymen 564
Phokeer, Amreesh 429 Wustrow, Eric 373
Poese, Ingmar 18
Pulls, Tobias 496 X
Xie, Gaogang 227
Q Xie, Jack 284
Qian, Feng 284 Xin, Rui 652
Xu, Lisong 257
R
Ramadan, Eman 284 Y
Rasaii, Ali 623 Yajima, Masanori 479
Yang, Xiaowei 652
S Ye, Wei 284
Saeed, Ahmed 313 Yoneya, Yoshiro 479
Sattler, Patrick 110, 525
Schmitt, Paul 3 Z
Sekar, Vyas 595 Zegura, Ellen 129
Singh, Shivani 623 Zhang, Zhi-Li 284
Sismis, Lukas 85 Zhauniarovich, Yury 461
Sosnowski, Markus 110 Zirngibl, Johannes 110