0% found this document useful (0 votes)
5 views13 pages

Optimizing RX Attacks on Simon/Simeck Ciphers

This paper enhances deep learning-based rotational-XOR (RX) attacks on lightweight block ciphers Simon32/64 and Simeck32/64 by optimizing neural distinguishers and presenting key-recovery attacks. The authors construct specialized data formats for training RX-neural distinguishers, achieving significant improvements in attack rounds and success rates compared to previous methods. Additionally, they introduce novel techniques for key recovery in related-key settings, marking the first successful neural-based key-recovery attacks for Simeck32/64.

Uploaded by

teknikshell01
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views13 pages

Optimizing RX Attacks on Simon/Simeck Ciphers

This paper enhances deep learning-based rotational-XOR (RX) attacks on lightweight block ciphers Simon32/64 and Simeck32/64 by optimizing neural distinguishers and presenting key-recovery attacks. The authors construct specialized data formats for training RX-neural distinguishers, achieving significant improvements in attack rounds and success rates compared to previous methods. Additionally, they introduce novel techniques for key recovery in related-key settings, marking the first successful neural-based key-recovery attacks for Simeck32/64.

Uploaded by

teknikshell01
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 1

Enhancing Deep Learning-Based Rotational-XOR Attacks on


Lightweight Block Ciphers Simon32/64 and Simeck32/64
Chengcai Liu, Siwei Chen , Zejun Xiang, Shasha Zhang, and Xiangyong Zeng

Abstract—At CRYPTO 2019, Gohr pioneered neural crypt- based cryptanalysis, like differential cryptanalysis [6], linear
analysis by introducing differential-based neural distinguishers cryptanalysis [7], rotational-XOR (RX) cryptanalysis [8] and
to attack Speck32/64, establishing a novel paradigm combin- so on, is not much.
ing deep learning with differential cryptanalysis. Since then,
constructing neural distinguishers has become a significant ap- At CRYPTO 2019, Gohr [9] combined deep learning and
proach to achieving the deep learning-based cryptanalysis for differential cryptanalysis to attack Speck32/64. He trained a
arXiv:2511.06336v1 [[Link]] 9 Nov 2025

block ciphers. This paper advances rotational-XOR (RX) attacks model with training datasets of ciphertext pairs and labels. The
through neural networks, focusing on optimizing distinguishers model, called the differential-neural distinguisher, can return
and presenting key-recovery attacks for the lightweight block a confidence level between 0 and 1 for a given ciphertext
ciphers Simon32/64 and Simeck32/64. In particular, we first
construct the fundamental data formats specially designed for pair, indicating whether the ciphertext pair is a real ciphertext
training RX-neural distinguishers by refining the existing data pair or a random ciphertext pair. Moreover, to mount key-
formats for differential-neural distinguishers. Based on these data recovery attacks using neural distinguishers, this work applies
formats, we systematically identify optimal RX-differences with a variant of Bayesian optimization to recovering information
Hamming weights 1 and 2 that develop high-accuracy RX-neural of key bits, known as the Bayesian key-recovery strategy. As
distinguishers. Then, through innovative application of the bit
sensitivity test, we achieve significant compression of data format a result, Gohr carried out the 11-round attacks on Speck32/64
without sacrificing the distinguisher accuracy. This optimization based on an 8-round differential-neural distinguisher. Since
enables us to add more multi-ciphertext pairs into the data then, studies on combining classical cryptanalysis methods
formats, further strengthening the performance of RX-neural with deep learning techniques have emerged, like linear-
distinguishers. As an application, we obtain 14- and 17-round neural cryptanalysis [10], integral-neural cryptanalysis [11],
RX-neural distinguishers for Simon32/64 and Simeck32/64, which
improves the previous ones by 3 and 2 rounds, respectively. In RX-neural cryptanalysis [12], etc. The common idea of these
addition, we propose two novel techniques, key bit sensitivity works is to construct the corresponding neural distinguishers
test and the joint wrong key response, to tackle the challenge according to the concrete attacking types. Similar to classical
of applying Bayesian’s key-recovery strategy to the target cipher cryptanalysis, good distinguishers can lead to good attacks.
that adopts nonlinear key schedule in the related-key setting That is to say, the higher the accuracy of the neural distin-
without considering of weak-key space. By this, we can straight-
forwardly mount a 17-round key-recovery attack on Simeck32/64 guishers are, the longer the attacking rounds or the higher the
based on the improved 16-round RX-nerual distinguisher. To the attacking success rates will be. Therefore, how to optimize the
best of our knowledge, the presented RX-neural distinguishers accuracy of neural distinguishers is extremely necessary in the
outperform the state-of-the-art neural-based distinguishers for deep learning based attacks.
both Simon32/64 and Simeck32/64, and this is the first successful
neural-based key-recovery attack for Simeck32/64 under the
related-key setting. A. Related Works
Index Terms—Deep Learning, Neural Distinguisher, In 2023, Chen et al. [13] used multi-ciphertext to improve
Rotational-XOR Attacks, Simon, Simeck, Key-recovery
Attack. the accuracy of differential-neural distinguishers. In the same
year, Gohr et al. [14] proposed a fair comparison scheme
between multi-ciphertext and other neural distinguishers. They
I. I NTRODUCTION indicated that it is feasible to improve neural distinguishers

W ITH the development of information technology, deep


learning has become a key technology driving the
intelligence of various industries. For example, deep learning
using multi-ciphertext because of the dependence between
ciphertext pairs. Taking 10-round Simon32/64 as an example,
the accuracy of the differential-neural distinguisher trained
is used for diagnosing diseases and analyzing medical images with 64 ciphertext pairs is 0.36 higher than that of the case
in the medical field and perceiving driving environments and without multi-ciphertext.
planning driving paths in the automotive industry. Moreover, At SAC 2023, Ebrahimi et al. [12] first introduced the
deep learning also plays an important role in strengthening RX-neural distinguisher. They utilized an evolutionary algo-
information security. For example, deep learning is an effective rithm framework and obtained 11- and 15-round RX-neural
technique frequently applied to the side-channel attacks [1]– distinguishers for Simon32/64 and Simeck32/64, respectively.
[4] in cryptography. In theory, cryptography and machine This study marks the first attempt to construct an RX-neural
learning are naturally linked fields [5] because many cryp- distinguisher, but key-recovery attacks based on neural dis-
tography tasks can be naturally framed as learning tasks. tinguishers were not discussed. We guess, the reason is that
Therefore, combining deep learning with classical cryptanal- RX cryptanalysis is a related-key attack method, while the
ysis technology has great potential. However, deep learning Bayesian key-recovery strategy is only applicable to single-key
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 2

attacks. This is because the crucial step in the Bayesian key- and Simeck32/64. Among these, the highest accuracy for 11-
recovery strategy involves performing the wrong key response round Simon32/64 and 13-round Simeck32/64 are 0.9215 and
(WKR), which is used to update the single guessed key. 0.6990 whereas the presented ones in [12] are 0.5445 and
However, how to perform WKR in a related-key setting 0.7057, respectively. Notably, our results are obtained without
remains unsolved. employing any optimization training strategy but the authors
At ASIACRYPT 2023, Bao et al. [15] successfully applied in [12] leveraged the multi-cipher technique, demonstrating
related-key attacks using the Bayesian key-recovery strategy that our strategy for good RX-differences selection is indeed
within the weak-key space to Speck32/64. This work provides effective.
new insights into the application of the Bayesian key-recovery
Presenting currently best neural distinguishers by lever-
strategy in a related-key setting. However, the weak-key space,
aging the multi-ciphertext and staged strategies. A critical
which requires the key to satisfy certain special conditions,
strategy for enhancing neural distinguisher accuracy involves
significantly reduces the key space. In other words, the re-
expanding data formats with multi-ciphertext. While incorpo-
quirements for this attack scenario are quite strict.
rating additional ciphertext pairs improves model discrimina-
tive power, this approach proportionally escalates computa-
B. Motivation of This Work tional demands. To make training feasible within limited com-
Indeed, the application of multi-ciphertext can significantly putational resources, we need to first cut down the fundamental
improve the accuracy of neural distinguishers, but the plat- data format as short as possible by identifying and eliminating
form is required to have extremely powerful computational the redundant components, then add into ciphertext pairs. For
resources. In order to enhance the accuracy under the limited this purpose, we utilize the bit sensitivity test (BST) technique,
computing resources, the data format needs to be cut down as introduced by Chen et al. [16], to analyze the influence of
short as possible. In addition, to our best knowledge, there is each ciphertext bit on RX-neural distinguishers. Surprisingly,
no related work about deep learning based RX attacks besides we find that the left branch of the ciphertext has negligible
Ebrahimi et al.’s work (SAC 2023). Moreover, Ebrahimi et effect on the performance of distinguishers. In other words,
al. presented the RX-neural distinguishers in a rather simple the components related to the left branch of ciphertext are
way, without optimizing them or mounting any key-recovery redundant. Thus, the number of components in the funda-
attacks. Furthermore, with regard to related-key attacks on mental data can be optimized from 8 to 5. Moreover, for the
ciphers with nonlinear key schedules, Bao et al.’s method remaining components, we use the method of controlling vari-
effectively addresses the challenges posed by key propagation ables to gradually identify and eliminate those with minimal
probabilities in the Bayesian key-recovery strategy. However, impact on accuracy. Consequently, we obtain two shortened
it also reduces the key space, which in turn impacts the success data formats only with 2 and 3 components, respectively.
rate of the final key-recovery attack. Therefore, expanding the Combining with the multi-ciphertext technique, we develop
key space or entirely avoiding the introduction of weak keys two multi-ciphertext data formats, successfully yielding 13-
is worth further investigation in the context of related-key and 16-round distinguishers for Simon32/64 and Simeck32/64,
settings. respectively. Furthermore, by leveraging the staged training
strategy, both the 13- and 16-round distinguishers are further
improved by one round, which extend the state-of-art neural-
C. Our Contributions based distinguishers from 13 to 14 for Simon32/64 and from
Motivated by the existing problems as mentioned above, 15 to 17 rounds for Simeck32/64. Our results, along with the
we in this paper focus on improving deep learning based various types of neural distinguishers are listed in Table I.
RX attacks on Simon32/64 and Simeck32/64, and give the
Mounting the first neural-based key-recovery attacks un-
following contributions:
der related-key setting. For Simon32/64, we can directly
Exploring the better RX-differences and fundamental data use Bayesian key-recovery strategy (BKS) as Simon adopts
formats for training RX-neural distinguishers. To train an a linear key schedule. As a result, we achieve 14- and 15-
RX-neural distinguisher, the first step is to collect cipher- round key-recovery attacks based on the presented RX-neural
text pairs using a fixed RX-difference and construct training distinguishers, with success rates 100% and 75% respectively.
dataset based on a given data format. Thus, it is quite crucial To mount key-recovery attacks for Simeck32/64 without con-
for training high-accuracy distinguisher to find good RX- sidering weak-key space, we introduce two novel techniques:
differences as well as appropriate data formats. Investigated the key bit sensitivity test (KBST) and the joint wrong key
from the previous works about the classical as well as deep response (JWKR). Using KBST, we identify the key bits that
learning based RX-cryptanalysis, the RX-difference with lower have a significant impact on neural distinguishers, which are
Hamming weight is more likely to derive the better results. called sensitive key bits. Then, we only need to construct
Besides, we adjust the existing data formats, which are proved JWKR for the sensitive key bits instead of all key bits, which
to be effective for training differential-neural distinguishers, makes the key-guessing process practical. Combining with
to fit RX-neural distinguishers training. Through exhaustively BKS, the key-recovery attacks can be achieved under the
evaluating the RX-nerual distinguishers trained using all pos- full key space in related-key setting. Based on the explored
sible RX-differences with Hamming weights of 1 and 2, we RX-nerual distinguishers, we conduct 16- and 17-round key-
respectively retain six good RX-differences for Simon32/64 recovery attacks for Simeck32/64, with success rates 98%
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 3

and 40%. Compared to the existing neural-based key-recovery II. P RELIMINARY


attacks (see Table II), our results indeed seem weak as they A. Notations and Concepts
make full use of the generic techniques of conventional attacks,
Simon and Simeck both adopt the Feistel structure, so we
which are not applicable to our attacks. Therefore, we can only
use CL and CR throughout this paper to denote the left and
consider to extend one round backward the distinguisher and
right branches of the ciphertext for ease. Also, some primary
directly utilize BKS. However, to our best knowledge, this is
notations are given in Table III.
the first time to present key-recovery attacks based on neural
distinguishers under related-key setting for Simon and Simeck. TABLE III: Notations of this paper
Also, the proposed KBST and JWKR are potential to neural-
based related-key attacks for other block cipher like Speck. Notations Descriptions
F2 A finite field only contains 2 elements, i.e. 0 and 1
Fn 2 An n-dimensional vectorial space defined over F2
TABLE I: Comparison of our work with other studies on ⊙ Bitwise AND
neural distinguishers for Simon32/64 and Simeck32/64 ⊕ Bitwise XOR
|| Concatenation of two bit-strings
x≪λ Circular left shift of x by λ bits
Cipher Round Accuracy TNR TPR Type† Ref. x≫λ Circular right shift of x by λ bits
CL r , Cr Left and right branches of r-round ciphertext
11 0.5445 - - RX [12] R
11 1.0000 1.0000 1.0000 RX Sect. IV-B ∆rL , ∆rR Left and right branches of r-round RX-difference
12 0.6477 0.6518 0.6435 RKD [17]
12 0.5225 - - SKD [18]
Simon32/64
13 0.5262 0.5437 0.5081 RKD [17] In addition, there are two key concepts, data format and
13 0.5810 0.5730 0.5890 RKD [19] multi-ciphertext, that will be frequently used in this paper.
13 0.7120 0.7075 0.7165 RX Sect. IV-B These concepts were mentioned in previous works but have
14 0.5241 0.5634 0.4848 RX Sect. IV-B
not been formally defined. To ensure a clearer understanding,
12 0.5161 0.4807 0.5504 SKD [20] we provide their formal definitions. Besides, note that Simon
12 1.0000 1.0000 1.0000 RX Sect. IV-B
15 0.5467 0.5173 0.5762 RKD [17] and Simeck both adopt the Feistel structure, we use CL and
Simeck32/64 15 0.5475 - - RX [12] CR throughout this paper to denote the left and right branches
15 0.5930 0.5950 0.5900 RX [19] of the ciphertext for ease.
15 0.6042 0.5330 0.6750 RX Sect. IV-B
16 0.5130 0.5255 0.5000 RX Sect. IV-B Definition 1 (Data Format). The data format is a data tuple
17 0.5040 0.5950 0.4130 RX Sect. IV-B
formed by N (N ≥ 1) elements D0 , D1 , ..., DN −1 used to
† RX: rotational-XOR; RKD: related-key differential; SKD: single-key generate the neural network training dataset, and denoted by
differential;
D = (D0 , D1 , ..., DN −1 ),
TABLE II: Comparison of our and previous deep learning- where each Di ∈ Fn2 for 0 ≤ i ≤ N − 1 corresponding to an
based key-recovery attacks n-bit data is called a component.
Cipher Round† Time Data Success rate Type Ref. For example, D = (CL , CR , CL ⊕ CL′ ) is a data format
14(13+1) 232 215.81 100% RX Sect. V that consists of three components where the corresponding
Simon32/64
15(14+1) 232 215.81 75% RX Sect. V plaintexts (P, P ′ ) of (CL , CL′ ) satisfies a deterministic rela-
16(4+11+1) 242.79 222 80% SKD [18] tion. Specifically, the first and second components are the left
17(5+11+1) 254.01 228 9% SKD [18]
and right branches of a ciphertext with size 2n, respectively.
15(3+10+2) 233.9 224 88% SKD [21] The third component is the difference on the left branch
16(4+11+1) 224 238.19 100% SKD [20]
Simeck32/64 16(15+1) 251 216.17 98% RX Sect. VI of a pair of ciphertexts. Therefore, if a dataset is generated
17(4+12+1) 226 245.04 30% SKD [20] using D, then each element in dataset has the same form as
17(16+1) 254 216.17 40% RX Sect. VI (CL , CR , CL ⊕CL′ ) and is generated by one pair of ciphertexts.
† The total attack rounds comprise the neural distinguisher rounds plus In order to improve the accuracy of neural distinguishers, in
any prepended or appended rounds. For example, the 17-round attack general, more pairs of ciphertexts can be appended into data
on Simon32/64 prepends 5 rounds and appends 1 round to the 11-round
neural distinguisher, denoted as 5+11+1. format.
Definition 2 (Multi-ciphertext Data Format). Given a data
format D = (D0 , ..., DN −1 ). If there are 2k (2k ≤ N ) differ-
D. Organization ent components Dki (0 ≤ i ≤ 2k − 1) among {D0 , ..., DN −1 },
and (Dk2i , Dk2i+1 ) for all 0 ≤ i ≤ k − 1 corresponds to a
Sect. II starts by giving some concepts and nota-
pair of ciphertexts, then we call D a k-multi-ciphertext data
tions throughout this paper, then introduces Simon32/64,
format and specially denote it by D k .
Simeck32/64 and conventional RX cryptanalysis in brief. We
demonstrate how to identify good RX-differences in Sect. III For example, D 3 = (C0 , C1 , C2 , C3 , C0′ , C1′ , C2′ ) is a 3-
and optimize the prepared RX-nerual distinguishers from multi-ciphertext data format. In particular, the pair of compo-
Sect. IV-A in Sect. IV. In Sect. V and Sect. VI, we present nents (Ci , Ci′ ) for i ∈ {0, 1, 2} represents a pair of ciphertexts,
the key-recovery attacks for Simon32/64 and Simeck32/64, where the corresponding pair of plaintexts (Pi , Pi′ ) satisfies a
repsectively. Finally, we conclude this paper in Sect. VII. deterministic relation.
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 4

B. Description of Simon and Simeck initial RX-difference with zero left branch and non-zero right
Simon [22] is a family of lightweight block ciphers branch. For the sake of convenience, we define this special
published by the National Security Agency based on the initial RX-difference as half RX-difference that is clarified as
Feistel structure. A member of the family is denoted by follows.
Simon2n/mn, where n is the branch size, 2n is the block Definition 4 (Half RX-difference). For a Feistel based block
size, and mn is the key length, n ∈ {16, 24, 32, 48, 64}, cipher with size of 2n, the half RX-difference [λ, ∆R ] is defined
m ∈ {2, 3, 4}. The round function consists of cyclic rotation as the input RX-difference, in which the left (resp. right)
(≪), bitwise AND (⊙), and bitwise XOR (⊕). The key branch ∆L ∈ Fn2 (∆R ∈ Fn2 ) is inactive (resp. active) and
schedule is recursive and divided into 3 types based on the the rotational offset is λ.
value of m. This paper only focuses on the key schedule when
m = 4 and its formula is defined as ki = ki−4 ⊕ ((ki−3 ⊕ III. F INDING G OOD RX- DIFFERENCES AND
(ki−1 ≫ 3)) ⊕ ((ki−3 ⊕ (ki−1 ≫ 3))) ≫ 1)) ⊕ c, where F UNDAMENTAL DATA F ORMATS
ci is a round constant. In 2015, Yang et al. [23] proposed For constructing the better RX-neural distinguishers, it is
the Simeck family. They chose the different rotation offsets of great significance to prepare better half RX-differences
in round function, and reuse the round function as its key firstly. In this section, we will explain how to experimentally
schedule which leads to better implementation in hareware finding the better half RX-differences. First, we will introduce
than Simon. Because we devote attention to Simon32/64 and two new fundamental data formats. Next, we exhaust all the
Simeck32/64, so we only list their round functions and key possible half RX-differences with Hamming weights of 1 and
schedule in Fig. 1. 2. Based on the half RX-differences and new data formats, we
then construct datasets and train the RX-neural distinguishers.
Finally, according to the accuracy of obtained RX-neural
distinguishers, we determine the half RX-differences.
We first provide a detailed explanation of two data formats
used in our experiments. The first one is inspired by Lu et al.’s
data format proposed in [17]. In particular, their data format
r r (r−1) (r−2)
is represented as (δL r
, δRr r
, CLr , CR , C ′ L , C ′ R , δR , δR )
where δ denotes the traditional difference. Specially, the r-
round ciphertexts (C r , C ′r ) are generated by a pair of plain-
(r−2)
texts that satisfy a given difference, and δ R that represents
the left branch of (r − 2)-round difference is derived by
decrypting (C r , C ′r ) for two rounds with the zero roundkeys.
The novelty of this data format is that it makes use of
the decrypted (r − 2)-round information. Note that this data
format is applied to differential-neural cryptanalysis. There-
fore, to fit RX cryptanalysis, we adjust this data format by
imposing rotation on C r and replacing traditional difference δ
(a) Simon32/64 (b) Simeck32/64
by RX-difference ∆. Consequently, the first data format is
r r (r−1) (r−2)
Fig. 1: The round functions and key schedules of Simon and (∆rL , ∆rR , CLr ≪ λ, CR r
≪ λ, C ′ L , C ′ R , ∆R , ∆R ),
(r−2) r ′r
Simeck where ∆R is derived by decrypting (C , C ) for two
rounds with the zero roundkeys. We denote it by D1 . The
second one is based on Gohr’s data format [9], which is
C. Rotational-XOR Cryptanalysis also proposed for differential-neural cryptanalysis and only
consists of a pair of r-round ciphertexts. This data format is
Rotational-XOR (RX) cryptanalysis, originally proposed by (CLr , CRr
, CL′r , CR′r
). Similar to the adjustment on Lu et al.’s
Ashur and Liu [8] at FSE 2016, is an improvement over data format, we impose rotation on C r and get our second
rotational cryptanalysis aiming to address the problem that data format as (CLr ≪ λ, CR r r r
≪ λ, C ′ L , C ′ R ). We denote it
the rotational-invariant property of internal state cannot hold by D2 .
if round constants are injected into key schedule. By investigating the previous good traditional RX distin-
Definition 3 (RX-difference [24]). The RX-difference of x and guishers in [24] as well as the RX-neural distinguishers in [12],
x′ = (x ≪ λ) ⊕ α is denoted by we found that the Hamming weight of initial RX-differences
are generally small, in particular, they are 1 or 2. Therefore,
∆λ (x, x′ ) = (x ≪ λ) ⊕ x′ = α we select all half RX-differences with the Hamming weight of
1 and 2 under the rotation offsets from
 1 to 15 as candidates.
where α ∈ Fn2 is a constant and λ is a rotational offset with
In total, there are 15 × 16 16

+ = 2040 candidates of
0 < λ < n, and (x, x′ ) is called an RX pair. 1 2
half RX-difference for Simon32/64 and Simeck32/64. For each
In this paper, our target ciphers are Simon and Simeck, both candidate, we first randomly choose 223 pairs of plaintexts
of which are with Feistel structure. Also, we only focus the satisfying the half RX-difference and encrypt them to generate
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 5

the pairs of ciphertexts, respectively for 11-round Simon32/64 accuracy, we here only list the top six candidates in IV. Also,
and 13-round Simeck32/64. Then we use the prepared data we compare the accuracy of these RX-neural distinguishers
format D1 and D2 to construct the training datasets. Moreover, with the previous results [12], all of ours are better than
the hyper-paramter of our experiments1 is set as: or consistent to that of [12], which demonstrates that our
• the size of validation dataset is 2 ;
18 proposed data formats and the method to construct half RX-
• the size of batch is 2 ;
15 differences are indeed effective. It is worth noting that the
• the number of iterations is 10; results of [12] listed in Table IV are the optimized one by
• the number of learning rates is 10, which vary from exploiting multi-ciphertext data format but ours do not. Hence,
0.0001 to 0.1 with the gap of 0.0111. we are sure that the accuracy of our presented RX-neural
Utilizing the distributed training strategy of Keras2 , each distinguisher can be further improved, which will be detailed
iteration takes 25 (resp. 20) seconds in average for D1 (resp. in the next section.
D2 ), and training one RX-neural network need to take about TABLE IV: The six half RX-differences for 11-round Si-
four minutes. As a result, it took about 23 days to accomplish mon32/64 and 13-round Simeck32/64 RX-neural distinguish-
the whole 2040 × 2 × 2 = 8160 experiments for Simon and ers using D1
Simeck.
In general, if the accuracy of a neural distinguisher is Simon Accuracy Simeck Accuracy Ref.
more than 0.5, then we regard it as a valid distinguisher [3, 0x2] 0.5445 [1, 0x2] 0.7057 [12]
for mounting attacks. Taking an accuracy of 0.51 as the [15, 0x3] 0.9215 [1, 0x4] 0.6990 Ours
[1, 0x6] 0.9203 [15, 0x2] 0.6989 Ours
standard, we consequently retain 1496 and 151 valid RX- [12, 0x2002] 0.8823 [15, 0x3] 0.6202 Ours
neural distinguishers3 for 11-round Simon32/64 and 13-round [4, 0x22] 0.8802 [1, 0x6] 0.6212 Ours
Simeck32/64, respectively. In particular, [3, 0x12] 0.8115 [6, 0x82] 0.6184 Ours
[13, 0x4002] 0.8097 [10, 0x802] 0.6168 Ours
• As for the 11-round Simon32/64, a total of 1496 valid
distinguishers are identified. When the Hamming weight
of the RX-difference is fixed to 1, the number of valid
distinguishers is 208. Among these, there are 143, 54 IV. E XPLORING B ETTER RX-N EURAL D ISTINGUISHERS
and 9 distinguishers with the accuracy distributed at 0.51- As illustrated in the last section, we find some good half
0.55, 0.55-0.60 and 0.60-0.65. The remaining two have RX-differences that can derive the high-accuracy RX-neural
the accuracy of 0.6550 and 0.6543, corresponding to the distinguishers even if we do not exploit the multi-ciphertext
half RX-differences [13, 0x4000] and [3, 0x2], respec- data formats. To obtain higher-accuracy or longer-round ones,
tively. Surprisingly, when the Hamming weight is 2, the it is of great significance to add multi-ciphertext into the
number of valid distinguishers is 1288. Additionally, there components of the two data formats. On the one hand, ac-
are a total of 6 distinguishers with an accuracy greater cording to the previous studies [13], [14], [17], [25], [26], we
than 0.80, as shown in Table IV. know that the more the number of multi-ciphertext components
• As for the 13-round Simeck32/64, when the Hamming contained in a data format, the higher the accuracy of the
weight of RX-difference is fixed to 1, the distribu- trained neural distinguisher. On the another hand, the number
tion of valid distinguishers exhibits a relatively con- of components of a data format has a directly influence on
centrated pattern. For rotation offsets {1, 4, 5, 11, 12, 15} the required memory space (it is GPU memory in general)
and {2, 3, 6, 10, 13, 14}, there are 6 and 5 distinguishers, for training the neural network. If the required memory
respectively. The remaining offsets {7, 8, 9} have no approaches or exceeds the total memory of the machine, the
valid distinguisher. When the Hamming weight is 2, the training process will be terminated. Therefore, we must strike
number of valid distinguishers becomes more widely a balance between the accuracy of the neural distinguishers
distributed. For rotation offsets {1, 15}, there are 18 and the memory requirements of the training process, under a
valid distinguishers. For offsets {11} and {4, 5, 12}, there limited computational resources. In other words, we need to
are 13 and 10 distinguishers, respectively. But for {2}, control the number of total components as much as possible
{14, 6, 10}, {3, 13}, {7}, and {9}, there are only 6, 5, while adding the multi-ciphertext into data formats.
3, 2, and 1 distinguishers, respectively. The offset {8} Regard D1 as the basic data format, we first remove some
has no valid distinguisher. Among these results, the best redundant components, which will be detailed in Sect. IV-A.
accuracy is closed to 0.7, with half RX-differences of Then in Sect. IV-B, we add multi-ciphertext into the simplified
[1, 0x4] and [15, 0x2]. data formats, and train distinguishers using multi-ciphertext
For more details, please refer to Sect. I of the online Sup- data formats. To explore RX-neural distinguishers that cover
plemental Material4 . By sorting these in descending order of more rounds, we in Sect. IV-C exploit the staged training
method to extend the trained ones from Sect. IV-B.
1 In this work, all experiments are conducted on a platform with four Intel
(R) Xeon (R) E5-2698 v4 @ 2.20GHz CPUs, four Tesla V100 DGXS 32GB
GPUs, and the Ubuntu 20.04 system. A. Removing Redundant Components from Data Format D1
2 The Keras Documentation is available via the link [Link]
3 Our experiments indicate that D is better than D . To remove components from D1 , we need to detect the
1 2
4 [Link] or [Link] impact of ciphertexts on the accuracy of neural disinguishers
anonymous ieee/anonymous [Link] and then determine the redundant components. Here we utilize
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 6

the bit sensitivity test (BST) method proposed in [16]. The By repeating the above process for the remaining 31 bit
main idea of this method is to randomly modify the value of positions, we can obtain the sensitivities for each ciphertext
a fixed bit position in the ciphertexts of each sample in the bit across the three types. For a more intuitive comparison, we
validation dataset by XORing with a random mask, and then present the results in Fig. 2, from which it can be observed
analyze the impact on the accuracy of neural distinguisher. We that all bits on the left branch of the ciphertext (note that 31
apply BST method to the six neural distinguishers in Table IV represents the most significant bit position) have the negligible
for Simon32/64 and Simeck32/64. Consequently, we conclude influcence on the accuracy of the neural distinguisher, as the
that ∆rL , CLr ≪ λ and CL′r are the redundant components of bit sensitivities are close to or equal to 0s. In other words, the
D1 . Next, we will give a detailed illustration about the process components related to the left branch of ciphertext in D1 , i.e.
of BST. ∆rL , CLr ≪ λ and CL′r are redundant. Similarly, we apply BST
As an example, we focus the 11-round neural distinguisher method to the 13-round neural disitinguisher of Simeck32/64
for Simon32/64 with the half RX-difference of [15, 0x3] and with the half RX-difference [1, 0x4] and the accuracy of 0.7
the accuracy of 0.9215. To detect the redundant components of and exhibit the bit sensitivities in Fig. 2b. Note that there
D1 , we first randomly choose 218 pairs of plaintexts satisfying are only five bits on the left branch of cipher can effect the
the half RX-difference of [15, 0x3] to generate an original accuracy. Besides, the corresponding bit sensitivities are quite
set that contains 218 pairs of ciphertexts. Taking BST on the small. Therefore, we can also regard ∆rL , CLr ≪ λ and CL′r
least significant bit (the bit position is 0) as an instance, we as the redundant components of D1 .
randomly choose 218 32-bit masks, where the value of the Thus, by removing the redundant components from D1
bit position 0 varies, and the values of other bit positions as mentioned above, we obtained the simplified data format
(r−1) (r−2)
are all 0s. Then, we XOR these masks with the elements in (∆rR , CLr ≪ λ, CR r
, ∆R , ∆R ), denoted by D3 . More-
the original pair-of-ciphertexts set one by one to generate a over, we wonder whether the data format D3 could be further
modified set. Here, we define three XORing types according simplified. Note that D3 contains two types of components: ci-
to the objective pair of ciphertexts (C, C ′ ) as XORing with phertexts and RX-differences. We try to remove the two types
the single ciphertext (C or C ′ ) or pair of ciphertexts might of components from D3 to respectively obtain the simplified
lead to different effects. The first type, denoted as TYPE1, data formats D4 and D5 . As shown in Table V, the accuracy
involves XORing the mask only with C, which corresponds of distinguisher trained using D4 has a significant decrease
to (CLr , CRr
) of the data format D1 . On the contrary, the compared to D3 (from 0.9221 to 0.5013). Nevertheless, D5
second type, denoted as TYPE2, XORs the mask only with has the nearly same accuracy as D3 . Therefore, the ciphertexts
′r
C ′ , corresponding to (CL′r , CR ) of D1 . The third type, denoted r
i.e., CR ≪ λ and CR ′r
, are the redundant components of D3 .
as TYPE3, involves XORing the mask with both C and C ′ . Similarly, regard D5 as a basic data format, we try to remove
For each of these types, a modified pair-of-ciphertexts set is each one of the three components respectively to obtain D6 ,
generated by applying the corresponding XOR operation to D7 , and D8 , and the corresponding trained accuracies are
the original one. Based on the data format D1 , three types of 0.7641, 0.9211, and 0.5912 as shown in Table V. This indicates
(r−1)
validation dataset can be generated from the modified pair-of- that ∆rR and ∆R are the necessary components, while
(r−2)
ciphertexts sets. Moreover, the validation dataset of each type ∆R is a redundant component in D5 .
is put into the 11-round neural distinguisher of Simon32/64 to
obtain a verified accuracy. Finally, the difference between the
original accuracy (i.e., 0.9215) and the verified accuracy can B. Training Distinguishers using Multi-ciphertext Data For-
be calculated, which is regarded as the bit sensitivity. By the mats
way, the smaller the value of bit sensitivity, the smaller the As mentioned earlier, the more multi-ciphertext a data
influence of the corresponding bit on the accuracy. format contains, the higher the corresponding accuracy will

0.12 TYPE1
TYPE2
0.14 TYPE1
TYPE2
TYPE3 TYPE3
0.12
0.10
0.10
0.08
Bit sensitivity

Bit sensitivity

0.08
0.06
0.06
0.04
0.04
0.02 0.02

0.00 0.00
31
30
29
28
27
26
25
24
23
22
21
20
19
18
17
16
15
14
13
12
11
10
9
8
7
6
5
4
3
2
1
0

31
30
29
28
27
26
25
24
23
22
21
20
19
18
17
16
15
14
13
12
11
10
9
8
7
6
5
4
3
2
1
0

Bit position Bit position

(a) Training using [15, 0x3] (b) Training using [1, 0x4]

Fig. 2: The bit sensitivity for 11-round Simon32/64 (left) and 13-round Simeck32/64 (right), both of which use the data format
D1 . For more bit sensitivity test, refer to Sect. II of Supplemental Material.
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 7

TABLE V: Data formats (written as DF in short) and the corresponding accuracy of RX-neural distinguishers for 11-round
Simon32/64. Note the half RX-difference is [15, 0x3], thus the value of λ is 15
DF Specification Accuracy

D1 r ≪ λ, C r ≪ λ, C ′ r , C ′ r , ∆(r−1) , ∆(r−2) )
(∆rL , ∆rR , CL 0.9215
R L R R R
(r−1) (r−2)
D3 (∆R , CR ≪ λ, C ′ rR , ∆R
r r , ∆R ) 0.9221
D4 r ≪ λ, C ′ r )
(CR 0.5013
R
(r−1) (r−2)
D5 (∆rR , ∆R , ∆R ) 0.9204
r−1 (r−2)
D6 (∆R , ∆R ) 0.7641
(r−1)
D7 (∆rR , ∆R ) 0.9211
(r−2)
D8 (∆rR , ∆R ) 0.5912

be. Due to the limitation of GPU memory of our platform, the to 13 rounds, with the best accuracy of 0.7120. Compared to
maximal number of multi-ciphertext that can be accommo- the previouly best RX-neural distinguisher presented in [12],
dated in D5 and D7 , determined by a series of experimentally our result not only extends the round (from 11 to 13) but also
tests, are 28 and 36 respectively. Consequently, we derive the improves the accuracy (from 0.5445 to 0.7120).
longest extended data formats from D5 and D7 , denoted by Similarly for Simeck32/64 (from 13 to 16 rounds), we
D528 and D736 , which adopt 28 and 36 pairs of ciphertexts. trained a series of RX-neural distinguishers using the three
Applying the two data formats as well as D5 to 11-round data formats (D5 , D528 and D736 ) and the six half RX-
Simon32/64 under the six half RX-differences (listed in Ta- differences. We list the accuracy in Table VII. It can be
ble IV), we have the trained results as presented in Table VI. seen that training with the two multi-ciphertext data formats
can actually derive the 16-round distinguishers with a stable
accuracy more than 0.51. Note that the currently best RX-
TABLE VI: Comparison on the accuracy of trained RX-neural
neural distinguisher, presented in [12], covers 15 rounds and
distinguishers for Simon32/64 using different data format.
has the accuracy of 0.5475. Our trained results improve
Note DF means data [Link] underlined items are used
the accuracy of 15-round one to 0.6022, notably extend the
for key-recovery attacks in Sect. V
number of round to 16.
#R DF [15, 0x3] [1, 0x6] [12, 0x2002] [4, 0x22] [13, 0x4002] [3, 0x12]
TABLE VII: Comparison on the accuracy of trained RX-neural
D5 0.9204 0.9204 0.8869 0.8905 0.8165 0.8170 distinguishers for Simeck32/64 using different data format.
11r D528 0.9999 0.9999 0.9997 0.9996 0.9999 0.9997
D736 1.0000 1.0000 0.9999 0.9998 0.9999 0.9999 Note DF means data format. The underlined items are used
for key-recovery attacks in Sect. VI
D5 0.5921 0.5946 0.6482 0.6437 0.5798 0.5750
12r D528 0.9236 0.9336 0.9744 0.9713 0.8991 0.8827
#R DF [1, 0x4] [15, 0x2] [15, 0x3] [1, 0x6] [6, 0x82] [10, 0x802]
D736 0.9176 0.9065 0.9821 0.9832 0.9075 0.9063
D5 0.6988 0.6997 0.6202 0.6214 0.6215 0.6198
D5 0.5100 0.5115 0.5249 0.5227 0.5091 0.5080
13r D528 0.9962 0.9999 0.9997 0.9997 0.9996 0.9997
13r D528 0.6300 0.6301 0.7120 0.7105 0.6837 0.6848
D736 0.9996 0.9995 0.9970 0.9967 0.9966 0.9968
D736 0.6086 0.5962 0.6893 0.6872 0.5889 0.6062
D5 0.5774 0.5789 0.5250 0.5262 0.5009 0.5014
D528 0.5006 0.5008 0.5009 0.5005 0.5004 0.5000 14r D528 0.9204
14r 0.9136 0.7317 0.7326 0.5000 0.5000
D736 0.5002 0.5001 0.5002 0.5003 0.5006 0.5004
D736 0.8982 0.8982 0.7095 0.7097 0.5019 0.5013
D5 0.5120 0.5136 0.5026 0.5016 0.5023 0.5016
15r D528 0.5905 0.5923 0.5191 0.5189 0.5000 0.5000
It can be seen that the accuracy corresponding to D5 under D736 0.6042 0.6022 0.5206 0.5250 0.5025 0.5000
each half RX-difference is actually identical to D1 (see in
D528 0.5066 0.5127 0.5004 0.5006 0.5005 0.5008
Table IV). Additionally, either D528 or D736 can achieve an ex- 16r
D736 0.5130 0.5057 0.5007 0.5008 0.5010 0.5009
ceptionally high accuracy, nearly reaching 100%. Apparently,
this appears that it is potential to leverage D528 or D736 to train
neural distinguishers covering longer rounds. Thus, we train
the neural distinguishers for 12 to 14 rounds of Simon32/64,
and give the corresponding accuracy in Table VI. The accura- C. Extending Distinguishers with Staged Training Method
cies of 14-round trained distinguishers shown in this table are By leveraging the multi-ciphertext data format, either the
more than 0.5, but in fact, they are not distinguishable. This is round or the accuracy of RX-neural distinguishers for Si-
because the accuracy recorded in the table is the maximum mon32/64 and Simeck32/64 are indeed improved. However, if
value among all epochs. Besides, we conducted prediction we only use the basic training scheme when training the neural
experiments on the trained distinguishers, and the actual distinguisher, it is difficult to break through 14 and 16 rounds
accuracies are unstable at 0.5 within the allowable range of for Simon32/64 and Simeck32/64, respectively. Therefore, we
statistical error, indicating that these distinguishers are invalid. apply the staged training scheme to enhancing the training
As a consequence, the longest valid RX-neural distinguisher process and hope to explore valid neural disinguishers cover-
for Simon32/64 trained using multi-ciphertext can reach up ing more rounds. The concept of staged training was originally
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 8

introduced by Gohr [9] at CRYPTO 2019 based on the idea V. K EY- RECOVERY ATTACKS ON S IMON 32/64
of reinforcement learning, and later has become a crucial
approach to enhance the neural-based distinguishers [17], [27]. In this section, we will give the details of recovering key bits
It turns an already trained (r−1)-round distinguisher into an r- of the round-reduced Simon32/64 based on the trained RX-
round distinguisher in several stages. In addition, the accuracy neural distinguishers. Different from the attacks based on con-
of derived r-round distinguisher relies on that of the (r − 1)- ventional distinguishers, the neural-based key-recovery attacks
round one. need to use the Bayesian key-recovery strategy (BKS) to guess
The application of staged training to Simon32/64 can be key bits. Previous works also combine with the techinque used
divided into three stages. In the first stage, we use the 13- in conventional distinguisher based key-recovery attacks, like
round disinguisher, which was already trained using D528 and extending round forward or backward on distinguishers, to
[4, 0x22] (resp. [12, 0x2002]) with accuracy 0.7105 (resp. mount attacks covering more rounds [9], [18], [20], [27]. In
0.7120), to recognize the 12-round encryption with the same this paper, we consider to extend the RX-neural distinguisher
data format and half RX-difference, i.e., D528 and [4, 0x22] backward by only one round, and straightforwardly use BKS
(resp. [12, 0x2002]). This stage was done on 223 training to achieve practical key-recovery attacks.
samples and 220 testing samples for 10 epochs. Note that all
the samples are derived from the 12-round multi-ciphertext. Algorithm 1 BayesianKeySearch: find a list of subkey
The batch size is 218 and learning rate is ranged from 0.0001 candidates
to 0.00001. In the second stage, we adopt the updated network Input: Ciphertext structure C = C0 , ..., Cm×k−1 ,
from the first stage to recognize the 14-round dataset using RX-neural distinguisher N D, set S :=
the same data format and half RX-difference. In this stage, {rk0 , ..., rkn−1 } composed of n candidates of
the hyper-parameter remain consistent with those from the guessed subkey, number of key search iteration
first stage, except that the samples are generated by the 14- l.
round multi-ciphertext and the number of training samples is Output: A List that contains l × n guessed candidates
increased to 224 . In the third stage, we fresh the 224 training of subkey.
samples and fed them to the network from the second stage. 1 S := {k0 , k1 , . . . , kn−1 } ← choose at random without
The number of epoch is set to 40 or more, and the learning rate replacement from the set of all subkey candidates.
is fixed to 0.00001. Finally, by this staged training strategy, we 2 L ← {};
retain two valid RX-neural distinguishers covering 14 rounds 3 for j ∈ {0, 1, . . . , m − 1} do
with the respective accuracies of 0.5190 and 0.5240, as listed 4 Pi,k ← Decrypt(Ci , k) for all i ∈ {0, 1, . . . , m −
in Table VIII. 1}, k ∈ S.
5 vi,k ← N (Pi,k ) for all i, k
However, for Simeck32/64, the staged training method 6 wi,k ← log2 (vi,k /(1 − vi,k )) for all i ∈ {0, . . . , m −
successfully applied to Simon32/64 failed. No valid 17-round 1}, k P∈S
RX-neural distinguisher could be explored. Thus, we skipped n
7 wk ← i=1 vi,k for all k ∈ S
the first stage and straightforwardly used the two 16-round 8 L ← L ∪ {(k, wk ) for k ∈ S}
distinguishers with the accuracies of 0.5130 and 0.5057 to 9
Pn−1
mk ← i=0 vi,k /n for k ∈ {k0 , . . . , kn−1 }
recognize the corresponding 17-round datasets in the second λk ←
Pn−1 2 2
10 i=0 (mki − µki⊕k ) /σki⊕k for k ∈
stage. After using a lower learning rate of 10−5 in the third 16
{0, 1, . . . , 2 − 1};
stage, we obtained a 17-round distinguisher with an accuracy 11 S ← argsortk (λ)[0 : n − 1]
of 0.5040, correspongding to the half RX-difference [1, 0x4] 12 end
(shown in Table VIII). This accuracy seems a little weak, but 13 return L
in fact, it is very stable and effective in the model prediction
experiments. We first describe the framework of BKS, which is used
to recover (r + 1)-round subkeys (we denote it by rk (r+1) )
for Simon32/64 based on an r-round RX-neural disinguisher
TABLE VIII: Enhanced RX-neural distinguishers for Si- trained using k-multi-ciphertext data format and a half RX-
mon32/64 and Simeck32/64 by staged training. DF: data for- difference [λ, ∆R ], as follows:
mat. HRXD: half [Link] underlined items are used
for [Link] underlined items are used for key-recovery 1) Randomly choose m×k pairs of plaintexts satisfying the
attacks in Sect. V and VI half RX-difference [λ, ∆R ], query for the corresponding
(r + 1)-round pairs of ciphertexts. A ciphertext structure
Cipher DF HRXD #R Accuracy TPR TNR Validity composed of m × k elements is generated.
Simon32/64 D528
[12, 0x2002] 14 0.5190 0.5176 0.5204 ✓ 2) Select n (n < 216 ) guessed candidates of rk (r+1) (there
[4, 0x22] 14 0.5241 0.5634 0.4848 ✓
are in total 216 candidates as the subkey is size of 16),
[1, 0x4] 17 0.5040 0.5950 0.4130 ✓ then use BayesianKeySearch algorithm, as depicted in
Simeck32/64 D736
[15, 0x2] 17 0.5000 - - ✗
Algorithm 1, to generate l × n guessed candidates for
rk (r+1) where l is the number of key search iteration.
Note that each candidate has a score.
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 9

3) Set the filtering threshold for rk (r+1) to c1 , retain the recovery framework, we can straightforwardly achieve a 14-
candidates whose scores are greater than c1 . If there is and 15-round key-recovery attacks on Simon32/64. Note that
no candidate left, then go back the step 2 and re-select in the attacking framework, we expect to find a pair of guessed
n guessed candidates for rk (r+1) . Otherwise, go to the subkeys (rk r , rk (r+1) ) as the correct one by repeating the
step 4. process from step 2. In theory, it can actually return a pair
4) For each of the n1 left candidates of rk (r+1) , generate as long as repeating for plenty of times. But it is not easy
m × n guessed candidates for rkr like doing in the step to evaluate the runtime in practice. Therefore, we need to set
2. There are in total m × n × n1 guessed candidate pairs an upper bound on the repeating times when applying the
of (rk r , rk (r+1) ), each one corresponds to a score. framework to practical attacks, we denote it by t. If there is
still no surviving candidate after t times of repeating, then
5) Set the filtering threshold for (rk r , rk (r+1) ) to c2 , retain we just regard the candidate pair, which has the highest score
the pairs whose scores are greater than c2 . If there among all t times of repeating, as the correct subkeys.
are n2 candidates left, rank them by their scores and
the pair corresponding to the highest score is regarded • To mount 14-round attack, we choose the 13-round
as the correct value of (rk r , rk (r+1) ). Otherwise, go RX-neural distinguisher with accuracy 0.7105, which
back to the step 2 and repeat the process till a pair of is trained using half RX-difference [4, 0x22] and data
(rk r , rk (r+1) ) is retained. format D528 . Under the same half RX-difference and data
format, the accuracy of 12-round one is 0.9713. The
In the above mentioned (r + 1)-round key-recovery frame- attack parameters are set as:
work based on BKS, BayesianKeySearch plays an important
role. The values of µ and σ in Line 9 of Algorithm 1 depend
m = 210 k = 28 c1 = 500 c2 = 1000
on the wrong key response (WKR) of the targeting subkey. As
the core step of BKS, the concept of WKR was introduced by t = 210 l=6 n = 32.
Gohr at CRYPTO 2019, to optimize the candidates of guessed The WKR profile for 14-round subkey is exhibited in
subkey such that the candidates that we generate cover the Fig. 3a. The results demonstrate that the mean response
correct subkey with a high probability. One can refer to [9] for of correct subkey is significantly different from the wrong
more detail, we will omit the WKR process and only exhibit ones. Notably, the correct subkey displays a mean re-
the WKR profile in the concrete attack on Simon32/64. It is sponse value exceeding 0.61, while the majority of wrong
worth noting that the success rate and runtime of our attacks candidates remain below the 0.40 threshold. Namely, the
depend on the values of c1 and c2 . In particular, if c1 and c2 recovered one possesses a high probability of matching
is too large, the success rate will be high but the attack will the correct subkey. In fact, based on the high-accuracy
take much more time. On the contrary, if c1 and c2 are small, RX-neural distinguisher, we can obtain a recovered pair
the runtime will be less but the success rate will decrease. of (rk 13 , rk 14 ) within four minutes using the key-
Therefore, to ensure that key-recovery attacks can be achieved recovery framework. To verify the effectiveness of our
in a reasonable time with a high success rate, we need to set attack, we then mount 100 key-recovery attacks under
c1 and c2 to the appropriate values by conducting a pre-attack 100 random master keys and get the 100% attacking suc-
process before starting attacks. The process involves setting the cess rate. The remaining 32-bit information of mask key
correct key of the last 2 rounds as the initial guessed candidate, can be recovered by brute-force attacks. Consequently,
then observing the pre-attack score and finally determining the time complexity of this 14-round key-recovery attack
threshold ranges through repeated experiments. is 232 and the data complexity is 210 × 28 × 2 = 215.81
Key-recovery attacks on 14- and 15-round Simon32/64. chosen plaintexts.
Based on the explored RX-neural distinguishers and the key- • For 15-round key-recovery attack, we use the 14-round

0.50 0.50
0.504
0.60

0.503
0.55

0.502
0.50
Mean response

Mean response

0.501
0.45

0.500
0.40

0.499
0.35

0.498
0 4096 8192 12288 16384 20480 24576 28672 32768 36864 40960 45056 49152 53248 57344 61440 65536 0 4096 8192 12288 16384 20480 24576 28672 32768 36864 40960 45056 49152 53248 57344 61440 65536
Difference to real key Difference to real key

(a) For 14-round subkey (b) For 15-round subkey

Fig. 3: The WKS profiles for 14- and 15-round Simon32/64. For more WKS for other rounds, one can refer to Sect. III of
Supplemental Material
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 10

RX-neural distinguisher as shown in Table VIII. It is take the bit position 0 of 16-round subkey of Simeck32/64 as
trained using half RX-difference [4, 0x22] and data for- an instance to illustrate the process of KBST. To launch KBST
mat D528 , with a accuracy of 0.5141. The corresponding for 16-round subkey, we need to prepare a 15-round RX-neural
13-round one has the acuracy of 0.7105. The parameters distinguisher. We choose the one with accuracy 0.6042 trained
of this attack are set as using half RX-difference [1, 0x4] and data format D736 . Firstly,
m = 210 k = 28 c1 = 16 c2 = 500 we randomly choose 106 × 36 pairs of plaintext satisfing the
half RX-difference [1, 0x4] and generate a set composed of
t = 210 l=6 n = 32. 106 × 36 pairs of ciphertext, denoted by C. Then we choose
Fig. 3b shows the WKR profile of the 15-round subkey. 106 masks where the value corresponding to the bit position
The analysis reveals that the correct key exhibits a 0 varies randomly, and the values of other bit positions are
mean response value of 0.504, closely aligning with the all 0s. We XOR these masks one by one on the fixed 16-
accuracy of 14-round RX-neural distinguisher. However, round subkey pair simultaneously to generate 106 modified
this differentiation is not as statistically significant as subkey pairs. Decrypt all elements in C using the modified
observed in the 14-round WKS, therefore theoretically subkey pairs and combine with data format D736 , we obtain a
yielding lower success rates compared to the 14-round at- validation dataset composed of 106 samples5 . Send this dataset
tack. Empirical verification via 100 key-recovery attacks to the prepared 15-round neural distinguisher, a validation
demonstrate an achievable success rate of 75%. The time accuracy is returned. The difference of this validation accuracy
and data complexities are the same as the 14-round attack. and the original accuracy of 15-round neural distinguisher is
regarded as the KTYPE1 sensitivity of the 0-th bit of 16-round
VI. K EY- RECOVERY ATTACKS ON S IMECK 32/64 subkey. As for the another type of bit sensitivity, denoted by
In related-key setting, Simon’s RX-difference of subkey pair KTYPE2, is derived using the constant mask where the value
for each round is uniquely determined as its key schedule is corresponding to 0-th bit is 1 and others are 0s. Considering
linear. If one subkey of the pair is guessed, then another one the other 15 bit positions, KBST of the 16-round subkey can
can be directly obtained using the deterministic RX-difference. be accomplished and the result is exhibited in Fig. 4. It appears
However, Simeck’s RX-difference is probabilistic due to its that only five sensitive key bits in 16-round subkey, namely, it
nonlinear key schedule, requiring simultaneous guessing both is sufficient to construct JWKR for these five bits. The KBST
of the two subkeys when mounting key-recovery attacks under for more rounds and data formats can be found in Sect. IV of
related-key setting. Therefore, BKS is not suitable for Simeck Supplemental Material. It is worth noting that KBST of 18-
as BKS only supports single-key guessing. round subkey is similar to that of other rounds subkeys at the
positions [4, 5, 8, 9, 10, 13, 14, 15] bits, confirming that our 17-
The joint wrong key response. In [15], Bao et al. used round RX-neural distinguisher for Simeck32/64 is valid, even
fixed key differences in related-key differential attacks against though it is very weak.
Speck, known as weak-key attacks. It makes sense but can only
success under a fixed key space (weak-key space) instead of Key-recovery attacks on 16- and 17-round Simeck32/64.
the full key space. To ensure that key-recovery attacks can be Key-recovery attacks for Simeck can also make use of the
achievable for the full key space under related-key setting, attack framework as described in Sect. V, but need to perform
we combine the idea of joint distribution to construct the JWKR instead of WKR in BayesianKeySearch algorithm.
wrong key response for key pairs, called the joint wrong key Therefore, we here reuse some notations denoted in Simon’s
response (JWKR). When attacking Simon32/64, the WKR pro- key-recovery attacks. All the RX-neural distinguishers utilized
file is constructed by systematically enumerating all possible in 16- and 17-round key-recovery attacks are trained using half
values of the target subkey (comprising 216 candidates) and RX-difference [1, 0x4] and data format D736 . The details of
computing their corresponding mean and standard deviation attacks are illustrated as follows:
responses. This process requires approximately 25 minutes. • For 16-round attack, we utilize the 14- and 15-round
Extending this methodology to JWKR profile construction, distinguishers with respective accuracies of 0.8982 and
which involves applying the same exhaustive approach to the 0.6042. The KBSTs for 15- and 16-round subkeys are
full 232 subkey pair space, yields a mathematically projected shown in Fig. 4. It indicates that we need to construct
runtime of 25 × 216 ≈ 220.6 minutes (equivalent to three JWKRs for bit positions [4, 5, 8, 9, 10, 13, 14, 15] of rk 15
years). Obviously, it is unfeasible. To tackle this problem, we and [5, 9, 10, 14, 15] of rk 16 , which are illustrated by
introduce key bits sensitivity test (KBST) method aiming to three-dimensional images in Fig 5. The JWKS for more
reduce the subkey pair space of JWKR. rounds can be found in Sect. V of Supplemental Material.
Key bits sensitivity test. Inspired by the application of BST By setting the attack parameter as
to detecting the redundant bits of ciphertext (as illustrated
in Sect. IV-A), we can use BST to detect the influence of
each bit of subkey on the accuracy of trained RX-neural
distinguisher. If the key bits have a negligible influence on 5 Note that C contains 106 × 36 pairs of ciphertext but there are only 106

the accuracy, then we call these bits as insensitive key bits, modified pairs of subkey, we have to divide C into 106 groups. Each group
containing 36 pairs of ciphertext is decrypted using one modified subkey pair,
otherwise as sensitive key bits. To distinguish from BST for and the corresponding decrypted group forms one sample using the 36-multi-
ciphertext, we call this key bits sensitivity test (KBST). We ciphertext data format D736 .
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 11

KTYPE1 0.10
KTYPE1
0.4 KTYPE2 KTYPE2
0.08
0.3
0.06
Bit sensitivity

Bit sensitivity
0.2
0.04

0.1 0.02

0.0 0.00
15
14
13
12
11
10
9
8
7
6
5
4
3
2
1
0

15
14
13
12
11
10
9
8
7
6
5
4
3
2
1
0
Bit positon Bit positon

(a) For 15-round subkey (b) For 16-round subkey

Fig. 4: The key bits sensitivity test for Simeck32/64

0.5100
0.8000 0.5000
0.6000 0.4900
Mean

Mean
0.4800
0.4000 0.4700
0.4600
0.2000 0.4500
0.4400
250 30
200 25
150 20
0 0 5 15
50 100 K1 10 15 10 K1
100 50
150 5
K0 200
250 0 K0 20 25 30 0
(a) For 15-round subkey (b) For 16-round subkey

Fig. 5: The joint wrong key response for sensitive bits of Simeck32/64

KTYPE1
KTYPE2 0.5008
0.008 0.5005
0.5002
Mean

0.006 0.5000
0.4998
Bit sensitivity

0.4995
0.004 0.4992

0.002 30
25
20
0 5 15
0.000
10 15 10 K1
5
K0 20 25 30 0
15
14
13
12
11
10
9
8
7
6
5
4
3
2
1
0

Bit positon

(a) KBST (b) JWKS

Fig. 6: The KBST and JWKS for 17-round Simeck32/64

m = 210 k = 36 c1 = 150 c2 = 1000 (rk 15 , rk 16 ) within seven minutes and 216.17 chosen
plaintexts, and the success rate is 98%. The remaining
t = 210 l=4 n = 32,
51-bit information can be recovered by brute-force attack.
we can straightforwardly recover the 13 bits of
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 12

Thus the time complexity is 251 . R EFERENCES


• For 17-round attack, we use the 15- and 16-round neural
[1] B. Timon, “Non-profiled deep learning-based side-channel attacks with
distinguisher with the accuracies of 0.6042 and 0.5130, sensitivity analysis,” IACR Trans. Cryptogr. Hardw. Embed. Syst., vol.
repsectively. According to KBST of rk 17 as shown in 2019, no. 2, pp. 107–131, 2019.
Fig. 6a, there are also five sensitive bits: [5, 9, 10, 14, 15], [2] J. Kim, S. Picek, A. Heuser, S. Bhasin, and A. Hanjalic, “Make some
noise. unleashing the power of convolutional neural networks for profiled
and the corresponding JWKS is exhibited in Fig 6b. The side-channel analysis,” IACR Trans. Cryptogr. Hardw. Embed. Syst., vol.
attack parameters are set as: 2019, no. 3, pp. 148–179, 2019.
[3] L. Wu, L. Weissbart, M. Krcek, H. Li, G. Perin, L. Batina, and S. Picek,
“Label correlation in deep learning-based side-channel analysis,” IEEE
m = 210 k = 36 c1 = 4 c2 = 100 Trans. Inf. Forensics Secur., vol. 18, pp. 3849–3861, 2023.
[4] H. Kim, C. Hahn, H. J. Kim, Y. Shin, and J. Hur, “Deep learning-
t = 210 l=4 n = 32. based detection for multiple cache side-channel attacks,” IEEE Trans.
Inf. Forensics Secur., vol. 19, pp. 1672–1686, 2024.
As a consequence, 10 bits of (rk 16 , rk 17 ) can be recov- [5] R. L. Rivest, “Cryptography and machine learning,” in Advances in
ered with data complexity 216.17 . The whole process takes Cryptology - ASIACRYPT ’91, International Conference on the Theory
about four hours, and the success rate is about 40%. The and Applications of Cryptology, Fujiyoshida, Japan, November 11-14,
1991, Proceedings. Springer, 1991, pp. 427–439.
total time complexity to recover 64-bit key is 254 . [6] E. Biham and A. Shamir, “Differential cryptanalysis of des-like cryp-
tosystems,” J. Cryptol., vol. 4, no. 1, pp. 3–72, 1991.
[7] M. Matsui, “Linear cryptanalysis method for DES cipher,” in Advances
VII. C ONCLUSION in Cryptology - EUROCRYPT ’93, Workshop on the Theory and Ap-
plication of of Cryptographic Techniques, Lofthus, Norway, May 23-27,
In this paper, we conducted a comprehensive study on 1993, Proceedings. Springer, 1993, pp. 386–397.
the deep learning-based rotational-XOR cryptanalysis for Si- [8] T. Ashur and Y. Liu, “Rotational cryptanalysis in the presence of
constants,” IACR Trans. Symmetric Cryptol., vol. 2016, no. 1, pp. 57–70,
mon32/64 and Simeck32/64. Specifically, inspired by previous 2016.
work about the differential-neural attack, we proposed two [9] A. Gohr, “Improving attacks on round-reduced speck32/64 using deep
fundamental data formats specially designed for training RX- learning,” in Advances in Cryptology - CRYPTO 2019 - 39th Annual
International Cryptology Conference, Santa Barbara, CA, USA, August
neural distinguishers. Through systematic exploration of initial 18-22, 2019, Proceedings, Part II. Springer, 2019, pp. 150–179.
RX-differences with Hamming weight ≤ 2, we identified some [10] B. Hou, Y. Li, H. Zhao, and B. Wu, “Linear attack on round-reduced
good input patterns that enable the development of high- DES using deep learning,” in Computer Security - ESORICS 2020 -
25th European Symposium on Research in Computer Security, ESORICS
accuracy RX-neural distinguishers. In addition, we further 2020, Guildford, UK, September 14-18, 2020, Proceedings, Part II.
enhanced the performance of the trained neural distinguish- Springer, 2020, pp. 131–145.
ers using multi-ciphertext data formats and staged training [11] B. Zahednejad and L. Lyu, “An improved integral distinguisher scheme
techniques. Finally, to verify effectiveness of the proposed based on neural networks,” Int. J. Intell. Syst., vol. 37, no. 10, pp. 7584–
7613, 2022.
RX-neural distinguishers, some round-reduced key-recovery [12] A. Ebrahimi, D. Gérault, and P. Palmieri, “Deep learning-based
attacks were presented. rotational-xor distinguishers for AND-RX block ciphers: Evaluations
Regarding the neural-based distinguishers, our results ex- on simeck and simon,” in Selected Areas in Cryptography - SAC 2023
- 30th International Conference, Fredericton, Canada, August 14-18,
tend the state-of-the-art from 13 to 14 rounds for Simon32/64 2023, Revised Selected Papers. Springer, 2023, pp. 429–450.
and from 15 to 17 rounds for Simeck32/64. Our key-recovery [13] Y. Chen, Y. Shen, H. Yu, and S. Yuan, “A new neural distinguisher
attacks under the related-key setting do not surpass exist- considering features derived from multiple ciphertext pairs,” Comput.
J., vol. 66, no. 6, pp. 1419–1433, 2023.
ing single-key differential-neural attacks [18] and [20] as [14] A. Gohr, G. Leander, and P. Neumann, “An assessment of differential-
we cannot prepend additional rounds in front of related-key neural distinguishers,” IACR Cryptol. ePrint Arch., p. 1521.
neural distinguisher, but this work establishes two notable [15] Z. Bao, J. Lu, Y. Yao, and L. Zhang, “More insight on deep learning-
aided cryptanalysis,” in Advances in Cryptology - ASIACRYPT 2023
contributions: 1) present the first key-recovery attacks for - 29th International Conference on the Theory and Application of
both cipher in related-key setting. 2) develop two novel Cryptology and Information Security, Guangzhou, China, December 4-8,
techniques, named key bits sensitivity test and joint wrong key 2023, Proceedings, Part III. Springer, 2023, pp. 436–467.
[16] Y. Chen, Y. Shen, and H. Yu, “Neural-aided statistical attack for
response, effectively addressing challenges in applying neural cryptanalysis,” Comput. J., vol. 66, no. 10, pp. 2480–2498, 2023.
distinguishers to related-key attacks without considering weak- [17] J. Lu, G. Liu, B. Sun, C. Li, and L. Liu, “Improved (related-key)
key space. These advancements provide new insights into RX- differential-based neural distinguishers for SIMON and SIMECK block
neural cryptanalysis, particularly in related-key attacks based ciphers,” Comput. J., vol. 67, no. 2, pp. 537–547, 2024.
[18] L. Zhang, Z. Wang, and B. Wang, “Improving differential-neural crypt-
on neural distinguishers, and it is also potential to explore analysis,” IACR Commun. Cryptol., vol. 1, no. 3, p. 13, 2024.
better attacks by combining with other types of related-key [19] X. Yuan and Q. Wang, “A multi-differential approach to enhance
neural attacks like related-key differential-neural attack. related-key neural distinguishers,” Cryptology ePrint Archive, Paper
2025/697, 2025. [Online]. Available: [Link]
[20] L. Zhang, J. Lu, Z. Wang, and C. Li, “Improved differential-neural
cryptanalysis for round-reduced SIMECK32/64,” Frontiers Comput. Sci.,
ACKNOWLEDGEMENT vol. 17, no. 6, p. 176817, 2023.
[21] L. Lyu, Y. Tu, and Y. Zhang, “Deep learning assisted key recovery
This work was supported by the National Natural Science attack for round-reduced simeck32/64,” in Information Security - 25th
Foundation of China (No. 62272147, 12471492, 62072161, International Conference, ISC 2022, Bali, Indonesia, December 18-22,
12401687) and the Innovation Group Project of the Natu- 2022, Proceedings. Springer, 2022, pp. 443–463.
[22] R. Beaulieu, D. Shors, J. Smith, S. Treatman-Clark, B. Weeks, and
ral Science Foundation of Hubei Province of China (No. L. Wingers, “The SIMON and SPECK families of lightweight block
2023AFA021). ciphers,” IACR Cryptol. ePrint Arch., p. 404.
SUBMIT TO IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY 13

[23] G. Yang, B. Zhu, V. Suder, M. D. Aagaard, and G. Gong, “The simeck


family of lightweight block ciphers,” in Cryptographic Hardware and
Embedded Systems - CHES 2015 - 17th International Workshop, Saint-
Malo, France, September 13-16, 2015, Proceedings. Springer, 2015,
pp. 307–329.
[24] J. Lu, Y. Liu, T. Ashur, B. Sun, and C. Li, “Improved rotational-xor
cryptanalysis of simon-like block ciphers,” IET Inf. Secur., vol. 16, no. 4,
pp. 282–300, 2022.
[25] Z. Hou, J. Ren, and S. Chen, “Improve neural distinguishers of simon
and speck,” Security and Communication Networks, vol. 2021, no. 1, p.
9288229.
[26] A. Benamira, D. Gérault, T. Peyrin, and Q. Q. Tan, “A deeper look
at machine learning-based cryptanalysis,” in Advances in Cryptology
- EUROCRYPT 2021 - 40th Annual International Conference on the
Theory and Applications of Cryptographic Techniques, Zagreb, Croatia,
October 17-21, 2021, Proceedings, Part I. Springer, 2021, pp. 805–835.
[27] Z. Bao, J. Guo, M. Liu, L. Ma, and Y. Tu, “Enhancing differential-
neural cryptanalysis,” in Advances in Cryptology - ASIACRYPT 2022
- 28th International Conference on the Theory and Application of
Cryptology and Information Security, Taipei, Taiwan, December 5-9,
2022, Proceedings, Part I. Springer, 2022, pp. 318–347.

Chengcai Liu received the M.E. degree with School of Cyber Science and
Technology in Hubei University, Wuhan, China. His research interest includes
cryptanalysis of block ciphers.

Siwei Chen received the Ph.D. degree from Hubei University in 2022. He is
currently a Lecturer with School of Cyber Science and Technology in Hubei
University, Wuhan, China. His current research interests include design and
cryptanalysis of symmetric ciphers.

Zejun Xiang received the Ph.D. degree from University of Chinese Academy
of Sciences, Beijing, China, in 2018. He is currently an Associate Professor
with School of Cyber Science and Technology in Hubei University, Wuhan,
China. His current research interests include design, cryptanalysis, classical
and quantum implementation of symmetric ciphers.

Shasha Zhang received the Ph.D. degree from Wuhan University, Wuhan,
China, in 2009. She is currently an Associate Professor with School of Cyber
Science and Technology in Hubei University, Wuhan, China. Her current
research interests include cryptanalysis, classical and quantum implementation
of symmetric ciphers and blockchain.

Xiangyong Zeng received the Ph.D. degree from Beijing Normal University,
Beijing, China, in 2002, then did the post-doctoral work from 2002 to 2004
in Wuhan University, Wuhan, China. He is currently a Professor with Faculty
of Mathematics and Statistics in Hubei University, Wuhan, China. His current
research interests include coding theory, cryptology, and blockchain.

You might also like