0% found this document useful (0 votes)
22 views18 pages

EPIPDLF: Predicting Enhancer-Promoter Interactions

The document presents EPIPDLF, a novel deep learning framework designed to predict enhancer-promoter interactions (EPIs) using genomic sequences. This approach addresses the limitations of traditional experimental methods by providing a cost-effective and efficient computational solution, demonstrating superior performance across benchmark datasets. The model incorporates interpretable analysis mechanisms to enhance understanding of the biological significance of identified sequences.

Uploaded by

23025076
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views18 pages

EPIPDLF: Predicting Enhancer-Promoter Interactions

The document presents EPIPDLF, a novel deep learning framework designed to predict enhancer-promoter interactions (EPIs) using genomic sequences. This approach addresses the limitations of traditional experimental methods by providing a cost-effective and efficient computational solution, demonstrating superior performance across benchmark datasets. The model incorporates interpretable analysis mechanisms to enhance understanding of the biological significance of identified sequences.

Uploaded by

23025076
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Page 1 of 18 Bioinformatics

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9 Sequence analysis
10
11
12
EPIPDLF: a pre-trained deep learning framework
13 for predicting enhancer-promoter interactions
14
15 Zhichao Xiao1, Yan Li2 , Yijie Ding3 and Liang Yu1,*
16 1School of Computer Science and Technology, Xidian University, Xian 710075, China, 2Xi'an Polytechnic
17 University No.19, Jinhua South Road, Xi'an, Shaanxi Province, China, 3 Yangtze Delta Region Institute
18 (Quzhou), University of Electronic Science and Technology of China, Quzhou, 324000, China
19
20 *To whom correspondence should be addressed.
21
Associate Editor: XXXXXXX
22
Received on XXXXX; revised on XXXXX; accepted on XXXXX
23
24
25 Abstract
26 Motivation: Enhancers and promoters, as regulatory DNA elements, play pivotal roles in gene
27 expression, homeostasis, and disease development across various biological processes. With
advancing research, it has been uncovered that distal enhancers may engage with nearby promoters
28
to modulate the expression of target genes. This discovery holds significant implications for
29 deepening our comprehension of various biological mechanisms. In recent years, numerous high-
30 throughput wet-lab techniques have been created to detect possible interactions between enhancers
31 and promoters. However, these experimental methods are often time-intensive and costly.
32 Results: To tackle this issue, we have created an innovative deep learning approach, EPIPDLF,
33 which utilizes advanced deep learning techniques to predict EPIs based solely on genomic
34 sequences in an interpretable manner. Comparative evaluations across six benchmark datasets
35 demonstrate that EPIPDLF consistently exhibits superior performance in EPI prediction. Additionally,
36 by incorporating interpretable analysis mechanisms, our model enables the elucidation of learned
37 features, aiding in the identification and biological analysis of important sequences.
38 Availability: The source code and data are available at: [Link]
Contact: lyu@[Link]
39
40
41
42
can be millions of base pairs apart, often not interacting with nearby
43
44 1 Introduction enhancers. Instead, most enhancers skip adjacent genes to connect with
distant promoters via long-range chromatin loops. The principles of
45 Enhancers and promoters are the two most important types of gene
chromatin interactions at the genome sequence level are unclear.
46 expression regulatory elements in mammals, especially humans. The
Therefore, it is crucial to establish an effective computational method for
47 efficient interaction between them ensures the accurate transcription of
identification and study EPI, and the large amount of data brought by
48 genes, thereby ensuring cell status and normal development. Erroneous
high-throughput sequencing technology makes this feasible (Wei, et al.,
49 associations between them can also lead to disease-related gene
2020).
50 expression abnormalities. Therefore, exploring enhancer-promoter
51 interactions is of great biological interest, but we know that genome-
52 wide chromatin interaction mechanisms are complex (Ni, et al., 2022). In
53 particular, the emergence of high-throughput sequencing technologies
54 such as Hi-C (Rao, et al., 2014) and ChIA-PET (Heidari, et al., 2014) has
55 enabled us to more clearly understand the complex mode of action of
56 EPI. In mammalian genomes, a gene's promoter and its distal enhancer
57
58
© The Author(s) 2025. Published by Oxford University Press.
59 This is an Open Access article distributed under the terms of the Creative Commons Attribution License
60 ([Link] which permits unrestricted reuse, distribution, and reproduction in any medium,
provided the original work is properly cited.
Bioinformatics Page 2 of 18

[Link] et al.
1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29 Fig. 1. The Workflow Framework of EPIPDLF
30 At present, many excellent calculation methods have been performance. In addition, more and more work is choosing to combine
31 developed to identify EPI. Due to the huge amount of data, most convolutional neural networks and recurrent neural networks on models.
32 calculation methods are based on deep learning technology(Liu, et al., For example, the deep learning model SPEID (Zhuang, et al., 2019)
33 2023; Qiao, et al., 2024). In the early development of computational proposed by Singh et al. combines convolutional neural networks (CNN)
34 methods, r researchers usually choose genomic features as input to the and long short-term memory (LSTM) (Chen, et al., 2022; Li, et al.,
35 model, such as TargetFinder (Whalen, et al., 2016), ChINN (Cao, et al., 2022). EPIVAN (Hong, et al., 2020) proposed by Liu et al. also
36 2021). TargetFinder proposed by Whalen et al. uses a large amount of combines CNN and Gated recurrent unit (GRU), and pre-trains the
37 genomic information, encompassing genomic peak data such as DNase- model. Recently, numerous studies have used sequence data to explore
38 seq, DNA methylation, transcription factor ChIP-seq, histone deep learning for EPI prediction, achieving notable improvements in
39 modifications, CAGE, and gene expression data, to select and use on the prediction performance. However, current deep learning predictors have
40 classifier random forest (RF) (Cox, 1958) and Support vector machine not fully leveraged feature representation learning, particularly in
41 (SVM) (Boser, et al., 1992; Wang, 2023) were used for training data. identifying key sequence patterns crucial for understanding EPI
42 Cao et al. proposed ChINN based on convolutional neural networks, mechanisms. Consequently, these models lack interpretability and fail to
43 which further achieved genome-wide prediction of chromatin harness the impact of sequence-based approaches in EPI prediction (Yin,
44 interactions. In recent years of research, more and more work has chosen et al., 2024).
45 to use sequence data to train models. The main reasons are: 1. Compared In order to solve the above problems, we refer to some advanced
46 with genome data, sequence data is out-of-the-box and does not require technologies in natural language processing that are developing rapidly
47 too many pre-processing steps; 2. The rapid development of the natural today, such as BERT (Devlin, et al., 2018; Ren, et al., 2024; Zhang, et
48 language field makes the process of training sequence data more al., 2024). Inspired by this, we treat DNA sequences as text data, and
49 complicated. technical means. Yang et al. developed PEP-Word (Yang, convert DNA sequences into "biological vocabulary" by building a
50 et al., 2017), which uses word embeddings to extract features directly vocabulary. Therefore, we propose a model pre-trained on large-scale
51 from sequences and trains a prediction algorithm for a boosted tree genome sequences to learn biological context semantics, converting EPI
52 ensemble model. The results of their work demonstrate that genome- 'biological vocabulary' into training data. In order to solve the lack of
53 wide EPIs can be reliably predicted based on sequence features alone. In interpretability of deep learning, we try to apply both genomic data and
54 the same year, Mao et al. proposed EPIANN (Mao, et al., 2017), which sequence data to the model. We employ adversarial training and transfer
55 is a neural network architecture that utilizes the attention mechanism, learning to boost prediction performance and enhance model robustness.
56 and introduced positional feature encoding to further improve Benchmark results from seven cell line datasets show that our model
57
58
59
60
Page 3 of 18 Bioinformatics

Article short title


1
2
3 substantially surpasses leading sequence-based approaches. Importantly, In this paper, we main employ Convolutional Neural Networks (CNNs)
4 it offers interpretable predictions and analysis at the sequence level by (Rakhlin, 2016; Zulfiqar, et al., 2024) and Recurrent Neural Networks
5 examining local features using attention mechanisms. The model (RNNs) (Rakhlin, 2016) to extract features, and to capture long-range
6 accurately and adaptively identifies sequence regions closely related to dependencies among the features within the sequences, we further
7 EPIs. Overall, our contributions can be summarized as follows: incorporate self-attention mechanisms, considering that the sequences
8 1. We introduce a novel deep learning method named EPIPDLF, are of considerable length(Li, et al., 2021). Combining CNN and RNN is

Downloaded from [Link] by guest on 05 March 2025


9 which is capable of training on pure sequence data and incorporates an a common technique in deep learning for extracting features from
10 additional gene data processing module. sequential data. CNN excels at capturing meaningful local features and
11 2. We propose transfer learning and adversarial learning strategies reducing the network's parameter count efficiently. On the other hand,
12 to enhance model performance during testing and cross-cell line RNN utilizes its recurrent structure to model temporal relationships
13 validation. within sequential data. Next, we will outline the methods used for feature
14 3. We utilize CNN modules and self-attention mechanisms to extraction. The complete framework of this research is depicted in
15 extract biologically meaningful motif sequences. Figure 1.
16 1D convolution layer: the model employs a 1D CNN layer is
17 employed to generate embedding features from the sequence.
2 Methods
18 Subsequently, MaxPooling is applied to perform down sampling and
19 further reduce the dimensionality of the features.
20 2.1 Dataset GRU layer: Next, the features flow into a Gated Recurrent Unit
21 In this study, we employed the same EPI dataset as TargetFinder to (GRU) (Chung, et al., 2014), which is an improvement over conventional
22 assess our model and compare it to existing approaches. The dataset RNNs and has the ability to address the issue of long-term dependencies.
23 comprises EPIs from six human cell lines: GM12878 (lymphoblastoid The GRU is a form of recurrent neural network (RNN) intended for
24 cells), HUVEC (umbilical vein endothelial cells), HeLa-S3 (cervical processing sequential data. It tackles the problem of long-term
25 carcinoma-derived cells), IMR90 (fetal lung fibroblasts), K562 dependencies by incorporating gating mechanisms that regulate
26 (leukemia-derived mesodermal cells), and NHEK (epidermal information flow. Here's how GRU works:
27 keratinocytes). TargetFinder utilized annotations from ENCODE and Consider an input sequence (or time step) represented as x and a
28 Roadmap Epigenomics to identify active enhancers and promoters within hidden state represented as h. GRU features two primary gates:
29 each cell. For the analysis of these enhancers and promoters, high- comprising the update gate and the reset gate.
30 resolution genome-wide measurements were performed for each cell line The reset gate (r) regulates the interaction between the previous
31 using Hi-C data. This enabled the classification of enhancer-promoter hidden state (h) and the current input (x), influencing how much old
32 pairs into interacting (positive samples) and non-interacting (negative information is discarded and how much new information is integrated.
33 samples). For each positive sample, 20 negative samples were selected, The formula for calculating the reset gate is provided below:
34 resulting in a ratio of 1:20 between positive and negative samples within r = s Wr ×  ht -1 , xt  (1)
35 each cell line. Additionally, care was taken to ensure that the positive Here, Wr represents the weight matrix associated with the reset gate,
36 and negative samples exhibited similar distributions of enhancer- s denotes the sigmoid function, and  ht -1 , xt  represents the
37 promoter distances. Table 1 provides detailed information on the datasets concatenation of the previous hidden state and the current input.
38 for each cell line. Moreover, we incorporated genomic information such The update gate (z) regulates the inclusion of the previous hidden state
39 as CTCF-binding sites, chromatin accessibility (DNase-I signals), and (h) and the current input (x) in the current hidden state. It is computed as
40 five histone marks (H3K27me3, H3K36me3, H3K4me1, H3K4me3, and follows:
41 H3K9me3). The processing of these genomic signals was conducted z = s Wz ×  ht -1 , xt  (2)
42 following the methods described in (Chen, et al., 2022) Here, Wz is the weight matrix corresponding to the update gate. Next,
43
Table 1. Sample distribution of all cell lines in the dataset
we can compute the current candidate hidden state h% :
44 h%= tanh W ×  r e ht -1 , xt  (3)
45 cell lines EPIs Non-EPIs Here, W is the weight matrix used to compute the candidate hidden
46 state, and e denotes element-wise multiplication. Finally, we can
NHEK 1291 25600
47 HUVEC 1524 30400
compute the updated hidden state (h) using the update gate (z):
48 GM12878 2113 42200 h = (1 - z ) e h + z e h%
t t -1 (4)
49 IMR90 1254 25000 Here, e denotes element-wise multiplication. This equation indicates
50 HeLa 1740 34800 that the current hidden state (h) is a weighted average of the previous
51 K562 1977 39500 
hidden state (h) and the candidate hidden state h% , with the update gate
52 (z) regulating the weights between them. By utilizing both the update
53 gate and the reset gate, GRU can determine which information to pass,
54 2.2 Description of the proposed EPIPDLF ignore, or update, enabling it to handle long-term dependencies more
55 effectively. This enables GRU to excel in a variety of sequence modeling
56 2.2.1 Feature extraction applications, such as speech recognition and natural language processing.
57
58
59
60
Bioinformatics Page 4 of 18

[Link] et al.
1
2
Attention layer: multi-head self-attention mechanism is an attention the batch. Then, by performing linear transformation and translation
3
mechanism used for sequence data modeling and is often used in natural operations on the input, the mean is adjusted to 0 and the standard
4
language processing tasks. Below I will use mathematical formulas to deviation is adjusted to 1. Finally, a learnable scaling factor and
5
describe the calculation process of the multi-head self-attention translation factor are used to restore the original distribution of the data.
6
mechanism in detail. Suppose we have an input sequence, denoted as
7
X =  x1 ,x 2 ,...,x n  , where x i represents the i-th element in the sequence, 2.2.2 Model training
8

Downloaded from [Link] by guest on 05 March 2025


with n denoting the length of the sequence. We need to calculate the
9
correlation between each element and other elements in the sequence in Loss functions and optimization method
10
order to obtain global contextual information. The multi-head self- The proposed model EPIPDLF employs the Binary Cross-Entropy
11
attention mechanism captures different attention representations by (BCE) loss function, often utilized in binary classification scenarios. It
12
introducing multiple attention heads. Suppose we have h attention heads, quantifies the discrepancy between predicted values and actual
13
each head has its own parameter matrix for calculating attention weights. outcomes, training the model by minimizing this difference, as described
14
The calculation process of each head is divided into three steps: linear below:
15
transformation, attention weight calculation and weighted summation. 1
16 BCE Loss = - å i =1 éë yi log  yˆ i  + 1 - yi  log 1 - yˆ i  ùû (10)
N

First, we linearly transform the input sequence to map it to different N


17
query, key, and value spaces. For each attention head i, we define three Among them, yi is the real binary classification label, yˆ i represents the
18
sets of parameter matrices: Wᵢⁿ_q ∈ ℝᵈⁿ×ᵈᵠ Win _q Î ¡ dn ´dj is query model's predicted output, while N denotes the number of samples in a
19
transformation matrix, Win _k Î ¡ dn ´dj is key transformation matrix and batch. The Adam optimizer (Diederik, 2014) is employed to modify the
20 d ´d
Win _v Î ¡ n j is value transformation matrix. Among them, d n is the learnable weights within the neural network. One of the benefits of using
21
dimension of the input sequence, d j is the dimension of the query key, the Adam optimizer is its adaptability to different parameters.
22
and d v is the dimension of the value.
23
q i = Win _q × x i (5) Pre-training strategy
24
k i = Win _k × x i (6)
25 To enhance the model's generalization performance and its efficacy
v i = Win _v × x i (7)
26 during cross-cell line validation, we employed a pre-training strategy. By
Next, we calculate the attention weight of each element relative to
27 integrating the training data from six cell lines, we constructed a
other elements. We measure the correlation between two elements using
28 substantial pre-training dataset, denoted as 𝐷 . The model was pre-
a query-key dot product and normalize it through a scaling operation.
29 trained on this dataset. Following the pre-training phase, the training
Then, we perform a weighted sum of relevance and value to obtain the
30 process of the model can be described as follows:
contextual representation of each element. For each attention head i, we
31 1. Construct the pre-training dataset 𝐷, incorporating the training data
compute the attention weight A i Î ¡ n´n :
32 from all cell lines.
33 
A i = soft max qi × k T dj  (8)
2. Train model EPIPDLF on dataset 𝐷 for 15 epochs, with a learning rate
Among them, softmax represents the normalization operation on the
34 set to 0.001.
attention weight, and d j is a scaling factor. Next, we perform a
35 3. Fine-tune the pre-trained model on the training data from each of the
weighted sum of attention weights and values to obtain the contextual
36 six cell lines, applying 10 epochs and maintaining a learning rate of
representation of each element. For each attention head i, we compute
37 0.001 during this process.
the context representation ci Î ¡ n´d v :
38 4. Conduct predictions for the specific cell line and perform cross-cell
ci = A i × vi (9)
39 line validation experiments.
Finally, we splice or average the context representation of each
40
attention head to obtain the final multi-head self-attention representation
41 Adversarial learning strategy
C.
42 Further, adversarial training is also incorporated into the model training
Regularization mechanism: To prevent overfitting during the
43 process to enhance its robustness. The basic idea is to augment the
training phase, we have employed a series of regularization techniques,
44 training set with adversarial samples, enabling the model to learn from
primarily including dropout and batch normalization. Dropout is a
45 them during training. The introduction of adversarial learning
regularization method employed to mitigate overfitting in neural network
46 necessitates the simultaneous fitting of adversarial samples during our
models. In each training batch, Dropout randomly sets the output values
47 standard training process, which may, to some extent, reduce the training
of some neurons to zero, that is, discards the contributions of these
48 speed. Adversarial examples are created using the Projected Gradient
neurons. The purpose of this is to force the model not to depend on
49 Descent (PGD) approach. The PGD algorithm finds the most deceptive
particular neurons, thereby enhancing the model's robustness and capable
50 adversarial sample by iteratively applying gradient ascent and projection
of generalization. During the prediction phase, all neurons are retained
51 operations in the input space. The formula is described as follows:
and scaled by a retention probability to maintain model consistency.
52 X* = clip X ,Î  X + a × sign  Ñ X Loss ( X , Y )   (11)
Batch normalization is a technique used in deep neural networks to
53 *
normalize mini-batch inputs, setting each feature's mean to near 0 and Among them, X is the original input sample, X is the generated
54
standard deviation to near 1. This accelerates training and enhances the adversarial sample , Loss ( X , Y ) represents the model's loss function,
55
network's generalization. Specifically, for each mini-batch of input, while denotes the true label of the original sample. In every iteration of
56
Batch normalization first calculates the mean and standard deviation of the PGD algorithm, the loss function Loss ( X , Y ) is computed to obtain
57
58
59
60
Page 5 of 18 Bioinformatics

Article short title


1
2
3 the gradient with regard to the input X . Then, the gradient is multiplied 3.1 The proposed EPIPDLF outperforms the state‑of‑the‑art
4 by the learning rate a and transformed into the direction of the gradient methods
5 using the sign function sign(⋅). Next, the generated adversarial sample To assess the performance of our proposed model EPIPDLF, we
6 X + a × sign  Ñ X Loss ( X , Y )  undergoes a projection operation to ensure compared it against four leading predictors: PEP-WORD, SPEID,
7 that it stays within the range of Î . The projection operation employs the SIMCNN, and EPIANN. Each model was trained and tested using the
8 function clip X ,Î  × to constrain the adversarial sample within the ϵ-range

Downloaded from [Link] by guest on 05 March 2025


same datasets for each cell line. The training process for each comparator
9 of the original sample X, thereby maintaining the acceptability of the predictor followed the methods described in their respective references.
10 adversarial sample. The AUROC and AUPR results for EPIPDLF and the four predictors
11 across six cell lines are presented in Table 2 and Table 3, respectively.
12 EPIPDLF achieved the highest AUROC values in the HUVEC, HeLa,
2.3 Evaluation metrics
13 K562, and NHEK cell lines, with exceptions in GM12878 and IMR90. In
14 The dataset used for performance evaluation in this study is highly
Table 2, we observe that SIMCNN, apart from EPIPDLF, performs
15 imbalanced. Therefore, we employ the Area Under the Receiver
optimally in terms of the AUROC metric. Constructed using a
16 Operating Characteristic Curve (AUROC) (Hanley and McNeil, 1982;
straightforward CNN architecture, SIMCNN demonstrates that CNNs
17 Li, et al., 2024; Liu, et al., 2019; Zou, et al., 2023) and the Area Under
can effectively extract features associated with EPIs, thereby validating
18 the Precision-Recall Curve (AUPR) (Ai, et al., 2023; Davis and
the appropriateness of utilizing CNN modules in our approach.
19 Goadrich, 2006; Tang, et al., 2021) as evaluation metrics. The Receiver
Specifically, our model outperformed the second-best predictor by 0.2%,
20 Operating Characteristic (ROC) curve illustrates the relationship
1.5%, and 3.1% in HUVEC, HeLa, and NHEK, respectively. Similarly,
21 between sensitivity (on the vertical axis) and the false positive rate (1 -
EPIPDLF demonstrated superior AUPR performance across the six cell
22 specificity, on the horizontal axis) across various thresholds. The area
lines, excelling in HeLa and NHEK with improvements of 4.6% and
23 under this curve, AUROC, indicates model performance, with values
4.3% over the second-best model. Notably, PEP-WORD also exhibited
24 closer to 1 (corresponding to the upper-left curve) reflecting better
high performance in AUPR. However, compared to all other models,
25 performance (Zhu, et al., 2023). Since the ROC curve is unaffected by
PEP-WORD exhibits the poorest performance in terms of the AUROC
26 the distribution of positive and negative samples, AUROC is suitable for
metric, indicating that it places greater emphasis on predicting positive
27 assessing models in imbalanced binary classification scenarios.
samples during evaluation. In summary, our model demonstrates
28 Conversely, the precision-recall curve depicts the trade-off between
superior performance compared to other predictors, achieving excellence
29 precision (vertical axis) and recall (horizontal axis), emphasizing the
in both AUROC and AUPR metrics.
30 balance between precision and recall for positive samples. The area
under this curve, AUPR, measures model performance, with values Table 4. Demonstrating the impact of pre-training strategy and
31
nearing 1 (associated with the upper-right curve) indicating superior adversarial learning on AUC across six cell line datasets
32
33 performance. strategy/cell K562 NHEK
GM12878 HUVEC HeLa IMR90
34 Table 2. Comparison of AUC performance between EPIPDLF and other line
35 leading models on different cell line datasets no strategy 0.912 0.923 0.953 0.891 0.930 0.973
36 pre-training 0.937 0.922 0.951 0.932 0.941 0.979
model/cell line GM12878 HUVEC HeLa IMR90 K562 NHEK
37 pre+adversarial 0.946 0.944 0.961 0.926 0.942 0.984
EPIANN 0.919 0.918 0.924 0.945 0.943 0.959
38
SIMCNN 0.941 0.933 0.949 0.951 0.943 0.962
39 Table 5. Demonstrating the impact of pre-training strategy and
PEP-WORD 0.842 0.845 0.843 0.898 0.883 0.917
40 adversarial learning on AUPR across six cell line datasets
SPEID 0.916 0.904 0.923 0.915 0.922 0.950
41
EPIPDLF 0.939 0.935 0.964 0.936 0.943 0.993 strategy/cell K562 NHEK
42 GM12878 HUVEC HeLa IMR90
line
43
44 Table 3. Comparison of AUPR performance between EPIPDLF and no strategy 0.715 0.661 0.805 0.708 0.741 0.874
other leading models on different cell line datasets pre-training 0.769 0.704 0.836 0.754 0.746 0.878
45
46 pre+adversarial 0.772 0.720 0.836 0.766 0.744 0.881
model/cell line GM12878 HUVEC HeLa IMR90 K562 NHEK
47 EPIANN 0.723 0.616 0.702 0.770 0.673 0.861
48 SIMCNN 0.706 0.640 0.737 0.737 0.679 0.882 3.2 Contributions of pre-trained strategy and adversarial
49 PEP-WORD 0.807 0.760 0.803 0.868 0.836 0.880 training
50 SPEID 0.773 0.523 0.797 0.732 0.771 0.852 Considering practical applications, EPI prediction models are often
51 EPIPDLF 0.788 0.730 0.849 0.779 0.755 0.925 required to generalize across different cell lines. Therefore, the ability to
52 predict EPIs across cell lines is particularly crucial. To improve the
53
54 3 Results prediction capability across different cell lines, we suggest two
strategies. The first strategy is pretraining, a common technique in
55 transfer learning (Han, et al., 2021) to improve model generalization. We
56 initially pretrained our model on sequence data from all cell lines and
57
58
59
60
Bioinformatics Page 6 of 18

[Link] et al.
1
2
subsequently fine-tuned the pretrained model on individual cell lines to To validate whether genomic information serves as an effective feature
3
derive the final models. In the previous section on model training, we input for EPIPDLF. Seven categories of genomic features were
4
provided a detailed description of the pre-training process. To further evaluated, including chromatin accessibility (DNase-I signals), CTCF-
5
enhance the model's cross-cell-line prediction capability, we integrated binding sites, and five histone modifications: H3K4me1, H3K9me3,
6
adversarial training (Madry, et al., 2017) into the pretraining process, H3K27me3, H3K4me3, and H3K36me3, following the methodology
7
which is a crucial component of our pretraining strategy. First, we aim to detailed in (Chen, et al., 2022). For the input genomic features, we
8

Downloaded from [Link] by guest on 05 March 2025


evaluate the effects of these two strategies on the validation of individual initially used a multilayer perceptron (MLP) for reconstruction, followed
9
cell lines. Tables 4 and 5 illustrate the impact of the pre-training strategy by the application of multi-head attention to extract distal features. These
10
and adversarial learning on a single cell line. The results indicate that the were then fused with sequence features for the final prediction. As
11
optimal performance is achieved when both strategies are employed shown in Table 6, incorporating genomic features resulted in
12
simultaneously. Additionally, pre-training alone significantly enhances performance improvements across six different cell lines. This
13
the AUC and AUPR values, demonstrating that both pre-training and demonstrates that genomic information can indeed enhance the
14
adversarial learning are beneficial for model performance. To further predictive performance of EPIs, provided that the genomic data is readily
15
validate the impact of these two strategies on cross-cell line validation, accessible.
16
we conducted experiments, with results presented in the heatmap shown
17
in Figure 3. The performance improvement due to pre-training in cross-
18 3.4 The biologically meaningful motifs learned by our CNN
cell line validation is anticipated; notably, the enhancement from
19 sequence learning module
adversarial learning is more pronounced in cross-cell line validation than
20 Next, we want to further explore whether the motifs learned by the
in the corresponding cell line validation. This is attributed to the
21 model from sequence data have biological significance. Referring to
substantial increase in robustness of the model following adversarial
22 reference (Jin, et al., 2022) , we extracted the motifs recognized by the
training.
23 model from the convolutional kernels corresponding to enhancers and
24 Table 6. Improvement of Model Performance by Gene Data
promoters, and matched them with the motifs in the JASPAR (Sandelin,
25 gene K562 NHEK et al., 2004) database. JASPAR is a well-known public transcription
26 cell line
data
GM12878 HUVEC HeLa IMR90
factor database. Furthermore, we used TOMTOM (Gupta, et al., 2007) to
27 Yes 0.939 0.935 0.964 0.936 0.943 0.993 calculate the similarity between the two motifs and use p-values to
28 AUC
No 0.946 0.944 0.961 0.926 0.942 0.983 measure the similarity. The lower the p-value, the higher the consistency
29 Yes 0.788 0.730 0.849 0.779 0.755 0.925 of the motifs. In Figures 3 and 4, the upper panel illustrates the motifs
30 AUPR
No 0.772 0.720 0.836 0.766 0.744 0.881 identified by EPIPDLF, while the lower panel displays similar motifs
31 from the database. A comparison reveals that the patterns learned by the
32 model from the sequences (enhancers and promoters) effectively match
33 certain functional motifs in the transcription factor database, which have
34 3.3 Genomic information effectively improves the previously been reported to be associated with EPI. This result
35 performance of EPIPDLF demonstrates that our model can effectively extract useful features from
36
37
38
39
40
41
42
43
44
45
a. AUC value without any strategies b. AUC value after pre-training c. AUC value after pre-training and adversarial learning
46
47
48
49
50
51
52
53
54
55
d. AUPR value without any strategies e. AUPR value after pre-training f. AUPR value after pre-training and adversarial learning
56
Fig. 2. The impact of pre-training strategy and adversarial learning on cross-cell line model validation
57
58
59
60
Page 7 of 18 Bioinformatics

Article short title


1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22 Fig. 3. Comparison between motifs identified by the model in enhancer sequences and those in JASPAR
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42 Fig. 4. Comparison between motifs identified by the model in promoter sequences and those in JASPAR
43
44
sequence data, and the high degree of matching with database motifs model training can be beneficial. In comparative experiments across six
45
proves that the motifs identified by the model have sufficient biological cell lines, EPIPDLF consistently outperformed other existing models,
46
significance. demonstrating its effectiveness. More importantly, we proposed two
47
strategies, pretraining and adversarial learning, to enhance the model's
48
capability for cross-cell-line EPI prediction. Experimental results showed
49 4 Conclusions
that pretraining significantly improves the model’s cross-cell-line
50 In this paper, we present EPIPDLF, an innovative method for identifying generalization ability, while adversarial learning enhances model
51 EPIs that relies solely on genomic sequence data. Compared to existing robustness, thereby improving its cross-cell-line prediction performance.
52 models, EPIPDLF incorporates pretrained DNABERT embedding Additionally, we validated whether the motifs learned by the
53 matrices and a multi-head self-attention mechanism, enhancing its ability convolutional neural networks (CNNs) layer of EPIPDLF possess
54 to capture latent features within sequences. Furthermore, we validated biological significance. Our findings indicate that the model can
55 that related genomic information effectively aids EPI identification, accurately identify functional sequence features, and the newly identified
56 suggesting that incorporating genomic data associated with EPIs into motifs also have biological significance.
57
58
59
60
Bioinformatics Page 8 of 18

[Link] et al.
1
2
We recognize the potential of pretrained large models in Hong, Z., et al. Identifying enhancer–promoter interactions with neural network
3
bioinformatics. However, we did not fully exploit the pre-trained based on pre-trained DNA vectors and attention mechanism. 2020;36(4):1037-
4
DNABETRT; instead, we only extracted its embedding matrix. In future 1043.
5
work, we aim to investigate whether more advanced pretrained models Jin, J., et al. iDNA-ABF: multi-scale deep biological language learning model for
6
can be employed for biological sequences to obtain more informative the interpretable prediction of DNA methylations. 2022;23(1):219.
7
pretrained DNA embedding matrices (Zou, et al., 2019). Notably, Li, H., Pang, Y. and Liu, B. BioSeq-BLM: a platform for analyzing DNA, RNA,
8

Downloaded from [Link] by guest on 05 March 2025


DNABERT-2 (Zhou, et al., 2023) has recently been published. EPIPDLF and protein sequences based on biological language models. Nucleic Acids
9
utilizes convolution and pooling to capture abstract local features Research 2021;49(22):e129.
10
without preserving the positional information of k-mers. Therefore, Li, Q., et al. Identification and classification of promoters using the attention
11
exploring how to incorporate additional features for EPI prediction mechanism based on long short-term memory. Frontiers Of Computer Science
12
remains a promising direction for future research. 2022;16(4).
13
Li, X., et al. TranSiam: Aggregating multi-modal visual features with locality for
14
medical image segmentation. Expert Systems with Applications 2024;237.
15 Funding
Liu, B., Gao, X. and Zhang, H. BioSeq-Analysis2.0: an updated platform for
16 This work was financially supported by the National Natural Science
analyzing DNA, RNA and protein sequences at sequence level and residue level
17 Foundation of China (Grant No. 62472344, 62072353, 62272065, 62172076
based on machine learning approaches. Nucleic Acids Research 2019;47(20):e127.
18 and U22A2038), Xidian University Specially Funded Project for
Liu, Y., et al. Sequence Alignment/Map format: a comprehensive review of
19 Interdisciplinary Exploration (No. TZJH2024027), the Municipal Government
approaches and applications. Briefings in Bioinformatics 2023;24(5):bbad320.
20 of Quzhou (Grant Number 2023D038) and the Zhejiang Provincial Natural
Science Foundation of China (Grant No. LY23F020003). Madry, A., et al. Towards deep learning models resistant to adversarial attacks.
21
2017.
22
Conflict of Interest: none declared. Mao, W., Kostka, D. and Chikina, M.J.b. Modeling enhancer-promoter interactions
23
with attention-based neural networks. 2017:219667.
24
Ni, P., Moe, J. and Su, Z.J.B.b. Accurate prediction of functional states of cis-
25 References
regulatory modules reveals common epigenetic rules in humans and mice. BMC
26 Ai, C., et al. Low Rank Matrix Factorization Algorithm Based on Multi-Graph biology 2022;20(1):221.
27 Regularization for Detecting Drug-Disease Association. Ieee-Acm Transactions on Qiao, J., et al. Towards Retraining-free RNA Modification Prediction with
28 Computational Biology and Bioinformatics 2023;20(5):3033-3043. Incremental Learning. Information Sciences 2024:120105.
29 Boser, B.E., Guyon, I.M. and Vapnik, V.N. A training algorithm for optimal Rakhlin, A.J.G. Convolutional neural networks for sentence classification.
30 margin classifiers. In, Proceedings of the fifth annual workshop on Computational 2016;6:25.
31 learning theory. 1992. p. 144-152. Rao, S.S., et al. A 3D map of the human genome at kilobase resolution reveals
32 Cao, F., et al. Chromatin interaction neural network (ChINN): a machine learning- principles of chromatin looping. 2014;159(7):1665-1680.
33 based method for predicting chromatin interactions from DNA sequences. Ren, X., et al. HydrogelFinder: A Foundation Model for Efficient Self ‐
34 2021;22:1-25. Assembling Peptide Discovery Guided by Non ‐ Peptidal Small Molecules.
35 Chen, J., Zou, Q. and Li, J. DeepM6ASeq-EL: Prediction of Human N6- Advanced Science 2024:2400829.
36 Methyladenosine (m6A) Sites with LSTM and Ensemble Learning. Frontiers of Sandelin, A., et al. JASPAR: an open‐access database for eukaryotic transcription
37 Computer Science 2022;16(2):162302. factor binding profiles. 2004;32(suppl_1):D91-D94.
38 Chen, K., Zhao, H. and Yang, Y.J.B.i.B. Capturing large genomic contexts for Tang, Y., Pang, Y. and Liu, B. IDP-Seq2Seq: identification of intrinsically
39 accurately predicting enhancer-promoter interactions. 2022;23(2):bbab577. disordered regions based on sequence to sequence learning. Bioinformatics
40 Chung, J., et al. Empirical evaluation of gated recurrent neural networks on 2021;36(21):5177-5186.
41 sequence modeling. 2014. Wang, Y., Zhai, Y., Ding, Y., Zou, Q. SBSM-Pro: Support Bio-sequence Machine
42 Cox, D.R.J.J.o.t.R.S.S.S.B.S.M. The regression analysis of binary sequences. for Proteins. arXiv preprint 2023:arXiv:2308.10275.
43 1958;20(2):215-232. Wei, L., et al. Computational prediction and interpretation of cell-specific
44 Davis, J. and Goadrich, M. The relationship between Precision-Recall and ROC replication origin sites from multiple eukaryotes by exploiting stacking framework.
45 curves. In, Proceedings of the 23rd international conference on Machine learning. Briefings in Bioinformatics 2020.
46 2006. p. 233-240. Whalen, S., Truty, R.M. and Pollard, K.S.J.N.g. Enhancer–promoter interactions
47 Devlin, J., et al. Bert: Pre-training of deep bidirectional transformers for language are encoded by complex genomic signatures on looping chromatin.
48 understanding. 2018. 2016;48(5):488-496.
49 Diederik, P.K.J. Adam: A method for stochastic optimization. 2014. Yang, Y., et al. Exploiting sequence-based features for predicting enhancer–
50 Gupta, S., et al. Quantifying similarity between motifs. 2007;8:1-9. promoter interactions. 2017;33(14):i252-i260.
51 Han, X., et al. Pre-trained models: Past, present and future. 2021;2:225-250. Yin, C., et al. NanoCon: contrastive learning-based deep hybrid network for
52 Hanley, J.A. and McNeil, B.J.J.R. The meaning and use of the area under a receiver nanopore methylation detection. Bioinformatics 2024;40(2):btae046.
53 operating characteristic (ROC) curve. 1982;143(1):29-36. Zhang, Z.Y., et al. A BERT-based model for the prediction of lncRNA subcellular
54 Heidari, N., et al. Genome-wide map of regulatory interactions in the human localization in Homo sapiens. International journal of biological macromolecules
55 genome. 2014;24(12):1905-1917. 2024;265(Pt 1):130659.
56
57
58
59
60
Page 9 of 18 Bioinformatics

Article short title


1
2
3 Zhou, Z., et al. Dnabert-2: Efficient foundation model and benchmark for multi-
4 species genome. 2023.
5 Zhu, W., et al. A First Computational Frame for Recognizing Heparin-Binding
6 Protein. Diagnostics (Basel) 2023;13(14).
7 Zhuang, Z., Shen, X. and Pan, W.J.B. A simple convolutional neural network for
8 prediction of enhancer–promoter interactions with DNA sequence data.

Downloaded from [Link] by guest on 05 March 2025


9 2019;35(17):2899-2906.
10 Zou, Q., et al. Gene2vec: Gene Subsequence Embedding for Prediction of
11 Mammalian N6‐Methyladenosine Sites from mRNA. RNA 2019;25(2):205-218.
12 Zou, X., et al. Accurately identifying hemagglutinin using sequence information
13 and machine learning methods. Front Med (Lausanne) 2023;10:1281880.
14 Zulfiqar, H., et al. Deep-STP: a deep learning-based approach to predict snake
15 toxin proteins by using word embeddings. Frontiers in Medicine 2024;10.
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Bioinformatics Page 10 of 18

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29 Fig. 1. The Workflow Framework of EPIPDLF
30
31 474x329mm (236 x 236 DPI)
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Page 11 of 18 Bioinformatics

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Fig. 2.(a. AUC value without any strategies)
33
34 645x516mm (236 x 236 DPI)
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Bioinformatics Page 12 of 18

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Fig. 2.(b. AUC value after pre-training)
33
34 645x516mm (236 x 236 DPI)
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Page 13 of 18 Bioinformatics

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Fig. 2.(c. AUC value after pre-training and adversarial learning)
33
34 645x516mm (236 x 236 DPI)
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Bioinformatics Page 14 of 18

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Fig. 2.(d. AUPR value without any strategies)
33
34 645x516mm (236 x 236 DPI)
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Page 15 of 18 Bioinformatics

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Fig. 2. (e. AUPR value after pre-training)
33
34 604x483mm (130 x 130 DPI)
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Bioinformatics Page 16 of 18

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Fig. 2. (f. AUPR value after pre-training and adversarial learning)
33
34 604x483mm (130 x 130 DPI)
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Page 17 of 18 Bioinformatics

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24 Fig. 3. Comparison between motifs identified by the model in enhancer sequences and those in JASPAR
25
26 641x345mm (157 x 157 DPI)
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Bioinformatics Page 18 of 18

1
2
3
4
5
6
7
8

Downloaded from [Link] by guest on 05 March 2025


9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24 Fig. 4. Comparison between motifs identified by the model in promoter sequences and those in JASPAR
25
26 640x345mm (157 x 157 DPI)
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60

You might also like