0% found this document useful (0 votes)
16 views12 pages

Predicting ECD in Drilling with ML

The paper presents a statistical machine learning model using support vector machines (SVM) and principal components analysis (PCA) to predict equivalent circulation density (ECD) while drilling. By analyzing actual field data, the model aims to improve the accuracy of ECD predictions by incorporating factors like bottom hole temperature and pipe rotation, which traditional mathematical equations often overlook. The results indicate that the PCA-based SVM model can effectively predict ECD, thereby enhancing drilling efficiency and safety.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views12 pages

Predicting ECD in Drilling with ML

The paper presents a statistical machine learning model using support vector machines (SVM) and principal components analysis (PCA) to predict equivalent circulation density (ECD) while drilling. By analyzing actual field data, the model aims to improve the accuracy of ECD predictions by incorporating factors like bottom hole temperature and pipe rotation, which traditional mathematical equations often overlook. The results indicate that the PCA-based SVM model can effectively predict ECD, thereby enhancing drilling efficiency and safety.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SPE/IADC-202101-MS

A Statistical Machine Learning Model to Predict Equivalent Circulation


Density ECD while Drilling, Based on Principal Components Analysis PCA

Ahmed AlSaihati, Salaheldin Elkatatny, Hani Gamal, and Abdulazeez Abdulraheem, KFUPM

Copyright 2021, SPE/IADC Middle East Drilling Technology Conference and Exhibition

This paper was prepared for presentation at the SPE/IADC Middle East Drilling Technology Conference and Exhibition held in Abu Dhabi, UAE, 25 - 27 May 2021.

This paper was selected for presentation by an SPE/IADC program committee following review of information contained in an abstract submitted by the author(s).
Contents of the paper have not been reviewed by the Society of Petroleum Engineers or the International Association of Drilling Contractors and are subject to correction
by the author(s). The material does not necessarily reflect any position of the Society of Petroleum Engineers or the International Association of Drilling Contractors,
its officers, or members. Electronic reproduction, distribution, or storage of any part of this paper without the written consent of the Society of Petroleum Engineers or
the International Association of Drilling Contractors is prohibited. Permission to reproduce in print is restricted to an abstract of not more than 300 words; illustrations
may not be copied. The abstract must contain conspicuous acknowledgment of SPE/IADC copyright.

Abstract
Mathematical equations, based on conservation of mass and momentum, are used to determine the ECD
at different depths in the wellbore. However, such equations do not consider important factors that have
a influence on the ECD such as: (i) bottom hole temperature, (ii) pipe rotation and eccentricity, and (iii)
wellbore roughness. Thus, discrepancy between the calculated ECDs and actual ones has been reported in
the literature.
This paper aims to explore how artificial intelligence (AI) and machine learning (ML) could provide
real-time accurate prediction of the ECD, to have more insight and management of wellbore downhole
conditions. For this purpose, a supervised ML algorithm, support vector machine (SVM), based on principal
components analysis (PCA), was developed.
Actual field data of Well-1 including drilling surface parameters and ECDs, measured by downhole
sensors, were collected to develop a classical SVM model. The dataset was split with an 80/20 training-
testing data ratio. Sensitivity analysis with different SVM parameters such as regularization parameter C,
gamma, kernel type (linear, radial basis function "RBF") was performed. The performance of the model was
assessed in terms of root mean square error (RMSE) and coefficient of determination (R2). Afterward, PCA
was applied to the dataset of Well-1 to develop an SVM model using the transformed dataset in PCA space.
The performance of the model while using different numbers of principal components was evaluated.
The results showed that the classical SVM with the linear kernel predicted the ECD with RMSE of 0.53
and R2 of 0.97 in the training set, while RMSE and R2 were 0.56 and 0.97 respectively in the testing set.
The PCA-based SVM model, with the linear kernel and four principal components (93.53% variation of the
dataset), predicted the ECD with RMSE 0.79 and R2 of 0.95 in the testing set.

Introduction
Drilling fluid hydraulics has an important role in well design when drilling a vertical or extended-reach well.
Therefore, accurate model and optimized drilling fluid hydraulics are essential to allow engineers to properly
design a well profile, thereby improving the drilling efficiency. Additionally, it empowers the drilling crew
to examine the downhole conditions and identify potential problems such as drill string washout, plugged
nozzles and the presence of a gas kick in the well.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
2 SPE/IADC-202101-MS

ECD is an important aspect of drilling fluid hydraulics which helps in avoiding kicks and drilling fluid
losses, particularly in deep HPHT wells where the temperature is significant and the margin between pore
pressure and fracture pressure is narrow (Rommeveit and Bjorkevoll,1997). The ECD can be expressed as:
(1)
where ECD is the equivalent circulation density (ppg), Ps is the frictional pressure drop in the annuals (psi),
TVD is the true vertical depth (ft), MW is the drilling fluid density (ppg).
Hydraulics programs have been widely used to evaluate the detailed hydraulic calculations that are
required in planning a well. The user has the option to select which rheological model (Bingham plastic,
Power-law, Herschel–Bulkley) the hydraulic calculations are based on. Each rheological model, however,
requires different inputs, which can be obtained from a full lab test. These parameters are then fed into the
software to compute the drilling fluid hydraulics and its associated parameters such as ECD. It has been
observed that there is a discrepancy between computed data and the ones recorded in the field (Maglion
and Robotti, 1996).
One causative factor to the inconsistency between calculated and actual drilling hydraulics is the
assumption that the rheological properties of the drilling fluid are independent of pressure and temperature
(Rommeveit and Bjorkevoll,1997). This can be valid in shallow wells, where temperature changes are
insignificant. Furthermore, the mathematical equations used to calculate the drilling hydraulics have a set
of assumptions such as (Maglion and Robotti, 1996; Dokhani at al., 2013; Erge et al., 2016): (i) concentric
annular and circular sections, (ii) laminar and turbulent flow, where plug flow is considered as laminar flow,
while transition flow is neglected, (iii) steady-state flow, where fluid properties at any single point in the
system do not change. Such potentially assumptions cannot be fulfilled in all drilling conditions (Rabia,
2001).
In HPHT wells, evaluations and analysis of the effect of temperature and pressure on drilling hydraulics
and kick probability are necessary (Rommeveit and Bjorkevoll,1997; Isambourg et al., 1996). Rommeveit
and Bjorkevoll, (1997) developed two models, a static ECD model and a dynamic ECD model, which
incorporate the temperature profile along the wellbore and allow the mud properties to be dependent on
pressure and temperature. They concluded that the models made it possible to make reliable evaluations
of any operational concerns during drilling. Scheid et al. (2009) performed an experiment to evaluate the
frictional pressure loss through pipes, annuli, and accessories such as pipe tool joint, and stabilizers to
estimate the drilling hydraulic calculations and the corresponding ECD with four types of drilling fluid.
They stated that the results obtained in the study can be used in the drilling industry for accurate drilling
hydraulic calculations.
Dokhani et al. (2013) concluded that eccentricity (Є) Eq. 2, affects frictional pressure loss and hence the
ECD for Herschel–Bulkley fluid.
(2)
where Є is the eccentricity (unitless), λ is the difference between the center position of the inner pipe and
the wellbore (inch), R0 is the wellbore radius (inch), Ri is the inner pipe diameter (inch).
Eccentricity reduces overall frictional pressure loss if є is > 0.1 and it will be neglected if Є is 0.8
(Cartlos and Dupuis, 1993; Escudier et al.,2000; Mokhtari et al., 2012). Moreover, Dokhani et al. (2013)
evaluated the effect of pipe rotations on frictional pressure loss. They observed that a drilling fluid's apparent
viscosity decreases as pipe rotation increases, hence overall frictional loss decreases. This phenomenon is
a result of the shear-thinning properties of drilling fluid (Hemphill and Ravi,2005). Dokhani et al. (2013)
also recommended considering the effect of roughness of the wellbore wall to have an accurate estimation
of drilling hydraulic calculations such as ECD.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
SPE/IADC-202101-MS 3

A more accurate way to evaluate the ECD is to use a downhole pressure sensor. The sensors provide
important real-time downhole pressure information which will allow the drilling crew to make a faster and
better decision (Halliburton, 2020). However, such sensors are expensive (Abdelgawad et al., 2018).
The literature review shows that major discrepancies exist between actual drilling hydraulic values and
those predicted by previously accepted mathematical equations. In addition, several studies have suggested
a consideration of other factors including pipe eccentricity, wellbore roughness, pressure and temperature,
and pipe rotation to improve the accuracy of drilling hydraulic calculations. This, however, would mandate
mathematical computations.
An alternative approach to estimating the ECD with higher accuracy is by using AI and ML (Abdelgawad
et al, 2018; Alkinani et al, 2019), which link the inputs and outputs based on a defined algorithm (Balaji et
al., 2018). AI has become an important subject in the drilling industry (Balaji et al., 2018), where streaming
data is continuously structured from surface and downhole sensors and logging. AI and ML allow for new
methods to learn from these data to mitigate drilling challenges and describe trends in real time (Al-Ghazal
and Vedpoathak, 2019).

Application of AI and ML in the Drilling Industry


There are some applications related to AI and ML in drilling operations such as (Bello et al., 2015): well
planning, rate of penetration (ROP) optimization, well integrity, detecting problems, procedure decision
making, and pattern recognition. The large number of publications on the application of AI and ML in the
drilling industry indicates that AI and ML can potentially reduce drilling cost and promote safety at the rig-
site (Bello et al., 2015; Al-Ghazal and Vedpoathak, 2019).
Abdelgawad et al. (2018) used standpipe pressure (SPP), ROP, and mud weight (MW) as input parameters
to predict the ECD using an artificial neural network (ANN) and adaptive neuro-fuzzy inference system
(ANFIS) in an 8-1/2-inch vertical hole section. The model predicted the ECD with a high correlation (R) of
0.99 and an absolute average percentage error (AAPE) of 0.22% for ANN and ANFIS respectively. Alkinani
et al. (2019) collected data from more than 2,000 wells located around the world and used an ANN to build a
model to predict the ECD. The input parameters for the model were flow rate, MW, nozzles total flow area,
plastic velocity, revolutions per minute, weight on bit (WOB), and yield point. Bayesian Regularization
(BR) was used as a training algorithm because it had the highest R2 (0.982).
Temizel et al. (2016) studied the factors which contribute to the performance of vertical and horizontal
wells using data-driven models. Al-Yami et al. (2016) used Bayesian Belief Network (BBN) to establish an
intelligent drilling system based on different fluid and reservoir properties. This tool can be utilized to train
young engineers in different drilling perspectives such as well control, underbalanced drilling, drilling fluid,
and cementing best practices. Ahmadi. (2016) simulated the performance of various types of drilling fluid's
rheology under different conditions using an SVM with good agreement between lab results and prediction.
Elkatatny et al. (2017) used an ANN, which incorporates drilling fluid and drilling mechanical properties,
to predict ROP with high accuracy. Additional work has been performed for providing new approaches
to optimize the rate of penetration using artificial neural network (Elkatatny 2018; Al-AbdulJabbar et al.
2019). Kamel et al. (2018) presented an adaptive and real-time optimal control of stick–slip and bit wear in
autonomous rotary steerable drilling. Al-Abduljabbar et al. (2018) predicted formation tops during drilling
based on actual field data with high accuracy. They concluded that the developed ANN model can potentially
replace any other expensive techniques to pick formation tops accurately.
Static Young's modulus is another rock parameter that was predicted using different AI tools (Elkatatny
et al. 2019), and new correlation was developed for estimating the static Young's modulus (Elkatatny et
al. 2018a). Tariq et al (2017a) presented a new technique to develop the rock strength correlation using
artificial intelligence tools. Elkatatny et al. (2018b) developed a new mathematical model for compressional
and shear sonic times from wireline log data using ANN and provided the mathematical correlation for
estimating the sonic times from the logs. Tariq et al (2017b) predicted the rock failure parameter of carbonate

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
4 SPE/IADC-202101-MS

rocks using AI tools. Moussa et al. (2018) presented a new permeability formulation to estimate the rock
permeability from well logs using artificial intelligence approaches.
This paper introduces a PCA-based SVM model to predict the ECD, which would give the drilling crew
the capability to monitor hole cleaning, maintain the wellbore pressures in horizontal extended-reach wells
in real time, and reduce the risk of formation fracture and collapse.

Principal Component Analysis


PCA is a useful statistical technique that has many applications in many fields, and is a common technique
for finding patterns in a dataset with high dimensions (smith, 2002). The objective of PCA is to reduce the
dimensionality of a large dataset of variables or features into a smaller one without loss of any important
information (Tharwat, 2016). PCA finds a lower dimensional space (W) to transform the data (X= [x1, x2,...,
xN]) from a higher dimensional space (RM) to a lower dimensional space (Rk),where N represents the total
number of observations (rows in a dataset) and xi represents ith observation.
PCA space for a dataset that contains a number of features K, has K principal components. These K
principal components are uncorrelated, orthonormal and represent the direction of the maximum variance
(Tharwat, 2016). The first component (PC1 ϵ RM) always represents the maximum variation of the data, (PC2
ϵ RM) represents the second-largest variation of the data, and so on (Tharwat, 2016; Smith, 2016). Figure 1
shows an example of a dataset containing two variables (x1, x2). The original dataset, before PCA is applied,
is on the left with initial coordinates x1, x2, while on the right is the original dataset projected on PC1 and PC2.

Figure. 1—Example of the two-dimensional dataset before and after applying PCA (Tharwat, 2016).

Support Vector Machines


SVMs are supervised ML models that analyze data for classification or regression problems (Bello et al.,
2015), and can lead to a high performance in particular applications (Hearst, 1989). Moreover, SVMs
have advantages over other ML algorithms including generalization capability, less learning time, and
strong interference capacity (Vapnik, 1995; Anifowose and Abdulraheem, 2011). SVMs move the data
from a low dimension to a high dimension, denoted as kernel space, to find a support vector classifier, i.e.
hyperplane, that minimizes the number of misclassified data points (Durka, 2011; Bello et al., 2015; Awad
and Khanna,2015). The hyperplane is defined as the plane with a maximal margin of separation between
two classes, and the distance between the hyperplane and the nearest data point from either set is known as
the margin. Support vectors are the data with the closest distance to the hyperplane. These are difficult to
classify (Durka, 2011), which is why the hyperplane with the maximum margin is the best separator.
To transform the data to kernel space, SVMs use kernel functions to systematically find the support vector
classifiers in higher dimensions. When kernels are used to transform the feature vectors from input space
to kernel space for linearly non-separable datasets, the kernel matrix computation requires computational

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
SPE/IADC-202101-MS 5

resources. The popular kernel functions are (Awad and Khanna,2015): linear kernel function, polynomial
kernel function, Gaussian RBF, and randomized blocks analysis of variance (ANOVA RB) kernel.
The selection of kernel functions is essentially dependent on the nature of the dataset. The linear kernel
ranks behind the polynomial kernel and it is useful in large sparse data vectors. On the other hand, the
polynomial kernel is commonly used in image processing. While the ANOVA RB kernel is generally
for regression tasks, the Gaussian RBF is mostly applied if the user lacks prior knowledge. (Awad and
Khanna,2015).

Method
Data Collection
Two types of actual data of Well-1 were collected from a 5-7/8-inch horizontal section to develop a classical
SVM model that helps in predicting ECD in real time: (i) drilling surface parameters, and (ii) ECDs. The
drilling surface parameters including flow rate (Q), hock-load (HL), ROP, rotary speed (RS), SPP and
surface drilling torque (T) were obtained from surface real-time transmitter sensors, while ECDs were
obtained from a pressure-while-drilling sensor, PWD. A total of 4,131 data points of ECD were obtained
at the same depth of the drilling surface parameters. Table 1 shows the statistical parameters of the whole
dataset. Q ranges from 250 to 296 GPM; HL ranges from 256 to 286 klbf; ROP ranges from 3.5 to 59.6 ft/
hr.; RS ranges from 50 to 141 RPM; SPP ranges from 2,354.8 to 3,656.5 psi; WOB ranges from 5.1 to 20.1
klbf; T ranges from 2.9 to 10 [Link]; and the ECD ranges from 83.4 to 95.53 pcf.

Table 1—Statistical parameters of the whole dataset (4,131 data points).

Statistical Q (GPM) HL (klbf) ROP (ft/hr.) RS (RPM) SPP (psi) WOB (klbf) T ([Link]) ECD (Pcf)
parameters

Min 250.0 256.0 3.5 50.0 2354.8 5.1 2.9 83.4

Max 296.0 286.0 59.6 141.0 3656.5 20.1 10.0 95.53

Range 47.2 30.1 56.1 91.3 1301.7 15.0 7.2 12.1

Mean 276.6 267.4 23.0 119.6 3034.7 15.2 6.9 90.39

Data Splitting
In ML, it is required to build a model that makes accurate predictions for future data. Thus, the dataset
is divided into two portions, training and testing sets. The training set is used to ensure that the machine
recognizes patterns in the dataset, while the testing set is used to evaluate how well the machine can predict
unseen data based on its training.
In this analysis, seven surface drilling parameters were used as independent variables (inputs): Q, HL,
ROP, RS, SPP, WOB, and T, while ECD was used as a dependent variable (output). The dataset was
randomly split, with a ratio of 80:20. The training set has 3,304 data points, while the testing set has 827
data points. Tables 2 and 3 show the statistical parameters of the training and testing sets respectively.

Table 2—Statistical parameters of the training set (3,304 data points).

Statistical Q (galUS/ HL (klbf) ROP (ft/h) RS (RPM) SPP (psi) WOB (klbf) T ([Link]) ECD (Pcf)
parameters min)

Min 249.41 256.13 3.50 50.00 2370.16 5.18 2.85 82.96

Max 296.58 286.80 59.64 141.33 3657.05 20.04 10.01 95.53

Range 47.17 30.67 56.14 91.33 1286.90 14.86 7.16 12.57

Mean 276.65 267.45 23.03 119.64 3038.11 15.20 6.91 90.44

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
6 SPE/IADC-202101-MS

Table 3—Statistical parameters of the testing set (827 data points).

Statistical Q (galUS/ HL (klbf) ROP (ft/h) RS (RPM) SPP (psi) WOB (klbf) T ([Link]) ECD (Pcf)
parameters min)

Min 250.29 256.07 4.93 59.00 2371.83 5.08 3.16 83.44

Max 296.04 286.15 56.71 140.20 3657.54 20.09 10.01 95.53

Range 45.75 30.08 51.78 81.20 1285.71 15.02 6.85 12.09

Mean 276.65 267.27 22.83 119.37 3019.58 15.13 6.81 90.21

Data Standardization
Standardizing a dataset refers to shifting the distribution of each variable to have a unit scale, i.e. a mean
of zero and a standard deviation of one, which is a necessary step for the PCA algorithm. The values of the
input parameters were standardized using the following equation:
(3)
where Y is the normalized input parameter, X is the input parameter to be normalized, σ is the standard
deviation of the input variable.

Building SVM Model


Python library's Scikit-Learn® was used to build the classical SVM model. The SVM parameters, known
as hyper-parameters (C, gamma, kernel type), were tuned using a built-in function in Scikit-Learn known
as GridSearchCV to evaluate the improvement and performance of the classical SVM. Two types of kernel,
RBF and linear, were tried while varying the value of C from 0.0001 to 1000 and the type of gamma (i.e.
auto and scale). Scikit-Learn® was used to apply the PCA algorithm on input variables of the dataset of
Well-1. The transformation to PCA space was completed in three steps: (i) instantiate the PCA by passing
the number of principal components to the constructor, (ii) call the fit, which will find the covariance matrix,
the eigenvectors and eigenvalues of the covariance matrix, and (iii) transform the dataset into the PCA
space. Then transformed dataset was fed into the SVM model for training and testing. The PCA-based SVM
model improvement and performance while using different numbers of PCs and different parameters was
evaluated and compared with the classical SVM model.

Results and Discussion


Model Assessment
The classical SVM model performance with the optimum parameters for each kernel type is presented
in Table 4. The optimal hyper-parameters were obtained using GridSearchCV. Table 4 shows that the
performance of the classical SVM model in terms of RMSE and R2 with the linear kernel was more robust
compared to the RBF kernel. The classical SVM with the linear kernel predicted the ECD with R2 of 0.97
and RMSE of 0.53 in the training set, while R2 and RMSE were 0.97 and 0.56 respectively in the testing set.
On the other hand, the classical SVM with the RBF kernel predicted the ECD with R2 of 0.97 and RMSE
of 0.53 in the training set, and with R2 and RMSE of 0.97 and 0.57 respectively in the testing set. Figures 2
and 3 are cross-plots of the actual and predicted ECD of the training and testing sets for the classical SVM
with the linear kernel type.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
SPE/IADC-202101-MS 7

Table 4—The performance of classical SVM with the optimum parameters for RBF and linear kernels.

Kernel Type C Gamma RMSE_Training R2_Training RMSE_Testing R2_Testing

RBF 500 Scale 0.53 0.97 0.57 0.97

Linear 0.01 Auto 0.53 0.97 0.56 0.97

Figure. 2—Cross-Plot of the actual and predicted ECD of the training


set using classical SVM with linear kernel (3,304 data points).

Figure. 3—Cross-Plot of the actual and predicted ECD of the


testing set using classical SVM with linear kernel (827 data points).

PCA for Dimensionality Reduction


The dataset was transformed to PCA space using Python library's Scikit-Learn®. Table 5 shows a sample
of the training set after transformation to PCA space.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
8 SPE/IADC-202101-MS

Table 5—A sample of training set after transformation to PCA space.

Sample PC1 PC2 PC3 PC4 PC5 PC6 PC7 ECD (Pcf)

1 -0.49519 4.461575 0.749265 3.937679 2.19006 2.351809 -0.69753 83.65

2 0.196929 5.553348 -1.06093 -1.19373 0.507604 0.569627 -0.5658 83.57

3 0.400576 5.55361 -1.01139 -0.99504 0.39389 0.328901 -0.51637 83.45

4 0.084996 5.255019 -0.01048 -1.58235 1.328692 0.021461 -0.09929 83.45

5 -0.44961 5.5468 0.237232 -1.55454 1.342653 0.718522 -0.21667 83.39

It is important to study the variation that each PC accounts for in the dataset to perform dimensionality
reduction. Figure 4 shows the percentages of the variation that each PC accounts for in the dataset. Principal
components that represent more than 90% of the variation in a dataset are often considered in the analysis.
The first thing to notice is that the first principal component PC1 accounts for 36.45% of the variation in
the dataset; PC2 accounts for 26.86%; PC3 accounts for 17.29%; PC4 accounts for 12.93%; PC5 accounts
for 4.41%; PC6 accounts for 1.29%; and PC7 accounts for 0.75%. This means that PC1, PC2, PC3, and
PC4 directions collectively explain 93.53% of the total variation in the dataset, while PC5, PC6, and PC7
combined explain only 6.45% of the total variation in the dataset. Thus, the dataset can be condensed to a
four-dimensional plane (R4) spanned by PC1, PC2, PC3, and PC4.

Figure. 4—The percentages of the variation that each PC accounts for in the whole dataset.

The contribution of each variable (Q, HL, ROP, RS, SPP, WOB, or T) in each principal component is
presented in Table 6. Now how is all of this to be interpreted? For example, in studying PC2, the fifth entry
"0.62", SPP, is the largest, which means a change in one unit of SPP affects the ECD more than a change
of one unit of Q, HL, ROP, RS, WOB, or T. The fourth entry "0.51", which corresponds to RS, is the next
most important factor in determining the ECD. On the other hand, the second entry "0.04", HL, is the least
important in determining the ECD. Similarly, the other PCs can be interpreted.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
SPE/IADC-202101-MS 9

Table 6—The contribution of variables in each principal component.

Variable PC1 PC2 PC3 PC4 PC5 PC6 PC7

Q 0.53 0.25 0.11 0.28 0.48 0.06 0.58

HL 0.58 0.04 0.27 0.01 0.12 0.45 0.61

ROP 0.21 0.28 0.07 0.89 0.16 0.23 0.01

RS 0.27 0.51 0.34 0.10 0.72 0.10 0.01

SPP 0.23 0.62 0.04 0.30 0.36 0.36 0.47

WOB 0.15 0.16 0.84 0.06 0.27 0.37 0.19

T 0.43 0.43 0.29 0.15 0.08 0.69 0.20

The performance of the PCA-based SVM model with a linear kernel while varying the numbers of PCs
is presented in Table 7. Table 7 shows that as the number of PCs for the dataset increases, R2 and RMSE
improve. However, when more than four principal components were considered, the improvement in the
RMSE was insignificant. The PCA-based SVM model, with the linear kernel and four principal components
(93.53% variation of the dataset), predicted the ECD with RMSE 0.79 and R2 of 0.95 in the testing set, while
RMSE and R2 were 0.56 and 0.97 respectively in the case of the classical SVM model, which was trained
with 100% variation of the dataset as discussed in section 3.1. The PCA-based SVM model performed
almost similar to the classical SVM and did not lose any important information, when only four dimensions
PC1, PC2, PC3, and PC4 were considered. The PCA-based SVM model reduced the dimensionality of the
dataset, which leads to a smaller model and possibly reduces the chance of model over-fitting.

Table 7—The performance of PCA-based SVM with different PCs compared to the classical SVM.

PCA-based SVM_Testing Classical SVM_Testing

NO. PC % Variation Kernel Type RMSE R 2 RMSE R2

PC1 36.45 2.36 0.47

PC1+ PC2 63.31 0.95 0.91

PC1+ PC2+PC3 80.60 0.94 0.92

PC1+ PC2+PC3+PC4 93.53 Linear 0.79 0.95 0.56 0.97

PC1+ PC2+PC3+PC4+PC5 97.94 0.78 0.95

PC1+ PC2+PC3+PC4+PC5+P6 99.23 0.75 0.95

PC1+ PC2+PC3+PC4+PC5+P6+PC7 100 0.74 0.95

Conclusions
A new SVM model based on principal components analysis was used to predict ECD while drilling HPHT
horizontal wells. The performance of PCA-based SVM model was compared with the classical SVM model.
Based on the results, the following can be concluded:

• The classical SVM with the linear kernel predicted the ECD with R2 of 0.97 and RMSE of 0.53 in
the training set, while R2 and RMSE were 0.97 and 0.56 respectively in the testing set.
• Four principal components PC1, PC2, PC3, and PC4 out of seven were sufficient to explain 93.53%
of the total variation in the dataset.
• The PCA-based SVM model, with the linear kernel and four principal components (93.53%
variation of the dataset), predicted the ECD with RMSE 0.79 and R2 of 0.95 in the testing set.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
10 SPE/IADC-202101-MS

References
1. Rommeveit, R., and Bjorkevoll, K. S. (1997). "Temperature and Pressure Effects on Drilling
Fluid Rheology and ECD in Very Deep Wells". Presented at the 1997 SPE/IADC Meddle East
Drilling Technology Conference, Manama, Bahrain, November. SPE/IADC 39282. pp.1–10.
10.2118/39282-MS.
2. Maglion, R., and Robotti, G. (1996). "A Computer Program to Predict Stand Pipe Pressure While
Drilling Using the Drilling Well as Viscometer". Presented at the Eleventh Petroleum Computer
Conference, Datta, Texas, USA, June. SPE 35994. pp. 1–11. 10.2118/35994-MS
3. Dokhani, V., Pordel, S. M., Karimi, M., and Salehi, S. (2013). "Evaluation of Annular Pressure
Losses While Casing drilling". Presented at the SPE Annual Technical Conference, and
Exhibition, New Orleans, USA, September. SPE 166103. pp. 1–15. 10.2118/166103-MS.
4. Erge, O., Vajargah, A. K., Ozbayoglu, M. E, and Van Oort, E. (2016). "Improved ECD Prediction
and Management in Horizontal and Extended Reach Wells with Eccentric Drill strings".
Presented at the SPE/IADC Drilling Conference and Exhibition held in, Fort Worth, Texas, USA,
1-3 March. IADC/SPE-178785-Ms. Pp. 1–21. 10.2118/178785-MS
5. Rabia, H. 2001. "Well Engineering and Construction. London, United Kingdome". Entrac
Petroleum.
6. Isambourg, P., Anfinsen, B. T., and Marken, C. (1996). "Volumetric Behavior of Drilling Muds
at High Pressure and High Temperature". Presented at the 1996 SPE European Petroleum
Conference held in Milan, Itally, 22-24 October. SPE 36830, Pp. 1–9. 10.2118/36830-MS.
7. Scheid, C. M, Cacada, L. M., Rocha, D. P., Aranha, P. E., Aragao, A. F., and Martins, A. L.
(2009). "Prediction of Pressure Losses in Drilling Fluid Flows in Circular and Annular Pipes
and Accessories". Presented at the Latin American and Caribbean Petroleum Engineering
Conference, Cartagena, Colombia, May. SPE 122072. Pp. 1–13. 10.2118/122072-MS.
8. Cartalos, U., and Dupuis, D. (1993). "An Analysis Accounting for The Combined Effect of
Drill string Rotation and Eccentricity on Pressure Losses in Slim Hole Drilling". Presented at
the SPE/IADC Drilling Conference, Amsterdam, Netherlands, February. SPE 25769. Pp. 1–11.
10.2118/25769-MS.
9. Escudier, M. P., Goulson, I. W., Oliveira, P. J., and Pinho, F. T. (2000). "Effects of inner cylinder
rotation on laminar flow of a Newtonian fluid through an eccentric annulus". International
Journal of Heat and Fluid Flow, 21 (1), 92–103. 10.1016/S0142-727X(99)00059-4.
10. Mokhtari, M., Ermila, M. A., Tutuncu, A. N., and Karimi, M. (2012). "Computational Modeling
of Drilling Fluids Dynamics in Casing Drilling". Presented at the SPE Eastern Regional Meeting,
Lexington, Kentucky, USA, October. SPE 161301. Pp. 1–13. 10.2118/161301-MS.
11. Hemphill, T., and Ravi, K. (2005). "Calculation of Drill Pipe Rotation Effects on Axial Flow:
An Engineering Approach". Presented at the SPE Annual Technical Conference and Exhibition,
Dallas, Texas, October. SPE 97158. pp. 1–4. 10.2118/97158-MS.
12. [Link]. (2020). [online] Available at: [Link]
public/ss/contents/Data_Sheets/web/[Link]?nav=en-
US_sperry_public [Accessed 7 Feb. 2020].
13. Abdelgawad, K. Z., Elzenary, M., Elkatatny, S., Mahmoud, M., Abdulraheem, A., and Patil, S.
(2018). "New Approach to Evaluate the Equivalent Circulating Density (ECD) Using Artificial
Intelligence Techniques". Journal of Petroleum Exploration and Production Technology, vol. 9,
no. 2, 2018, Pp. 1569–1578. 10.1007/s13202-018-0572-y.
14. Alkinani, H. H., Al-Hameedi, A., Dunn-Noman, S., Al-Alwani, M. A., Mutar, R. A., and Al-
Bazzaz, W. H. (2019). "Data-Driven Neural Network Model to Predict Equivalent Circulation

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
SPE/IADC-202101-MS 11

Density ECD". Presented at the SPE Gas and Oil Technology Showcase and Conference held in
Dubai, UAE, 21-23 October. SPE-198612. Pp. 1–9. 10.2118/198612-MS.
15. Balaji, K., Rabiei, K., Suicmez, V., Canbaz, C., Agharzeyva, Z., Tek, S., Bulut, U., and Temizel,
C. (2018). "Status of Data-Driven Methods and Their Applications in Oil and Gas Industry".
In Proceeding of the SPE Europe featured at 80th EAGE Conference and Exhibition held in
Copenhagen, Denemark, 11-14 June. SPE-190812-MS. 10.2118/190812-MS.
16. Al-Ghazal, M., and Vedpoathak, V. (2019). "A Novel Machine Learning Models for Early
operational Anomaly Detection Using LWD/MWD Data". Presented at the international
Petroleum Technology Conference held in Beijing, China, 26-28 March. [Link].
1–7.
17. Bello, O., Holzmann, J., Yaqoob, T. and Teodoriu, C. (2015). "Application of Artificial
Intelligence Methods in Drilling System Design and Operations: A Review of The State of
The Art". Journal of Artificial Intelligence and Soft Computing Research, 5 (2), Pp. 121–139.
10.1515/jaiscr-2015-0024.
18. Temizel, C., Aktas, S., Kirmaci, H., Susuz, O., Zhu, Y., Balaji, K., Ranjith, R., Tahir, S.,
Aminzadeh, F., and Yegin, C. (2016). "Turning Data into Knowledge: Data-Driven Surveillance
and Optimization in Mature Fields". Presented at the SPE Annual Technical Conference and
Exhibition held in Dubai, UAE, 26-28 September. SPE-181881. Pp. 1–32. 10.2118/181881-MS.
19. Al-Yami, A. S., Al-Shaarawi, A., Al-Bahrani, H., Wagle, V. B., Al-Gharbi, S., and Al-Khudairi,
M. B. (2016). "Using Bayesian Network to Develop Drilling Expert Systems". Presented at the
SPE Heavy Oil Conference and Exhibition, Kuwait City, Kuwait, 6-8 December. SPE-184168.
Pp. 1–21. 10.2118/184168-MS.
20. Ahmedi, M. A. (2016). "Toward Reliable Model for Prediction Drilling Fluid Density at Wellbore
Conditions: A LS SVM model". Neurocomputing 211, 143–149. 10.1016/[Link].2016.01.106.
21. Elkatatny, S. M., Tariq, Z., Mahmoud, M. A., Al-Abduljabbar, A. (2017). "Optimization of Rate
of Penetration Using Artificial Intelligence Technique". This paper was prepared for presentation
at the 51st US Rock Mechanics /Geomechanics Symposium held in San Francisco, California,
USA, 25-28 June. ARMA-17-771. Pp. 1–8.
22. Elkatatny, S., 2018. New approach to optimize the rate of penetration using artificial neural
network. Arabian Journal for Science and Engineering, 43 (11), pp.6297–6304.
23. Al-AbdulJabbar, A., Elkatatny, S., Mahmoud, M., Abdelgawad, K. and Al-Majed, A., 2019.
A Robust Rate of Penetration Model for Carbonate Formation. Journal of Energy Resources
Technology, 141 (4), p.042903.
24. Kamel, M. A., Elkatatny, S., Mysorewala, M. F., Al-Majed, A. and Elshafei, M., 2018. Adaptive
and real-time optimal control of stick-slip and bit wear in autonomous rotary steerable drilling.
Journal of Energy Resources Technology, 140 (3), p.032908.
25. Al-Abduljabbar, A., Elkatatny, S., Mahmoud, M., and Abdulraheem, A. (2018). "Prediction
Formation Tops While Drilling Using Artificial Intelligence". In Proceeding of the SPE
Kingdome Annual Technical Symposium and Exhibition held in Dammam, Saudi Arabia, 23-26
April. SPE-192345-MS. 10.2118/192345-MS.
26. Elkatatny, S., Tariq, Z., Mahmoud, M., Abdulraheem, A. and Mohamed, I., 2019. An integrated
approach for estimating static Young's modulus using artificial intelligence tools. Neural
Computing and Applications, 31 (8), pp. 4123–4135.
27. Elkatatny, S., Mahmoud, M., Mohamed, I. and Abdulraheem, A., 2018a. Development of a
new correlation to determine the static Young's modulus. Journal of Petroleum Exploration and
Production Technology, 8 (1), pp. 17–30.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025
12 SPE/IADC-202101-MS

28. Tariq, Z., Elkatatny, S., Mahmoud, M., Ali, A. Z. and Abdulraheem, A., 2017a, May. A new
technique to develop rock strength correlation using artificial intelligence tools. Paper presented
at the SPE Reservoir Characterisation and Simulation Conference and Exhibition, Abu Dhabi,
UAE, 8-10 May. SPE-186062-MS. 10.2118/186062-MS
29. Elkatatny, S., Tariq, Z., Mahmoud, M., Mohamed, I. and Abdulraheem, A., 2018b. Development
of new mathematical model for compressional and shear sonic times from wireline log data using
artificial intelligence neural networks (white box). Arabian Journal for Science and Engineering,
43 (11), pp. 6375–6389.
30. Tariq, Z., Elkatatny, S., Mahmoud, M., Ali, A. Z. and Abdulraheem, A., 2017b. A new approach
to predict failure parameters of carbonate rocks using artificial intelligence tools. Paper presented
at the SPE Kingdom of Saudi Arabia Annual Technical Symposium and Exhibition, Dammam,
Saudi Arabia, 24-27 April. SPE-187974-MS. 10.2118/187974-MS
31. Moussa, T., Elkatatny, S., Mahmoud, M. and Abdulraheem, A., 2018. Development of new
permeability formulation from well log data using artificial intelligence approaches. Journal of
Energy Resources Technology, 140 (7), p. 072903.
32. Lindsay. I, Smith. "A Tutorial on Principal Component Analysis". [Link]
cosc453/studenttutorials/[Link], February 26, 2002.
33. Tharwat, A. (2016). "Principal component analysis - a tutorial". International Journal of Applied
Pattern Recognition, 3 (3), p.197. 10.1504/IJAPR.2016.079733.
34. Wold, S., Esbensen, K., and Geladi, P. (1987). " Principal component analysis". Chemometrics
and Intelligent laboratory systems, Vol. 2, No. 1, Pp. 37–52. 10.1016/0169-7439(87)80084-9.
35. Hearst, M. A. (1998). "Support Vector Machines". IEEE Intelligent Systems and Their
Applications, July/August 1998, Vol. 13, Issue 4, pp. 18–28, ISSN. 10.1109/5254.708428.
36. Vapnik. V, The Nature of Statistical Learning The-Ory.2nd ED, New York, Springer p.1–314,
1995.
37. Anifowose, F., Abdulraheem, A. (2011). "Fuzzy Logic Driven and SVM-Driven Hybrid
Computational Intelligence Models Applied to Oil and Gas Reservoir Characterization". Journal
of Natural Gas Science and Engineering. 10.1016/[Link].2011.05.002.
38. Durka, B. (2011). "A Classification Algorithm Using Mahalanobis Distance Clustering of Data
with Applications onBiomedical Data Sets". Msc dissertation, Middle East Technical University,
Ankara, Turkey.
39. Awad M., Khanna R. (2015). "Support Vector Machines for Classification". In: Efficient
Learning Machines. Apress, Berkeley, CA. 10.1007/978-1-4302-5990-9_3.

Downloaded from [Link] by The University of West Indies - St Augustine user on 08 November 2025

You might also like