Petroleum Geostatistics Overview
Petroleum Geostatistics Overview
Modeling Module
– Introduction to Spatial Statistics
– Review of Basic Statistics
– Spatial Statistics (Semivariogram),
Random Variables, and Random Functions
– Model Building using Estimation Algorithms
such as Kriging
– Model Building using Conditional
Simulation Algorithms such as Sequential
Gaussian Simulation and Indicator-based
Methods
– Additional Topics in Model Building -
Object-based Methods, Simulated
Annealing, and Multi-Point Statistic (MPS)
based Techniques to Build Facies Models
• Population:
– Collection of a finite number of
measurements or virtually infinitely large
collection of data about something of
interest.
• Sample:
– Representative subset selected from the
population. A good sample must reflect the
essential features of the population from
which it is drawn.
Mean
Mode
0.08 Median
0.07
Frequency
0.06
0.05
0.04
0.03
0.02
0.01
0
Porosity
Mode
2
1.8 Median
1.6
Frequency
1.4 Arithmetic
1.2 Mean
1
0.8
0.6
0.4
0.2
0
Permeability
σ = ∑ ( z(αi ) − m)
n
12
n i =1
where z(αi) = sample value and n = number of data
values
• Note that the Unit of Variance is the Data Value Unit
Squared, i.e. %2. Also note that Variance
Magnitude Is “Sensitive” to the Unit Magnitude, i.e.
Millidarcys Vs Darcys. To Get Around these
Problems, the Standard Deviation is Used.
• The Standard Deviation, Usually Denoted by σ, is
the Square Root of the Variance, that is
σ = σ 2
12 20
10
15
8
6 10
4
5
2
0 0
0.02
0.06
0.14
0.18
0.22
0.26
0.1
• Example of Histogram
Histogram
Histogram, 3
3
2.5
Plots for
1.5
1.5
1
1
Porosity
0.5
0.5
0
0 0 2 4 6 8 10 12 14 16 18 20
Values in 0 2 4 6 8 10
Porosity
Porosity
12 14 16 18 20
Example Data
Set #1
PDF
PDF
0.16
0.16
0.14
"Normalized" Frequency
0.14
"Normalized" Frequency
0.12
0.12
0.1
0.1
0.08
0.08
0.06
0.06
0.04
0.04
0.02
0.02
0
0 0 2 4 6 8 10 12 14 16 18 20
0 2 4 6 8 10 12 14 16 18 20
Porosity
Porosity
CDF
CDF
1
1
0.8
Cumulative Frequency
0.8
Cumulative Frequency
0.6
0.6
0.4
0.4
0.2
0.2
0
0
0 2 4 6 8 10 12 14 16 18 20
0 2 4 6 8 10 12 14 16 18 20
Porosity
Porosity
• Selection of Bin
Histogram - Class Size = 5
or Class Size Is Histogram - Class Size = 5
a Function of
6
6
5
Raw Frequency
5
Data Range,
Raw Frequency
4
4
3
Number of Data
3
2
2
1
Values, and 0
1
0
0 5 10 15 20 25
Distribution 0 5 10
Porosity
Porosity
15 20 25
Information that
Is Needed. Histogram - Class Size = 2
Histogram - Class Size = 2
Same Data 3
2.5
Raw Frequency
2
Histograms. 1.5
2
1.5
1
1
0.5
0.5
0
0
10
12
14
16
18
20
22
24
26
0
10
12
14
16
18
20
22
24
26
0
Porosity
Porosity
2
2
1.5
Raw Frequency
1.5
Raw Frequency
1
1
0.5
0.5
0
0
11
13
15
17
19
21
23
25
1
11
13
15
17
19
21
23
25
1
Porosity
Porosity
2701 11
2702 5
2703 6
Statistics 2704
2705
2706
2707
3
5
8
12
2708 13
2709 256
2710 390
2711 44
Values
2719 2
2720 5
2721 7
2722 8
16 1
0.9
14
0.8
12
Number of Values 40.00
Cumulative Frequency
0.7
Raw Frequency
8
Geometric Average 9.12 0.5
Variance 12778.04
0.4
6 Standard Deviation 113.04
0.3
4 Median 6.50
0.2
2
0.1
0 0
100
125
150
175
200
225
250
275
300
325
350
375
400
425
450
475
500
25
50
75
0
Permeability (md)
• Outlier
LOW Permeability (md) HIGH Permeability (md)
Handling 11
5
– Delete
6
3
5
Extreme
8
12
13
Values 44
256
390
11
– Treat 2
1
2
Separately 4
5
4
– Transform 2
5
7
(next page) 8
14
59
389
17
452
12
11
5
6
8
6
3
2
1
5
7
6
8
1 1
0.5
X Y XNS YNS
0 0
0 15 30 -3 0 +3
1 1
0.5
X Y XNS YNS
0 0
0 15 30 -3 0 +3
• Types of Distributions
– Parametric Distributions Can Be
Completely Described by a Few
Parameters such as Mean and Variance
• Normal or Gaussian
– Porosity (Sometime)
– Saturation (Usually)
• Log-Normal
– Permeability (Sometime)
– Non-Parametric Distributions Can Not
Easily Be Described by Parameters.
Frequently the Result of a Mixture of
Populations.
• Bi-Modal
• Multi-Modal
• Normal Distributions
– Example Histogram and CDF
– CDF for a Normal (Gaussian) Distribution
Has an “S” Shape
Histogra m a nd CDF for Norm a l Distribution
25 1
0.9
20 0.8
Cumulative Frequency
0.7
Raw Frequency
15 0.6
0.5
10 0.4
Mean = 17.1 0.3
5 0.2
0.1
0 0
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
Po r o s ity
• Log-Normal Distribution
– Example Histogram and CDF
– Note Skewed “S” Shape of CDF
45 1
40
0.8
Cumulative Frequency
35
Raw Frequency
30
25 Mean = 4.2 0.6
20
15
Geometric Mean = 3.6 0.4
10 0.2
5
0 0
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
Porosity
• Non-Parametric Distribution
– Example Histogram and cdf for Bi-Modal
Distribution
– Note “Step” Shape of CDF
16 1
0.9
14
Mean = 13.8 0.8
12
0.7 Cumulative Frequency
Raw Frequency
10 0.6
8 0.5
6 0.4
0.3
4
0.2
2 0.1
0 0
11
13
15
17
19
21
23
25
1
Porosity
• Non-Parametric Distributions
(continued)
– Separating a Bi-Modal or Multi-Modal
Mixture (for Example, Effective Porosity in
a Sand-Shale Sequence) May Result in
Individual Distributions that are Normal or
near-Normal
– Indicator Transform Often Provides a
Convenient Way of Splitting Mixed
Populations
• Summary
– Measures of Location and Spread
• Mean - Measure of Central Value
(Arithmetic Vs Geometric)
• Median - Another Measure of Central Value
• Variance and Standard Deviation - Measure
of Spread
– Common Displays
• Histogram and Population Density Function
(pdf)
• Cumulative Density Function (cdf)
– Outlier Handling – Use Cautiously, if at all!
• Exclusion
• Separate Population
• Transform
– Types of Distribution
• Parametric (Normal, Log-Normal)
• Non-parametric (Bi-Modal)
• Learning Objectives
Review of Measures of Location and
Spread
Mean
Variance
Standard Deviation
Review of Univariate Plots
Histogram
Probability Density Function - pdf
Cumulative Density Function - cdf
Handling Outliers
Review of Types of Distributions
Parametric
Normal (Gaussian)
Log-Normal
Non-Parametric
• Learning Objectives
9 Review of Measures of Location and
Spread
9 Mean
9 Variance
9 Standard Deviation
9 Review of Univariate Plots
9 Histogram
9 Probability Density Function - pdf
9 Cumulative Density Function - cdf
9 Handling Outliers
9 Review of Types of Distributions
9 Parametric
9 Normal (Gaussian)
9 Log-Normal
9 Non-Parametric
• Learning Objectives
– Bivariate Data Display
• Scatterplot or Crossplot
– Bivariate Measures
• Covariance
• Correlation Coefficient
– Brief Review of Linear Regression
• Procedure
• Example
• Limitations
Relationship Between
900
800
700
Two Variables?
600
500
400
• Scattergram or
300
200
100
– Used to Examine
Relationship
10
Between Two 1
0 5 10 15 20 25 30
Variables 1000
– Used to Examine
Data for Extreme
100
Semi-Log, or Log-
1 10 100
Log Scales
• Bivariate Measures
– The Covariance, Usually Denoted by C or σXY
is the Measure of Joint Variation of Two
Variables, X and Y, About Their Respective
Means, mX and mY. That is
σ xy = ∑ ( xi − mx ) yi − my
1 n
n i =1
( )
where n = number of data pairs
– Covariance = Variance if X = Y
– Covariance Has a Magnitude and Units
“Problem” Similar to Variance.
– Therefore, We Use the Correlation Coefficient,
Usually Denoted by ρXY.. The Correlation
Coefficient Is Given by
σ xy
ρxy =
σ xσ y
5
ρ = −0.9
0
0 5 10
10
5 ρ = 0.08
0
0 5 10
10
0
ρ = +0.9
0 5 10
20.000
15.000
10.000
5.000
0.000
0.000 10.000 20.000 30.000 40.000 50.000 60.000
• Linear Regression
– Goal is to Summarize a Relationship
Between Two Variables such as Porosity
and Permeability with a Linear Equation of
the Form
y = a + bx
where y = y value, x = x value, b = slope,
and a = y intercept
∑ ⎝ y i − y ⎞⎟⎠
⎛⎜
n ∧ 2
i =1
r2 = 1−
∑(y )
n 2
i − my
i =1
y = a + bx
my
(y1 - m y)
( y1 ^- y )
y(1)
x(1)
Number of Values 20 20 20
Column Sum 218.00 4.09 4088.0
Arithmetic Average 10.90 0.20 204.4
Geometric Average 8.19 0.07 74.0
Median 9.50 0.11 107.0
Variance 52.83 0.05 53566.8
Standard Deviation 7.27 0.23 231.4
Covariance Bewteen
Porosity (X) and Permeability (Y) 1.537 1536.990
Correlation Coefficient 0.962 0.962
b 0.029 29.092
a -0.113 -112.706
r2 0.925 0.925
900
800
700
600
500
400
300
200
Y = 29.02X - 112.7
100
-100
-200
0 5 10 15 20 25 30
300
200 y = - 1108 + 0.66x
100
0
1900 1950 2000
• Summary
– Scatterplot Is Basic Bivariate Display
• Relationship Between Variables
• Extreme Values Easily Spotted
• May Be Linear, Semi-Log, or Log-Log
– Bivariate Measures
• Covariance
• Correlation Coefficient
– Unit-Less
– Varies Between -1 and +1
– Linear Regression – Quantitative
Description of the Linear Relationship
between Two Data Types
• Learning Objectives
Bivariate Data Display
Scatterplot or Crossplot
Bivariate Measures
Covariance
Correlation Coefficient
Brief Review of Linear Regression
Procedure
Example
Limitations
• Learning Objectives
9 Bivariate Data Display
9 Scatterplot or Crossplot
9 Bivariate Measures
9 Covariance
9 Correlation Coefficient
9 Brief Review of Linear Regression
9 Procedure
9 Example
9 Limitations
• Declustering weights Ai
are taken inversely
( p)
w =
∑
i n
proportional to the Ai
areas i =1
• Histogram:
A Normal Score Transform consists of mapping log-normal distribution data points to a normal distribution's corresponding points in a standard score space. This process brings all data to a Gaussian format, essential because many geostatistical algorithms assume a normal distribution of input data. In log-normal distributions, this transformation is akin to a logarithmic transform, helping to manage vastly different scales of data within the same dataset .
Extreme values can skew the analysis by affecting measures like mean, variance, and standard deviation, potentially leading to incorrect conclusions about the data. To mitigate their influence, methods such as deleting extreme values, treating them separately as a distinct statistical population, or applying transformations (e.g., log transformations) can be utilized. For example, using a Normal Score Transform can help by mapping extreme values to a standard normal distribution, minimizing their impact on statistical algorithms that assume Gaussian distribution .
The arithmetic mean is the sum of values divided by the number of values, sensitive to extreme values which can skew it upwards or downwards. The geometric mean is the nth root of the product of values and dampens the influence of extreme values, providing a more central tendency reflecting the dataset's typical behavior. In datasets with outliers, the geometric mean offers better insights into central tendency, whereas the arithmetic mean might misrepresent it .
Declustering adjusts for biases in data collection to ensure representative statistics. It addresses data where wells are drilled in high-probability areas or cores are taken from high-quality rocks, potentially skewing summary statistics. Techniques such as polygonal/cell declustering, declustering by estimation or kriging, and histogram smoothing can rectify these biases by redistributing data to reflect the entire study area's true conditions .
A bi-modal distribution is characterized by two distinct peaks or modes, often indicating a mixture of two different populations. In contrast, a normal distribution has a single peak and symmetrical shape. Visualizations like histograms and cumulative distribution functions (CDFs) highlight these differences; while a normal distribution's CDF has an "S" shape, a bi-modal distribution's CDF might show a "step" shape, indicating the presence of two separate modes .
The correlation coefficient quantifies the linear relationship between two variables by providing a value between -1 and +1. Values near zero suggest no significant linear correlation. However, it is limited as it is sensitive to outliers and does not account for non-linear relationships. When extreme values are present, the rank correlation coefficient may offer additional insights by indicating possible non-linear relationships or extreme data points .
Parametric distributions can be completely described by parameters such as mean and variance. Examples include Normal (Gaussian) distributions often used for porosity and Log-Normal distributions for permeability. Non-parametric distributions cannot be easily described by parameters and usually result from a mix of populations, leading to bi-modal or multi-modal distributions. An example of a non-parametric distribution is effective porosity in a sand-shale sequence, which can be split using an Indicator Transform .
Linear regression is suitable for linear relationships as it models them with an equation of the form y = a + bx, where a is the y-intercept and b is the slope. Extrapolating beyond the data range can be deceptive because the relationship may not remain linear outside the observed data. A subset of the data might present a linear trend, while the entire set might not, leading to misleading predictions if extrapolated .
Transforming a distribution to a standard normal distribution is crucial because many statistical algorithms assume data are normally distributed. This transformation allows such algorithms to work effectively. It is achieved by using a Normal Score Transform, which involves mapping each point on a cumulative distribution function (CDF) to the corresponding point on a CDF for a normal distribution. This mapping helps in achieving a mean of zero and a standard deviation range of -3 to +3 .
Measures of location, such as the mean and median, provide central values of a dataset, indicating where most data points are grouped. Measures of spread, like variance and standard deviation, indicate the variability or dispersion of data points around the central value. These measures are essential for understanding data distribution and are typically represented in plots like histograms and CDFs, facilitating visual comprehension of data behavior .