0% found this document useful (0 votes)
81 views52 pages

Petroleum Geostatistics Overview

The document discusses univariate and bivariate statistics used in geostatistics and reservoir modeling. It introduces topics such as measures of central tendency and spread like the mean, median, variance and standard deviation. It also covers univariate distribution plots including histograms, probability density functions and cumulative density functions. The document uses examples to illustrate key statistical concepts.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
81 views52 pages

Petroleum Geostatistics Overview

The document discusses univariate and bivariate statistics used in geostatistics and reservoir modeling. It introduces topics such as measures of central tendency and spread like the mean, median, variance and standard deviation. It also covers univariate distribution plots including histograms, probability density functions and cumulative density functions. The document uses examples to illustrate key statistical concepts.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Geostatistics and Reservoir

Modeling Module
– Introduction to Spatial Statistics
– Review of Basic Statistics
– Spatial Statistics (Semivariogram),
Random Variables, and Random Functions
– Model Building using Estimation Algorithms
such as Kriging
– Model Building using Conditional
Simulation Algorithms such as Sequential
Gaussian Simulation and Indicator-based
Methods
– Additional Topics in Model Building -
Object-based Methods, Simulated
Annealing, and Multi-Point Statistic (MPS)
based Techniques to Build Facies Models

Petroleum Geostatistics, Review of Basic Statistics. 1


Institute of Petroleum Engineering,
University of Tehran, 2010.
Statistics Review

• Although Our Knowledge of Most


Reservoirs is Limited, the Amount of
Data is Often Difficult to Manage and
Communicate.

• Univariate (Single Variable) Statistics


• Bivariate (Two Variable) Statistics

Petroleum Geostatistics, Review of Basic Statistics. 2


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Statement of Learning Objectives


– Review of Measures of Location and
Spread
• Mean
• Variance
• Standard Deviation
– Review of Univariate Plots
• Histogram
• Probability Density Function - pdf
• Cumulative Density Function - cdf
– Handling Outliers
– Review of Types of Distributions
• Parametric
– Normal (Gaussian)
– Log-Normal
• Non-Parametric

Petroleum Geostatistics, Review of Basic Statistics. 3


Institute of Petroleum Engineering,
University of Tehran, 2010.
Additional Vocabulary

• Population:
– Collection of a finite number of
measurements or virtually infinitely large
collection of data about something of
interest.
• Sample:
– Representative subset selected from the
population. A good sample must reflect the
essential features of the population from
which it is drawn.

Petroleum Geostatistics, Review of Basic Statistics. 4


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• The Mean, Usually Denoted by m, is the Central


Value of a Distribution. The Arithmetic Mean is
Given by
n
1
m =
n
∑i =1
z (α i )
where z(αi) = sample value and n = number of data
values
• For Log-Normal Distributions, the Geometric Mean
is a Better Measure of the Central Value. The
Geometric Mean Given by

log m = ∑ log( z(αi ))


1 n
n i =1
• The Median Is Also Frequently Used As an
Appropriate Central Value for Distributions that are
Highly Asymmetric. The Median, Usually Denoted
by M, is the Value that Splits a Distribution in Half,
i.e., 50% of the Values are Less than the Median
and 50% are Greater.

Petroleum Geostatistics, Review of Basic Statistics. 5


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Measures of Location (continued)


– Normal Distribution
– Mode = Median = Mean

Mean
Mode
0.08 Median
0.07
Frequency

0.06
0.05
0.04
0.03
0.02
0.01
0

Porosity

Petroleum Geostatistics, Review of Basic Statistics. 6


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Measures of Location (continued)


– Lognormal Distribution
– Mode < Median < Mean

Mode
2
1.8 Median
1.6
Frequency

1.4 Arithmetic
1.2 Mean
1
0.8
0.6
0.4
0.2
0
Permeability

Petroleum Geostatistics, Review of Basic Statistics. 7


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• The Variance, Usually Denoted by σ2, Is a Measure


of the Spread of the Data Values Around the Mean.
The Variance is Given by
2

σ = ∑ ( z(αi ) − m)
n
12
n i =1
where z(αi) = sample value and n = number of data
values
• Note that the Unit of Variance is the Data Value Unit
Squared, i.e. %2. Also note that Variance
Magnitude Is “Sensitive” to the Unit Magnitude, i.e.
Millidarcys Vs Darcys. To Get Around these
Problems, the Standard Deviation is Used.
• The Standard Deviation, Usually Denoted by σ, is
the Square Root of the Variance, that is

σ = σ 2

Petroleum Geostatistics, Review of Basic Statistics. 8


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Measures of Spread and Shape


(continued)
m = 0.128

12 20
10
15
8
6 10
4
5
2
0 0
0.02

0.06

0.14

0.18

0.22

0.26
0.1

Two Distributions with Same Mean Value.


Standard Deviation(s) of Distribution Shown
by Line with Solid Boxes is Larger than
Standard Deviation of Distribution Shown by
Line with Crosses.

Petroleum Geostatistics, Review of Basic Statistics. 9


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Univariate Distribution Plots


– Histogram
• Plot of the Data Value (X-axis) Against
Actual Frequency of Occurrence of the Data
Value (Y-axis)
• The X-axis is Normally Divided Into a
Number of Bins or Classes.
– Probability Density Function (pdf)
• Plot of the Data Value (X-axis) Against
“Normalized” Frequency of Occurrence of
the Data Value (Y-axis)
• The “Normalized” Frequency of Occurrence
is Obtained by Dividing the Actual
Frequency of Occurrence by Total Number
of Sample Values

Petroleum Geostatistics, Review of Basic Statistics. 10


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Univariate Distribution Plots


(continued)
– Cumulative Density Function (cdf)
• Plot of the Data Value (X-axis) Against
“Cumulative” Frequency of Occurrence of
the Data Value (Y-axis)
• The Cumulative Frequency of Occurrence is
Obtained by Summing the “Normalized”
Frequency of Occurrence of the Current Bin
or Class and all Lower Bins or Classes.
• The cdf Y-axis Must Range Between 0 and
1 (or 0% and 100%).

Petroleum Geostatistics, Review of Basic Statistics. 11


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Example Data Set #1

Depth Porosity (%) Permeability (darcy) Permeability (md)


Depth Porosity (%) Permeability (darcy) Permeability (md)
1001 16 0.222 222
1001
1002 1116 0.222
0.345 345222
1002
1003 1011 0.345
0.124 124345
1003
1004 910 0.124
0.076 76124
1004
1005 79 0.076
0.050 5076
1005
1006 57 0.050
0.020 2050
1006
1007 15 0.020
0.002 220
1007
1008 21 0.002
0.005 52
1008
1009 72 0.005
0.013 135
1009
1010 47 0.013
0.001 113
1010
1011 84 0.001
0.033 331
1011
1012 138 0.033
0.045 4533
1012
1013 1913 0.045
0.055 5545
1013
1014 2019 0.055
0.513 513 55
1014
1015 2120 0.513
0.456 456513
1015
1016 1721 0.456
0.345 345456
1016
1017 1717 0.345
0.411 411345
1017
1018 1517 0.411
0.250 250411
1018
1019 615 0.250
0.004 4250
1019
1020 26 0.004
0.006 64
1020 2 0.006 6
Number of Values 20 20 20
Number of Values 20 20 20
Column Sum 210.00 2.98 2976.2
Column Sum 210.00 2.98 2976.2
Arithmetic Average 10.50 0.15 148.8
Arithmetic Average 10.50 0.15 148.8
Geometric Average 8.06 0.05 46.0
Geometric Average 8.06 0.05 46.0
Median 9.50 0.05 52.5
Median 9.50 0.05 52.5
Variance 40.79 0.03 30390.2
Variance 40.79 0.03 30390.2
Standard Deviation 6.39 0.17 174.3
Standard Deviation 6.39 0.17 174.3

Petroleum Geostatistics, Review of Basic Statistics. 12


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Example of Histogram, pdf, and cdf


Calculations for Example Data Set #1
(Porosity)
– Basic Steps Are:
• Set Up Bins (Classes)
• Determine Number of Data Values in Each
Bin
• Normalized Frequency = Number of Values
in Each Bin Divided by Total Number of
Values
• Cumulative Frequency is Obtained by
Summing the “Normalized” Frequency of
Occurrence of the Current Bin or Class and
all Lower Bins or Classes.
Bin Raw Frequency Normalized Frequency Cumulative Frequency
Bin Raw Frequency Normalized Frequency Cumulative Frequency
0 1 0.05 0.05
0 1 0.05 0.05
2 3 0.15 0.2
2 3 0.15 0.2
4 2 0.1 0.3
4 2 0.1 0.3
6 2 0.1 0.4
6 2 0.1 0.4
8 1 0.05 0.45
8 1 0.05 0.45
10 2 0.1 0.55
10 2 0.1 0.55
12 1 0.05 0.6
12 1 0.05 0.6
14 1 0.05 0.65
14 1 0.05 0.65
16 2 0.1 0.75
16 2 0.1 0.75
18 3 0.15 0.9
18 3 0.15 0.9
20 2 0.1 1
20 2 0.1 1

Petroleum Geostatistics, Review of Basic Statistics. 13


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Example of Histogram
Histogram
Histogram, 3
3
2.5

pdf, and cdf Raw Frequency


2.5
2
Raw Frequency
2

Plots for
1.5
1.5
1
1

Porosity
0.5
0.5
0
0 0 2 4 6 8 10 12 14 16 18 20

Values in 0 2 4 6 8 10
Porosity
Porosity
12 14 16 18 20

Example Data
Set #1
PDF
PDF
0.16
0.16
0.14
"Normalized" Frequency

0.14
"Normalized" Frequency

0.12
0.12
0.1
0.1
0.08
0.08
0.06
0.06
0.04
0.04
0.02
0.02
0
0 0 2 4 6 8 10 12 14 16 18 20
0 2 4 6 8 10 12 14 16 18 20
Porosity
Porosity

CDF
CDF
1
1

0.8
Cumulative Frequency

0.8
Cumulative Frequency

0.6
0.6

0.4
0.4

0.2
0.2

0
0
0 2 4 6 8 10 12 14 16 18 20
0 2 4 6 8 10 12 14 16 18 20
Porosity
Porosity

Petroleum Geostatistics, Review of Basic Statistics. 14


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Selection of Bin
Histogram - Class Size = 5
or Class Size Is Histogram - Class Size = 5

a Function of
6
6
5

Raw Frequency
5

Data Range,
Raw Frequency
4
4
3

Number of Data
3
2
2
1
Values, and 0
1
0
0 5 10 15 20 25
Distribution 0 5 10
Porosity
Porosity
15 20 25

Information that
Is Needed. Histogram - Class Size = 2
Histogram - Class Size = 2
Same Data 3

Used to Build All


3
2.5
Raw Frequency

2.5
Raw Frequency

2
Histograms. 1.5
2
1.5
1
1
0.5
0.5
0
0
10

12

14

16

18

20

22

24

26
0

10

12

14

16

18

20

22

24

26
0

Porosity
Porosity

Histogram - Class Size = 1


Histogram - Class Size = 1

2
2

1.5
Raw Frequency

1.5
Raw Frequency

1
1

0.5
0.5

0
0
11

13

15

17

19

21

23

25
1

11

13

15

17

19

21

23

25
1

Porosity
Porosity

Petroleum Geostatistics, Review of Basic Statistics. 15


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Histograms, pdf’s, and cdf’s are Also


Used to Spot Outliers
– Outliers, Either Extreme Low Values or
Extreme High Values, May Strongly Affect
Summary Univariate Statistics Such as
Mean, Variance, Bivariate Statistics Such
as Covariance and the Correlation
Coefficient, or Spatial Statistics such as the
Semi-Variogram
– Outliers May Be Handled by
• Declare the Extreme Values to be
Erroneous and Discard. This Approach
Should be Used Judiciously as the Extreme
Values Are Often the Most Interesting
• Transform Data to Minimize the Influence of
Extreme Values
• Classify Extremes Into a Separate Statistical
Population

Petroleum Geostatistics, Review of Basic Statistics. 16


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate
Depth Permeability (md)

2701 11
2702 5
2703 6

Statistics 2704
2705
2706
2707
3
5
8
12
2708 13
2709 256
2710 390
2711 44

• Example Data Set


2712 11
2713 2
2714 1
2715 2

with Extreme 2716


2717
2718
4
5
4

Values
2719 2
2720 5
2721 7
2722 8

– Summary Statistics 2723


2724
2725
14
59
389
2726 17
2727 452
2728 12
2729 11
2730 5
2731 6
2732 8
2733 6
2734 3
2735 2
2736 1
2737 5
2738 7
2739 6
2740 8

Histogram and CDF (All Data)

16 1

0.9
14
0.8
12
Number of Values 40.00
Cumulative Frequency
0.7
Raw Frequency

10 Arithmetic Average 45.38 0.6

8
Geometric Average 9.12 0.5
Variance 12778.04
0.4
6 Standard Deviation 113.04
0.3
4 Median 6.50
0.2
2
0.1

0 0
100

125

150

175

200

225

250

275

300

325

350

375

400

425

450

475

500
25

50

75
0

Permeability (md)

Petroleum Geostatistics, Review of Basic Statistics. 17


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Outlier
LOW Permeability (md) HIGH Permeability (md)
Handling 11
5

– Delete
6
3
5

Extreme
8
12
13

Values 44
256
390

11

– Treat 2
1
2
Separately 4
5
4

– Transform 2
5
7

(next page) 8
14
59
389
17
452
12
11
5
6
8
6
3
2
1
5
7
6
8

36.00 4.00 Number of Values


9.11 371.75 Arithmetic Average
6.06 364.00 Geometric Average
127.13 6822.92 Variance
11.28 82.60 Standard Deviation
6.00 389.50 Median

Petroleum Geostatistics, Review of Basic Statistics. 18


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Normal Score Transform


– Gaussian Anamorphosis
– Essentially Involves Transforming any
Distribution a Standard Normal or
Gaussian Distribution with a mean of zero
and a range of -3 to +3 (standard
deviations) by “Mapping” Each Point on a
CDF to the Corresponding Point on a CDF
for a Normal Distribution
– Important as Many Algorithms Assume a
Gaussian Data Distribution
– Often, this Transformation Is Done
Automatically
– Note that for Log-Normal Distributions the
Normal Score Transform is the Same as
Doing a Logarithmic Transform.

Petroleum Geostatistics, Review of Basic Statistics. 19


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Normal Score Transform


– Forward Transform

1 1

0.5

X Y XNS YNS
0 0
0 15 30 -3 0 +3

16.5 in “real” space becomes 0 in normal score space

28 in “real” space becomes +1.25 in normal score space

Process of mapping from “real” to normal score space


defines the transform

Petroleum Geostatistics, Review of Basic Statistics. 20


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Normal Score Transform


– Back Transform

1 1

0.5

X Y XNS YNS
0 0
0 15 30 -3 0 +3

0 in normal score space becomes 16.5 in “real” space

+1.25 in normal score space becomes 28 in “real” space

Given knowledge of the forward transform “mapping”

Petroleum Geostatistics, Review of Basic Statistics. 21


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Types of Distributions
– Parametric Distributions Can Be
Completely Described by a Few
Parameters such as Mean and Variance
• Normal or Gaussian
– Porosity (Sometime)
– Saturation (Usually)
• Log-Normal
– Permeability (Sometime)
– Non-Parametric Distributions Can Not
Easily Be Described by Parameters.
Frequently the Result of a Mixture of
Populations.
• Bi-Modal
• Multi-Modal

Petroleum Geostatistics, Review of Basic Statistics. 22


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Theoretical Normal (Gaussian)


Distribution
– Some Distributions Have a Concise
Mathematical Description
– A Normal or Gaussian Distribution is
Completely Described by
⎡ 1 ⎛ z − m ⎞2 ⎤
−⎢ ⎜ 2 ⎟ ⎥
1
g( z ) = e ⎣ 2⎝ σ ⎠ ⎦
σ 2π

• where g(z) = frequency, m = arithmetic


mean and
• σ2 is the variance
– The Standard Normal Distribution Sets m =
0 and σ2 = 1

Petroleum Geostatistics, Review of Basic Statistics. 23


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Normal Distributions
– Example Histogram and CDF
– CDF for a Normal (Gaussian) Distribution
Has an “S” Shape
Histogra m a nd CDF for Norm a l Distribution

25 1
0.9
20 0.8

Cumulative Frequency
0.7
Raw Frequency

15 0.6
0.5
10 0.4
Mean = 17.1 0.3
5 0.2
0.1
0 0
10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

25

Po r o s ity

– Unit Normal Distribution Has Mean = 0 and


Standard Deviation = 1

Petroleum Geostatistics, Review of Basic Statistics. 24


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Log-Normal Distribution
– Example Histogram and CDF
– Note Skewed “S” Shape of CDF

Histogram and CDF for Log-Normal Distribution

45 1
40
0.8

Cumulative Frequency
35
Raw Frequency

30
25 Mean = 4.2 0.6

20
15
Geometric Mean = 3.6 0.4

10 0.2
5
0 0
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16

Porosity

Petroleum Geostatistics, Review of Basic Statistics. 25


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Non-Parametric Distribution
– Example Histogram and cdf for Bi-Modal
Distribution
– Note “Step” Shape of CDF

Histogram and CDF For Bi-Modal Distribution

16 1
0.9
14
Mean = 13.8 0.8
12
0.7 Cumulative Frequency
Raw Frequency

10 0.6
8 0.5

6 0.4
0.3
4
0.2
2 0.1
0 0
11

13

15

17

19

21

23

25
1

Porosity

Petroleum Geostatistics, Review of Basic Statistics. 26


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Non-Parametric Distributions
(continued)
– Separating a Bi-Modal or Multi-Modal
Mixture (for Example, Effective Porosity in
a Sand-Shale Sequence) May Result in
Individual Distributions that are Normal or
near-Normal
– Indicator Transform Often Provides a
Convenient Way of Splitting Mixed
Populations

Petroleum Geostatistics, Review of Basic Statistics. 27


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Summary
– Measures of Location and Spread
• Mean - Measure of Central Value
(Arithmetic Vs Geometric)
• Median - Another Measure of Central Value
• Variance and Standard Deviation - Measure
of Spread
– Common Displays
• Histogram and Population Density Function
(pdf)
• Cumulative Density Function (cdf)
– Outlier Handling – Use Cautiously, if at all!
• Exclusion
• Separate Population
• Transform
– Types of Distribution
• Parametric (Normal, Log-Normal)
• Non-parametric (Bi-Modal)

Petroleum Geostatistics, Review of Basic Statistics. 28


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Learning Objectives
‰ Review of Measures of Location and
Spread
‰ Mean
‰ Variance
‰ Standard Deviation
‰ Review of Univariate Plots
‰ Histogram
‰ Probability Density Function - pdf
‰ Cumulative Density Function - cdf
‰ Handling Outliers
‰ Review of Types of Distributions
‰ Parametric
‰ Normal (Gaussian)
‰ Log-Normal
‰ Non-Parametric

Petroleum Geostatistics, Review of Basic Statistics. 29


Institute of Petroleum Engineering,
University of Tehran, 2010.
Univariate Statistics

• Learning Objectives
9 Review of Measures of Location and
Spread
9 Mean
9 Variance
9 Standard Deviation
9 Review of Univariate Plots
9 Histogram
9 Probability Density Function - pdf
9 Cumulative Density Function - cdf
9 Handling Outliers
9 Review of Types of Distributions
9 Parametric
9 Normal (Gaussian)
9 Log-Normal
9 Non-Parametric

Petroleum Geostatistics, Review of Basic Statistics. 30


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Learning Objectives
– Bivariate Data Display
• Scatterplot or Crossplot
– Bivariate Measures
• Covariance
• Correlation Coefficient
– Brief Review of Linear Regression
• Procedure
• Example
• Limitations

Petroleum Geostatistics, Review of Basic Statistics. 31


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Basic Question that


Bivariate Statistics
Seeks to Answer is
What Is the 1000

Relationship Between
900

800

700

Two Variables?
600

500

400

• Scattergram or
300

200

100

Crossplot Is Basic Plot 0


0 5 10 15 20 25 30

Used to Summarize 1000

Bivariate Data 100

– Used to Examine
Relationship
10

Between Two 1
0 5 10 15 20 25 30

Variables 1000

– Used to Examine
Data for Extreme
100

Values (High or Low) 10

– May Use Linear, 1

Semi-Log, or Log-
1 10 100

Log Scales

Petroleum Geostatistics, Review of Basic Statistics. 32


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Bivariate Measures
– The Covariance, Usually Denoted by C or σXY
is the Measure of Joint Variation of Two
Variables, X and Y, About Their Respective
Means, mX and mY. That is

σ xy = ∑ ( xi − mx ) yi − my
1 n
n i =1
( )
where n = number of data pairs
– Covariance = Variance if X = Y
– Covariance Has a Magnitude and Units
“Problem” Similar to Variance.
– Therefore, We Use the Correlation Coefficient,
Usually Denoted by ρXY.. The Correlation
Coefficient Is Given by

σ xy
ρxy =
σ xσ y

Petroleum Geostatistics, Review of Basic Statistics. 33


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Bivariate Measures (continued)


– The Correlation Coefficient Is Unit-less and
Varies Between -1 and +1.
– Values Near Zero Indicate No Significant
Linear Correlation.
10

5
ρ = −0.9
0
0 5 10

10

5 ρ = 0.08
0
0 5 10

10

0
ρ = +0.9
0 5 10

Note: Variance (s2 2) of X and Y


Note: Variance (s ) of X and Y
Values is Same for All Plots
Values is Same for All Plots

Petroleum Geostatistics, Review of Basic Statistics. 34


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Bivariate Measures (continued)


– Correlation Coefficient is Extremely
Sensitive to Extreme Values (Outliers)
25.000

20.000

15.000

10.000

5.000

0.000
0.000 10.000 20.000 30.000 40.000 50.000 60.000

ρ (all points) = 0.96


ρ (excluding extreme) = 0.23

– May Also Calculate the Rank Correlation


Coefficient, ρR. If ρR Significantly Different
than ρ, then Extremes are Present or a
Non-Linear Relationship Exists.

Petroleum Geostatistics, Review of Basic Statistics. 35


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Linear Regression
– Goal is to Summarize a Relationship
Between Two Variables such as Porosity
and Permeability with a Linear Equation of
the Form

y = a + bx
where y = y value, x = x value, b = slope,
and a = y intercept

– In Addition to Obtaining the Model


Parameters (i.e. a and b), Regression
Analysis Also Gives Information About the
Dependency of the Variables and the
Source of Variability

Petroleum Geostatistics, Review of Basic Statistics. 36


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Linear Regression (continued)


– When Two Variables Have a Non-Zero
Covariance, Measurements of One
Variable Can Be Used To Estimate the
Other
– Assume a Deterministic Model of the Form
y = a + bx
– The Best Estimates of a and b are
Obtained by Finding the Unbiased,
Minimum Variance Estimates of a and b
– This Process Gives:
σ xy
b= 2 a = my − b(mx )
σx
where mi = mean value

Petroleum Geostatistics, Review of Basic Statistics. 37


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Linear Regression (continued)


– The Coefficient of Determination, r2, is
Used to Estimate How Much of the Overall
Variation in Y is Explained by the Linear
Model Shown Below
– That is,

r2= Explained Variation of Y


Total Variation of Y
– Mathematically, r2 is Given by

∑ ⎝ y i − y ⎞⎟⎠
⎛⎜
n ∧ 2

i =1
r2 = 1−
∑(y )
n 2
i − my
i =1

Petroleum Geostatistics, Review of Basic Statistics. 38


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Linear Regression (continued)


– Sample Plot Showing
• (yi - my)
• (yi - ^
y)

y = a + bx
my
(y1 - m y)
( y1 ^- y )
y(1)

x(1)

Petroleum Geostatistics, Review of Basic Statistics. 39


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Linear Regression Example


Depth Porosity Permeability Permeability
% darcy md
1001 16 0.288 288
1002 11 0.276 276
1003 10 0.124 124
1004 9 0.076 76
1005 7 0.050 50
1006 5 0.020 20
1007 1 0.001 1
1008 2 0.005 5
1009 7 0.090 90
1010 4 0.010 10
1011 8 0.033 33
1012 13 0.222 222
1013 19 0.459 459
1014 20 0.513 513
1015 29 0.887 887
1016 17 0.345 345
1017 17 0.411 411
1018 15 0.250 250
1019 6 0.022 22
1020 2 0.006 6

Number of Values 20 20 20
Column Sum 218.00 4.09 4088.0
Arithmetic Average 10.90 0.20 204.4
Geometric Average 8.19 0.07 74.0
Median 9.50 0.11 107.0
Variance 52.83 0.05 53566.8
Standard Deviation 7.27 0.23 231.4

Covariance Bewteen
Porosity (X) and Permeability (Y) 1.537 1536.990
Correlation Coefficient 0.962 0.962

b 0.029 29.092
a -0.113 -112.706
r2 0.925 0.925

900

800

700

600

500

400

300

200

Y = 29.02X - 112.7
100

-100

-200
0 5 10 15 20 25 30

Petroleum Geostatistics, Review of Basic Statistics. 40


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

Average Weight of Linemen


(American Football) at the
University of Texas

300
200 y = - 1108 + 0.66x
100
0
1900 1950 2000

Given the Best-Fit Line Above, What Will Be the


Average Weight of Linemen at the University of
Texas in the Year
3000?

What About in the Year 1000?

(Data and Example from Larson & Marx, 1986)

Petroleum Geostatistics, Review of Basic Statistics. 41


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Linear Regression (continued)


– Limitations
• Linear Regression Only Useful if Data Trend
is Linear.
• Extrapolation Using Best-Fit Line Outside of
Data Range is Often Outright Wrong or
Misleading.
• Examine Your Data Carefully. A Subset of
the Data May Have a Useful Linear Trend
Even if the Entire Data Set Does Not.

Petroleum Geostatistics, Review of Basic Statistics. 42


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Summary
– Scatterplot Is Basic Bivariate Display
• Relationship Between Variables
• Extreme Values Easily Spotted
• May Be Linear, Semi-Log, or Log-Log
– Bivariate Measures
• Covariance
• Correlation Coefficient
– Unit-Less
– Varies Between -1 and +1
– Linear Regression – Quantitative
Description of the Linear Relationship
between Two Data Types

Petroleum Geostatistics, Review of Basic Statistics. 43


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Learning Objectives
‰ Bivariate Data Display
‰ Scatterplot or Crossplot
‰ Bivariate Measures
‰ Covariance
‰ Correlation Coefficient
‰ Brief Review of Linear Regression
‰ Procedure
‰ Example
‰ Limitations

Petroleum Geostatistics, Review of Basic Statistics. 44


Institute of Petroleum Engineering,
University of Tehran, 2010.
Bivariate Statistics

• Learning Objectives
9 Bivariate Data Display
9 Scatterplot or Crossplot
9 Bivariate Measures
9 Covariance
9 Correlation Coefficient
9 Brief Review of Linear Regression
9 Procedure
9 Example
9 Limitations

Petroleum Geostatistics, Review of Basic Statistics. 45


Institute of Petroleum Engineering,
University of Tehran, 2010.
Declustering

• Data are rarely collected for their statistical


“representivity”
– Wells are drilled in areas with the greatest
probability of high production rates
– Core measurements are taken preferentially
from good reservoir quality rock
• These “data collection” practices should not be
changed as they lead to the best economics
and the greatest number of data in portions of
the study area that are the most important
• There is a need, however, to adjust the
histograms and summary statistics to be
representative of the entire volume of interest.
• Declustering techniques
– Polygonal / Cell Declustering
– Declustering by Estimation / Kriging
– Correcting Non-Representative Data
– Histogram Smoothing

Petroleum Geostatistics, Review of Basic Statistics. 46


Institute of Petroleum Engineering,
University of Tehran, 2010.
Example of Declustering

• Location map of 122


wells. The gray scale
code shows the
underlying “true”
distribution of porosity
(inaccessible in
practice)
• Histogram of 122 well
data with the true
reference histogram
shown as the black
line (inaccessible in
practice). Note the
greater proportion of
data between 25-30%
and the sparsity of
data in the 0-20%
porosity range

Petroleum Geostatistics, Review of Basic Statistics. 47


Institute of Petroleum Engineering,
University of Tehran, 2010.
Declustering by Area of
Influence
• Location map of 122
wells with polygonal
areas of influence

• Declustering weights Ai
are taken inversely
( p)
w =

i n
proportional to the Ai
areas i =1

• Weight between 25-


30% has been
decreased and the
weight in the 0-20%
porosity range has
been increased

Petroleum Geostatistics, Review of Basic Statistics. 48


Institute of Petroleum Engineering,
University of Tehran, 2010.
Cell Declustering

• Another technique, called Cell


Declustering, is more robust in 3-D and
when the polygon limits are poorly
defined or not applicable

• Histogram with cell declustering


weights shows the proportion of data
between 25-30% has largely been
corrected and the weight given to data
in the 0-20% porosity range has been
increased

Petroleum Geostatistics, Review of Basic Statistics. 49


Institute of Petroleum Engineering,
University of Tehran, 2010.
Correcting for Non-
Representative Data

• What do we do when there are too few


data or the data are not
representative?
• Nothing, unless there is some
secondary information

Petroleum Geostatistics, Review of Basic Statistics. 50


Institute of Petroleum Engineering,
University of Tehran, 2010.
Non-Representative Data -
Example

• Reservoir with 63 wells:

• Histogram:

Petroleum Geostatistics, Review of Basic Statistics. 51


Institute of Petroleum Engineering,
University of Tehran, 2010.
Non-Representative Data -
Example
• Could use polygonal or cell declustering but a
seismic attribute (low frequency) at the top of
the reservoir interval provides valuable
information

Petroleum Geostatistics, Review of Basic Statistics. 52


Institute of Petroleum Engineering,
University of Tehran, 2010.

Common questions

Powered by AI

A Normal Score Transform consists of mapping log-normal distribution data points to a normal distribution's corresponding points in a standard score space. This process brings all data to a Gaussian format, essential because many geostatistical algorithms assume a normal distribution of input data. In log-normal distributions, this transformation is akin to a logarithmic transform, helping to manage vastly different scales of data within the same dataset .

Extreme values can skew the analysis by affecting measures like mean, variance, and standard deviation, potentially leading to incorrect conclusions about the data. To mitigate their influence, methods such as deleting extreme values, treating them separately as a distinct statistical population, or applying transformations (e.g., log transformations) can be utilized. For example, using a Normal Score Transform can help by mapping extreme values to a standard normal distribution, minimizing their impact on statistical algorithms that assume Gaussian distribution .

The arithmetic mean is the sum of values divided by the number of values, sensitive to extreme values which can skew it upwards or downwards. The geometric mean is the nth root of the product of values and dampens the influence of extreme values, providing a more central tendency reflecting the dataset's typical behavior. In datasets with outliers, the geometric mean offers better insights into central tendency, whereas the arithmetic mean might misrepresent it .

Declustering adjusts for biases in data collection to ensure representative statistics. It addresses data where wells are drilled in high-probability areas or cores are taken from high-quality rocks, potentially skewing summary statistics. Techniques such as polygonal/cell declustering, declustering by estimation or kriging, and histogram smoothing can rectify these biases by redistributing data to reflect the entire study area's true conditions .

A bi-modal distribution is characterized by two distinct peaks or modes, often indicating a mixture of two different populations. In contrast, a normal distribution has a single peak and symmetrical shape. Visualizations like histograms and cumulative distribution functions (CDFs) highlight these differences; while a normal distribution's CDF has an "S" shape, a bi-modal distribution's CDF might show a "step" shape, indicating the presence of two separate modes .

The correlation coefficient quantifies the linear relationship between two variables by providing a value between -1 and +1. Values near zero suggest no significant linear correlation. However, it is limited as it is sensitive to outliers and does not account for non-linear relationships. When extreme values are present, the rank correlation coefficient may offer additional insights by indicating possible non-linear relationships or extreme data points .

Parametric distributions can be completely described by parameters such as mean and variance. Examples include Normal (Gaussian) distributions often used for porosity and Log-Normal distributions for permeability. Non-parametric distributions cannot be easily described by parameters and usually result from a mix of populations, leading to bi-modal or multi-modal distributions. An example of a non-parametric distribution is effective porosity in a sand-shale sequence, which can be split using an Indicator Transform .

Linear regression is suitable for linear relationships as it models them with an equation of the form y = a + bx, where a is the y-intercept and b is the slope. Extrapolating beyond the data range can be deceptive because the relationship may not remain linear outside the observed data. A subset of the data might present a linear trend, while the entire set might not, leading to misleading predictions if extrapolated .

Transforming a distribution to a standard normal distribution is crucial because many statistical algorithms assume data are normally distributed. This transformation allows such algorithms to work effectively. It is achieved by using a Normal Score Transform, which involves mapping each point on a cumulative distribution function (CDF) to the corresponding point on a CDF for a normal distribution. This mapping helps in achieving a mean of zero and a standard deviation range of -3 to +3 .

Measures of location, such as the mean and median, provide central values of a dataset, indicating where most data points are grouped. Measures of spread, like variance and standard deviation, indicate the variability or dispersion of data points around the central value. These measures are essential for understanding data distribution and are typically represented in plots like histograms and CDFs, facilitating visual comprehension of data behavior .

You might also like