0% found this document useful (0 votes)
9 views14 pages

Simple Descriptive Statistics Overview

The document provides an overview of simple descriptive statistics, including definitions of qualitative and quantitative characteristics, statistical series, and graphical representations. It explains various statistical concepts such as frequencies, statistical tables, position parameters (like mode and average), dispersion parameters (like variance and standard deviation), and shape parameters (like skewness and kurtosis). Examples are included to illustrate the application of these concepts in analyzing data.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views14 pages

Simple Descriptive Statistics Overview

The document provides an overview of simple descriptive statistics, including definitions of qualitative and quantitative characteristics, statistical series, and graphical representations. It explains various statistical concepts such as frequencies, statistical tables, position parameters (like mode and average), dispersion parameters (like variance and standard deviation), and shape parameters (like skewness and kurtosis). Examples are included to illustrate the application of these concepts in analyzing data.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

USTHB-Faculty of Mathematics.

2017/2018

MI–Section5.

SIMPLE DESCRIPTIVE STATISTICS

I / Definitions:

We want to study the data related to the characteristics of a set of individuals.


or objects called population. It is difficult to observe all the data when the
the number of individuals in the entire population is high. We examine a number
restricted from the population that we call a sample. Each individual can be
studied in relation to one or more characters. A character is called a variable.
statistic denoted by X. The set of observed data related to a
a characteristic constitutes a statistical series. A characteristic can present several
terms. It is assumed that there exists noted arrangements { . }

1. Qualitative characteristic: If its modalities are categories and are not


measurable. There are two types of qualitative characters:

Ordinals can be ordered, like degree of burn (1er, 2èmeor 3emedegree) or


level of education.

Nominals do not get sorted, like profession or eye color.

2. Quantitative character: If its modalities are measurable, that is to say, one can
correspond a number to each modality. We distinguish two types:

a. Discrete quantitative character: If it can only take isolated values of a


certain interval, for example the number of children in a family .{ }

b. Continuous quantitative character: If it can take any value belonging to a


interval (the possible values of the character are real numbers), for example the
size or weight of newborns .

II/ Statistical series:

We call a statistical series the sequence of values taken (measured or observed) by


a variable X on individuals. The number of individuals is noted n, and the values
observed values of the variable X are noted { . }

1
Examples:
1/We are interested in the variable civil status noted X and in the statistical series of values
prices for 10 people. The terms are: (C: single, M: married, W: widower, D:
divorced). Let the observed statistical series be:
M M V C C M C C M C here and .
2/The character X: the number of children per family. The statistical series:
320113325022134 . .

1/ Effective, effective cumulative frequency and cumulative frequency:

The frequency of a modality is the number of times that modality appears.


in the statistical series. We note the number of the modality , .

The frequency of a modality is the corresponding proportion. , .

We note the cumulative staff the cumulative frequency of the modality .

and .

2/ Statistical table: In a statistical table, we summarize the data in


important the corresponding terms and frequencies.
Examples In the qualitative case, we take the example of the marital status characteristic:

civil status
C 5 0.5
M 4 0.4
V 1 0.1
D 0 0
Totals 10 1

In the discrete quantitative case, we take the example of the number of children per family.
Number of children

0 2 2 2/15 2/15
1 3 5 3/15 5/15
2 4 9 4/15 9/15
3 4 13 4/15 13/15
4 1 14 1/15 14/15
5 1 15 1/15 1
Totals 15 / 1 /

2
- Continuous case: Since the modalities are infinite, we then make a distribution of
th
data in intervals called classes and is noted by the center of the
class. If the number of classes is not given, we take . √

Let it be the thclass, We note extent of the series.


So the amplitude of this class is: and

Note: Choose the range so that all the values of the series are
included in the table.

Example: We measure the height in cm of 50 students in a class.

– , [ ] . √

we take .

Classes
[152, 155[ 153.5 8 8
[155, 158[ 156.5 10 18
[158, 161[ 159.5 9 27
[161, 164[ 162.5 6 33
[164, 167[ 165.5 6 39
[167, 170[ 168.5 6 45
[170, 173[ 171.5 5 50
Totals / 50 /

3
III/ Graphical representation: The statistical data is primarily presented
treatment in a disordered form. After sorting and presenting the
these data in tabular form, graphical representations allow to
visualize the distribution.

1/ Qualitative variable: The principle consists of representing the data by


diagrams where the different parts have areas proportional to the sizes
or to the frequencies. (The and the ).

Bar chart or organ pipes:

In the qualitative case, the modalities are placed on a horizontal line that is not
not oriented because they are not measurable. The numbers are on a vertical axis and
the height of each band is proportional to or .

modalities
Pie chart (pie chart)

It is the most used. Each modality is represented in a circular sector.


don't the angle .

Example :

Civil status
C 5 0.5
M 4 0.4
V 1 0.1
D 0 0

Totals 10 1

4
2/ quantitative variable :

a/ discrete case:

Bar chart: The different modalities of are plotted on the abscissa.


character and on the vertical axis the numbers (or the frequencies ).

Polygon of workforce

0 1 2 3 4 5 Number of children

Bar chart
Polygon of frequencies: It is the broken line that connects the
tops of sticks.
Cumulative diagram: The values are plotted on the x-axis. of the character and in
arranged the counts (or frequencies) of individuals for which the character
is less than or equal to We obtain a stepwise curve.

0 1 2 3 4 5 Nombre d’enfants

Cumulative diagram
5
b/ Continuous case:

Histogram: To each class, we associate a rectangle whose width is


the amplitude of the class and whose height is the size (or frequency) that
correspond.

Height of students
Histogram
Polygon of the frequencies: it is the broken line that passes through the midpoints of the
tops of rectangles. Two fictional classes have been added. and
where the numbers are zero, for the polygon to join the axis of
abscissas.
Cumulative frequency polygon: (or cumulative frequencies). It is the line
broken line joining the points obtained by plotting on the ordinate on the right side of
chaque classe (limite supérieure en abscisse), l’effectif cumulé (ou fréquence
cumulative).

Size of students
Cumulative frequency polygon 6
ulcumulés
Distribution function: We denote by the associated cumulative distribution function
to a value , which is an application of in .
Discrete case: It is the cumulative frequency of observations less than .

{ for

Case continues:

{ ∑ for

Example: Size of 50 students. Calculate the distribution function for .

IV/ position parameters:

These parameters are characteristic values that allow for a representation


condensed information contained in the statistical series.

1/ Mode ( It is the value taken by the studied character that is the largest.
effective. A statistical series can be unimodal or multimodal.
Discrete case: The mode is the modality of the characteristic that corresponds to the largest.
effective. Example: (number of children).

Case continues: This is the modal class. It is the class of maximum frequency, and it
corresponds to the peak of the histogram. We can have several modal classes.
All the values of the class can be realized, we determine a single value that
will represent the mode. For simplicity, we can choose the center of the modal class,
however, it is preferable to perform an interpolation to take the classes into account
adjacent.
7
Size of the modal class.

Strength of the previous class.

Class size of the following.

Amplitude of the class.

and .

So the mode is:

Example: (size of 50 students).

2/ Order quantiles ( ) :

Let it be It is called quantile of order , I noticed the value as it is (or


observations that are lower than it in an ordered series of sizes.
For we obtain respectively 1er, 2th, 3thquartile, noted
The second quartile is called the median. .

Discreet case:

is the value which corresponds to . [.] designates the integer part.

One can determine from the statistical table. It corresponds to the modality
for which the cumulative total is equal to (or cumulative frequency equal ).

Example:

8
Continuing case:

We will talk about a class containing It is the class as such (ou ).


is determined by interpolation. Let us be:

class containing .

the cumulative workforce of and that of the previous class.

cumulative frequency of and that of the previous class.

So: or

Example: (size of 50 students).

Graphical representation: On the graph of cumulative frequencies, is


the abscissa of the point of ordinate (or if we used the (or the ) to plot
the graph. In the discrete case we have:

0 1 2 3 4 5 Number of children

Cumulative diagram

9
In the continuous case, we have:

Height of the students

Cumulative curve

3/ The average:

̅
Arithmetic mean ( It is generally the characteristic that represents the
better the center of the distribution of the statistical series. It is the characteristic of
central tendency the most used. If the are the modalities of a discrete variable
or the centers of classes of a continuous variable, the arithmetic mean is:
̅ ∑ or well ̅ ∑
Yes so we have a simple arithmetic mean and if the are
different then we have a weighted arithmetic mean.
The calculations are summarized in the statistical table.

Classes
[152, 155[ 153.5 8 8
[155, 158[ 156.5 10 18
[158, 161[ 159.5 9 27
[161, 164[ 162.5 6 33
[164, 167[ 165.5 6 39
[167, 170[ 168.5 6 45
[170, 173[ 171.5 5 50
Totals / 50 /
10
V/ Dispersion parameters:

Summarize a statistical distribution by a single characteristic such as the mode,


the mean or the median, which represent a central value, is insufficient. The
Statistics begins where there is variability; it is therefore necessary to define at least one
central value and a measure of dispersion around this central value.

1/ The extent ( It is the difference between the highest recorded observation. and
the lowest rated So the extent is:

2/ Interquartile range (decile, centile):

The interval between the third and first quartile , or the difference
among them is an indicator of dispersion around the median .

Indeed this indicator corresponds to an interval that groups 50% of the


observations around the median.

The gap between the ninth and the first decile corresponds to a
interval that groups 80% of the observations around the median.
The gap between the ninety-ninth and the first percentile contains 98%
observations around the median.

3/ The moments:

These are algebraic quantities that allow us to describe the characteristics of


distributions statistiques : forme, symétrie, aplatissement, tendance centrale,
dispersion. The arithmetic mean and variance are the most commonly used moments.

a/ Order moment :

Let the modalities of a statistical variable discreet or the


centers of classes in the continuous case. The moment of order noted of the variable
is defined by:

∑ (If .̅ )
11
b/ Central moment of order :

The centered moment of order noted in relation to ̅ is defined by:

∑ ̅

Yes .
Variance and standard deviation:

It is the most commonly used characteristic to measure dispersion or spread of


data around the average. It is rated by or and it is the moment
center of order 2, .

∑ ̅ or ∑ ̅

Note: It shows that ∑ ̅ ∑ ̅

On the statistical table, we add a column to calculate the variance:

Yes is a constant, then we have the following two properties:

-
-

The standard deviation is the square root of the variance and is denoted by .

Change of variable:

The calculation of the mean and variance can sometimes be laborious due to the values.
raised from so from their squares, but the situation becomes particularly
different if we make the following change of variable:

Yes so and ̅. ̅

Where class amplitude and middle of the middle class (if the number of
even classes, 2 central classes, choose the one with the largest population.

On the statistical table, we calculate the values then̅ and of the


variable knowing that:

̅ ∑ ∑ ̅

We will deduce the average. ̅ and the variance of .


12
Example: (size of 50 students).

Classes
[152, 155[ 153.5 8 8
[155, 158[ 156.5 10 18
[158, 161[ 159.5 9 27
[161, 164[ 162.5 6 33
[164, 167[ 165.5 6 39
[167, 170[ 168.5 6 45
[170, 173[ 171.5 5 50

4/ Coefficient of variation:

It is the number which̅ allows to relativize the standard deviation based on size
values. It thus allows for comparing the dispersion of series of measurements expressed
in different units because it has no unit.

La série avec le plus petit coefficient de variation serait la moins dispersée, c'est-à-
She would have her values located more around the average than the others.

Example: (size of 50 students).

1̅ 61.3 and √ so ̅

VI/ Shape parameters:

1/ Fisher's skewness coefficient:

One can visually judge if a distribution is more spread out to the right or to the left.
left from a bar chart or a histogram. The coefficient
Fisher's asymmetry is a characteristic that allows measuring asymmetry.
of a distribution: ⁄

It is a dimensionless number (independent of measurement units of the ).


13
If the distribution is symmetrical around the mean .̅

Yes the distribution is more spread out to the right.

Yes , the distribution is more spread out on the left.

2/ Fisher's flattening coefficient:

It is the number (dimensionless) following:

For a known Normal distribution variable in statistics, Such a


The bell-shaped distribution is often considered ideal, hence:

If , the series is normal.


If the series is less flattened than a normal statistical series of the same
average and the same variance. The curve is sharper and has
longer queues.
Yes , the series is flatter than a normal statistical series of the same.
average and the same variance. The curve is rounder and has
shorter queues.
Example of two distributions with the same mean and the same variance.
The sharpest distribution has a thicker tail.

Example: (size of 50 students).

14

You might also like