AMITY INSTITUTE OF PSYCHOLOGY AND ALLIED SCIENCE, AIPS
WHAT IS A STATISTIC?
• Statistics are part of our everyday life.
• Kuzma (1984) provides a formal definition: A body of techniques and procedures dealing with the
collection, organization, analysis, interpretation, and presentation of information that can be stated
numerically.
• Perhaps an example will clarify this definition. Say, for example, we wanted to know the level of job
satisfaction nurses experience working on various units within a particular hospital (ie. psychiatric,
cardiac care, obstetrics, etc.). The first thing we would need to do is collect some data. We might have
all the nurses on a particular day complete a job satisfaction questionnaire. We could ask such
questions as "On a scale of 1 (not satisfied) to 10 (highly satisfied), how satisfied are you with your
job?". We might examine employee turnover rates for each unit during the past year. We could also
examine absentee records for a two month period of time as decreased job satisfaction is correlated
with higher absenteeism. Once we have collected the data, we would then organize it. This entire
process of organizing and interpreting data will be called as statistics.
AMITY INSTITUTE OF PSYCHOLOGY AND ALLIED SCIENCE, AIPS
TYPES OF STATISTICS
• Descriptive statistics are used to organize or summarize a particular set of measurements. In other words, a
descriptive statistic will describe that set of measurements. For example, in our study above, the mean
described the absenteeism rates of five nurses on each unit. The Indian census represents another example of
descriptive statistics. In this case, the information that is gathered concerning gender, race, income, etc. is
compiled to describe the population of Indiaat a given point in time. A Batsman batting average is another
example of a descriptive statistic. It describes the player's past ability to Score runs at any point in time. What
these three examples have in common is that they organize, summarize, and describe a set of measurements.
• Inferential statistics use data gathered from a sample to make inferences about the larger population from
which the sample was drawn. For example, we could take the information gained from our nursing
satisfaction study and make inferences to all hospital nurses. We might infer that cardiac care nurses as a
group are less satisfied with their jobs as indicated by absenteeism rates. Opinion polls and television ratings
systems represent other uses of inferential statistics. For example, a limited number of people are polled
during an election and then this information is used to describe voters as a whole
AMITY INSTITUTE OF PSYCHOLOGY AND ALLIED SCIENCE, AIPS
WHAT IS MEASUREMENT?
• Normally, when one hears the term measurement, they may think in terms of measuring the length of
something (ie. the length of a piece of wood) or measuring a quantity of something (ie. a cup of flour). This
represents a limited use of the term measurement. In statistics, the term measurement is used more broadly
and is more appropriately termed scales of measurement. Scales of measurement refer to ways in which
variables/numbers are defined and categorized. Each scale of measurement has certain properties which in
turn determines the appropriateness for use of certain statistical analyses. The four scales of measurement
are nominal, ordinal, interval, and ratio.
• Nominal: Categorical data and numbers that are simply used as identifiers or names represent a nominal scale
of measurement. Numbers on the back of a cricket player and your social security number are examples of
nominal data. If I conduct a study and I'm including gender as a variable, I will code Female as 1 and Male as 2
or visa versa when I enter my data into the computer. Thus, I am using the numbers 1 and 2 to represent
categories of data.
• Ordinal: An ordinal scale of measurement represents an ordered series of relationships or rank order.
Individuals competing in a contest may be fortunate to achieve first, second, or third place. First, second, and
third place represent ordinal data. If Rahul takes first and Anjali takes second, we do not know if the
competition was close; we only know that Rahul outperformed Anjali.
AMITY INSTITUTE OF PSYCHOLOGY AND ALLIED SCIENCE, AIPS
• Likert-type scales (such as "On a scale of 1 to 10 with one being no pain and ten being high pain, how
much pain are you in today?") also represent ordinal data.. Fundamentally, these scales do not represent a
measurable quantity. An individual may respond 8 to this question and be in less pain than someone else
who responded 5. A person may not be in half as much pain if they responded 4 than if they responded 8.
All we know from this data is that an individual who responds 6 is in less pain than if they responded 8 and
in more pain than if they responded 4. Therefore, Likert-type scales only represent a rank ordering
• Interval: A scale which represents quantity and has equal units but for which zero represents simply an
additional point of measurement is an interval scale. The Fahrenheit scale is a clear example of the interval
scale of measurement. Thus, 60 degree Fahrenheit or -10 degrees Fahrenheit are interval data.
Measurement of Sea Level is another example of an interval scale. With each of these scales there is
direct, measurable quantity with equality of units. In addition, zero does not represent the absolute lowest
value. Rather, it is point on the scale with numbers both above and below it (for example, -10 degrees
Fahrenheit).
• Ratio: The ratio scale of measurement is similar to the interval scale in that it also represents quantity and
has equality of units. However, this scale also has an absolute zero (no numbers exist below the zero). Very
often, physical measures will represent ratio data (for example, height and weight). If one is measuring the
length of a piece of wood in centimeters, there is quantity, equal units, and that measure can not go below
zero centimeters. A negative length is not possible
AMITY INSTITUTE OF PSYCHOLOGY AND ALLIED SCIENCE, AIPS
• As stated above, many measures (ie. personality, intelligence, psycho-social, etc.) within the behavioral and social
sciences represent ordinal data. IQ scores may be computed for a group of individuals. They will represent differences
between individuals and the direction of those differences but they lack the property of indicating the amount of the
differences. Psychologists have no way of truly measuring and quantifying intelligence. An individual with an IQ of 70
does not have exactly half of the intelligence of an individual with an IQ of 140. Therefore, IQ scales should theoretically
be treated as ordinal data.
CENTRAL TENDENCY
• It is that one representative value which can-be used to locate and summaries the entire set of varying values.
• This one value can be used to make many decisions concerning the entire set.
• We can define measures of central tendency (or location) to find some central value around which the data tend to
cluster.
• Measures of central tendency i.e. condensing the mass of data in one single value, enable us to get an idea of the
entire data. For example, it is impossible to remember the individual incomes of millions of earning people of India. But
if the average income is obtained, we get one single value that represents the entire population.
• Measures of central tendency also enable us to compare two or more sets of data to facilitate comparison.
• Suppose you have data, for instance, marks in psychology obtained by students in 12th standard and you want to
analyse it . You can of course organise the data with the help of classification and tabulation but if you want to further
analyse the data then you can compute the average marks obtained by the whole class or find the midpoint for marks
above and below which will lie half of the students or you can also find out most frequent marks obtained by the
students. The techniques you are employing here are mean, median and mode. These are called measures of central
tendency and can be categorised under descriptive statistics
FUNCTIONS OF MEASURE OF CENTRAL TENDENCY
The main functions of measures of central tendency are as follows:
1) They provide a summary figure with the help of which the central location of the whole data can be explained. When we
compute an average of a certain group we get an idea about the whole data.
2) Large amount of data can be easily reduced to a single figure. Mean, median and mode can be computed for a large data
and a single figure can be derived
3) When mean is computed for a certain sample, it will help gauge the population mean.
4) The results obtained from computing measures of central tendency will help in making certain decisions. This holds true
not only to decisions with regard to research but could have applications in varied areas like policy making, marketing and
sales and so on.
5) Comparison can be carried out based on single figures computed with the help of measures of central tendency. For
example, with regard to performance of students in mathematics test, the mean marks obtained by girls and the mean
marks obtained by boys can be compared.
PROPERTIES OF GOOD MEASURE OF CENTRAL TENDENCY
i) It should he easy to understand.
ii) It should he simple to compute. It should not involve elaborate mathematical calculations.
iii) It should be based on all observations.
iv) It should be uniquely defined. It should not be subject to varied interpretations and needs to be unaffected by any
individual bias.
v) It should be capable of further algebraic treatment.
vi) It should not be unduly affected by extreme values.
vii)The data needs to be collected from a sample that truly represents the population. The sample thus needs to be
randomly selected.
ARITHMETIC MEAN
• The arithmetic mean (or mean or average) is the most commonly used and readily understood measure of central
tendency. In statistics, the term average refers to any of the measures of central tendency. The arithmetic mean is
defined as being equal to the sum of the numerical values of each and every observation divided by the total number of
observations. Symbolically, it can be represented as( Refer to formula at the end of the PPT)
• Mean for sample is denoted by symbol ‘M or x̅ (‘x-bar’)’ and mean for population is denoted by ‘µ’ (mu). It is one of the
most commonly used measures of central tendency and is often referred to as average.
• It can also be termed as one of the most sensitive measure of central tendency as all the scores in a data are taken in to
consideration when it is computed (Bordens and Abbott, 2011). Further statistical techniques can be computed based on
mean, thus, making it even more useful.
• Mean is a total of all the scores in data divided by the total number of scores. For example, if there are 100 students in a
class and we want to find mean or average marks obtained by them in a psychology test, we will add all their marks and
divide by 100, (that is the number of students) to obtain mean.
• When the observations are classified into a frequency distribution, the midpoint of the class interval would be treated as
the representative average value of that class. Therefore, for grouped data; the arithmetic mean is stated as ( Refer to
formula at the end of the PPT)
• The arithmetic mean has the great advantages of being easily computed and readily understood. It is due to the fact
that it possesses almost all the properties of a good measure of central tendency.
• No other measure of central tendency possesses so many properties. However, the arithmetic mean has some
disadvantages.
• The major disadvantage is that its value may be distorted by the presence of extreme values in a given set of data. A
minor disadvantage is. When it is used for open-end distribution since it is difficult to assign a midpoint value to the
open-end class.
MEDIAN
A second measure of central tendency is the median. Median is that value which divides the distribution into two equal parts.
Fifty per cent of the observations in the distribution are above the value of median and other fifty per cent of the
observations are below this value of median. The median is the value of the middle observation when the series is arranged
in order of size or magnitude. ( Refer to formula at the end of the PPT)
Quantiles are the related positional measures of central tendency. These are useful and frequently employed measures of
non-central location. The most familiar quantiles are the quartiles, deciles, and percentiles.
Quartiles: Quartiles are those values which divide the total data into four equal parts. Since three points divide the
distribution into four equal parts, we shall have three quartiles. Let us call them Q1, Q2, and Q3. The first quartile, Q1, is the
value such that 25% of the observations are smaller and 75% of the observations are larger. The second quartile, Q2, is the
median, i.e., 50% of the observations are smaller and 50% are larger. The third quartile, Q3, is the value such that 75% of the
observations are smaller and 25% of the observations are larger.
MODE
The mode is the typical or commonly observed value in a set of data. It is defined as the value which occurs most often
or with the greatest frequency. The dictionary meaning of the term mode is most usual'. For example, in the series of
numbers 3, 4, 5, 5, 6, 7, 8, 8, 8, 9, the mode is 8 because it occurs the maximum number of times.
Mode is easier to compute than mean and media. But it is not used often because of lack of stability from one sample to
another and also because a single set of data may possibly have more than one mode. Also, when there is more than
one mode, then the modes cannot be termed to adequately measure central location.
Mode is not affected by outliers or extreme scores
( Refer to formula at the end of the PPT)
HOW TO CHOOSE A MEASURE OF CENTRAL TENDENCY?
• The choice of a measure of central tendency will depend on first of all, the scales of measurement
• For nominal scales one can compute mode but not mean or median. For example, in case of males and females, the males
can be coded as 1 and females can be coded as 2 (or vice versa) in such a case, we can compute frequently occurring score,
that will provide us information whether there are more males or more females. However, it is not possible to compute mean
or median.
• With regard to ordinal scale, median or mode can be used. For example, if we rank the students based on their performance
in mathematics test, it is possible to find median below and above which lie half of the ranks. Mode can also be computed if
more than one student gets same rank.
• With regard to interval scale and ratio scale mean can be computed.
• Yet another aspect that is important while making a choice with regard to which measure of central tendency to use is,
whether the data is normally distributed or not. If the data is normally distributed we will compute mean and if it is not
normally distributed, we will compute median or mode. This is because mean may not adequately represent the data when
the data is not normally distributed
The Graphical Presentation of Ungrouped Data
For the data which is not grouped into a frequency distribution we use the following common graphs or diagrams.
(i) Pictographs or Pictograms.
Pictographs or Pictograms Pictographs or pictograms are the graphs or diagrams used for presenting an ungrouped
statistical data in pectoral (picture like) form: A picture is said to be worth more than 100 words spoken or written. Thereby
the pictorial representation of the data is always considered better than its description in the words and figures.
Let us illustrate this fact through an example
In a data collection process it was found that there are 100 students in class VI; 85 students in Class VII, 80 Students in VIII,
90 Students in IX and 70 Student in X. Present this data first into a tabular form and then in pectoral form.
(ii) Bargraphs on Bar Diagrams.
Bar Graphs or Bar Diagrams In stead of using pictures we can: use bars (rectangles of similar breadth) for the
representation of numerical data. this mode of presentation of statistical data through bars is known as bargraphs or bar
diagrams.
(iii) Circle or pie graphs/diagrams.
Circle graph or Pie diagrams Circle or pie graphs/diagrams provide us an opportunity to represent statistical data through
the figure or a circle and its constituents Le. proportionate sub-divisions. These are specifically helpful in the case for which
the question of proportion is of much interest. To construct them requires a working knowledge of angle measurement and
percentages.
(iv) Line graphs
Line Graphs Line graphs can be better used in describing the Con-committed relationships between two variables by
plotting their respective values on the x and y axes of a graph paper (After choosing appropriate scales).
THE GRAPHICAL PRESENTATION OF FREQUENCY DISTRIBUTION (GROUPED DATA)
1. The Histogram or column diagram
The Histogram–A histogram or column diagram is essentially a bar graph of a frequency distribution
2. The Frequency Polygon.
The Frequency Polygon–A frequency polygon is essentially, a line graph for the graphical representation of the
frequency distribution. We can get a frequency polygon from a histogram, if the mid points of the upper bases of the
rectangles are connected. But it is not essential a plot histogram first to draw a frequency polygon. We can construct it
directly from a given frequency distribution.
3. The Cumulative Frequency Graph.
Cumulative frequency graphs (or cumulative frequency diagrams) are useful when representing or analysing the
distribution of large grouped data sets. They can also be used to find estimates for the median value, the lower quartile
and the upper quartile for the data set.
Comparison Between the Histogram and the Frequency Polygon
Although Histogram and Frequency polygon-both are used for the graphic representation of the frequency distribution
and are alike in many aspects yet they posses points of differences. Some of these differences can be cited as below–
1. Where Histogram is essentially the bar graph of the given frequency distribution, the Frequency polygon is a line
graph of this distribution.
2. In Frequency polygon we assume frequencies to be concentrated at the mid-points of the class interval. It points
out merely the graphical relationship between mid-points and frequencies and thus is unable to show the distribution
of frequencies within each class interval. But the Histogram gives a very clear as well as accurate picture of the relative
proportions of frequency from interval to interval.
3. In comparing two or more distributions by plotting two or more graphs on the same axes, Frequency polygon is
more useful and practicable than the Histogram as in such cases vertical and horizontal lines in the histogram tend to
coincide.
MEASURES OF VARIABILITY
• Dispersion actually refers to the variations that exist within and amongst the scores obtained by a group. In average there
is a convergence of scores towards a mid-point in a normal distribution.
• In dispersion, we try and see how each score in the group varies from the mean or the average score. The larger the
dispersion, less is the homogeneity of the group concerned and if the dispersion is less it means that the group is
homogeneous.
• Dispersion is an important statistic which helps us to know how far the sample population varies from the universe
population. It tells us about the standard error of the mean
• Variability in statistics means deviation of scores in a group or series, from their mean scores. It actually refers to the
spread of scores in the group in relation to the mean. It is also known as dispersion. For instance, in a group of 10
participants who have scored differently on a mathematics test, each individual varies from the other in terms of the
marks that he/she has scored.
• These variations can be measured with the help of measure of variability, that measure the dispersion of different values
for the average value or average score. Variability or dispersion also means the scatter of the values in a group. High
variability in the distribution means that scores are widely spread and are not homogeneous. Low variability means that
the scores are similar and homogeneous and are concentrated in the middle
• Measures of variability fall under descriptive statistics that describe how similar a set of scores are to each other. The
greater the similarity of the scores to each other, lower would be the measure of variability or dispersion. The less the
similarity of the scores are to each other, higher will be the measure of variability or dispersion.
Importance if measure of variability
Measures of variability are used to test the extent to which an average represents the characteristics of a data.
If the variation is small then it indicates high uniformity of values in the distribution and the average represents the
characteristics of the data. On the other hand, if variation is large then it indicates lower degree of uniformity and
unreliable average.
Measures of variability help in identifying the nature and cause of variation. Such information can be useful to control
the variation. Measures of variability help in the comparison of the spread in two or more sets of data with respect to
their uniformity or consistency.
Measures of variability facilitate the use of other statistical techniques such as correlation, regression analysis, and so on.
FUNCTIONS OF VARIABILITY
• The major functions of dispersion or variability are as follows:
•
• It is used for calculating other statistics such as analysis of variance, degree of correlation, regression etc.
• It is also used for comparing the variability in the data obtained as in the case of Socio-Economic Status, income,
education etc.
• To find out if the average or the mean/median/mode worked out is reliable. If the variation is small then we could state
that the average calculated is reliable, but if variation is too large, then the average may be erroneous.
• Dispersion gives us an idea if the variability is adversely affecting the data and thus helps in controlling the variability
TYPES OF MEASURES OF DISPERSION OR VARIABILITY
The measures of variability most commonly used in psychological statistics are
1) Range 2) Quartile Deviation 3) Average Deviation or Mean Deviation 4) Standard Deviation 5) Variance
The Average Deviation (AD) or Mean Deviation (MD)
• To measure the variation, as a degree to which values within a data deviate from their mean, we use average deviation.
• According to Garrett (1971, as cited in Mangal 2002, page 70)
• “The average deviation is the mean of the deviation of all of the separate scores is a series taken from their mean”.
According to Guilford (1963) average deviation can be described as an average or mean of all the deviations when the
algebraic signs are not taken in to the account.
• Average is a central value and thus, some deviations will be positive (+) and some may be negative (-).
• Mean deviation ignores the signs of the deviations, and it considers all the deviations to be positive. This is so because
the algebraic sum of all the deviations from the mean equals to zero. MD or AD is arithmetic mean of the difference of
the values from the average.
• The average is either the arithmetic mean or the median. It is a measure of variability that takes into account the
variations of all the scores in the data. It is an absolute measure of dispersion and is expressed in the same unit as the
raw scores
• AD or MD is easy to understand and compute.
( Refer to formula at the end of the PPT)
• It is based on all observations, unlike R or QD.
• It is an accurate measure of variability since it averages the absolute deviations.
• It is less affected by extreme observations.
• It is based on average thus, it is a better measure to compare about the formations of different distributions.
The main limitations of average deviation are as follows:
• While calculating average deviation we ignore the plus minus sign and consider all values as plus. Because of this
mathematical property, it is not used in inferential statistics.
• AD cannot be computed for open-end classes.
• It tends to increase with the size of the sample
THE STANDARD DEVIATION (SD)
• The term standard deviation was first used in writing by Karl Pearson in 1894.
• The standard deviation of population is denoted by ‘σ’ (Greek letter sigma) and that for a sample is ‘s’.
• A useful property of SD is that unlike variance it is expressed in the same unit as the data.
• This is most widely used method of variability.
• The standard deviation indicates the average of distance of all the scores around the mean. It is the positive square root
of the mean of squared deviations of all the scores from the mean. It is the positive square root of variance. It is also
called as ‘root mean square deviation’.
• Mangal (2002, page 71) defined standard deviation as “as the square root of the average of the squares of the deviations
of each score from the mean”.
• SD is an absolute measure of dispersion and it is the most stable and reliable measure of variability. Standard deviation
shows how much variation there is, from the mean.
( Refer to formula at the end of the PPT)
• SD is calculated from the mean only. If standard deviation is low it means that the data is close to the mean.
• A high standard deviation indicates that the data is spread out over a large range of values. Standard deviation may serve
as a measure of uncertainty.
• If you want to test the theory or in other word, want to decide whether measurements agree with a theoretical
prediction, the standard deviation provides the information.
• If the difference between mean and standard deviation is very large then the theory being tested probably needs to be
revised.
• The mean with smaller standard deviation is more reliable than mean with large standard deviation.
• A smaller SD shows the homogeneity of the data.
• The value of standard deviation is based on every observation in a set of data. It is the only measure of dispersion
capable of algebraic treatment therefore, SD is used in further statistical analysis.
MERITS AND LIMITATIONS OF THE STANDARD DEVIATION
The main merits of using standard deviation are as follows:
1) It is widely used because it is the best measure of variation by virtue of its mathematical characteristics.
2) It is based on all the observations of the data.
3) It gives an accurate estimate of population parameter when compared with other measures of variation.
4) SD is least affected by sample fluctuations
5) It is also possible to calculate combined SD, that is not possible with other measures.
6) Further statistics can be applied on the basis of SD like, correlation, regression, tests of significance, etc.
7) Coefficient of variation is based on mean and SD. It is the most appropriate method to compare variability of two or more
distributions.
The limitation of SD
While calculating standard deviation more weight is given to extreme values and less to those, near the means. When we
calculate SD, we take deviation from mean (X-M) and square these obtained deviations. Therefore, large deviations, when
squared are proportionally more than small deviations. For example, the deviations 2 and 10 are in the ratio of 1:5 but their
square 4 and l00 are in the ratio 1:25.
2) It is difficult to compute as compared to other measures of dispersion
Uses of Standard Deviation The uses of standard deviation are as follows:
1) SD is used when one requires a more reliable and accurate measure of variability but it is recommended when the
distribution is normal or near to normal.
2) It is used when further statistics like, correlation, regression, tests of significance, etc. have to be computed.
The Quartile Deviation (QD)
• Since a large number of values in the data lie in the middle of the frequencies distribution and range depends on
the extreme (outliers) of a distribution, we need another measure of variability.
• The Quartile deviation, is a measure that depends on the relatively stable central portion of a distribution.
• According to Garret (1966), the Quartile deviation is half the scale distance between 75th and 25th per cent in a
frequency distribution.
• The entire data is divided into four equal parts and each part contains 25% of the values.
• According to Guilford (1963) the Semi-Interquartile range is the one half the range of the middle 50 percent of the
cases. On the basis of above definitions, it can be said that quartile deviation is half the distance between Q1 and
Q3.
• Quartile deviation is closely related to the median because median is responsive to the number of scores lying
below it rather than to their exact positions and Q1 and Q3 are defined in a same manner.
( Refer to formula at the end of the PPT)
• The median and quartile deviation have common properties. Both median and quartile deviation are not affected by
extreme values. In a symmetrical distribution, the two quartiles Q1 and Q3 are at equal distance from the median or
Q1 =Q3- Median.
• Thus, like median, quartile deviation covers exactly 50 per cent of observed values in the data. In normal distribution,
quartile deviation is called the Probable Error or PE. If the distribution is open-class, then quartile deviation is the only
measure of variability that is reasonable to compute.
Merits and Limitations of Quartile Deviation
• Quartile deviation is easy to understand and compute.
• Quartile deviation is a better measure of dispersion than range because it takes into account 50 per cent of the data,
unlike the range which is based on two values of the data, that is highest value and the lowest value.
• Secondly, quartile deviation is not affected by extreme scores since it does not consider 25 per cent data from the
beginning and 25 per cent from the end of the data.
• Lastly, quartile deviation is the only measure of dispersion which can be computed from the frequency distribution
with open-end class.
Despite the major merits of quartile deviation, there are limitations to it as well.
1) The value of quartile deviation is based on the middle 50 percent values, it is not based on all the observations. Thus, it
is not regarded as a stable measure of variability
2) The value of quartile deviation is affected by sampling fluctuation.
3) The value of quartile deviation is not affected by the distribution of the individual values within the intervals of middle
50 percent observed values.
Uses of Quartile Deviation
1) The distribution contains few and very extreme scores.
2) When the median is the measure of central tendency.
3) When our primary interest is to determine the concentration around the median.
DIFFERENCE BETWEEN PARAMETRIC AND NON PARAMETRIC TEST
Parametric test Non Parametric test
One which has the complete information about A non parametric test is one for which the
the population parameter or that can make researcher has no knowledge about the
assumptions about the parameters of population population parameter neither he can make
distribution from which the samples have been assumptions about the distribution of population
drawn
Uses a normal probability distribution The distribution is not normal
Uses interval and ration scale of measurement Uses nominal and ordinal scale of measurement
Preferred measure of central tendency is mean Preferred measure of central tendency is median
Homogeneity of variance exists There is no homogeneity of variance
Example: Pearson Correlation Coefficients, t test Example: Spearman Rank correlation coefficient,
Mann Whitney test