Statistics
Statistics
Statistics is the science of collecting, organizing, presenting, analyzing, and interpreting numerical data to assist
in making more effective decisions. Statistical techniques are used extensively by marketing, accounting, quality
control, consumers, professional sports people, hospital administrators, educators, politicians, physicians, and
many others.
Descriptive statistics
Descriptive statistics is a branch of statistics that deals with the collection, analysis, presentation, and
interpretation of data. It is used to summarize and describe the main features of a data set, such as its central
tendency, variability, and distribution. Descriptive statistics can be used to identify patterns and relationships in
the data, and to make inferences about the population from which the sample was drawn.
Inferential statistics
Inferential statistics is a branch of statistics that deals with drawing conclusions about populations based on
data collected from samples. It uses statistical methods to estimate population parameters, test hypotheses,
and make predictions.
• Population
A population in statistics is the entire group of individuals, objects, or events that you want to draw
conclusions about. It can be any size, but it is often too large and impractical to study the entire population.
In these cases, researchers will select a sample to study instead.
• Sample
A sample in statistics is a subset of the population that is selected to represent the entire population.
The sample should be representative of the population in terms of all relevant characteristics, such as
age, gender, location, and income.
Variable
A variable is a quantity that can take on different values. It is a quantity that “varies.”
Qualitative or Attribute variable
A qualitative or attribute variable is a variable that describes a characteristic that can't be easily measured, but
can be observed subjectively—such as smells, tastes, textures, attractiveness, and color. Broadly speaking, when
you measure something and give it a numerical value, you create quantitative data. A data set (or dataset) is a
collection of data.
Discrete variable
A discrete variable is a variable that can only take on a finite number of values. The values of a discrete variable
can be counted, rather than measured. For example, the number of students in a class is a discrete variable,
because there can only be a certain number of students in a class.
Continuous
A continuous variable is a variable that can take on any value within a range. A continuous variable takes on an
infinite number of possible values within a given range. The values of a continuous variable can be measured,
rather than counted. For example, the height of a person is a continuous variable, because it can be measured
to any number of decimal places.
Levels of data
The four levels of data in statistics are nominal, ordinal, interval, and ratio. Each level of data has its own unique
characteristics and limitations.
• Nominal data is the most basic level of data. It is used to classify data into categories. Nominal data has
no intrinsic order. For example, eye color, hair color, and nationality are all nominal variables. Nominal
level variables must be mutually exclusive and exhaustive.
Mutually exclusive means that no observation can belong to more than one category. For example, if
you are classifying eye color, no person can have both blue eyes and brown eyes.
Exhaustive means that all observations must belong to at least one category. For example, if you are
classifying eye color, all people must have either blue eyes, brown eyes, green eyes, or some other color
of eyes.
• Ordinal data is data that has been ranked or ordered. Ordinal data has an intrinsic order, but the intervals
between the ranks are not necessarily equal. For example, customer satisfaction ratings (e.g., very
satisfied, satisfied, neutral, dissatisfied, very dissatisfied) are ordinal data.
• Interval data is data that has been measured on a scale with equal intervals between the units. Interval
data has an intrinsic order and the intervals between the units are equal. For example, temperature in
degrees Celsius or Fahrenheit is interval data.
• Ratio data is the highest level of data. It is data that has been measured on a scale with equal intervals
between the units and a true zero point. Ratio data has an intrinsic order, the intervals between the units
are equal, and there is a true zero point. For example, height and weight are ratio data.
Ungrouped data is data that is presented in its raw form, without being grouped or summarized. It is a list of all
the individual values in the data set. For example, the following is ungrouped data showing the heights of 10
students:
160, 165, 170, 175, 180, 185, 190, 195, 200, 205
Grouped data is data that has been grouped into categories or intervals. This can make the data easier to analyze
and visualize. For example, the following is grouped data showing the heights of the same 10 students:
Height range Frequency
160-169 2
170-179 4
180-189 3
190-199 1
Function of Statistics
Here are some of the key functions of statistics:
• To describe data: Statistics can be used to describe data in a way that is easy to understand and interpret.
This can be done by calculating summary statistics, such as the mean, median, and mode, or by creating
visualizations such as histograms and bar charts.
• To make inferences about populations: Statistics can be used to make inferences about populations based
on samples. This is done by using statistical methods such as hypothesis testing and confidence intervals.
• To predict future events: Statistics can be used to predict future events based on historical data. This is
done by using statistical methods such as time series analysis and machine learning.
• To make informed decisions: Statistics can be used to make informed decisions by weighing the risks and
benefits of different options. This is done by using statistical methods such as decision analysis and cost-
benefit analysis.
Scope of Statistics
The scope of statistics is vast and ever-expanding. It is used in almost every field of human endeavor, including
science, medicine, business, government, and education.
Here are some examples of how statistics is used in different fields:
• Science: Statistics is used by scientists to design experiments, analyze data, and draw conclusions about
the natural world. For example, a scientist might use statistics to determine whether a new drug is
effective or to study the impact of climate change on the environment.
• Medicine: Medical researchers use statistics to develop new drugs and treatments, design clinical trials,
and analyze medical data to identify trends and patterns. For example, medical researchers might use
statistics to determine the efficacy of a new drug or to study the risk factors for a particular disease.
• Business: Businesses use statistics to make decisions about everything from product development to
marketing campaigns to financial planning. For example, a business might use statistics to determine which
products to sell, how to price those products, and how to reach potential customers.
• Government: Governments use statistics to inform public policy decisions. For example, the government
might use statistics to determine how to allocate funding for education or healthcare, or to develop
policies to reduce crime or poverty.
• Education: Statistics is used by educators to assess student learning, evaluate teaching methods, and
develop educational programs. For example, an educator might use statistics to identify students who are
at risk of falling behind or to develop a new curriculum for a particular subject.
Limitation of Statistics
Statistics is a powerful tool, but it is important to be aware of its limitations. Here are some of the key limitations
of statistics:
• Statistics is only as good as the data it is based on. If the data is inaccurate or incomplete, the statistical
results will also be inaccurate or incomplete.
• Statistics can be misused to mislead people. For example, someone could cherry-pick data to make a
particular point, or they could use misleading charts and graphs to distort the data.
• Statistics cannot prove causation. Statistics can show that two variables are correlated, but they cannot
prove that one variable causes the other.
• Statistics is not a magic bullet. Statistics cannot solve all problems, and it is important to use common
sense and judgment when interpreting statistical results.
Hypothesis
A hypothesis is a statement about a phenomenon that can be tested with data. It is a proposed explanation for
something that is based on known facts but has not yet been proven.
Prediction
A prediction is a statement about a future event or outcome. It is a forecast or prophesy based on current
knowledge or past experience. Predictions can be made about anything, from the weather to the stock market
to the outcome of a sporting event.
Pie chart
A pie chart is a circular chart that is divided into slices. The size of each slice is proportional to the value of the
category it represents. Pie charts are used to show the relative contributions of different categories to a whole.
Line chart
A line chart is a type of chart that shows how a variable changes over time. The independent variable is plotted
on the x-axis and the dependent variable is plotted on the y-axis. Line charts are used to identify trends and
patterns in data.
Column chart
A column chart is a type of bar chart that displays data in vertical bars. The height of each bar represents the
value of the category it represents. Column charts are used to compare data across different categories.
Bar chart
A bar chart is a type of chart that displays data in horizontal bars. The length of each bar represents the value
of the category it represents. Bar charts are used to compare data across different categories.
Area chart
An area chart is a type of line chart that fills the area under the line with color. Area charts are used to show the
cumulative contribution of different categories to a whole.
Scatter plot
A scatter plot is a type of chart that shows the relationship between two variables. Each point on the scatter
plot represents a pair of values for the two variables. Scatter plots are used to identify patterns and relationships
between variables.
Combo chart
A combo chart is a type of chart that combines two or more different chart types. For example, a combo chart
might combine a line chart with a column chart or a scatter plot with an area chart. Combo charts are used to
visualize complex data in a single chart.
Pivot chart
A pivot chart is a type of interactive table that allows you to summarize and analyze data. Pivot charts can be
used to create a variety of different chart types, including pie charts, line charts, column charts, and bar charts.
Pivot charts are a powerful tool for data analysis and visualization.
Frequency Distribution
A frequency distribution is a table or chart that shows the number of times each value in a data set occurs. It is
a way of organizing and summarizing data so that it is easier to understand and analyze. Step of Frequency
Distribution:
• Step 1: Determine the range of the raw data.
• Step 2: Determine how many classes it will contain.
• Step 3: Determine the width of the class interval.
• Step 4: Start at a value equal to or lower than the lowest number of the ungrouped data and end at a
value equal to or higher than the highest number.
• Step 5: Class endpoints are selected so that value of the data can’t fit into more than one class.
Class Midpoint
The class midpoint is the middle value between the lower- and upper-class limits. It is a way of representing the
center of a class interval in a frequency distribution. The class midpoint is important, because it becomes the
representative value for each class in most group statistics calculations.
Frequency
Frequency is the number of times an observation or event occurs in a data set.
Relative frequency
Relative frequency is the proportion of observations in a data set that fall into a particular category or class. It
is calculated by dividing the frequency of the category by the total number of observations in the data set.
Cumulative frequency
Cumulative frequency is the total number of observations in a data set that fall below or at a certain value. It is
calculated by adding up the frequencies of all the categories up to and including the given value.
Measures of Central
Measures of central tendency are statistics that are used to describe the center of a data set. There are three
main measures of central tendency: mean, median, and mode.
Mean
The mean, also known as the arithmetic average, is the sum of all the values in a data set divided by the number
of values.
Properties of Arithmetic Mean
• The sum of deviations of the items from the arithmetic mean is always zero i.e. ∑(X–X) =0.
• The Sum of the squared deviations of the items from A.M. is minimum, which is less than the sum of the
squared deviations of the items from any other values.
• If each item in the series is replaced by the mean, then the sum of these substitutions will be equal to
the sum of the individual items.
Weighted Arithmetic mean
The weighted arithmetic mean is a type of mean that is calculated by multiplying each value in a data set by a
corresponding weight and then dividing the sum of the products by the sum of the weights. It is a more general
form of the arithmetic mean, which is calculated by simply averaging all of the values in a data set.
Formula:
Weighted mean = (sum of (value X weight)) / (sum of weights)
Median
The median is the middle value in a data set when the values are arranged in ascending or descending order. If
a data set has an even number of values, the median is the mean of the two middle values.
Mode
The mode is the most frequent value in a data set.
Range
The range is the difference between the largest and smallest values in a data set.
Objectives of Average
• To get one single value that describe the characteristics of the entire data.
• To facilitate comparison.
Good Average
• Understandable
• Simple
• Based on all the observation
• Capable of further algebraic treatment
• Sampling stability
• Not unduly affected by the presence of extreme values.
Importance of Measures of Central Tendency
• Businesses use measures of central tendency to track customer satisfaction, employee productivity, and
sales performance.
• Governments use measures of central tendency to track economic indicators, such as unemployment rates
and inflation rates.
• Researchers use measures of central tendency to summarize the findings of their studies and to draw
conclusions about populations based on samples.
• Individuals use measures of central tendency to make informed decisions about their personal finances,
health, and careers.
Percentile
A percentile is a measure that divides a set of data into 100 equal parts. To determine the location of a percentile
in a data set, you can follow these steps:
1. Order the data set from smallest to largest.
2. Calculate the percentile. You can do this using the following formula:
Percentile = (number of values less than the given value) / (total number of values) * 100
3. Round the percentile to the nearest integer.
4. Find the value in the data set that corresponds to the percentile. You can do this by counting the number
of values in the data set from smallest to largest until you reach the percentile. For example, if the
percentile is 75, then the 75th percentile is the value in the data set that is 75th in order from smallest
to largest.
Quartile
A quartile is a type of percentile that divides a set of data into four equal parts.
Pythagorean means
A Pythagorean mean is a type of mean that is calculated using the Pythagorean theorem. The three classical
Pythagorean means are the arithmetic mean, the geometric mean, and the harmonic mean.
• Arithmetic mean: The arithmetic mean is the simplest type of mean. It is calculated by adding all of the
values in a data set and then dividing by the number of values.
• Geometric mean: The geometric mean is calculated by multiplying all of the values in a data set and then
taking the nth root, where n is the number of values.
• Harmonic mean: The harmonic mean is calculated by taking the reciprocal of the arithmetic mean of the
reciprocals of the values in a data set.
Measure of Dispersion
A measure of dispersion is a statistical measure that quantifies the spread of data around a central value. It is
used to describe how much the data values vary from each other and from the central value. There are many
different measures of dispersion, but some of the most common include:
• Range: The range is the difference between the largest and smallest values in a data set. It is a simple
but effective measure of dispersion.
• Standard deviation: The standard deviation is the most common measure of dispersion. It is calculated
by taking the square root of the variance. The standard deviation is a good measure of dispersion
because it takes into account all of the data values.
• Variance: The variance is the average squared deviation from the mean. It is a good measure of
dispersion because it is additive. This means that the variance of a combined data set is equal to the sum
of the variances of the individual data sets.
Measures of dispersion have many advantages:
• They provide a better understanding of the distribution of data. Measures of dispersion, such as the range,
IQR, and standard deviation, can be used to identify outliers, skewness, and other patterns in the data.
This information can be used to make more informed decisions about how to use and interpret the data.
• They allow for comparisons between different data sets. Measures of dispersion can be used to compare
the variability of two or more data sets. This can be useful for identifying differences in the distribution of
data between different groups or populations.
• They can be used to assess the quality of data collection and analysis. Measures of dispersion can be used
to identify errors in data collection and analysis. For example, if the standard deviation is too high, it may
indicate that there are errors in the data or that the data collection method is not reliable.
Mean of a population
The mean of a population is the sum of all the values in the population divided by the number of values in the
population. It is a measure of the central tendency of the population.
Mean of a sample
The mean of a sample is the sum of all the values in the sample divided by the number of values in the sample.
It is an estimate of the mean of the population from which the sample was drawn.
population variance
Population variance is a measure of how spread out the values in a population are. It is calculated by taking the
squared deviations of each value from the mean, averaging them, and then taking the square root.
Sample variance
Sample variance is a measure of how spread out the values in a sample are. It is calculated by taking the squared
deviations of each value from the sample mean, averaging them, and then taking the square root.
The normal rule of Standard deviation
The normal rule, also known as the 68-95-99.7 rule, is a statistical rule that states that for normally distributed
data, almost all of the values will fall within three standard deviations of the mean. Specifically, the normal rule
predicts that:
• 68% of the values will fall within one standard deviation of the mean
• 95% of the values will fall within two standard deviations of the mean
• 99.7% of the values will fall within three standard deviations of the mean
Geometric mean
The geometric mean is a type of mean that is calculated by multiplying all of the values in a data set and then
taking the nth root, where n is the number of values.
Coefficient of variation
The coefficient of variation (CV) is a statistical measure of the dispersion of data points around the mean relative
to the mean. It is calculated by dividing the standard deviation by the mean and expressing the result as a
percentage.
Moments
Moments are measures of the central tendency and variability of a data set. The first moment is the mean, the
second moment is the variance, the third moment is the skewness, and the fourth moment is the kurtosis.
Skewness
Skewness is a measure of the asymmetry of a data distribution. A positive skewness indicates that the data is
skewed to the right, with a longer tail on the right side of the distribution. A negative skewness indicates that
the data is skewed to the left, with a longer tail on the left side of the distribution. A skewness of zero indicates
that the data is symmetric.
Kurtosis
Kurtosis is a measure of the peakedness of a data distribution. A high kurtosis indicates that the data is more
peaked, with more values concentrated around the mean. A low kurtosis indicates that the data is less peaked,
with the values more spread out. A kurtosis of three indicates that the data is normally distributed.
Linear regression
Linear regression is a simpler method than nonlinear regression, and it is easier to interpret the results of linear
regression models. However, linear regression can only model linear relationships between variables.
Nonlinear regression
Nonlinear regression can model more complex relationships between variables than linear regression, but it is
also a more complex method. Nonlinear regression models can be difficult to fit to data, and the results of
nonlinear regression models can be difficult to interpret.
Probability
Probability is the measure of how likely an event is to occur. It is a number between 0 and 1, with 0 representing
impossibility and 1 representing certainty.
Birth of Probability
The concept of probability first emerged in the 17th century, when mathematicians such as Girolamo Cardano
and Pierre de Fermat began to study games of chance. They developed early probability models to predict the
outcome of dice rolls and other gambling games.
In the 18th century, the mathematician Abraham de Moivre wrote the first book on probability theory, The
Doctrine of Chances. De Moivre's work helped to establish probability as a legitimate field of mathematical
study.
Three Approaches of probability
There are three main approaches to probability:
• Classical probability: This approach assumes that all outcomes of an experiment are equally likely. For
example, if you roll a die, the probability of getting any given number is 1/6.
• Relative frequency probability: This approach is based on the results of repeated experiments. For
example, if you roll a die 100 times and get a 6 20 times, the relative frequency probability of getting a
6 is 0.2.
• Subjective probability: This approach is based on personal beliefs and judgments. For example, if you
believe that it is more likely to rain tomorrow than it is not, your subjective probability of rain is greater
than 0.5.
Venn diagram
A Venn diagram is a graphical representation of the relationships between different sets. It is often used to
illustrate probability concepts. In a Venn diagram, each set is represented by a circle. The intersection of two
circles represents the elements that are in both sets. The union of two circles represents the elements that are
in either set or both sets.
Tree diagram
A tree diagram is a graphical representation of the possible outcomes of a sequence of events. It is often used
to calculate the probability of a compound event. In a tree diagram, each event is represented by a branch. The
branches that connect to a given branch represent the possible outcomes of the previous event.
Addition rule
The addition rule states that the probability of the union of two disjoint events is equal to the sum of the
probabilities of the two events. In other words, if event A and event B are disjoint (meaning that they cannot
both occur at the same time), then the probability that either event A or event B will occur is equal to the
probability of event A plus the probability of event B.
Multiplication rule
The multiplication rule states that the probability of the intersection of two independent events is equal to the
product of the probabilities of the two events. In other words, if event A and event B are independent (meaning
that the outcome of one event does not affect the outcome of the other event), then the probability that both
event A and event B will occur is equal to the probability of event A multiplied by the probability of event B.
Classical approach of probability
The classical approach to probability is based on the assumption that all outcomes of an experiment are equally
likely. This means that each outcome has the same chance of occurring.
To calculate the probability of an event using the classical approach, we use the following formula:
Probability of event = Number of favorable outcomes / Total number of outcomes
For example, if we flip a coin, there are two possible outcomes: heads or tails. Since we assume that all outcomes
are equally likely, the probability of getting heads is 1/2 and the probability of getting tails is 1/2.
Experiment
An experiment is an activity that is carried out to observe the result of a process under controlled conditions.
Outcome
An outcome is a possible result of an experiment.
Exhaustive
An exhaustive set of events is a set of events that includes all possible outcomes of an experiment.
Mutually exclusive
Mutually exclusive events are events that cannot occur at the same time.
Examples:
• Experiment: Flipping a coin
• Outcomes: Heads or tails
• Exhaustive set of events: {heads, tails}
• Mutually exclusive events: {heads}, {tails}
Sample space
The sample space of an experiment is the set of all possible outcomes of that experiment.
Sample point
A sample point is a single possible outcome of an experiment.
Event
An event is a subset of the sample space. In other words, an event is a collection of one or more sample points.
Examples:
• Experiment: Flipping a coin
• Sample space: {heads, tails}
• Sample point: heads
• Event: {heads}
• Experiment: Rolling a die
• Sample space: {1, 2, 3, 4, 5, 6}
• Sample point: 3
Index Number Definition
An index number is a statistical measure of the relative change in a variable or group of variables over time.
Index numbers are used to track changes in prices, quantities, wages, and other economic variables.
Uses of Index Numbers
Index numbers are used in a variety of ways, including:
• Measuring inflation: Price indices are used to measure the rate of inflation, which is the rate at which
prices are rising over time.
• Deflating economic data: Index numbers can be used to deflate economic data, such as GDP, to account
for inflation.
• Making international comparisons: Index numbers can be used to make international comparisons of
economic variables, such as prices and wages.
• Setting wages and salaries: Index numbers are often used to set wages and salaries, such as through cost-
of-living adjustments.
Types of Index Numbers
There are two main types of index numbers:
Price indices: Price indices measure changes in prices over time.
Quantity indices: Quantity indices measure changes in quantities over time.
Steps in Computation of Fishers Ideal Index Number
The Fisher's ideal index number is a weighted average of the Laspeyres and Paasche price indices. It is considered
to be a more accurate measure of inflation than either the Laspeyres or Paasche indices alone.
To calculate the Fisher's ideal index number, the following steps are taken:
• Calculate the Laspeyres price index.
• Calculate the Paasche price index.
• Take the geometric mean of the Laspeyres and Paasche indices.
Merits of Fisher's Ideal Index Number
The Fisher's ideal index number has several merits, including:
• It is a more accurate measure of inflation than either the Laspeyres or Paasche indices alone.
• It satisfies the time reversal test and the factor reversal test.
• It is a weighted average of the Laspeyres and Paasche indices, which gives it the advantages of both
indices.
Demerits of Fisher's Ideal Index Number
The Fisher's ideal index number has one main demerit: it is more complex to calculate than the Laspeyres and
Paasche indices.
Why is it called an ideal index number?
The Fisher's ideal index number is called an ideal index number because it satisfies several important properties,
including the time reversal test, the factor reversal test, and the circular test.
Time Reversal Test
The time reversal test states that if an index number is reversed in time, the result should be equal to 1. The
Fisher's ideal index number satisfies the time reversal test.
Factor Reversal Test
The factor reversal test states that if an index number is calculated for two different sets of weights, the result
should be the same. The Fisher's ideal index number satisfies the factor reversal test.
Circular Test
The circular test states that if an index number is calculated for three or more time periods, the result should
be the same as if the index number is calculated directly between the first and last time periods. The Fisher's
ideal index number satisfies the circular test.
Price Index
A price index is a statistical measure of the relative change in prices over time. Price indices are used to track
changes in the cost of living, the cost of production, and other economic variables.
Quantity Index
A quantity index is a statistical measure of the relative change in quantities over time. Quantity indices are used
to track changes in the volume of production, the volume of consumption, and other economic variables.
Value Index
A value index is a statistical measure of the relative change in the value of goods and services over time. Value
indices are calculated by multiplying price indices and quantity indices.
Unweighted Composite/Aggregate Index
An unweighted composite/aggregate index is a statistical measure of the relative change in a group of variables
over time. Unweighted composite/aggregate indices are calculated by taking the average of the relative changes
in the individual variables.
Weighted Composite/Aggregate Index
A weighted composite/aggregate index is a statistical measure of the relative change in a group of variables
over time. Weighted composite/aggregate indices are calculated by taking the weighted average of the relative
changes in the individual variables.
Joint Probability
The joint probability of two events A and B is the probability that both events A and B occur. It is denoted by
P(A and B).
Marginal Probability
The marginal probability of an event A is the probability that event A occurs, regardless of whether any other
events occur. It is denoted by P(A).
Conditional Probability
The conditional probability of event A given event B, denoted by P(A|B), is the probability that event A occurs,
given that event B has already occurred.
Dependent Event
Dependent events are events whose outcomes are not independent of each other. This means that the outcome
of one event affects the outcome of the other event.
Independent Event
Independent events are events whose outcomes are independent of each other. This means that the outcome
of one event does not affect the outcome of the other event.