MODULE 1
INTRODUCTION TO BUSINESS STATISTICS
Business Statistics is the branch of statistics that deals with collecting, organizing,
presenting, analyzing, and interpreting numerical data for business decision-making.
It helps businesses understand trends, forecast future outcomes, and make informed
decisions in areas like sales, production, finance, and marketing.
Applications of Business Statistics
Business statistics helps businesses make better decisions by using numbers and
data. Its applications are:
1. Forecasting and Planning
Business statistics helps a company predict future demand, sales, and
expenses.
Example: A shop estimates how many products will sell next month and plans
stock accordingly.
2. Policy Formulation and Decision Making
Statistics helps management make important business decisions and create
policies.
Example: A company decides whether to launch a new product based on
customer data.
3. Market Research and Customer Analysis
Businesses study customer preferences, buying habits, and market trends.
Example: A company conducts surveys to know which product customers like
most.
Sales and Revenue Forecasting
4.
Statistics helps estimate future sales and income using past records.
Example: A company predicts festival-season sales based on previous years.
5. Inventory Management
Statistics helps maintain the right amount of stock.
Example: A store keeps enough goods to avoid shortage or excess stock.
6. Production and Quality Control
Statistical methods help check product quality and improve production.
Example: A factory checks whether products meet quality standards.
7. Risk Assessment and Control
Statistics helps identify possible business risks and reduce losses.
Example: A bank studies loan repayment data before giving loans.
8. Financial Analysis
Statistics helps analyze profit, loss, expenses, and financial condition.
Example: A company studies financial statements to know its performance.
Functions of Business Statistics
Business statistics helps businesses understand data and make better decisions. Its
main functions are:
1. Informed Decision-Making
Statistics helps managers make decisions based on facts and numbers instead
of guessing.
Example: A company checks sales data before launching a new product.
2. Performance Evaluation
It helps measure how well different departments are performing.
Example: Comparing monthly sales to see whether business is improving.
3. Risk Assessment and Control
Statistics helps identify possible risks and reduce losses.
Example: A bank studies customer repayment history before giving loans.
4. Market Understanding
Businesses use statistics to understand customer needs and market trends.
Example: A company studies which products are popular among customers.
5. Resource Optimization
Statistics helps use money, labour, and materials efficiently.
Example: A factory decides how many workers are needed for production.
6. Quality Improvement
It helps improve product quality by checking defects and errors.
Example: A company checks whether products meet quality standards.
7. Forecasting
Statistics helps predict future sales, demand, or profits.
Example: A store predicts festival-season demand.
Statistical Methods
Statistical Methods are techniques used to collect, organize, analyze, and interpret
data. They help convert raw data into useful information.
Two Main Types of Statistical Methods
1. Descriptive Statistics
Descriptive statistics describes and summarizes data already collected.
It explains “What the data shows.”
Example: Finding the average marks of students in a class.
Types of Descriptive Statistics
a) Frequency Distribution
Shows how often values occur.
Example: Number of students scoring different marks.
b) Measures of Central Tendency
Shows the middle or average value.
● Mean = Average
● Median = Middle value
●
Mode = Most repeated value
Example: Average salary of employees.
c) Measures of Dispersion
Shows how spread out the data is.
Example: Difference between highest and lowest marks.
2. Inferential Statistics
Inferential statistics uses a sample to make conclusions about a larger group.
It explains “What we can predict or conclude.”
Example:
Surveying 100 customers to understand opinions of all
customers.
Main Methods of Inferential Statistics
a) Estimation
Used to estimate population values from sample data.
● Point Estimation: Gives one value estimate.
● Interval Estimation: Gives a range of values.
Example: Estimating average income of people in a city.
b) Hypothesis Testing
Used to test whether a statement is true or false.
Example: Checking whether a new teaching method improves student marks.
Types of Data Analysis – Detailed Simple Explanation
Data analysis means studying data to understand information and make decisions.
Based on the number of variables used, data analysis is divided into three types:
1. Univariate Analysis
Meaning:
“Uni” means one.
Univariate analysis studies only one variable at a time.
It helps to describe and summarize data.
Purpose:
● To understand one characteristic of data.
● To find average, highest, lowest, or spread of data.
Examples:
● Marks of students in a class.
● Age of employees.
● Salary of workers.
Tools Used:
● Mean (Average)
● Median
● Mode
● Range
● Charts and Graphs
2. Bivariate Analysis
Meaning:
“Bi” means two.
Bivariate analysis studies two variables together.
It helps to know whether one variable affects another.
Purpose:
● To find relationship between two variables.
● To understand cause and effect.
Examples:
● Study hours and exam marks.
● Advertising expense and sales.
● Price and demand.
Tools Used:
• Correlation
• Regression
• Scatter Diagram
3. Multivariate Analysis
Meaning:
“Multi” means many.
Multivariate analysis studies more than two variables at the same time.
Purpose:
● To understand complex relationships.
● To analyze many factors together.
Examples:
● Income, education, and spending habits.
● Sales affected by price, advertisement, and season.
● Student performance affected by study hours, attendance, and family
background.
Tools Used:
● Multiple Regression
● Factor Analysis
● MANOVA
Module 2
CORRELATION AND REGRESSION ANALYSIS
Meaning of Correlation
Correlation means the relationship between two variables. It shows whether changes
in one variable are associated with changes in another variable.
In simple words, correlation tells us how two things are connected.
Definition of Correlation
Correlation is a statistical technique used to measure the degree and direction of
relationship between two variables.
Assumptions of Correlation
1. Data should be normal
Most values should be near the average.
2. Relationship should be straight
The relationship between variables should form a straight line.
3. Variables should be measurable
Data must be numbers like marks, income, height, etc.
4. Equal number of values
Both variables should have the same number of observations.
5. No extreme values
Very high or very low values should be avoided.
6. Variables should be related
One variable should have some connection with the other.
Types of Correlation
Correlation shows the relationship between two variables. It is classified into
different types.
1. Positive and Negative Correlation
Positive Correlation
When both variables move in the same direction.
• If one increases, the other also increases.
• If one decreases, the other also decreases.
Examples:
●Income and Savings
● Advertisement and Sales
Example:
More income → More savings
Negative Correlation
When variables move in opposite directions.
● If one increases, the other decreases.
Examples:
●Price and Demand
● Speed and Time taken
Example:
Higher price → Lower demand
Zero Correlation
When there is no relationship between variables.
Example:
Shoe size and intelligence
2. Perfect and Imperfect Correlation
Perfect Correlation
When two variables change in exact proportion.
● Perfect Positive = +1
● Perfect Negative = –1
Example:
Temperature in Celsius and Fahrenheit.
Imperfect Correlation
When relationship exists but not exactly proportional.
Value lies between –1 and +1.
Example:
Income and expenditure.
3. Linear and Non-linear Correlation
Linear Correlation
When change in variables happens at a constant rate.
Graph forms a straight line.
Example:
Hours worked and wages earned.
Non-linear Correlation
When change is not constant.
Graph forms a curve.
Also called Curvilinear Correlation.
Example:
Stress and performance.
4. Simple, Partial and Multiple Correlation
Simple Correlation
Relationship between two variables only.
Examples:
● Height and Weight
● Price and Demand
Multiple Correlation
Relationship between more than two variables.
Example:
Sales affected by:
● Price
● Advertisement
● Income level
Partial Correlation
Relationship between two variables while keeping other variables constant.
Example:
Studying relation between demand and price while ignoring income.
Uses of Correlation
Correlation is useful in many areas:
[Link] economists study relationships like:
● Price and Demand
● Income and Consumption
[Link] businesses analyze:
● Advertisement and Sales
● Wages and Productivity
[Link] doctors study health conditions.
[Link] the strength of relationship between variables.
[Link] as a base for Regression Analysis.
Limitations of Correlation
1. Correlation shows relationship but not cause and effect.
2. It assumes a linear relationship.
3. Extreme values may affect results.
4. Sometimes correlation may be false (Spurious Correlation).
Spurious Correlation
A relationship appears but actually has no connection.
Example:
Height of people and amount of savings.
Methods of Measuring Correlation
Two main methods:
1. Graphic Methods
a) Correlation Graph
● Two curves are drawn.
● Same direction → Positive correlation
● Opposite direction → Negative correlation
c) Scatter Diagram
● Dots are plotted on a graph.
● Helps understand relationship visually.
2. Statistical Methods
1. Karl Pearson’s Correlation
2. Spearman’s Rank Correlation
3. Concurrent Deviation Method
Scatter Diagram
Scatter diagram is a graph showing relationship between two variables using dots.
● X-axis → One variable
● Y-axis → Another variable
Each pair of values is shown as one dot.
Interpretation of Scatter Diagram
Positive Correlation
Dots move upward from left to right.
Negative Correlation
Dots move downward from left to right.
No Correlation
Dots are spread randomly.
Perfect Positive
Dots form a straight upward line.
Perfect Negative
Dots form a straight downward line.
Advantages of Scatter Diagram
● Easy to draw
● Easy to understand
● Extreme values do not affect much
Disadvantages of Scatter Diagram
● Gives rough idea only
● Not mathematically exact
Karl Pearson’s Correlation Method
It is a statistical method used to calculate the degree of relationship between two
variables. Formula gives value of r
● r = +1 → Perfect Positive
● r = –1 → Perfect Negative
● r = 0 → No Correlation
EQUATIONS :
● DIRECT METHOD (IMPORTANT)
● ARITHMETIC MEAN METHOD
● ASSUMED MEAN METHOD
Probable Error (P.E.) – Definition
Probable Error is the amount by which the actual correlation coefficient is expected
to differ from the calculated correlation coefficient.
Probable Error is used to check accuracy of the correlation coefficient (r).
Formula of Probable Error
Where r = Correlation coefficient
n = Number of observations
0.6745 is constant
Interpretation of Probable Error
If r < P.E. → Correlation is not significant
If r > 6 × P.E. → Correlation is highly significant
If P.E. is small → Correlation is reliable
If P.E. is large → Correlation is less reliable
Standard Error
Standard Error is the standard deviation of a sample mean or estimate.
It measures the amount of variation between sample results.
Formula of Standard Error in Correlation
Where:
● r = Correlation coefficient
● N = Number of observations
● A small standard error means the estimate is more accurate.
● A large standard error means the estimate is less accurate.
Advantages of Standard Error
1. Helps Measure Accuracy
It shows how accurate the estimate is.
2. Helps Reduce Errors
It helps identify sampling and measurement errors.
3. Useful for Decision Making
Businesses and researchers use it to trust statistical results.
4. Improves Reliability
It helps judge whether data results are dependable.
Merits of Karl Pearson’s Coefficient of Correlation
1. Karl Pearson’s method shows whether correlation exists and also measures
the strength of relationship between variables.
2. It indicates the direction of correlation as positive or negative.
3. It helps in predicting the value of one variable from another variable.
4. It is useful for further mathematical and statistical analysis.
Demerits of Karl Pearson’s Coefficient of Correlation
1. Karl Pearson’s method is comparatively difficult to calculate.
2. The result may be affected by extreme values in the data.
3. It is based on certain assumptions which may not always be correct.
4. The results may sometimes be misinterpreted.
5. Time consuming
6. The method is subject to probable error, so reliability should be checked.
Spearman’s Rank Correlation
Definition
Spearman’s Rank Correlation is a statistical method used to measure the degree of
relationship between two ranked variables.
EQUATION :
Equation when ranks are repeated
Regression Analysis – Meaning
Regression Analysis is a statistical method used to find the relationship between
variables and predict one variable from another.
Definition
Regression Analysis is a technique that estimates the value of one variable
(dependent variable) based on the value of another variable (independent variable).
Difference Between Correlation and Regression
Correlation Regression
Correlation shows the relationship Regression predicts one variable using
between two variables. another variable.
It tells the direction and strength of It tells the effect of one variable on
relationship. another.
Correlation coefficient remains same Regression coefficients change when
even if variables are interchanged. variables are interchanged.
No need to identify dependent and Dependent and independent variables
independent variables. must be identified.
Correlation shows only relationship, Regression shows cause-and-effect
not cause and effect. relationship.
Correlation mainly studies linear Regression can study both linear and
relationship. non linear relationships
Correlation may sometimes be It avoid spurios relationship
spurious(false relationship)
Line of Regression – Meaning
A Line of Regression is a straight line that shows the relationship between two
variables and helps predict the value of one variable from another.
It is also called the best fit line.
Types of Regression Lines
1. Regression Line of Y on X
Used to predict Y based on X.
2. Regression Line of X on Y
Used to predict X based on Y.
Regression Coefficient – Meaning
A Regression Coefficient shows the amount of change in one variable due to a
change in another variable.
It tells the slope or direction of the regression line.
Types of Regression Coefficients
1. Regression Coefficient of Y on X (bᵧₓ)
Shows change in Y when X changes.
2. Regression Coefficient of X on Y (bₓᵧ)
Shows change in X when Y changes.
Methods of Regression Analysis
Regression analysis can be done mainly by two methods:
1. Graphic Method (Scatter Diagram)
2. Algebraic Method (Regression Equation)
1. Graphic Method – Scatter Diagram
A Scatter Diagram is a graph used to show the relationship between two variables.
It helps to draw a regression line (best fit line).
2. Algebraic Method – Regression Equation
This method uses a mathematical equation to show the relationship between
variables.
Regression Equation
Y = a + bX
Where:
● Y = Dependent variable
● X = Independent variable
● a = Constant (intercept)
● b = Regression coefficient (slope)
Properties of Regression Coefficient
1. Both regression coefficients have the same sign
If one coefficient is positive, the other will also be [Link] one is negative,
the other will also be negative.
2. Correlation and regression coefficients have the same sign
Positive correlation gives positive regression coefficients and vice versa.
3. Correlation is related to regression coefficients
Correlation coefficient is the geometric mean of regression coefficients.
4. If one regression coefficient is greater than 1, the other will be less than 1
Both cannot be greater than 1 at the same time.
5. Regression coefficient is not affected by origin
Changing the starting point of data does not affect it.
6. Average of regression coefficients is greater than or equal to correlation
coefficient.
7. If r = 0
There is no correlation, and regression lines become perpendicular.
8. If r = ±1
Regression lines overlap or become the same line.
9. Angle between regression lines shows relationship strength
Smaller angle → Stronger relationship.
Limitations of Regression Analysis
1. Calculation is lengthy and difficult
Regression involves many calculations.
2. Not suitable for qualitative data
It cannot measure qualities like honesty, beauty, or character.
3. Affected by extreme values
Very high or very low values can change results.
4. Cause and effect may not be clear
Regression assumes cause-and-effect relationship, but it may not always exist.
5. Assumes linear relationship
It assumes variables move in a straight-line relationship.
6. Relationship may change over time
Future conditions may differ from past data.
7. Limited data may give wrong prediction
Results based on small data may not always be accurate.
Coefficient of Determination (R²)
Coefficient of Determination (R²) shows how much variation in one variable is
explained by another variable.
It is the square of correlation coefficient (r²).
Module 3
SET THEORY
Set – Meaning
A Set is a collection of well-defined objects, numbers, or items grouped together.
These objects are called elements or members of the set.
Set Theory – Meaning
Set Theory is a branch of mathematics that studies sets and their relationships.
It deals with grouping objects and understanding how different sets are related.
Methods of Describing a Set
A set can be described in two methods:
1. Tabular Method (Roster Method)
2. Selector Method (Rule Method / Set Builder Method)
1. Tabular Method (Roster Method)
In this method, all elements of the set are written inside [Link] list all members
of the set one by one.
Example:
Set of vowels:
A= \{a, e, i, o, u\}
Set of even numbers below 10:
B= \{2, 4, 6, 8\}
Important Points:
● Elements are written inside { }
● Order does not matter.
● Repeating elements makes no difference.
2. Selector Method (Rule Method / Set Builder Method)
In this method, a rule or condition is used to describe the [Link] describe the set by
stating a common property of elements.
Example:
Set of even numbers less than 10:
A = \{x : x \text{ is an even number less than 10}\}
Meaning: A contains all values of x that are even and less than 10.
Types of sets
1. Finite Set
A set that has a limited number of elements.
Example:
A = {1, 2, 3, 4}
2. Infinite Set
A set that has unlimited elements.
Example:
Even numbers = {2, 4, 6, 8, …}
3. Empty Set (Null Set)
A set with no elements.
Symbol: ∅ or {}
Example:
Set of months with 32 days = ∅
4. Equal Sets
Two sets are equal when they contain exactly the same elements.
Example:
A= {1, 2, 3}
B= {3, 2, 1}
A=B
5. Equivalent Sets
Two sets are equivalent when they have the same number of elements.
Example:
A= {1, 2, 3}
B= {a, b, c}
Both have 3 elements.
6. Subset
A set is a subset if all its elements are included in another set.
Example:
A= {1, 2}
B= {1, 2, 3, 4}
A⊆B
7. Universal Set
The set containing all elements under discussion.
Example:
U = {1, 2, 3, 4, 5, 6}
8. Cardinality of a Set
Cardinality means number of elements in a set.
Example:
A = {2, 4, 6}
n(A) = 3
Operations on Sets
1. Intersection of Sets (A ∩ B)
Common elements present in both sets.
Example:
A= {1, 2, 3}
B= {2, 3, 4}
A ∩ B = {2, 3}
2. Union of Sets (A ∪ B)
All elements from both sets without repetition.
Example:
A= {1, 2}
B= {2, 3}
A ∪ B = {1, 2, 3}
3. Complement of a Set (Aᶜ)
Elements in universal set but not in A.
Example:
U = {1,2,3,4,5}
A = {1,2}
Aᶜ = {3,4,5}
4. Difference of Sets (A − B)
Elements in A but not in B.
Example:
A= {1,2,3,4}
B= {3,4}
A − B = {1,2}
5. Symmetric Difference (A Δ B)
Elements in either A or B but not common.
Example:
A= {1,2,3}
B= {3,4,5}
A Δ B = {1,2,4,5}
Venn Diagram
A Venn Diagram uses circles to represent sets and show relationships between them.
PROBABILITY
Basic Terms in Probability
1. Experiment
An experiment is an action that gives a result.
Examples:
● Tossing a coin
● Rolling a dice
● Drawing a card
2. Random Experiment
A random experiment gives an outcome that cannot be predicted exactly.
Examples:
● Tossing a coin → Head or Tail
● Rolling a die → 1 to 6
3. Trial
One performance of a random experiment.
Example:
Throwing a die once = One trial
4. Event
An event is the result of an experiment.
Examples:
● Getting Head in a coin toss
● Getting 4 on a die
5. Sample Space (S)
The sample space is the set of all possible outcomes.
Example:
Rolling a die:
S = \{1,2,3,4,5,6\}
6. Exhaustive Events
All possible outcomes together.
Example:
For a die → 1,2,3,4,5,6
7. Favorable Events
Outcomes that satisfy a condition.
Example:
Getting an even number on a die:
Favorable outcomes = {2,4,6}
8. Equally Likely Events
Events having the same chance of occurring.
Example:
Head or Tail in a coin toss.
9. Mutually Exclusive Events
Two events cannot happen together.
Example:
Getting Head and Tail in one toss.
10. Non-Mutually Exclusive Events
Two events can happen together.
Example:
Drawing a King and a Red card.
11. Simple Event
Event with only one outcome.
Example:
Getting 3 on a die.
12. Compound Event
Event with more than one outcome.
Example:
Getting an odd number = {1,3,5}
13. Complementary Event
Opposite of an event.
Example:
Head ↔ Not Head (Tail)
14. Dependent Event
One event depends on another.
Example:
Drawing cards without replacement.
15. Independent Event
One event does not affect another.
Example:
Tossing a coin twice.
16. Impossible Event
An event that cannot happen.
Probability = 0
Example:
Getting 7 on a six-faced die.
17. Odds of an Event
Ratio of favorable outcomes to unfavorable outcomes.
Approaches to Probability
1. Classical Approach
Probability based on equal chances.
Example:
Probability of Head = 1/2
2. Relative Frequency Approach
Probability based on past experience or data.
Example:
4 defective products in 100 items → Probability = 4%
3. Subjective Approach
Probability based on personal opinion or experience.
Example:
Doctor estimating success of surgery.
Theorems of Probability
1. Addition Theorem (OR Rule)
Used when finding probability of A or B.
Example:Probability of getting Head or Tail.
2. Multiplication Theorem (AND Rule)
Used when finding probability of A and B together.
Inverse Probability
Inverse Probability means finding the probability of a cause after knowing the result.
Bayes’ Theorem
Bayes’ Theorem is a method used to calculate inverse probability.
It helps revise probability using new information.
Combination
Combination means selecting objects without considering order.
Here, order is not important.
Permutation
Meaning
Permutation means arranging objects in a particular order.
Here, order is important.
Module 4
THEORETICAL DISTRIBUTION
Theoretical Distribution – Meaning
A Theoretical Distribution is a distribution based on mathematical rules and
probability rather than actual collected data.
Probability Distribution – Meaning
A Probability Distribution shows how probabilities are distributed among possible
outcomes of a random variable.
Random Variable – Meaning
A Random Variable is a variable whose value depends on the outcome of a random
experiment.
Types of Theoretical Probability Distribution
There are many theoretical distributions, but the three most common are:
1. Binomial Distribution
2. Poisson Distribution
3. Normal Distribution
Binomial Distribution
– Meaning
A binomial distribution is a type of probability distribution used when an experiment
has only two possible outcomes.
These outcomes are usually:
● Success or Failure
● Yes or No
● Pass or Fail
● Head or Tail
Simple Terms Related to Binomial Distribution
1. Bernoulli Trial
A Bernoulli trial is an experiment with only two possible outcomes.
Example:
● Head or Tail
● Pass or Fail
● Yes or No
[Link] distribution: A discrete probability distribution for a Bernoulli trial.
[Link] random variable: Random variable having a Bernoulli distribution.
[Link] experiment: An experiment involving repeated independent Bernoulli
trials.
Properties of Binomial Distribution
1. Binomial Distribution is a discrete probability distribution because it deals
with countable outcomes.
2. It is based on two parameters: n (number of trials) and p (probability of
success).
3. For n trials, the distribution has (n + 1) possible outcomes.
4. The mean of binomial distribution is np.
5. The variance is npq and the standard deviation is √npq.
6. The shape of the distribution changes when the probability value
changes.
7. A binomial distribution may have one peak or two peaks depending on
the values of n and p.
8. The variance is usually smaller than the mean.
9. The distribution is symmetric when p = 0.5.
Important Constants of Binomial Distribution
● Mean = np
● Variance = npq
● Standard Deviation = √npq
Poisson distribution
Poisson distribution is a type of probability distribution used to measure how many
times an event occurs in a fixed period of time, area, or [Link] is mainly used for
rare or uncommon events.
Properties of Poisson Distribution
1. Poisson Distribution is a discrete distribution because it counts events like 0,
1, 2, 3, etc.
2. It has no fixed upper limit, so the number of events can continue increasing.
3. It is mainly used for rare events, so the distribution is usually positively
skewed.
4. In Poisson distribution, mean and variance are equal.
5. It depends only on one value called mean (λ), so it is called a single-
parameter distribution.
6. If the mean is known, all probabilities can be calculated.
7. Events are assumed to happen randomly and evenly over time or space.
8. The probability of an event in a very small interval is very small.
9. As the mean increases, the distribution becomes less skewed.
10. The spread of the distribution increases when the mean increases.
Normal distribution
Normal distribution is a probability distribution in which data values are spread
evenly around an average [Link] forms a bell-shaped [Link] distribution is a
continuous probability distribution
Properties of Normal Distribution
1. The normal distribution is a smooth and continuous curve.
2. It has a bell shape and is perfectly symmetrical.
3. It has only one highest point, so it is called unimodal.
4. The highest point of the curve is at the mean, which is also equal to
median and mode.
5. Half of the data lies above the mean and half lies below the mean.
6. The curve looks the same on both sides of the mean.
7. The curve gets closer to the horizontal line but never touches it.
8. The total area under the curve is always equal to 1.
9. Standard deviation shows how far values spread from the mean.
10. Most values are close to the mean, and very few are far away.
11. It follows the Empirical Rule:
• About 68% values lie within 1 standard deviation
• About 95% values lie within 2 standard deviations
• About 99.7% values lie within 3 standard deviations.