Stat
Stat
Meaning of Statistics
Statistics is the branch of mathematics that deals with numerical data. It helps us to
understand, organize, and draw conclusions from data.
1. Collection of Data
This is the first step of statistics. It means gathering information from different sources.
2. Presentation of Data
• Diagrams
3. Analysis of Data
4. Interpretation of Data
This is the final step where we draw conclusions from the analyzed data.
Example: If average marks are low, we may conclude that students need better preparation.
Characteristics of Statistics
Statistics has some important characteristics that define its nature and scope. These are
discussed below:
1. Aggregates of Facts
• Example: It studies marks of all students in a class, not just one student.
So, individual data is not statistics, but a collection is.
2. Numerically Expressed
Statistics plays a very important role in modern life. It is used in almost every field for decision-
making, analysis, and planning. Its importance and uses are described below:
1. Importance of Statistics
Statistics provides facts and figures that help people and organizations make correct decisions.
Large and complex data can be summarized into simple forms using statistics.
2. Uses of Statistics
(i) In Economics
• Population census
(iv) In Education
• Educational research
• Hospital management
• Risk estimation
• Loan analysis
1. d Define with examples: (i) Population (ii) Sample (iii) Variable (iv) Constant
Here are the definitions with simple examples:
1. Population
Population means the entire group of people, objects, or observations about which we
want information.
Example:
All students of a university are a population.
2. Sample
A sample is a part or subset taken from the population for study.
Example:
100 students selected from the university for a survey are a sample.
3. Variable
A variable is a characteristic that can change or take different values.
Example:
Height, age, marks, and income are variables because their values differ from person to person.
4. Constant
A constant is a quantity or characteristic that remains fixed and does not change.
Example:
In a class, if every student studies in the same university, then the university name is a constant.
What is a Variable?
A variable is a characteristic or attribute that varies from one unit to another within a
population. For example, if you are studying a group of students, "height" is a variable because
it differs from person to person.
Classification of Variables
Variables are broadly classified into two categories based on how they are measured or
described:
1. Qualitative Variable
These describe attributes or qualities that cannot be expressed numerically. They represent
categories or labels.
2. Quantitative Variable
• Discrete Variable: A variable that takes on isolated or whole numbers (integers). There
are no values in between the points.
o Examples: Number of siblings, family size, or the number of road accidents per
day.
• Continuous Variable: A variable that can take any value within a specific range or limit,
including decimals and fractions.
Origin
The origin is the starting or reference point from which measurements are taken.
It indicates the zero point of a scale.
Example:
In Celsius temperature scale, 0°C is an origin.
Scale
Example:
In a ruler, each centimeter division represents a fixed scale.
Uses of Origin and Scale
1. Measurement of Data
They help in measuring and expressing numerical data properly.
2. Comparison
They make comparison between observations easy and meaningful.
4. Graphical Representation
Graphs and charts are drawn correctly by choosing suitable origin and scale.
5. Transformation of Data
They are used in changing units or simplifying calculations in statistics and mathematics.
1. Qualitative Data
This refers to information that cannot be measured in numerical form. It describes qualities,
attributes, or categories rather than quantities.
• Examples from the text: Religion, economic condition, color, and gender.
2. Quantitative Data
This refers to information that can be expressed in numerical form or numbers. It represents
counts or measurements.
• Examples from the text: Family size, population size, height, weight, and monthly
income.
• Key characteristic: It answers questions like "how many," "how much," or "how often."
Quick Comparison
Measurement scales are the systems or methods used to classify and measure data.
Graphical Representation
Graphical representation means showing data by graphs or charts so that it becomes easy to
understand.
Example:
Showing students’ marks by a graph.
1. Histogram
• Y-axis → frequencies
Example:
Marks: 0–10, 10–20, 20–30, etc.
Use:
Shows how data are distributed.
2. Frequency Polygon
It is made by joining the middle points of histogram bars with straight lines.
Use:
Easy to compare data.
3. Frequency Curve
Use:
Shows the general shape of data.
Two types:
Use:
Used to find median and quartiles.
Central tendency means the value that represents the center or typical value of a group of data.
It shows around which value most of the data are gathered.
• Mean
• Median
• Mode
The arithmetic mean is obtained by dividing the sum of all observations by the total number of
observations.
Formula
∑𝑥
𝑥ˉ =
𝑛
Where,
• x = observations
• n = number of observations
Example
For data: 2, 4, 6
Mean = (2 + 4 + 6)/3 = 4
Uses
• Easy to calculate
2. Median
Example
For data: 2, 4, 6, 8, 10
Median = 6
Uses
3. Mode
Example
For data: 2, 3, 3, 5, 7
Mode = 3
Uses
2.b What do you mean by measures of central tendency? Describe the various
measures of central tendency with their merits, demerits, and uses.
A Measure of Central Tendency is a single value that attempts to describe a set of data by
identifying the central position within that set of data. They are also known as measures of
location or center.
The Arithmetic Mean is the sum of all observations divided by the total number of observations.
• Demerits: Highly affected by outliers (extreme values); not suitable for qualitative data.
• Uses: Used in schools for average marks, economics for per capita income, and general
statistics.
The Geometric Mean is the $n^{th}$ root of the product of $n$ observations.
• Merits: Less affected by extreme values than AM; useful for calculating ratios and
percentages.
• Uses: Used to find the average growth rate (population growth) or compound interest.
The Harmonic Mean is the reciprocal of the arithmetic mean of the reciprocals of the data
values.
• Uses: Ideal for finding the average speed or rates involving time and distance.
4. Median
The Median is the middle value of a dataset when the observations are arranged in ascending or
descending order.
• Demerits: Does not use all observations (it only cares about the middle position); not
suitable for further algebraic treatment.
• Uses: Used for skewed data, such as calculating the median household income.
5. Mode
• Merits: Very easy to identify; the only measure that can be used for qualitative data
(e.g., favorite color).
• Demerits: A dataset can have no mode or multiple modes (bimodal/multimodal); it is
not based on all observations.
• Uses: Used in business and market research to identify the most popular product or
size.
2. c What are the essential characteristics of an ideal average? Which measure of central
tendency is the best and why?
1. Easy to Understand
It should be simple and clear.
2. Easy to Calculate
Its calculation should not be difficult.
5. Rigidly Defined
Its value should be definite and unique.
7. Stable in Sampling
Small changes in data should not change it greatly.
Reasons
However, when data contain extreme values, the median may be more suitable
2. e Describe the calculation method of median and mode from the graph.
Based on the images you provided from your textbook, here is a detailed description of how to
determine the Median and Mode graphically.
The median is found by plotting a Cumulative Frequency Curve, commonly known as an Ogive.
Step-by-Step Method:
1. Preparation: First, calculate the cumulative frequencies from the given frequency
distribution.
2. Plotting: * Plot the Upper Class Limits along the X-axis.
3. Drawing the Curve: Join all the plotted points with a smooth freehand curve to obtain
the Ogive curve (labeled AB in your image).
4. Finding the Median Position: Identify the point on the Y-axis that represents N/2 (where
N is the total number of observations).
5. Intersection: From the N/2 position on the Y-axis, draw a horizontal line parallel to the X-
axis until it intersects the Ogive curve at Point C.
6. Final Result: Draw a perpendicular line (CM) from Point C down to the X-axis. The
distance OM on the X-axis indicates the Median value
The mode is determined using a Histogram, specifically by looking at the highest rectangle
which represents the Modal Class.
Step-by-Step Method:
1. Draw the Histogram: Represent the frequency distribution as a histogram with class
intervals on the X-axis and frequency on the Y-axis.
2. Identify the Modal Class: Locate the highest rectangle in the histogram.
3. Drawing Connecting Lines: To pinpoint the mode within the highest bar:
o Draw a straight line from the top-left corner of the highest bar (A) to the top-left
corner of the next adjacent bar (C).
o Draw another straight line from the top-right corner of the highest bar (B) to the
top-right corner of the previous adjacent bar (D).
4. Find the Intersection: These two lines will intersect at a specific point (Q).
5. Final Result: Draw a perpendicular line from the intersection point Q down to the X-axis
(Point P). The value at point P (the abscissa) on the X-axis is the Mode.
Example:
Income distribution where a few people are extremely rich.
Example:
“50 and above”.
2. g Establish the relationship between arithmetic mean, geometric mean, and harmonic mean.
2. h When are the arithmetic mean, geometric mean, and harmonic mean equal?
3. a Discuss the concept of correlation. Write down the importance and uses of correlation.
Concept of Correlation
Correlation is a statistical tool used to measure the strength and direction of the relationship
between two variables. It determines how closely two variables move together.
Types:
• Positive: Both variables increase or decrease together (e.g., Height and Weight).
• Negative: One increases while the other decreases (e.g., Price and Demand).
1. Measuring Association: It quantifies how strongly two variables are linked, helping to
simplify complex data.
2. Basis for Prediction: It is the foundation for Regression Analysis, allowing us to estimate
unknown values based on known data.
3. Business Decisions: Helps managers understand the impact of specific factors, like how
advertising affects sales volume.
4. Policy Making: Governments use it to study social links, such as the correlation between
education levels and poverty.
6. Reliability Check: Helps in verifying the consistency and accuracy of collected data.
• Definition: When both variables move in the same direction at an equal rate.
• Value: r = +1
• Value: r = -1
• Definition: When variables move in opposite directions, but the rate of change is not
equal.
5. Zero Correlation
• Value: r = 0
3. c What is scatter diagram? Discuss the different nature of correlation with the help of scatter
diagram.
A Scatter Diagram (or Scatter Plot) is a graphical method used to study the relationship between
two variables. In this method, the values of two variables are plotted as points (dots) on a graph
sheet with an X-axis and a Y-axis. By observing the pattern or "scatter" of these points, we can
visually identify the nature and strength of the correlation between the variables.
Pearson’s coefficient of correlation measures the degree of linear relationship between two
variables. It is denoted by r.
−1 ≤ 𝑟 ≤ +1
Properties
• r = 0 → no correlation
5. Symmetric Property
𝑟𝑥𝑦 = 𝑟𝑦𝑥
Measures of Dispersion
Measures of dispersion are the statistical methods used to show how much the data values are
spread out or scattered around a central value (average).
In simple words, dispersion tells us whether the data are close together or widely spread.
• Range
• Quartile Deviation
• Mean Deviation
• Standard Deviation
2. To Compare Variability
Two datasets may have the same average but different spreads.
Measures of dispersion help compare their consistency.
3. To Control Variations
In manufacturing, business, and medicine, dispersion helps detect unwanted variations and
maintain quality.
4. Basis for Further Statistical Analysis
Many statistical methods like correlation, regression, and hypothesis testing depend on
measures of dispersion.
Dispersion helps us know whether the data are homogeneous (similar) or heterogeneous
(different).
3. g Describe the various measures of dispersion. Discuss the following measures of dispersion
indicating their relative merits and demerits:
(i) Range
Measures of Dispersion
Measures of dispersion show how much the data values are spread around the average.
1. Range
2. Mean Deviation
3. Standard Deviation
4. Quartile Deviation
(i) Range
𝑅𝑎𝑛𝑔𝑒 = 𝐿 − 𝑆
Where,
• L = largest value
• S = smallest value
Merits
• Easy to understand
Demerits
• Not reliable
Mean deviation is the average of the absolute deviations of observations from a central value.
Merits
• Easy to understand
Demerits
Standard deviation is the square root of the average of squared deviations from the arithmetic
mean.
∑(𝑥 − 𝑥ˉ )2
𝜎=√
𝑁
Merits
Demerits
Quartile deviation is half of the difference between the third quartile and first quartile.
𝑄3 − 𝑄1
𝑄. 𝐷. =
2
Merits
• Simple to calculate
Demerits
1. Independent of origin
Adding or subtracting the same number from all observations does not change S.D.
2. Dependent on scale
Multiplying or dividing all observations changes S.D. in the same ratio.
𝑀. 𝐷. ≤ 𝑆. 𝐷.
𝑛2 − 1
𝑆. 𝐷. = √
12
𝜎<𝑅
5. a Define the following terms with examples: experiment, random experiment, event,
composite event, probability.
1. Experiment
An experiment is an action that gives a result.
2. Random Experiment
A random experiment is an experiment where the result cannot be known before it
happens.
3. Event
An event is a result or group of results of a random experiment.
Example: Getting a number greater than 4 when rolling a die → {5, 6}.
5. Probability
Probability means the chance that an event will happen.
Its value is between 0 and 1.
Formula:
Favorable outcomes
𝑃(𝐸) =
Total outcomes
Example:
Probability of getting an Ace from 52 cards = 4/52 = 1/13.
5. b Write down the difference between a priori and a posteriori definition of probability.
Permutation and Combination are the two fundamental ways we count and arrange objects in
mathematics. The simplest way to distinguish them is to remember that in Permutations,
order matters, while in Combinations, order does not matter.
1. Permutation (Arrangement)
A permutation is an arrangement of objects in a specific order. If you change the order of the
objects, you create a new permutation.
Formula:
𝒏
𝒏!
𝑷𝒓 =
(𝒏 − 𝒓)!
Math Example:
How many ways can you arrange the letters in the word "CAT"?
Possible ways:
CAT, CTA, ACT, ATC, TCA, TAC
2. Combination (Selection)
A combination is a selection of objects where the order does not matter. We are only
interested in which items are picked, not the sequence in which they are picked.
Key Idea: Order is irrelevant.
Formula:
𝒏
𝒏!
𝑪𝒓 =
𝒓! (𝒏 − 𝒓)!
Math Example:
If you have 3 friends (A, B, and C) and you can only invite 2 of them to dinner, how many
groups can you form?
Possible groups:
{A, B}, {A, C}, {B, C}
Note that {A, B} is the same as {B, A}, so we only count it once. There are 3 combinations.
5. e Prove the additive law of probability with statement for two mutually disjoint events.
5. f State and prove the additive law of probability for two non-mutually exclusive events.
1. Classical Approach
Formula:
Favorable outcomes
𝑃(𝐸) =
Total possible outcomes
Example:
When a die is rolled, probability of getting 2 = 1/6.
2. Empirical (Statistical) Approach
Formula:
Number of times the event occurs
𝑃(𝐸) =
Total number of trials
Example:
If it rains on 20 days out of 100 days, then probability of rain = 20/100 = 0.2.
3. Axiomatic Approach
Main axioms:
Formula:
Example:
If P(A) = 0.3 and P(B) = 0.4, then
P(A ∪ B) = 0.7.
6. a Define the following terms with examples: mutually exclusive events, exhaustive events,
complementary events, sample space, favourable outcomes
Definition:
If one event happens, the other event cannot happen at the same time. These are called
mutually exclusive events.
Example:
When a coin is tossed, “Head” and “Tail” cannot come together.
So, Head and Tail are mutually exclusive events.
2. Exhaustive Events
Definition:
All possible outcomes of an experiment together are called exhaustive events.
Example:
When a die is thrown, the possible outcomes are:
{1, 2, 3, 4, 5, 6}
These six outcomes are exhaustive events.
3. Complementary Events
Definition:
If event A does not happen, it is called the complement of A.
It is written as Ā or Aᶜ.
𝑃(𝐴) + 𝑃(𝐴‾) = 1
Example:
If A = “Bangladesh wins the match”,
then Ā = “Bangladesh does not win the match”.
4. Sample Space
Definition:
The set of all possible outcomes of a random experiment is called the sample space.
It is usually written by S.
Example:
For throwing a die:
S = {1, 2, 3, 4, 5, 6}
Definition:
The outcomes which help an event happen are called favourable outcomes.
Example:
A box has 10 balls and 5 are red.
If you want a red ball, then the favourable outcomes are 5.
6.b Prove the multiplicative law of probability with statement for two independent events.
6. c Prove the multiplicative law of probability with statement for two dependent events.
6.d Prove the Bayes’ formula with statement.
Bayes’ Theorem খুব সহজে বুঝজে হজে আজে Conditional Probability বুঝজে হজব।
এবং
𝑃(𝐴 ∩ 𝐵)
𝑃(𝐵 ∣ 𝐴) =
𝑃(𝐴)
এখাজি দুজটাজেই 𝑃(𝐴 ∩ 𝐵)আজে।
𝑃(𝐵 ∣ 𝐴) × 𝑃(𝐴)
𝑃(𝐴 ∣ 𝐵) =
𝑃(𝐵)
6. e If A and B are two independent events then prove that 𝐴 and 𝐵 ̅ independent.
It is used for discrete random variables. It is used for continuous random variables.
It gives the probability of an exact value. It gives the density of probability over an interval.
Random Variable
Definition
A random variable is a variable that takes different numerical values according to the outcomes
of a random experiment.
It is usually represented by X, Y, Z etc.
Example
• If 1 comes, X = 1
• If 2 comes, X = 2
• …
• If 6 comes, X = 6
A random variable that takes countable values is called a discrete random variable.
Example
A random variable that can take any value within an interval is called a continuous random
variable.
Example
• Height of students
• Weight of a person
• Temperature
Definition
The expected value (or mean) of a random variable is the average value we expect from a
random experiment in the long run.
It is denoted by E(X).
𝐸(𝑋) = ∑𝑥𝑃(𝑥)
Where:
• P(x) = probability of x
Example
Suppose:
X P(X)
1 0.2
2 0.5
3 0.3
Then,
1. Expectation of a Constant
𝐸(𝑐) = 𝑐
Where c is a constant.
𝐸(𝑐𝑋) = 𝑐𝐸(𝑋)
3. Addition Rule
The expectation of the sum of two random variables equals the sum of their expectations.
The expectation of the difference of two random variables equals the difference of their
expectations.
𝐸(𝑋𝑌) = 𝐸(𝑋)𝐸(𝑌)
𝐸(𝑎𝑋 + 𝑏) = 𝑎𝐸(𝑋) + 𝑏
These are the basic rules and laws of expectation used in probability and statistics.
7. e Define the variance of the random variable. State the properties of variance of random
variable
Definition
Variance is a measure of dispersion or spread of a random variable from its mean (expected
value).
It shows how far the values are spread from the average value.
Formula of Variance
𝑉𝑎𝑟(𝑋) = 𝜎 2 ≈ 1.96
μ-σ+σVar(X) ≈ 1.96
Properties of Variance
𝑉𝑎𝑟(𝑐) = 0
where c is a constant.
𝑉𝑎𝑟(𝑋) ≥ 0
𝑉𝑎𝑟(𝑐𝑋) = 𝑐 2 𝑉𝑎𝑟(𝑋)
where c is a constant.
𝑉𝑎𝑟(𝑋 + 𝑐) = 𝑉𝑎𝑟(𝑋)
7. g What is a joint probability function? What conditions must a function satisfy to quality as a
joint probability function?
Definition
A joint probability function gives the probability of two discrete random variables occurring
together.
If X and Y are two random variables, then the joint probability function is written as:
𝑃(𝑋 = 𝑥, 𝑌 = 𝑦) = 𝑓(𝑥, 𝑦)
Example
Suppose:
Then,
1
𝑃(𝑋 = 1, 𝑌 = 2) =
36
because there are 36 possible outcomes when two dice are thrown.
A function must satisfy the following conditions to be a valid joint probability function:
1. Non-negative Condition
𝑓(𝑥, 𝑦) ≥ 0
∑ ∑ 𝑓(𝑥, 𝑦) = 1
𝑦
𝑥
These two conditions are necessary for a function to qualify as a joint probability function.
7. h What do you mean by moment generating function and state its uses, limitations and
properties?
Definition
The moment generating function (MGF) of a random variable is a function used to find the
moments such as mean, variance, etc. of a probability distribution.
It is denoted by Mₓ(t).
Formula of MGF
𝑀𝑋 (𝑡) = 𝐸(𝑒 𝑡𝑋 )
𝑀𝑋 (𝑡) = ∑𝑒 𝑡𝑥 𝑃(𝑥)
Uses of MGF
1. Finding Mean
2. Finding Variance
3. Identifying Distribution
4. Simplifying Calculations
Limitations of MGF
2. Difficult Calculation
Properties of MGF
1. Value at Zero
𝑀𝑋 (0) = 1
4. Unique Property